Quick answer: A voice AI pilot runs a pre-live build (4 weeks in the one vendor plan we use as a placeholder), then as long as your call volume takes to produce a readable result. Confirming a lift from 2% to 4% needs 2,274 leads. At 200 eligible leads a day, that is 2 weeks of live traffic and a decision around day 44. At 50 a day, it is 7 weeks and about day 79.
This page covers the pilot itself: the stages between contract signature and a go/no-go decision, which of them actually sets the date, and what makes each one slip. Two things it deliberately leaves to other pages. The US carrier critical path (A2P 10DLC brand and campaign review, toll-free verification) is on our US AI voice agent go-live timeline and its Registration Clock. How many test runs a scenario needs before you trust it is on our guide to testing an AI voice agent before go-live. Both feed into the clock below as inputs.
How long does a typical voice AI platform pilot take before go-live?
A pilot has two clocks. The first is engineering and approvals: scoping, CRM integration (HubSpot, Salesforce or whatever holds your pipeline), script build, testing and number registration. The second is evidence: how many live conversations the agent must handle before its difference from your current process stops being noise.
We call the combination the Pilot Clock:
Days from contract to decision = pre-live build + live-traffic days, rounded up to whole weeks + outcome lag
Live-traffic days = total leads the test needs ÷ eligible leads per day. It is rounded up to whole weeks so that every weekday appears in the pilot equally often; a pilot that runs Monday to Wednesday of its last week has over-sampled the start of the week. The outcome lag is how long after the call the outcome is known. For a booked meeting, we use the same illustrative lag as our post on sample sizes: up to 2 days.
For the pre-live build we use 28 days. That is not our number. It comes from one vendor’s own published implementation table: Assort Health, an inbound healthcare scheduling vendor, lays out illustrative timings in which discovery, AI customisation, EHR integration and testing run from week one to week four, with “go-live and improvement” in weeks five to six. An outbound sales deployment is a different job, so treat 28 days as a placeholder and replace it with your own vendor’s plan.
| Eligible leads per day | Confirm 2% to 3% (7,638 leads) | Confirm 2% to 4% (2,274 leads) | Confirm 2% to 8% (442 leads) | Longest stage when confirming 2% to 4% |
|---|---|---|---|---|
| 50 | 22 weeks live; decision day 184 | 7 weeks live; decision day 79 | 2 weeks live; decision day 44 | Live traffic (49 days against 28) |
| 100 | 11 weeks live; decision day 107 | 4 weeks live; decision day 58 | 1 week live; decision day 37 | Tie (28 days each) |
| 200 | 6 weeks live; decision day 72 | 2 weeks live; decision day 44 | 1 week live; decision day 37 | Pre-live build (28 days against 14) |
| 500 | 3 weeks live; decision day 51 | 1 week live; decision day 37 | 1 week live; decision day 37 | Pre-live build |
| 1,000 | 2 weeks live; decision day 44 | 1 week live; decision day 37 | 1 week live; decision day 37 | Pre-live build |
Each cell is 28 + (7 × whole weeks of live traffic) + 2. At 50 leads a day, 2,274 ÷ 50 = 45.5, so 46 days of traffic, which rounds up to 7 weeks (49 days); 28 + 49 + 2 = 79. The quotable version: for a pilot that must confirm a doubling from a 2% booking rate after a four-week build, below 82 eligible leads a day the pilot’s length is set by statistics, not by engineering.
The six stages from contract signature to a go/no-go decision
A voice AI pilot breaks into six stages. What differs between deployments is which ones overlap, and what has to be true before each can close.
| Stage | Exit condition | What makes it slip | Can run in parallel with | Assort Health’s illustrative timing (inbound healthcare scheduling) |
|---|---|---|---|---|
| 1. Scoping | One workflow, one success metric, the lift the pilot must detect, and the daily eligible volume, all written down | Nobody will commit to a lift threshold, so the sample size cannot be set | Security and procurement paperwork | Week one (discovery and workflow mapping) |
| 2. Integration and CRM | A test booking written back to the system of record on the same call, with the lead ID and script variant attached | Write access, API credentials and security review sit with a different team from the pilot sponsor | Script build, number registration | Weeks two to three (EHR/PMS integration) |
| 3. Script build | Compliance-locked lines (identification, consent, opt-out) signed off; variable lines drafted | Legal review of disclosure wording queues behind other work | Integration, number registration | Weeks one to two (AI customisation) |
| 4. Testing | Scenario suite passing at the run counts your test plan requires | Failures found in testing send work back to stage 2 or 3 | Number registration | Weeks three to four (testing against real data) |
| 5. Controlled live traffic | The pre-committed number of leads has been contacted in both arms | Eligible volume below forecast; a script change mid-pilot that restarts the count | Nothing: this stage is sequential | Weeks five to six (go-live and improvement, under monitoring) |
| 6. Read-out and decision | Outcomes for the last leads have arrived, and the result is compared with the threshold agreed in stage 1 | The threshold is renegotiated after the numbers are seen | Nothing | Not shown as a separate phase |
Two inputs sit outside the six stages because they are set by third parties, not by you or the vendor. In the US, if the pilot sends any SMS, the carrier registration queue runs in parallel with stages 2 to 4 and can end up being the longest pre-live item; the Registration Clock sets out that critical path. In the UK and Australia, the equivalent is regulatory-bundle review for the phone numbers, covered in our note on UK and AU number KYC bundle approval time. Start both on the day the contract is signed: nobody can speed them up by adding people.
Stages 4 and 5 answer different questions. A scenario that passes testing has not been measured on live traffic, and a pilot that passes stage 5 has not been tested against the edge cases in stage 4.
Which stage takes the longest? Below 82 leads a day, it is the live traffic
Stage 5 is the only stage whose length you can calculate in advance, and the only one nobody can shorten by working harder. It depends on the lift you need to confirm, your daily eligible volume, and your traffic split. Set live-traffic days equal to your pre-live build and solve for volume, and you get the crossover below which the live phase is the longest stage.
| Lift the pilot must confirm | Total leads needed | If pre-live build is 2 weeks | If pre-live build is 4 weeks | If pre-live build is 6 weeks |
|---|---|---|---|---|
| 2% to 3% (50% relative) | 7,638 | 546 a day | 273 a day | 182 a day |
| 2% to 4% (100% relative) | 2,274 | 163 a day | 82 a day | 55 a day |
| 2% to 8% (300% relative) | 442 | 32 a day | 16 a day | 11 a day |
Each cell is total leads ÷ build days, rounded up: 2,274 ÷ 28 = 81.2, so 82 a day. Read the middle row. If your pre-live build takes four weeks and the pilot has to confirm that the agent doubles a 2% booking rate, then with fewer than 82 eligible leads a day you will spend longer waiting for evidence than you spent building the agent.
This is also why published pilot durations disagree so widely: they rarely say which lift, at what volume, or whether stage 5 is included.
How many calls does a voice AI pilot need before its result means anything?
The total-leads column above comes from the standard large-sample formula for comparing two independent proportions, as set out by Hae-Young Kim in Statistical notes for clinical researchers: Sample size calculation 2 (Restorative Dentistry & Endodontics, 2016):
n₂ = (zα/2 + zβ)² ÷ (p₁ − p₂)² × [p₁(1 − p₁) ÷ κ + p₂(1 − p₂)], and n₁ = κ × n₂
Here p₁ is the control rate (your current process), p₂ is the agent’s rate you want to be able to detect, κ is the allocation ratio n₁ ÷ n₂, zα/2 = 1.96 for a two-sided test at 95% confidence and zβ = 0.84 for 80% power. Those are the paper’s values. Our post on why call volume alone won’t improve an AI sales agent uses the same formula for script tests. The formula’s per-arm figures below (1,137, 3,819 and 203) match that post, and we recomputed each one before reusing it; the Fisher’s exact figure of 221 per arm for 2% against 8% (step 4) is new to this page.
Step 1: check the formula against the source. Kim’s example, 20% against 30% with κ = 1: 7.84 × (0.16 + 0.21) ÷ 0.01 = 290.08. The paper prints 290.08.
Step 2: 2% against 4%, equal split. (1.96 + 0.84)² = 7.84. Variance terms: 0.02 × 0.98 + 0.04 × 0.96 = 0.0196 + 0.0384 = 0.058. Squared difference: 0.02² = 0.0004. So n = 7.84 × 0.058 ÷ 0.0004 = 1,136.8, rounded up to 1,137 per arm, or 2,274 in total.
Step 3: 2% against 3%. 7.84 × (0.0196 + 0.0291) ÷ 0.0001 = 3,818.08, so 3,819 per arm, or 7,638 in total.
Step 4: 2% against 8%, and why the formula is not enough here. The formula gives 7.84 × (0.0196 + 0.0736) ÷ 0.0036 = 202.97, so 203 per arm. But Kim states the large-sample approximation is appropriate when the sample is “large enough (e.g., n1p1 > 5, and n2p2 > 5)”, and recommends Fisher’s exact test when any expected cell is below 5. At 203 leads and 2%, the control arm expects 203 × 0.02 = 4.06 bookings. That is below the threshold. We therefore computed the exact power of a two-sided Fisher’s exact test by enumerating every possible pair of outcomes (no simulation). At 203 per arm it gives 75.9% power, not 80%. The first sample size that reaches 80% is 221 per arm, or 442 in total, and that is the figure used on this page.
Step 5: convert leads into calendar days. Divide total leads by eligible leads per day, round up to a whole day, then up to a whole week. At 100 a day: 2,274 ÷ 100 = 22.74, so 23 days, then 4 weeks.
Step 6: check the approximation. We also computed the exact power of the pooled two-proportion z-test at each designed sample size. At 50/50 splits the formula is well calibrated: 80.2% for 2% against 3%, and 80.8% for 2% against 4%.
This is the core of what we ran, in Python, in our working directory for this page. It prints 290.08, 1137, 3819 and 203:
from math import ceil
def n2(p1, p2, k=1.0, za=1.96, zb=0.84):
return (za + zb)**2 / (p1 - p2)**2 * (p1*(1 - p1)/k + p2*(1 - p2))
print(round(n2(0.20, 0.30), 2)) # 290.08, Kim's worked example
print(ceil(n2(0.02, 0.04))) # 1137 per arm
print(ceil(n2(0.02, 0.03))) # 3819 per arm
print(ceil(n2(0.02, 0.08))) # 203 per arm: below Kim's n*p > 5 rule, so use Fisher (221)
Substitute your own control rate and the smallest lift that would change your decision. The result, divided by your eligible daily volume, is the minimum length of stage 5.
Giving the agent a small share of traffic makes the pilot longer
Sending only 10% or 25% of leads to the agent lowers exposure. It also lengthens stage 5, because the smaller arm holds the test back.
Kim’s formula handles this through κ. With the agent on share s of traffic, κ = (1 − s) ÷ s. The formula is conservative for lopsided splits, though. When we computed exact power at its output, the 25% split gave 87.8% and the 10% split gave 90.8%, well above the 80% they were designed for. So we searched for the smallest agent arm that reaches 80% exact power at each split and report both.
| Agent’s share of traffic | κ | Total leads, Kim formula | Total leads, exact power search | Live days at 200 a day (exact) | Multiple of the 50/50 pilot’s leads |
|---|---|---|---|---|---|
| 50% | 1 | 2,274 | 2,274 (formula gives 80.8%) | 12 | 1.0 |
| 25% | 3 | 3,524 | 2,804 (701 agent, 2,103 control) | 15 | 1.2 |
| 10% | 9 | 7,960 | 5,520 (552 agent, 4,968 control) | 28 | 2.4 |
Arithmetic for the 10% row: 5,520 ÷ 200 = 27.6, so 28 days; 5,520 ÷ 2,274 = 2.43. For a pilot confirming 2% to 8% with Fisher’s exact test, the same search gives 504 total leads at a 25% share and 960 at 10%, against 442 at 50/50.
The decision rule: a 10% agent share costs about 2.4 times the live-traffic calendar of a 50/50 split for the same answer. If exposure is the worry, a shorter 50/50 pilot on a narrow, low-risk segment usually beats a long 10% pilot on everything.
What a two-week monitored window can and cannot tell you
Many go-live plans end with a short monitored period. That is right for operations: it catches misrouted transfers, broken write-backs and scenarios the test suite missed. It is not a measurement. Turn the formula around: for a fixed window, what rate must the agent reach before its difference from a 2% control is detectable?
| Eligible leads per day | 2-week window | 4-week window | 8-week window |
|---|---|---|---|
| 50 | 6.17% (350 per arm) | 4.68% (700 per arm) | 3.77% (1,400 per arm) |
| 100 | 4.68% (700 per arm) | 3.77% (1,400 per arm) | 3.19% (2,800 per arm) |
| 200 | 3.77% (1,400 per arm) | 3.19% (2,800 per arm) | 2.81% (5,600 per arm) |
| 500 | 3.05% (3,500 per arm) | 2.72% (7,000 per arm) | 2.50% (14,000 per arm) |
| 1,000 | 2.72% (7,000 per arm) | 2.50% (14,000 per arm) | 2.34% (28,000 per arm) |
Per arm = leads per day × days ÷ 2; the rate is the p₂ at which Kim’s formula returns exactly that per-arm number. At 100 leads a day for two weeks, the agent would need to book at 4.68%, more than double the control, before the pilot could tell it apart from chance. A real improvement from 2% to 3% in that window has 22.4% power (exact z-test enumeration), so about 78% of the time it would be reported as “no significant difference” and could kill a good deployment. The 200-a-day, four-week cell (3.19%) matches the figure in our call-volume post, which is a useful check that the two pages use the same arithmetic.
What’s a realistic go-live date from contract signature for a mid-size voice AI deployment?
“Mid-size” has no standard definition. For this worked example we take it to mean 100 to 500 eligible leads a day, which is our working assumption, not an industry category. We also take a pilot that must confirm the agent at least doubles a 2% booking rate. With the 28-day pre-live placeholder and a 50/50 split:
- Day 0: contract signed. Start number registration (US) or regulatory bundles (UK and AU), security review and CRM access requests the same day.
- Days 1 to 28: scoping, integration, script build and testing, overlapping where the stage table allows.
- Day 29: controlled live traffic starts.
- At 100 a day: 2,274 leads take 23 days, rounded up to 4 weeks. Traffic ends day 56, the last outcomes land by day 58, and the go/no-go meeting can be held from day 58, about week 9.
- At 200 a day: 12 days, rounded up to 2 weeks. Decision from day 44, about week 7.
- At 500 a day: 5 days, rounded up to 1 week. Decision from day 37, about week 6.
So for a mid-size deployment, a realistic go/no-go date is day 37 to day 58 after contract signature, in week 6 to week 9, when the pilot has to confirm a doubling. Full go-live follows that decision. If the bar is a 50% lift (2% to 3%), the same volumes give decisions from day 107 (100 a day), day 72 (200 a day) and day 51 (500 a day). None of this includes slippage in the pre-live build, and that is where the next section starts.
What makes a voice AI pilot slip
In order of where they hit the calendar:
Approvals outside the pilot team. Moveworks, an enterprise AI-agent vendor, puts it plainly on its own blog: “Unclear goals, conflicting timelines, or stalled approvals from legal, procurement, or security can stretch a four-week rollout into four months.” Start them on day 0, alongside scoping, rather than when integration asks for access.
Write access to the system of record. An agent that can read the CRM but not write back produces a pilot whose outcomes live in a spreadsheet. Stage 6 then becomes a reconciliation project. The exit condition in the stage table (a test booking written back on the same call, with the lead ID and variant attached) is the check that prevents it.
Eligible volume below forecast. In stage 5, slip is proportional to the shortfall. Consent flags, do-not-call scrubbing, business-hours windows and segment filters all shrink the list the pilot may contact. If the plan assumed 200 a day and 120 are eligible, a 2-week live phase becomes 19 days (2,274 ÷ 120 = 18.95), which rounds up to 3 weeks. Measure eligible volume in stage 1, after the filters, not before.
Changing the script mid-pilot. A pilot measures one agent configuration. Changing the opener in week two means the leads from week one no longer measure the thing you are now running, and the count restarts. Save changes for after the read-out.
Reading the result early. Stopping the moment the agent looks ahead inflates the false-positive rate above the 5% the test was designed for. Fix the lead count in stage 1 and do not act before it is reached.
Moving the threshold after the numbers arrive. A threshold agreed after the fact slides towards whatever the pilot achieved. Writing it into the stage 1 scope keeps stage 6 short; our Discover, Deploy, Scale AI agent playbook covers setting that success gate before the pilot starts.
What this means when you compare platforms
The lift a pilot is built to confirm is the biggest lever on its length, so put this to any vendor, including us: what lift is this pilot sized to detect, at what eligible daily volume, and on what split? A bare number of weeks does not tell you whether stage 5 is in it.
For scale: Zian has taken accounts from roughly 2% conversion to around 8%. That is Zian’s own first-party figure, not an independent measurement. A lift of that size is the 2% to 8% column of the Pilot Clock table: 442 leads, one or two weeks of live traffic at 50 or more eligible leads a day, and a pilot whose length is set almost entirely by the pre-live build. A smaller expected lift puts you in the 2% to 4% or 2% to 3% column, with a longer stage 5. Zian is in partnership-application beta, and PrecisionPitch AI™ keeps split-testing scripts after go-live, the method set out in our guide to split-testing sales scripts with AI. Each of those later tests carries the same arithmetic, which is why the pilot should be sized to answer one question, cleanly, rather than several at once.
Running the pilot yourself needs the method on this page, plus a shared key between your voice platform and CRM, a variant ID on every call, and someone who holds the line on stage 5 when the first week looks good.
Frequently asked questions
How long does a typical voice AI platform pilot take before go-live?
A pre-live build, four weeks in the one published vendor plan used here as a placeholder, then a live-traffic phase whose length depends on call volume. To confirm a lift from 2% to 4% on a 50/50 split you need 2,274 leads. That is 2 weeks at 200 eligible leads a day and 7 weeks at 50 a day, so a go/no-go decision around day 44 or day 79.
Which stage of a voice AI pilot takes the longest?
It depends on volume. With a four-week build and a pilot that must confirm a doubling from 2% to 4%, the live-traffic phase is the longest stage below 82 eligible leads a day. From 82 to 108 a day the two take the same four weeks, once live traffic is rounded to whole weeks; above 108, the build and approvals take longer. For a 50% lift from 2% to 3%, the crossover is 273 leads a day.
How many calls does a voice AI pilot need?
For 95% confidence and 80% power, the two-proportion formula in Hae-Young Kim’s 2016 statistical note gives 1,137 leads per arm to tell 2% from 4% and 3,819 per arm to tell 2% from 3%. For 2% against 8%, the control arm expects fewer than 5 bookings, so Kim recommends Fisher’s exact test, which needs 221 per arm.
Can I run the pilot on 10% of my leads to reduce risk?
You can, but the pilot takes longer. For 2% against 4%, a 10% agent share needs 5,520 leads in total against 2,274 for a 50/50 split, about 2.4 times the live-traffic calendar. A shorter 50/50 pilot on a narrow, low-risk segment usually gives the same answer sooner.
Is a two-week monitored launch long enough to prove the agent works?
It is long enough to catch operational faults, not to measure a lift at modest volume. At 100 eligible leads a day for two weeks on a 50/50 split, the agent would have to book at 4.68% against a 2% control before the difference was detectable.
What delays a voice AI pilot most often?
Before go-live: approvals from legal, procurement or security, CRM write access, and third-party number registration. During the pilot: fewer eligible leads than forecast, script changes that restart the count, and reading the result early. Starting approvals and registration on the day the contract is signed removes most of the pre-live slip.
Where every figure on this page comes from
| Figure | Who published it | Link | Date read |
|---|---|---|---|
| Two-proportion sample-size formula with allocation ratio κ; z = 1.96 and z = 0.84; worked example 290.08 for 20% against 30%; large-sample condition n₁p₁ > 5 and n₂p₂ > 5, and Fisher’s exact test for small cells | Hae-Young Kim, Restorative Dentistry & Endodontics 2016;41(2):154–156 | rde.ac | 2026-09-26 |
| Illustrative phase timings: discovery week one, AI customisation weeks one to two, EHR/PMS integration weeks two to three, testing weeks three to four, go-live and improvement weeks five to six; “Assort Health go-live runs 30 to 45 days” (the basis of this page’s 28-day pre-live placeholder) | Assort Health (page metadata dated August 2026) | assorthealth.com | 2026-09-26 |
| “Unclear goals, conflicting timelines, or stalled approvals from legal, procurement, or security can stretch a four-week rollout into four months.” | Moveworks (page dated 6 August 2025) | moveworks.com | 2026-09-26 |
| 1,137 / 3,819 / 203 per arm; 221 per arm by Fisher’s exact test (75.9% power at 203); exact z-test power 80.2%, 80.8%, 87.8%, 90.8% and 22.4%; 2,804 and 5,520 total leads at 25% and 10% agent shares; 504 and 960 for 2% to 8%; all Pilot Clock, crossover and detectable-rate cells | This page’s own calculation from Kim’s formula, run in Python on 2026-09-26 (exact binomial enumeration for power, no simulation) | Formula and code shown on this page | 2026-09-26 |
| 2% control rate, 2-day outcome lag, 28-day pre-live build, 100 to 500 leads a day as “mid-size”, the daily volumes in every table | Illustrative assumptions chosen for this page, not measurements | Not applicable | 2026-09-26 |
| Roughly 2% conversion to around 8% | Zian AI, first-party figure released by the company | First-party; no external URL | 2026-09-26 |
Want the pilot sized before the contract, not after it? Apply For Partnership