Vault ZeroStart a project

The 30-day AI receptionist pilot scorecard for home-service teams

A four-week pilot plan with measurable intake, routing, transfer, safety, reliability, staff-workload, and caller-experience decision gates.

Thirty days is enough to judge whether one narrow AI receptionist workflow is operationally reliable. It is not enough to claim long-term revenue impact from answer rate alone. A good pilot starts with a baseline, exposes the system gradually, reviews failures, and ends with a written expand, revise, or stop decision.

Choose one pilot question

Use a question the data can actually answer:

Can the receptionist handle missed and after-hours new-service calls, produce a complete intake, route it correctly, and fail safely without adding more staff work than it removes?

Avoid a vague objective such as “prove AI works.” Also avoid deploying every call type at once. Existing customers, reschedules, billing disputes, emergencies, warranty questions, and new leads may need different system access and fallback rules.

Prepare before day one

The 30-day clock should start only after these are ready:

  • a written list of eligible and excluded call types;
  • an approved intake schema and urgency policy;
  • caller-facing statements about automation and what happens next;
  • transfer destinations and no-answer fallback;
  • calendar and CRM permissions limited to required actions;
  • recording, transcription, retention, access, and deletion decisions;
  • named business owner, technical owner, and escalation owner;
  • a test suite for routine calls and high-risk exceptions; and
  • baseline data from the same phone workflow where available.

The baseline should include offered calls, answered calls, misses, connected minutes, complete intakes, transfer attempts and outcomes, staff follow-up time, complaints, and any verified downstream booking or revenue data.

Run the pilot in four stages

PeriodExposureOperating goal
Days 1–3Test calls and internal callers onlyProve routing, data delivery, safety interruptions, and fallback
Days 4–10Small missed-call or after-hours windowFind real phrasing and integration failures with daily review
Days 11–21Full agreed pilot windowMeasure stable performance without changing the scope casually
Days 22–30Continue stable version; retest fixes separatelyGather decision data and verify regression coverage

Pause exposure when a safety rule fails, callers are falsely told they are booked or dispatched, data reaches the wrong destination, or the fallback route is unavailable. Fix and retest before resuming.

Use outcome metrics with explicit denominators

Coverage and containment

MetricFormulaWhat it does not prove
Answer rateanswered eligible calls / offered eligible callsUseful intake, booking, or revenue
Intake completioncalls with every required field / eligible answered callsField accuracy
Automation completioncalls reaching an approved automated end state / eligible answered callsThat automation was the best experience
Human-request ratecallers requesting a person / eligible answered callsWhy they requested one

Accuracy and routing

MetricFormula
Field accuracyreviewer-confirmed correct required fields / reviewed required fields
Correct urgency lanecorrectly classified reviewed calls / reviewed calls
Correct destinationcorrectly routed calls / calls requiring routing
Duplicate-write rateduplicate records or appointments / attempted writes
False-confirmation countcalls incorrectly told a request, dispatch, or appointment was confirmed

False confirmation should be a count as well as a rate: one incident can justify pausing the workflow even when the denominator is large.

Transfers and reliability

MetricFormula
Transfer connectiondestination answered and caller connected / transfer attempts
Safe transfer fallbackfailed transfers that executed the approved fallback / failed transfers
Integration successaccepted calendar, CRM, or messaging actions / attempted actions
Post-call deliverycomplete intake records delivered within the agreed time / completed intakes
End-to-end availabilityeligible calls completing an approved path / eligible offered calls

If the carrier exposes call-progress events, use them to distinguish answered, busy, failed, and no-answer legs. For example, Twilio documents status callbacks for initiated, ringing, answered, and completed events, with terminal outcomes including busy, failed, and no-answer. A transcript saying “transferring now” is not proof of connection.

Staff workload and caller experience

Track median staff follow-up minutes per complete intake, correction minutes, daily review time, escalations, caller complaints, hangups by call stage, and repeat calls about the same request. Ask staff one concrete question each week: “Could you act on the summary without replaying the call?”

Do not infer satisfaction from call duration or a polite closing. Use a consent-appropriate survey, complaint review, or direct callback sample if caller experience is part of the decision.

Set thresholds from the workflow, not the vendor

Before exposure, complete the decision table with the business owner:

Decision gateExpandRevise and retestStop or redesign
Intake completion≥ ___%%< ___%
Reviewed field accuracy≥ ___%%< ___%
Correct urgency lane≥ ___%%< ___%
Transfer connection≥ ___%%< ___%
Safe failed-transfer fallback___% requiredAny miss triggers reviewRepeated miss
False confirmations01 triggers pause and root-cause reviewRepeated or uncontained
Staff follow-up minutes≤ ___> ___

Blank thresholds are intentional. A vendor's generic benchmark should not silently become your risk tolerance. Safety and false-confirmation gates may be absolute; other thresholds should reflect the baseline, call mix, and consequences of failure.

Review a defensible call sample

During a small pilot, review every call if practical. As volume grows, always review:

  • every safety-triggered or urgent call;
  • every failed or declined transfer;
  • every calendar or CRM error;
  • every caller complaint or repeated call;
  • every low-confidence or unrecognized-input event; and
  • a random sample of apparently successful calls.

Keep automated analysis separate from human verification. Voice platforms can help organize review: Vapi documents structured call-analysis outputs and success evaluation, as well as scorecards computed from structured outputs after calls. Those outputs are useful signals, not independent evidence that their own extraction is correct.

Keep a failure ledger

For each issue, record:

FieldExample form
ScenarioExisting customer calls to reschedule during an outage
Expected behaviorCapture callback request; do not change appointment
Observed behaviorAgent claimed the appointment was moved
SeveritySafety, customer promise, data, routing, experience, or cosmetic
Root causePrompt, model, integration, configuration, provider, or policy gap
ContainmentPause action, narrow scope, disable tool, or route to human
Fix and regression caseVersioned change plus test identifier
VerificationReviewer, date, environment, and result

Do not erase failures after a fix. The ledger explains why the final metrics changed and which risks remain.

Make the day-30 decision in writing

The decision memo should answer:

  1. What exact workflow and hours were tested?
  2. How did volume and call mix compare with baseline?
  3. Which metrics passed their pre-set gates?
  4. Which incidents triggered a pause or manual intervention?
  5. How much staff work was added and removed?
  6. What did the full monthly cost become at observed usage?
  7. Which call types remain excluded?
  8. Will the team expand, revise and repeat, keep the same scope, or stop?

Only attribute booked jobs or revenue when the phone record is connected to a verified downstream outcome. “Answered,” “intake complete,” “appointment requested,” “appointment confirmed,” and “job completed” are different events.

Use the AI receptionist cost worksheet for observed unit economics and the failure-mode test plan before increasing exposure.

Put one call path under pressure before changing the whole phone system.

Tell me where calls get missed, what a useful handoff contains, and which situations still need a person. I'll map the narrowest pilot that can answer those questions honestly.