The 30-day AI receptionist pilot scorecard for home-service teams
A four-week pilot plan with measurable intake, routing, transfer, safety, reliability, staff-workload, and caller-experience decision gates.
Field note
By Vault Zero
Thirty days is enough to judge whether one narrow AI receptionist workflow is operationally reliable. It is not enough to claim long-term revenue impact from answer rate alone. A good pilot starts with a baseline, exposes the system gradually, reviews failures, and ends with a written expand, revise, or stop decision.
Choose one pilot question
Use a question the data can actually answer:
Can the receptionist handle missed and after-hours new-service calls, produce a complete intake, route it correctly, and fail safely without adding more staff work than it removes?
Avoid a vague objective such as “prove AI works.” Also avoid deploying every call type at once. Existing customers, reschedules, billing disputes, emergencies, warranty questions, and new leads may need different system access and fallback rules.
Prepare before day one
The 30-day clock should start only after these are ready:
- a written list of eligible and excluded call types;
- an approved intake schema and urgency policy;
- caller-facing statements about automation and what happens next;
- transfer destinations and no-answer fallback;
- calendar and CRM permissions limited to required actions;
- recording, transcription, retention, access, and deletion decisions;
- named business owner, technical owner, and escalation owner;
- a test suite for routine calls and high-risk exceptions; and
- baseline data from the same phone workflow where available.
The baseline should include offered calls, answered calls, misses, connected minutes, complete intakes, transfer attempts and outcomes, staff follow-up time, complaints, and any verified downstream booking or revenue data.
Run the pilot in four stages
| Period | Exposure | Operating goal |
|---|---|---|
| Days 1–3 | Test calls and internal callers only | Prove routing, data delivery, safety interruptions, and fallback |
| Days 4–10 | Small missed-call or after-hours window | Find real phrasing and integration failures with daily review |
| Days 11–21 | Full agreed pilot window | Measure stable performance without changing the scope casually |
| Days 22–30 | Continue stable version; retest fixes separately | Gather decision data and verify regression coverage |
Pause exposure when a safety rule fails, callers are falsely told they are booked or dispatched, data reaches the wrong destination, or the fallback route is unavailable. Fix and retest before resuming.
Use outcome metrics with explicit denominators
Coverage and containment
| Metric | Formula | What it does not prove |
|---|---|---|
| Answer rate | answered eligible calls / offered eligible calls | Useful intake, booking, or revenue |
| Intake completion | calls with every required field / eligible answered calls | Field accuracy |
| Automation completion | calls reaching an approved automated end state / eligible answered calls | That automation was the best experience |
| Human-request rate | callers requesting a person / eligible answered calls | Why they requested one |
Accuracy and routing
| Metric | Formula |
|---|---|
| Field accuracy | reviewer-confirmed correct required fields / reviewed required fields |
| Correct urgency lane | correctly classified reviewed calls / reviewed calls |
| Correct destination | correctly routed calls / calls requiring routing |
| Duplicate-write rate | duplicate records or appointments / attempted writes |
| False-confirmation count | calls incorrectly told a request, dispatch, or appointment was confirmed |
False confirmation should be a count as well as a rate: one incident can justify pausing the workflow even when the denominator is large.
Transfers and reliability
| Metric | Formula |
|---|---|
| Transfer connection | destination answered and caller connected / transfer attempts |
| Safe transfer fallback | failed transfers that executed the approved fallback / failed transfers |
| Integration success | accepted calendar, CRM, or messaging actions / attempted actions |
| Post-call delivery | complete intake records delivered within the agreed time / completed intakes |
| End-to-end availability | eligible calls completing an approved path / eligible offered calls |
If the carrier exposes call-progress events, use them to distinguish answered, busy, failed, and no-answer legs. For example, Twilio documents status callbacks for initiated, ringing, answered, and completed events, with terminal outcomes including busy, failed, and no-answer. A transcript saying “transferring now” is not proof of connection.
Staff workload and caller experience
Track median staff follow-up minutes per complete intake, correction minutes, daily review time, escalations, caller complaints, hangups by call stage, and repeat calls about the same request. Ask staff one concrete question each week: “Could you act on the summary without replaying the call?”
Do not infer satisfaction from call duration or a polite closing. Use a consent-appropriate survey, complaint review, or direct callback sample if caller experience is part of the decision.
Set thresholds from the workflow, not the vendor
Before exposure, complete the decision table with the business owner:
| Decision gate | Expand | Revise and retest | Stop or redesign |
|---|---|---|---|
| Intake completion | ≥ ___% | –% | < ___% |
| Reviewed field accuracy | ≥ ___% | –% | < ___% |
| Correct urgency lane | ≥ ___% | –% | < ___% |
| Transfer connection | ≥ ___% | –% | < ___% |
| Safe failed-transfer fallback | ___% required | Any miss triggers review | Repeated miss |
| False confirmations | 0 | 1 triggers pause and root-cause review | Repeated or uncontained |
| Staff follow-up minutes | ≤ ___ | – | > ___ |
Blank thresholds are intentional. A vendor's generic benchmark should not silently become your risk tolerance. Safety and false-confirmation gates may be absolute; other thresholds should reflect the baseline, call mix, and consequences of failure.
Review a defensible call sample
During a small pilot, review every call if practical. As volume grows, always review:
- every safety-triggered or urgent call;
- every failed or declined transfer;
- every calendar or CRM error;
- every caller complaint or repeated call;
- every low-confidence or unrecognized-input event; and
- a random sample of apparently successful calls.
Keep automated analysis separate from human verification. Voice platforms can help organize review: Vapi documents structured call-analysis outputs and success evaluation, as well as scorecards computed from structured outputs after calls. Those outputs are useful signals, not independent evidence that their own extraction is correct.
Keep a failure ledger
For each issue, record:
| Field | Example form |
|---|---|
| Scenario | Existing customer calls to reschedule during an outage |
| Expected behavior | Capture callback request; do not change appointment |
| Observed behavior | Agent claimed the appointment was moved |
| Severity | Safety, customer promise, data, routing, experience, or cosmetic |
| Root cause | Prompt, model, integration, configuration, provider, or policy gap |
| Containment | Pause action, narrow scope, disable tool, or route to human |
| Fix and regression case | Versioned change plus test identifier |
| Verification | Reviewer, date, environment, and result |
Do not erase failures after a fix. The ledger explains why the final metrics changed and which risks remain.
Make the day-30 decision in writing
The decision memo should answer:
- What exact workflow and hours were tested?
- How did volume and call mix compare with baseline?
- Which metrics passed their pre-set gates?
- Which incidents triggered a pause or manual intervention?
- How much staff work was added and removed?
- What did the full monthly cost become at observed usage?
- Which call types remain excluded?
- Will the team expand, revise and repeat, keep the same scope, or stop?
Only attribute booked jobs or revenue when the phone record is connected to a verified downstream outcome. “Answered,” “intake complete,” “appointment requested,” “appointment confirmed,” and “job completed” are different events.
Use the AI receptionist cost worksheet for observed unit economics and the failure-mode test plan before increasing exposure.
Put one call path under pressure before changing the whole phone system.
Tell me where calls get missed, what a useful handoff contains, and which situations still need a person. I'll map the narrowest pilot that can answer those questions honestly.