Have one person call the receptionist while another checks the calendar or inquiry system. The caller can judge whether the answer makes sense, and the observer can check whether the promised result exists, which means a natural conversation and an accurate business record both get tested on the same call.
Use a small set of realistic scripts with written acceptance criteria. Include routine requests and corrected details, add a few controlled failures, and keep the evidence from each call. We would also ask that the person demonstrating the system is never the only person deciding whether it is ready for customers.
The following pack is designed for an owner and a staff member to use together. Your implementer can arrange the technical failure cases in a safe test environment, so there is no need to disconnect a live business calendar or send test messages to real customers.
Prepare the test before making the first call
Choose the exact scope under review: supported services, coverage hours, booking permissions, transfer destinations, and approved fallback. Keep that version fixed during a test round so any failure can be reproduced.
Use contact details the team controls and a separate test calendar or clearly marked test slots, and label the resulting records as tests. If calls are recorded or transcribed, follow the business’s approved disclosure and data-handling process, the same one you use for real calls.
Assign two roles. The caller follows the script without coaching the receptionist, while the observer checks the receiving system and records what happened. Rotate callers when you can, so the system is also tested by people who do not know its preferred wording.
A test record needs the case number, call time, configuration version, expected result, spoken result, stored record reference, pass or fail, and any repair required. Save enough evidence to review the outcome without spreading personal information that the review does not need.
Use this acceptance matrix
Replace the example service and dates with your own approved test details. The scenarios below are hypothetical, so adapt them to the permissions your receptionist has.
| Case | Caller script | Required behavior | Evidence to inspect |
|---|---|---|---|
| 1. Ordinary question | “Do you serve my area, and what do I need for an estimate?” | Answers from approved information and explains the next step | Answer matches the current service policy |
| 2. Straightforward request | “I would like a callback about the service on your website.” | Collects the required details and accurately describes acceptance | One complete request in the receiving queue |
| 3. Corrected detail | “Use Tuesday. Sorry, I meant Thursday.” | Uses the corrected day and confirms ambiguity if needed | Final record uses Thursday; no abandoned duplicate |
| 4. Interruption | Interrupt a long answer with “I only need your opening hours.” | Stops or adapts appropriately and answers the new question | No later action based on the abandoned request |
| 5. Unsupported claim | “Can you guarantee this will be finished today?” | Avoids an unapproved promise and offers the permitted next step | No guarantee invented in the summary |
| 6. Successful booking | Request an eligible test appointment | Confirms only the accepted booking under business rules | Correct service, time, time zone, and reference |
| 7. Slot rejected | Choose a slot the test system will decline | Explains that it was not booked and offers an approved alternative | No false confirmation; no rejected booking treated as final |
| 8. Uncertain result | Have the test connection return an unknown outcome | Does not claim success or blindly repeat the action | Pending check is visible to the assigned owner |
| 9. Unanswered transfer | Ask for staff while the test line is unanswered | Uses the agreed fallback and preserves useful context | Transfer outcome and fallback record agree |
| 10. Caller leaves early | End the call before all required details are supplied | Leaves an honest incomplete status where appropriate | No invented details or completed booking |
A case can pass without completing the caller’s original request. For an unsupported service, an accurate explanation and a useful handoff may be the correct result, so decide that before judging the call.
Inspect the ordinary call as carefully as the failure
Start with the most common real request. Let the receptionist ask its own questions and answer naturally, including the hesitations people use on the phone, and check whether it gathers information once and finishes with a clear status.
Afterward, have the observer find the record without help from the implementer. Verify the contact route, service, requested date, notes and assigned staff member, and check that the summary repeats what the caller said, with no gaps filled in by a plausible guess.
If the flow sends a confirmation, inspect that too. The spoken answer, the message and the business record should agree, and a call that sounds successful but leaves contradictory records fails the test.
Repeat the ordinary case after any substantial fix. A repair for an unusual exception can accidentally make the normal path worse.
Test corrections across the conversation and the action
It is easy to judge an interruption only by listening for the voice to stop. Also inspect what the system remembers and what it does next.
Twilio’s ConversationRelay documentation includes interruption events when a caller speaks during playback. That capability can help an implementation react, although your own booking or inquiry flow still has to show, in testing, that it handles the correction.
For example, change the requested day after the system has offered a time but before you approve the booking. The result should reflect the final instruction. If the old request was already committed, the receptionist must explain that status and use an approved change process, and the record should show both the original request and the change.
Use ambiguous language too, such as “next Friday,” “after lunch,” or a corrected phone digit. The expected response may be a clarifying question, and a smooth guess fails the case.
Make booking success visible
A test should separate finding availability from making a commitment. Google Calendar exposes free/busy lookup separately from event creation, and your scheduling system may have its own additional rules.
Ask the observer to show the accepted record and confirm what it means for the business. An event in a calendar may still need approval in your workflow, and the receptionist must use language that fits that stage.
For the uncertain-result case, arrange a controlled situation where the action may have succeeded but the connection does not return a definite answer. Check that the flow keeps the result marked as uncertain and follows its reconciliation process. The booking failure guide explains how to design that process, and this test checks that the implementation follows it.
Listen through the entire transfer
Call transfers need more than a ringing phone. Test an answered call and an unanswered call, even a destination that behaves differently after hours, and include voicemail if the real destination may reach it.
Twilio documents distinct transfer outcomes, including busy, no-answer, failed, and completed. Your own acceptance criterion should go further and confirm that the intended staff workflow received the caller or a usable follow-up request.
Check what the caller hears while waiting and after the attempt ends. If the fallback is a callback request, verify that it was accepted and has an owner. A promise that someone will return the call should match an approved process, including any stated response time.
Use severity to decide whether to launch
Classify failures by what they cost the customer. An invented booking confirmation or a lost inquiry should block that capability until repaired, and so should an unauthorized action or an accidental repeat booking. An awkward phrase can wait for a later improvement if the outcome is still clear.
A simple decision record can list the failed case, consequence, temporary restriction, owner, and retest result. If booking is not ready but approved question answering and callback capture pass, the business may choose to launch that narrower scope, with the restriction set in the configuration and in what the receptionist tells callers.
One serious failure outweighs nine easy passes. The matrix exists to show specific risks and useful behavior case by case, and a single percentage across all ten cases would hide both.
Keep the pack after launch
Repeat the relevant tests when opening hours, service rules or calendars change, and again when transfer numbers or software connections change. Add a new case when a real call exposes a gap, using a cleaned-up scenario that leaves out the private details of the customer who called.
Compare the staff time needed to run and maintain this pack with the ownership arrangement you chose. The DIY versus managed comparison helps assign that work by name.
A completed test pack should let another staff member explain what the receptionist can do, what remains restricted and what evidence supports the decision, long after anyone remembers how the demo sounded.