Test customer-facing AI against approved knowledge, prohibited claims, ambiguous questions, sensitive situations, data capture, real actions, human handover and technical failure. Use realistic conversations and record what must change before public use.
The strongest test is not “Can it answer this question?” It is “Can it respond safely and usefully across the situations our customers will actually create?” That includes interruptions, incomplete information, frustration and requests the system should refuse.
1. Test the knowledge boundary
Start with questions the assistant should answer, questions it might answer and questions it must not answer without a person. Check prices, policies, opening hours, availability, integrations and any regulated or sensitive subjects. A confident invention is a failure, even when it sounds fluent.
2. Test natural language, not a script
People will use regional phrases, vague descriptions and several questions at once. They will change their mind mid-sentence. Test the same intent in different words and listen for whether the assistant asks a useful clarifying question rather than guessing.
3. Test every real action
If the assistant can book, capture an enquiry, send a message or update a record, verify the complete result. Check the correct calendar, timezone, customer record, notification and confirmation. A convincing sentence saying an action happened is not evidence that it did.
- Approved answers are accurate and current.
- Unknown information is acknowledged rather than invented.
- Prices, guarantees and availability stay within approved boundaries.
- Bookings and CRM updates are verified outside the conversation.
- Sensitive or upset customers reach a suitable human route.
- Consent and privacy wording appear before optional data use.
- A network, microphone or provider failure has a clear fallback.
4. Test tone under pressure
An assistant may sound warm in its greeting but become repetitive or defensive when challenged. Test corrections, complaints, silence and requests to speak to a person. For voice AI, also review pace, interruption handling, pronunciation and whether the disclosure is clear.
5. Test the handover with context
A handover should not simply tell the customer to start again elsewhere. Check what information is passed, who receives it and what expectation is set. The assistant should be clear about whether someone will reply, the visitor should call, or a booking route is available.
6. Review the evidence after launch
Public use will reveal language and edge cases no test plan can predict. Review unanswered questions, failed actions, handovers and complaints. Update the approved knowledge and retest material changes. Improvement needs an owner and a cadence.
A practical release decision
Do not ask whether the AI is perfect. Ask whether its supported scope is clear, its failure modes are controlled and the team can see what happened. A smaller, well-tested scope is usually safer than a broad promise that the assistant cannot consistently keep.
