The problem
A voice ordering agent has to resolve noisy speech, remember changes, select valid menu items, and submit the right order while a customer is waiting. A correct answer delivered too late can still be a failed interaction.
The engineering challenge was to make the whole conversation reliable and identify which part of the system had failed when it wasn’t.
What I owned
I led core engineering for the agent and owned orchestration, conversation quality, and the evaluation function. Speech and telephony capabilities came from vendors; my work connected those components to the agent’s behavior and business workflow.
The scope was a single-location production pilot. Checkmate later discontinued the voice product. The outcomes here describe that pilot, not a multi-location deployment.
The engineering decisions
Separate intent from execution. The model interprets what the customer wants. Deterministic logic validates actions and controls what gets submitted. Failure paths provide a human handoff.
Evaluate complete conversations. I built simulation suites against the production pipeline, covering reasoning, tool calling, state management, interruption handling, entity extraction, latency, and task completion.
Trace the system at the right level. Prompt versions, tool calls, and conversation traces made it possible to compare iterations and assign failures to model, infrastructure, or data.
A recurring failure looked like a model problem. Evaluation traced it to transcription. Fixing the speech-related layer cut that error class about 50%.
How I measured it
Simulation suites became a release gate. Production conversations, speech outputs, annotations, and tool traces fed evaluation sets, while LLM graders were checked against human labels.
- Outcome: task completion and order accuracy.
- Conversation: latency, interruption recovery, and workflow adherence.
- Understanding: entity extraction and tool-call correctness.
- Speech: word, character, and semantic transcription errors.
The key was actionable attribution: a failure should identify the next engineering change, not simply lower an aggregate score.
What changed
Task completion improved from the low 70s to over 90%, with sub-second p95 turn latency reported for the pilot. A persistent transcription error class was reduced about 50%.
Simulation-led evaluation and AI-native development also contributed to approximately 3× engineering iteration velocity.
These are reported pilot outcomes. Call counts and measurement windows are not included here, so the figures should not be interpreted as a controlled benchmark or a claim of statistical significance.
A public engineering summary. Results are scoped to the project and period described; implementation details are summarized at a high level.