Faisal RehmanAI ENGINEER
← All selected work

Checkmate · Case study 01

Making a voice agent reliable enough to order.

VoiceBite: core engineering and evaluation for a real-time restaurant ordering agent, tested in a single-location production pilot.

My role
Led core engineering; owned orchestration, conversation quality, and evaluation
Period
2025–2026
Technology
Python asyncio · Pydantic AI · Coval · Langfuse · Redis · Real-time voice
Low 70s → 90%+Task completion in the pilot
~50% reductionA persistent transcription error class
Sub-secondReported p95 turn latency

The problem

A voice ordering agent has to resolve noisy speech, remember changes, select valid menu items, and submit the right order while a customer is waiting. A correct answer delivered too late can still be a failed interaction.

The engineering challenge was to make the whole conversation reliable and identify which part of the system had failed when it wasn’t.

What I owned

I led core engineering for the agent and owned orchestration, conversation quality, and the evaluation function. Speech and telephony capabilities came from vendors; my work connected those components to the agent’s behavior and business workflow.

The scope was a single-location production pilot. Checkmate later discontinued the voice product. The outcomes here describe that pilot, not a multi-location deployment.

The engineering decisions

Customer speech→Agent reasoning→Validated actions→Order or handoff

Separate intent from execution. The model interprets what the customer wants. Deterministic logic validates actions and controls what gets submitted. Failure paths provide a human handoff.

Evaluate complete conversations. I built simulation suites against the production pipeline, covering reasoning, tool calling, state management, interruption handling, entity extraction, latency, and task completion.

Trace the system at the right level. Prompt versions, tool calls, and conversation traces made it possible to compare iterations and assign failures to model, infrastructure, or data.

A decision that changed the outcome

A recurring failure looked like a model problem. Evaluation traced it to transcription. Fixing the speech-related layer cut that error class about 50%.

How I measured it

Simulation suites became a release gate. Production conversations, speech outputs, annotations, and tool traces fed evaluation sets, while LLM graders were checked against human labels.

  • Outcome: task completion and order accuracy.
  • Conversation: latency, interruption recovery, and workflow adherence.
  • Understanding: entity extraction and tool-call correctness.
  • Speech: word, character, and semantic transcription errors.

The key was actionable attribution: a failure should identify the next engineering change, not simply lower an aggregate score.

What changed

Task completion improved from the low 70s to over 90%, with sub-second p95 turn latency reported for the pilot. A persistent transcription error class was reduced about 50%.

Simulation-led evaluation and AI-native development also contributed to approximately 3× engineering iteration velocity.

These are reported pilot outcomes. Call counts and measurement windows are not included here, so the figures should not be interpreted as a controlled benchmark or a claim of statistical significance.

A public engineering summary. Results are scoped to the project and period described; implementation details are summarized at a high level.

NEXT CASE STUDY

From business questions to actions inside Checkmate.

Continue reading
07 WHAT’S NEXT

Hard problems.
Meaningful work.

I’m interested in teams bringing capable AI into the real world. Let’s talk about AI engineering, agent reliability, and products worth building.