01Agents change when nothing changed
A prompt tweak, a new tool or a model update can shift an agent's behaviour in ways no type checker will catch. Without evaluations, you find out from users.
02Scenarios for Gobbles' ordering assistant
Gobbles' voice waiter works through 12 tools: finding and describing dishes, managing the cart, notifying the kitchen, calling a waiter. Its expected behaviour is written down as 19 scripted scenarios that can be re-run whenever the agent changes.
19 / 19 agent scenarios passing
Voice waiter
Simulated conversation
12 tools · 19 eval scenarios · English, Hindi & Hinglish
03Pass, flaky or fail
Models aren't deterministic, so a single run proves little. Zabber's coach evaluation runs golden cases against the live model several times each and classifies them as passing, flaky or failing. Treating flaky as its own result keeps inconsistent behaviour from hiding behind a lucky pass.
04Built to recover, not just to pass
- Tools make behaviour observable: you can see exactly what the agent looked up before it answered.
- A watchdog recovers voice responses that stall instead of leaving a diner in silence.
- Connection limits are per device, not per IP, because diners share restaurant Wi-Fi.
- All of it sits on top of 2,000+ conventional automated tests across both products.