Field note · 4 min read

    Testing AI agents like software

    Agents regress quietly. Scripted scenarios, repeatable golden cases and observable tools are how we catch it.

    4 min read4 sectionsWritten from shipped code

    01Agents change when nothing changed

    A prompt tweak, a new tool or a model update can shift an agent's behaviour in ways no type checker will catch. Without evaluations, you find out from users.

    02Scenarios for Gobbles' ordering assistant

    Gobbles' voice waiter works through 12 tools: finding and describing dishes, managing the cart, notifying the kitchen, calling a waiter. Its expected behaviour is written down as 19 scripted scenarios that can be re-run whenever the agent changes.

    19 / 19 agent scenarios passing

    Live demo · Every scenario re-checked on change

    Voice waiter

    Simulated conversation

    Is the paneer tikka spicy? I'm allergic to nuts.
    describe_itemitem: "Paneer Tikka"
    spice_level: 2 · medium · allergens: [dairy]
    It's medium spicy, and there are no nuts in it. Just so you know, it does contain dairy.
    Kuch halka aur bina dairy wala batao?
    find_dishesexclude: [dairy] · richness ≤ 1
    Chana Masala · allergens: [] · richness: 1 · moderate
    Chana masala try kijiye. Halka hai, aur is menu par dairy-free tagged hai.

    12 tools · 19 eval scenarios · English, Hindi & Hinglish

    Live demo · Simulated conversation: answers come from tool calls on enriched data

    03Pass, flaky or fail

    Models aren't deterministic, so a single run proves little. Zabber's coach evaluation runs golden cases against the live model several times each and classifies them as passing, flaky or failing. Treating flaky as its own result keeps inconsistent behaviour from hiding behind a lucky pass.

    04Built to recover, not just to pass

    • Tools make behaviour observable: you can see exactly what the agent looked up before it answered.
    • A watchdog recovers voice responses that stall instead of leaving a diner in silence.
    • Connection limits are per device, not per IP, because diners share restaurant Wi-Fi.
    • All of it sits on top of 2,000+ conventional automated tests across both products.

    AI data for safety-critical products: unknown is never safe

    When an AI agent answers "is this nut-free?" or "is this safe in pregnancy?", the data behind it decides whether someone gets hurt. Five rules for building data an agent can be trusted with.

    Read next