Field note · 4 min read

    Why we took the names out of our embeddings

    Similar-sounding exercises kept ranking next to each other. The fix was to stop letting names speak for meaning.

    4 min read4 sectionsWritten from shipped code

    01The symptom

    Zabber's AI coach finds exercises by meaning. Every one of the 1,324 exercises in its catalog is embedded, and the coach searches that space when it builds or adjusts a plan.

    Early on, the neighbours were sometimes wrong in a very specific way. Exercises with similar names ranked close together even when they worked completely different parts of the body. Our own code comment records the example that made it obvious: "arm circles" landing next to "ankle circles".

    structure only
    arm circles≈shoulder mobility moves
    Live demo · What changed when the name left the embedding text

    02Why it happens

    An embedding reflects the text you feed it. When that text starts with the exercise name, the name carries a lot of weight, and two names that share most of their words look similar no matter what the movements actually do.

    For a coach, that's more than a ranking quirk. A shoulder drill and an ankle drill are not substitutes, and suggesting one for the other is the kind of mistake that erodes trust fast.

    03The fix: describe the movement, not the label

    We rebuilt the embedding text from the enriched structure instead: movement pattern and type, angle, whether it's one-sided, difficulty, impact, target and secondary muscles with their emphasis, equipment, suitable goals and the joints it loads. The name field no longer leads the text or defines it.

    • Structure first: the fields the enrichment pipeline already produces become the description.
    • Versioned: the change shipped as a new enrichment version, so a re-run only touches rows below it.
    • Resumable: interrupted runs pick up where they stopped, so re-embedding a whole catalog is safe to start.

    04Proving it worked

    Retrieval changes are easy to get wrong silently, so Zabber has a dedicated retrieval evaluation harness. It runs without calling a language model and exits with a failure if expectations aren't met, which makes it cheap enough to run every time the search code changes.

    A catalog that can't forget an allergen

    How a shared, cross-restaurant dish catalog saves repeated AI work, and why its allergen data is only ever allowed to grow.

    Read next