About the role
The claim this company makes is that its data is correct. This role is the one that checks. You would build the evaluation sets, measure precision and recall per field, and be able to say what a contract change did to quality rather than whether it felt better.
The second half is ranking. Semantic search here has a specific failure mode we design against: in an earlier system, 'arm curl' matched 'ankle curl' because the name dominated the embedding. Embeddings are now built from structure rather than names, similar-item search is gated on anchor fields, and every release runs ranking checks before it can be published. You would own that measurement and push it further.
You would also own the prompt and model choices where they are still choices: what to send, what schema to constrain it to, and what it costs per thousand records.
What you'd do
- Build and maintain evaluation sets per field, and report precision, recall and abstention honestly
- Own the ranking quality story: structure-first embeddings, anchor gating, golden pairs that run before a release
- Measure the effect of a contract or prompt change, with a result someone can disagree with
- Tune model choice, prompt and schema against cost and accuracy, not vibes
- Investigate the records the pipeline got wrong and turn each into a rule, a test or a contract change
What we need
- Three or more years in data science or applied ML, with work that shipped rather than sat in a notebook
- You can design an evaluation that survives someone trying to poke holes in it
- Strong Python and SQL
- Practical experience with embeddings, retrieval or classification quality
- You are comfortable reporting a result that says your idea did not work
Nice to have, not required
- Information retrieval or search relevance background
- Experience with LLM evaluation, structured output or constrained decoding
- You have worked where a wrong prediction had a real-world consequence
How we work
- Small team, thin slices, shipped weekly. Nothing sits on a branch for a month.
- Reviews are real. Every feature ships with its tests, and a bug found in review is cheaper than one found by a customer.
- Safety-critical means the boring answer usually wins: fail closed, keep the receipt, don't guess.
- We only claim what ships. That applies to the product, the roadmap and the offer letter.
How hiring works
- 1A 30-minute call about something you've built and what was hard about it.
- 2A working session on a real problem from this codebase. You can drive, or we can pair.
- 3A conversation about how you make decisions when the evidence is thin.
- 4References, then an offer. We aim to answer within a week at every stage.