About the role
Customers point live agents at tools we serve. That makes deploys, migrations and observability a product concern rather than an internal one: if a released tool starts answering slowly or wrongly, we need to know before the customer's agent does.
You would own CI, the deploy path, database migrations (including renames that must not drop a table) and the monitoring around the served endpoints: latency, error rates, and how rate limits and quotas actually behave under load.
You would also own the unglamorous half of security: secrets out of the codebase and out of logs, credential rotation as a routine task rather than an incident, and least privilege on the database roles.
What you'd do
- Own CI: typecheck, lint, the full test suite against a real database, and a build that fails loudly rather than shipping broken
- Run the deploy path and the database migrations, including renames that must not drop a table
- Set up logging, metrics and alerting for the served tools: latency, error rates, quota and rate-limit behaviour
- Keep secrets out of the codebase and out of logs, and make credential rotation a routine task
- Write the runbook for the failure you just fixed, so the next person does not start from nothing
What we need
- Three or more years keeping production systems running, including being on call for them
- Comfortable operating Postgres: backups, restores, migrations, and reading what it is doing under load
- Fluent in the shell and in at least one scripting language; you automate the thing you did twice
- You have set up monitoring that caught a real problem, and can describe the alert that did it
- You treat a rollback as a normal operation, not a last resort
Nice to have, not required
- Experience with containerised deploys and infrastructure as code
- You have run a cost-reduction exercise and know where the money actually went
- Security hardening or a compliance review you helped pass
How we work
- Small team, thin slices, shipped weekly. Nothing sits on a branch for a month.
- Reviews are real. Every feature ships with its tests, and a bug found in review is cheaper than one found by a customer.
- Safety-critical means the boring answer usually wins: fail closed, keep the receipt, don't guess.
- We only claim what ships. That applies to the product, the roadmap and the offer letter.
How hiring works
- 1A 30-minute call about something you've built and what was hard about it.
- 2A working session on a real problem from this codebase. You can drive, or we can pair.
- 3A conversation about how you make decisions when the evidence is thin.
- 4References, then an offer. We aim to answer within a week at every stage.