Careers
Work on the hard part of shipping agents
Judgeworthy works with a small network of senior US-based engineers on a contract basis. The work is evals, reliability engineering and red-teaming for LLM agents: reading traces, building datasets and graders, calibrating judges, wiring suites into CI, and running structured adversarial campaigns against systems that can act.
Engagements are project-shaped and remote, typically a few weeks to a few months, with clients who are shipping real agents into production. You would work directly with the client's engineers, in their repo and on their tooling, and everything you build ships as reproducible code and data. There is no bench and no filler work between engagements; we reach out when a project fits what you do.
If you have shipped an agent to production, built the harness that tested it, or broken one on purpose, we would like to hear from you.
Senior peers only
Every engagement is staffed by engineers who have built and run agent harnesses themselves. No layers, no juniors to supervise.
Work that ships as code
Datasets, graders, CI jobs and findings land in the client repo. Nothing you build ends life as a slide deck.
Remote, US-based
Contracts run remote from anywhere in the United States, with occasional onsite time when a client engagement calls for it.
Honest scoping
Scope is agreed in writing before a project starts, and raised early when it moves. You will not be asked to eat the difference.
Senior AI Evals Engineer - Contract, Remote (US)
RemoteBuild eval datasets, trajectory and tool-call graders, calibrated LLM-as-judge scoring and CI gating for client agents. You have shipped an LLM agent to production and built the harness that tested it. Project-based; hours vary by engagement.
Apply Online ▸AI Red-Team Engineer - Contract, Remote (US)
RemoteRun structured adversarial campaigns against client agents: indirect prompt injection, authorization probes, exfiltration paths, unsafe tool use. A security background helps; writing reproducible findings is the job.
Apply Online ▸Join The Team. Apply Now.
Tell us who you are and what you have shipped: the agents you have built or tested, the eval tooling you know (Langfuse, Promptfoo, Braintrust, or homegrown), and a link to code or writing if you have one. We read every note and reply to the ones that fit.