How We Work

Senior engineers, hourly billing, code in your repo

Engagement model

We bill hourly on a time-and-materials basis. There are two ways to engage: a project - a readiness assessment, an eval suite build, a red-team campaign - with the scope, staffing and estimate agreed in writing before work starts; or staff augmentation, with a senior engineer embedded in your team under your direction. If intake shows the work is larger than the estimate assumes (many tools, heavy fixtures for external systems, multiple agents), we say so before starting, never partway through.

Each engagement is staffed by senior US-based engineers who have built and run agent evaluation harnesses themselves. There is no hand-off to a junior bench or an offshore team. You will talk to the people doing the work.

A typical engagement at a glance

EngagementTypical durationCadence
Readiness assessment~2 weeksKickoff, mid-point check, readout
Eval suite build4-6 weeksWeekly demo of the working suite
Red-team campaign2-3 weeksThreat-model review, findings readout, retest
Ongoing reliabilitySteady cadenceMonthly report, quarterly red-team re-run

How deliverables land

  • Code and data in your repo. Datasets, graders, CI configuration and runbooks are committed to a branch in your repository. Reports are markdown in the same branch, with a PDF if you want one for leadership.
  • Your tooling. We build on Langfuse, Promptfoo, Braintrust, OpenTelemetry or plain pytest, whichever you already run. We will recommend a tool if you have none, and we have no commercial relationship with any of them.
  • Reproducibility. Every risk, finding or regression we report comes with a case you can re-run. If we cannot reproduce it, it does not go in the report.

Defining "correct"

The first thing we do on any engagement is write down what a correct trajectory and a correct outcome look like for your top user journeys, with your product owner in the room. Graders are built against that definition, and LLM-as-judge graders are calibrated against your reviewers' labels, with the agreement number reported. We do not import a generic benchmark and call it your quality bar.

Security and data handling

  • Least-privilege access. Read-only access to repositories and trace stores wherever possible; write access limited to a feature branch. Staging environments preferred over production for red-team work.
  • Your data stays yours. Traces, datasets and findings are stored in your systems. Anything we hold during an engagement for analysis is deleted at close-out unless you ask us to retain it.
  • No training, no reuse. We never use client traces, prompts or findings to train models or to inform another client's engagement.
  • Model providers. Graders run against whichever model provider you already use, under your accounts and data-handling terms. We do not route your data through our own accounts unless you ask.
  • Responsible disclosure. Red-team findings are shared only with your named contacts, in your systems. We do not publish findings, anonymized or otherwise, without written permission.
  • Agreements. Mutual NDA on request before intake; our standard services agreement covers IP assignment of all deliverables to you.

After the engagement

You own everything. The suite runs in your CI without us. Many teams go on to ongoing reliability engineering so someone keeps the judges calibrated and the dataset current; that is a choice, not a lock-in.

Get in touch to talk through fit.

Ask. We would rather answer a scoping question by email than have you guess.
Questions about the process?