Private evals
Task contracts, graders and rubrics written with your experts
Control · ChoiceTo get good results from a model, a company has to show it how the company works. That is the reverse-information paradox: the more you reveal, the more of your know-how sits in prompts, traces and corrections on someone else’s platform. We help you battle it with the 5Cs.
We write private evals with the people who do the work, because they define what good looks like in your company. Every agent run, tool call, correction and approval is recorded in your own database, and model calls go through a gateway you control.
We turn your workflows into learning environments with verifiable tasks, and fine-tune open-weight models against them. Training and serving run on dedicated GPUs from our neocloud partners or in your own cloud.
Stablehand, our self-hosted agent runtime, runs Claude Agents, LangGraph and PI harnesses behind a gateway across frontier APIs and open-weight models. When a new model arrives, your evals decide whether it takes over, and what you have built stays with you.
Budgets per agent and per team, with routing decided by eval results. In our document-extraction study, a fine-tuned 27B open-weight model scored higher than a frontier model on the sealed test set and cost a ninth as much to run.
We extract learnings from frontier-model attempts on your work (the corrections, approvals and outcomes) and feed them back into your evals and your own models, so all the tokens you pay for compound.
Self-hosted agent runtime: harnesses, one sandbox per session, tools, approvals, budgets and the event log.
Task contracts, graders and rubrics written with your experts
Control · ChoiceEvery run, tool call, correction and approval, in your database
Control · CompoundingRoutes each task to a frontier API or your own open-weight model
Choice · CostVerifiable tasks cut from your workflows, for training and testing
CapabilityFine-tuning and reinforcement learning on open-weight models, with versioned adapters
Neocloud partner GPUsCapability · CostYour adapted models, behind the same gateway
Neocloud partner GPUsControl · CostDe-identification, rights and provenance records for each dataset
ControlCorrections, approvals and outcomes from frontier-model attempts, kept as training and eval signal
CompoundingDefine the tasks, the evidence a good answer needs and the failure cases, with the people who do the work.
Deploy Stablehand in your infrastructure, connect your tools and put approvals where judgment is needed.
Learnings from frontier-model attempts, the corrections and the outcomes become eval cases and training examples.
Train an open-weight model for the gap, and promote it only when it wins on your evals.
Reference: “The Reverse Information Paradox”, July 2026.
Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.