Stablehand

Run agents inside your infrastructure, on any harness, and keep every run as learning signal.

Stablehand is the runtime at the centre of our stack. Each session runs in its own sandbox, the harness you choose drives the model, and every turn, tool call, approval, correction and dollar spent lands in your own Postgres. Swap the harness or the model; the record of how your work gets done stays with you.

Status
Early access
Language
TypeScript
Stack
Fastify, PostgreSQL, React
Licence
MIT
3agent harnesses on one control plane
17/17model and harness combinations certified across seven providers
2,430turns in a ten-wave endurance run, with zero failures
300/300benchmark tasks run and graded in one campaign

Why we built it

The best agent framework changes every quarter. The record of how your company does its work should outlast all of them. Every team we built agents for was rebuilding the same runtime underneath (sandboxes, approvals, streaming, budgets), and every hosted platform kept the most valuable output, the traces and corrections, in someone else’s cloud.

We built Stablehand to keep that record inside your infrastructure. A small core owns runs, events, budgets, approvals and artifacts, and everything else, from evals to domain tools, plugs in as an extension.

What it does

Where it’s going

Who it’s for

Agents inside a SaaS product

Specialist agents that use company context and tools, with human review, per-customer budgets and a run history behind every output.

Back-office workflows

Claims intake, finance close and document review as chained agents, with approvals before anything is filed and the event log as the audit trail.

Model and harness bake-offs

Run the same private tasks across harness and model combinations, grade them, and route each task to the cheapest configuration that passes.

Work with us

Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.