Fine-tuning open models for specialist work
We fine-tune open-weight models for one task and measure quality, cost and latency against frontier models.

How we fine-tune a model
Specify the task
Define inputs, outputs, evidence requirements, abstention and failure conditions. Learnings from frontier-model attempts on the same work become examples and test cases.
Curate the data
Build source-linked examples. Separate matters across training, calibration and test splits.
Adapt the model
Choose the base, training recipe and checkpoint against the evaluation goal. Training runs on dedicated GPUs from our neocloud partners or in your own cloud.
Test the result
Compare with baselines, inspect task-level failures and check capability outside the training domain.
Two fine-tuned open models beat every frontier model we tested
We fine-tuned open-weight models on structured extraction from public legal decisions, each answer tied to its source passage. Both models, trained from the same data with the same recipe, scored above all twelve frontier configurations on a sealed test set of 496 tasks, at a fraction of the cost.
Sealed test set, September 2026. Supervised fine-tuning with low-rank adapters. Cost is inference per 1,000 tasks at hosted prices.
A frontier model solves 42% of our clinical environments within five tries
Hundreds of environments built from longitudinal patient histories, with five attempts per environment. The other frontier model we ran solves 7%. That gap is the training signal for reinforcement learning.
Work with us
Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.