Fine-tuning open models for specialist work

We fine-tune open-weight models for one task and measure quality, cost and latency against frontier models.

Fine parallel striations within a pale mineral grain.

How we fine-tune a model

  1. Specify the task

    Define inputs, outputs, evidence requirements, abstention and failure conditions. Learnings from frontier-model attempts on the same work become examples and test cases.

  2. Curate the data

    Build source-linked examples. Separate matters across training, calibration and test splits.

  3. Adapt the model

    Choose the base, training recipe and checkpoint against the evaluation goal. Training runs on dedicated GPUs from our neocloud partners or in your own cloud.

  4. Test the result

    Compare with baselines, inspect task-level failures and check capability outside the training domain.

Two fine-tuned open models beat every frontier model we tested

We fine-tuned open-weight models on structured extraction from public legal decisions, each answer tied to its source passage. Both models, trained from the same data with the same recipe, scored above all twelve frontier configurations on a sealed test set of 496 tasks, at a fraction of the cost.

Sealed test set, September 2026. Supervised fine-tuning with low-rank adapters. Cost is inference per 1,000 tasks at hosted prices.

Quality against inference costTwelve frontier-model configurations score between 29% and 67% complete answers. The two fine-tuned open models score 82.7% and 84.5% at $1.62 and $5.71 per thousand tasks.10%20%30%40%50%60%70%80%90%$0.50$1$2$5$10$20$50Inference cost per 1,000 tasks (log scale)Complete answersBest frontier: 66.9%, $50.90Before fine-tuningFine-tuned 35B MoE82.7%, $1.62Fine-tuned 27B84.5%, $5.71Quality against inference costTwelve frontier-model configurations score between 29% and 67% complete answers. The two fine-tuned open models score 82.7% and 84.5% at $1.62 and $5.71 per thousand tasks.10%30%50%70%90%$0.50$2$10$50Cost per 1,000 tasks (log scale)Best frontier: 66.9%Before fine-tuning35B 82.7%27B 84.5%
Fine-tuned open models: 27B (84.5%, $5.71) and 35B MoE (82.7%, $1.62)Same models before fine-tuningFrontier-model configurationsBest frontier score at each cost

A frontier model solves 42% of our clinical environments within five tries

Hundreds of environments built from longitudinal patient histories, with five attempts per environment. The other frontier model we ran solves 7%. That gap is the training signal for reinforcement learning.

Pass rate by number of attemptsAcross our clinical environments, frontier model A passes 22% on one attempt and 42% within five; frontier model B passes 4% and 7%.0%10%20%30%40%50%60%12345Attempts per environment (k)Pass@kFrontier model A42% within 5Frontier model B7% within 5Pass rate by number of attemptsAcross our clinical environments, frontier model A passes 22% on one attempt and 42% within five; frontier model B passes 4% and 7%.0%20%40%60%12345Attempts per environment (k)A 42%B 7%
Frontier model AFrontier model BBands: 95% bootstrap intervals

Work with us

Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.