Insurance
Private evals built around claims tasks, and open-weight models fine-tuned where the evals show a gap. In a study on public legal decisions, a fine-tuned 27B model beat a frontier model at a ninth of the inference cost.
How we develop models
