iofold auto-evals

Your reviewers rate a handful of traces. iofold writes the evals, and the evals decide which agent version goes live.

iofold imports your agent traces from Langfuse, collects quick ratings from your team, and has an agent write Python eval functions that must agree with those ratings before they score anything. The evals run in a sandbox on every trace, and GEPA uses them to propose a better prompt as a candidate version for you to promote.

Language
TypeScript
Evals
Python
Licence
MIT
80%minimum agreement with your reviewers before an eval is kept
5candidate evals written for every request

Why we built it

Teams ship an agent, users complain, and the feedback never reaches the prompts. Hand-written eval scripts rot as the agent changes.

We built iofold so the judgement in your reviewers’ heads becomes code that runs on every trace. It was our first build of the learning loop that now sits at the centre of how we work with design partners.

What it does

Where it’s going

Who it’s for

Support and ops agents

Turn a pile of thumbs-downs into regression evals, and see exactly where the evals and your reviewers disagree.

Tool-using agents

Catch wrong tool calls, wrong parameters and answers that invent facts missing from the tool results.

Prompt changes

Run an optimised candidate against the current version on a task set, then promote it by hand.

Get started

Run it locally
git clone https://github.com/iofold/iofold && cd iofold
pnpm install
cp .dev.vars.example .dev.vars   # Langfuse keys and Cloudflare AI Gateway settings
docker compose up -d             # backend :8787, frontend :3000, Python sandbox :9999
curl http://localhost:8787/health

Full documentation is in the repository README.

Work with us

Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.