Support and ops agents
Turn a pile of thumbs-downs into regression evals, and see exactly where the evals and your reviewers disagree.
Your reviewers rate a handful of traces. iofold writes the evals, and the evals decide which agent version goes live.
iofold imports your agent traces from Langfuse, collects quick ratings from your team, and has an agent write Python eval functions that must agree with those ratings before they score anything. The evals run in a sandbox on every trace, and GEPA uses them to propose a better prompt as a candidate version for you to promote.
Teams ship an agent, users complain, and the feedback never reaches the prompts. Hand-written eval scripts rot as the agent changes.
We built iofold so the judgement in your reviewers’ heads becomes code that runs on every trace. It was our first build of the learning loop that now sits at the centre of how we work with design partners.
Turn a pile of thumbs-downs into regression evals, and see exactly where the evals and your reviewers disagree.
Catch wrong tool calls, wrong parameters and answers that invent facts missing from the tool results.
Run an optimised candidate against the current version on a task set, then promote it by hand.
git clone https://github.com/iofold/iofold && cd iofold
pnpm install
cp .dev.vars.example .dev.vars # Langfuse keys and Cloudflare AI Gateway settings
docker compose up -d # backend :8787, frontend :3000, Python sandbox :9999
curl http://localhost:8787/healthFull documentation is in the repository README.
Tell us what the work is, what data it touches and what constraints apply. We’ll work out what to build and how to measure it.