evaluation2 MIN READ
What to measure before you launch an agent
Separate task success, mistakes, time and cost before you build a leaderboard.
Alex Vale / Sep 20, 2026↗
Loading…

We turn complicated data and promising ideas into AI systems people can actually use.
Read the insights ↓Field notes from
problem to practice.
Separate task success, mistakes, time and cost before you build a leaderboard.
Completion should be observable, and retries should have limits.
Keep inputs and evaluation consistent when choosing a model.