AI & automation · June 2026 · 5 min read
Evaluation sets are cheaper than incidents
Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.
Prompt tuning without a labelled set feels productive until production disagrees. The cheaper habit is writing forty examples in an afternoon and running them on every change.
Incidents teach the same lessons more slowly and in public. Evaluation sets teach them before customers notice.
Why a small set beats endless tuning
Tuning optimises for the last conversation. An eval set freezes the failure modes you already care about so a 'better' prompt cannot silently erase them.
- 01Capture known failures firstStart with the cases that already embarrassed you. Perfect coverage comes later; regression coverage comes now.
- 02Score what operators scoreIf reviewers care about field accuracy, score fields. If they care about tone, score tone. Do not invent metrics nobody uses.
- 03Run on every changeModel bump, prompt edit, retrieval tweak — same suite. If it is optional, it will be skipped under deadline.
How small is useful
Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.
Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.
When the suite should run
On every pull request that touches prompts, tools, or retrieval. Weekly against a larger holdout. After every incident, add the case that escaped.
A short test
Change one line of the system prompt and ask whether anyone would notice before a customer does. If the answer is no, you do not have an evaluation set yet.
