AI9 MIN READ31 JUL 2026

Evaluation harnesses: the boring half of an LLM project

Why your model demo passes and your production deployment fails and the testing infrastructure that closes the gap.

Igor Stajić

Igor Stajić

CBO

If you cannot answer "is the new prompt better than the old prompt" with a number, you do not have an LLM product. You have a vibes-based product that happens to involve an LLM.

What an eval harness actually does

Curated test set + scoring function + versioned prompts + leaderboard. Run on every prompt change, every model change, every retrieval change. Without it you are guessing.

About the author

Igor Stajić

Igor Stajić

CBO

Keep talking

Talk to the engineering team

If you're shipping something into production and want a sanity check from people who've done it, we'd be glad to compare notes.