If you cannot answer "is the new prompt better than the old prompt" with a number, you do not have an LLM product. You have a vibes-based product that happens to involve an LLM.
What an eval harness actually does
Curated test set + scoring function + versioned prompts + leaderboard. Run on every prompt change, every model change, every retrieval change. Without it you are guessing.
About the author

Keep reading
Related articles.
Keep talking
Talk to the engineering team
If you're shipping something into production and want a sanity check from people who've done it, we'd be glad to compare notes.