Here is a scenario that plays out in AI teams everywhere, probably including yours… Someone notices the chatbot is not handling a specific type of question well. They tweak the prompt. They test it on that specific question. It works now. They ship it.
Three days later a different type of question that was working perfectly before is now broken. Nobody connects it to the prompt change. It gets filed as a bug. Someone spends two days trying to figure out what happened. Eventually someone checks the git history for the prompt file, sees the change, tests the broken case against both versions, and finds the regression.
Two days of debugging for a problem that a fifteen minute test suite would have caught before it ever shipped.
This happens constantly in AI development. And the reason it keeps happening is that most teams treat prompts as text files rather than as software, which means they deploy prompt changes with none of the testing discipline they would apply to actual code changes.
Why prompt changes are more dangerous than code changes
When you change a line of code in a traditional software system, the blast radius of that change is usually bounded and somewhat predictable. You changed a function, so things that call that function might be affected. You changed a database query, so results from that query might be different. The relationships are traceable.
When you change a prompt, the blast radius is the entire distribution of inputs your system will ever receive. A small change in wording can shift the model's behavior in ways that affect certain input types dramatically and leave others completely unchanged. And because model outputs are probabilistic, you might not even see the regression on the first few tests. It might only show up reliably on a specific class of inputs that you did not think to test.
As Eugene Yan wrote in his guide on prompting fundamentals and how to apply them effectively, reliable evals are a prerequisite for any serious prompt engineering work, because without them you have no way of knowing whether a change improved things, made them worse, or broke something entirely that you were not even looking at.
The word prerequisite is doing a lot of work in that sentence. Not nice to have. Not best practice. Prerequisite. Before you change the prompt. Not after.
What prompt regression actually looks like in practice
A prompt regression is when a change you made to a prompt to fix one problem causes a different problem to appear, usually on inputs that were working fine before.
The insidious part is that regressions almost never show up on the inputs you tested the change against. If you changed the prompt to fix a specific edge case, you tested that edge case. It works now. What you did not test is the hundred other input types that your system handles, because you were focused on the thing you were fixing, not on the things you might have broken.
This is exactly the same problem as traditional software regression, and the solution is exactly the same: a test suite that runs against every prompt change before it ships, covering not just the case you are fixing but the full range of inputs your system handles.
The difference is that software engineers have had decades to build the tools and culture around regression testing, and it is now table stakes. Nobody ships a code change without running the test suite. AI teams are still in the phase where most of them do not have a test suite for their prompts at all, which means every prompt change is a gamble.
What a minimal prompt test suite looks like
You do not need a sophisticated evaluation framework to get most of the benefit. A minimal prompt test suite for a production AI feature is a set of representative inputs with expected outputs or expected output characteristics, run against every version of the prompt.
The inputs should cover your happy path, your most common edge cases, and the specific cases that have caused problems in the past. Twenty to fifty inputs is enough to catch most regressions. A hundred is better. The point is not exhaustive coverage. The point is enough coverage that a change which breaks something meaningful cannot slip through without showing up in at least one test case.
The expected outputs do not need to be exact strings, because model outputs are not deterministic and exact matching will produce too many false positives. What works better is checking for expected characteristics: the response contains a certain element, the response does not contain a certain phrase, the response is under a certain length, the response follows a certain format. These checks are robust to the natural variation in model outputs while still catching meaningful regressions.
Hamel Husain covers this in detail in his work on AI evals for production systems, noting that binary pass-fail checks based on specific behaviors are more reliable and more actionable than scoring systems, because they tell you clearly whether a specific property holds rather than giving you an ambiguous number to interpret.
The token counter use case nobody mentions
Here is a practical connection between prompt testing and token costs that most people miss.
Every time you change a prompt, the token count changes. Sometimes it goes up because you added instructions. Sometimes it goes down because you cleaned something up. Either way, the new token count is now the token count for every request your system processes from that point forward.
If your prompt change accidentally added two hundred tokens, you just increased your API costs by the cost of two hundred tokens per request, permanently, until someone notices and fixes it. At high volumes that is not trivial.
Running your prompts through the Token Counter on Prompt Toolbox before and after every change takes thirty seconds and tells you immediately if the change affected your token budget in ways you did not intend. It is the kind of check that should be part of your prompt change review process the same way a diff review is part of your code change process.
The cultural problem underneath the technical one
The reason most teams do not test prompt changes properly is not that they do not know they should. It is that the culture around AI development has not yet caught up to the maturity level that makes testing feel non-negotiable.
In software engineering, the norm is that you do not ship untested changes. That norm exists because enough things broke badly enough times that the industry as a whole decided testing was worth the overhead. AI development is still in the phase before that norm solidifies, where moving fast feels more important than moving carefully and the consequences of a broken prompt feel less severe than the consequences of a broken feature.
They are not less severe. A chatbot that gives confidently wrong answers to a class of inputs is a trust problem that is harder to recover from than a bug that throws an error. Errors are obvious. Bad AI outputs are invisible until a user notices, and by then the damage is done.
Build the test suite before you need it. It is significantly cheaper than building it after the first serious regression reaches your users.




