When you build workflows or agents, a small tweak to a prompt or skill can have effects well beyond the behavior you meant to change. You add an instruction, try it a couple of times, and move on. Later, you discover that the agent now skips an approval step or mishandles a task it used to complete correctly.
These regressions can be hard to spot through everyday use. Each run looks plausible on its own, and the failure may only appear under particular conditions.
Anthropic’s recent work on automating evaluations prompted me to examine how I was checking for these problems in the agent skills I maintain.
Changing those instructions changes how the product works. You need to check that the change helps without breaking existing behavior, and understand any added cost.
Decide what success means
An evaluation, or eval, gives an agent a defined task and checks what happens against explicit expectations. For a workflow, that includes the actions taken along the way. A correct final answer does not excuse making an unauthorized change before asking for approval.
Choose tasks that reflect how people use the workflow, and write down what the agent should do in each case. Include actions it must never take, along with ordinary cases where the new instruction should make no difference.
Tools for building and running evals
Claude’s /claude-api build-eval helps design test cases and graders, with your review. /claude-api hillclimb proposes changes against those tests and checks examples kept out of the editing process for improvement. You choose the goal and what it may edit. Anthropic’s walkthrough explains both workflows.
The first helps turn expectations into tests. The second can help tune a prompt or reduce cost while preserving quality, once you trust the tests. If the score rewards the wrong behavior, automated improvements to that score may make the workflow worse.
For plugin authors, claude plugin eval provides a runner with isolated sessions, simulated tools, and grading. Its default comparison tests with and without the plugin. That helps establish whether the plugin contributes anything, but answers a different question from whether an edited version improves on the previous one.
Using Claude’s built-in runner saves you from maintaining your own test machinery. I would start there and build a custom runner only when a test requires more control.
Compare against a baseline
Recently, I updated workflow skills that read issues from GitHub, Linear, Jira, and Azure DevOps to also read relevant issue comments. Before modifying the skills, I ran regression evals across two models to record their existing behavior. That baseline let me compare the edited skills against the originals and check for regressions.
Within each model’s comparison, I held the model, effort setting, tasks, and simulated tool responses constant between versions, and ran each case three times.
For this work, I used a custom runner because I needed to simulate user replies and check that the agent waited for approval before acting. I also had to verify that the runner itself handled those interactions correctly.
Some failures were in my tests. An agent correctly said it couldn’t finish reading the comments, but my check expected different wording. Changing the agent’s instructions would have addressed the wrong problem. I corrected the check and kept the earlier results in the record.
The edited skills passed more runs on both models, but still made mistakes. I reviewed the failures individually before deciding to ship. In one case, the agent correctly ignored irrelevant comments but unnecessarily listed what it had ignored. I accepted that extra wording and left the test marked as failed. A higher overall score did not make the failure disappear.
Start small, inspect the results
For teams starting out, I would keep the first set of tests small enough to read every result. When a test fails, examine what the agent actually did before changing its instructions. Check important requirements individually; a better average score can still hide a skipped approval.
Those tests are there for the next edit too. Each time I change a prompt or skill, I can check the tasks I already rely on and investigate anything that behaves differently. A passing suite cannot cover every situation, but it gives me a firmer basis for shipping than trying the change a couple of times. Over time, that becomes part of maintaining the workflow: keeping evidence that it still does the work I expect of it.