How Do We Know the Agent Got Better?

What I learned from using evals to test changes to agent workflows.

When you build workflows or agents, a small tweak to a prompt or skill can have effects well beyond the behavior you meant to change. You add an instruction, try it a couple of times, and move on. Later, you discover that the agent now skips an approval step or mishandles a task it used to complete correctly.

These regressions can be hard to spot through everyday use. Each run looks plausible on its own, and the failure may only appear under particular conditions.

Anthropic’s recent work on automating evaluations prompted me to examine how I was checking for these problems in the agent skills I maintain.

Changing those instructions changes how the product works. You need to check that the change helps without breaking existing behavior, and understand any added cost.

Decide what success means

An evaluation, or eval, gives an agent a defined task and checks what happens against explicit expectations. For a workflow, that includes the actions taken along the way. A correct final answer does not excuse making an unauthorized change before asking for approval.

Choose tasks that reflect how people use the workflow, and write down what the agent should do in each case. Include actions it must never take, along with ordinary cases where the new instruction should make no difference.

Tools for building and running evals

Claude’s /claude-api build-eval helps design test cases and graders, with your review. /claude-api hillclimb proposes changes against those tests and checks examples kept out of the editing process for improvement. You choose the goal and what it may edit. Anthropic’s walkthrough explains both workflows.

The first helps turn expectations into tests. The second can help tune a prompt or reduce cost while preserving quality, once you trust the tests. If the score rewards the wrong behavior, automated improvements to that score may make the workflow worse.

For plugin authors, claude plugin eval provides a runner with isolated sessions, simulated tools, and grading. Its default comparison tests with and without the plugin. That helps establish whether the plugin contributes anything, but answers a different question from whether an edited version improves on the previous one.

Using Claude’s built-in runner saves you from maintaining your own test machinery. I would start there and build a custom runner only when a test requires more control.

Compare against a baseline

Recently, I updated workflow skills that read issues from GitHub, Linear, Jira, and Azure DevOps to also read relevant issue comments. Before modifying the skills, I ran regression evals across two models to record their existing behavior. That baseline let me compare the edited skills against the originals and check for regressions.

Within each model’s comparison, I held the model, effort setting, tasks, and simulated tool responses constant between versions, and ran each case three times.

For this work, I used a custom runner because I needed to simulate user replies and check that the agent waited for approval before acting. I also had to verify that the runner itself handled those interactions correctly.

Some failures were in my tests. An agent correctly said it couldn’t finish reading the comments, but my check expected different wording. Changing the agent’s instructions would have addressed the wrong problem. I corrected the check and kept the earlier results in the record.

The edited skills passed more runs on both models, but still made mistakes. I reviewed the failures individually before deciding to ship. In one case, the agent correctly ignored irrelevant comments but unnecessarily listed what it had ignored. I accepted that extra wording and left the test marked as failed. A higher overall score did not make the failure disappear.

Start small, inspect the results

For teams starting out, I would keep the first set of tests small enough to read every result. When a test fails, examine what the agent actually did before changing its instructions. Check important requirements individually; a better average score can still hide a skipped approval.

Those tests are there for the next edit too. Each time I change a prompt or skill, I can check the tasks I already rely on and investigate anything that behaves differently. A passing suite cannot cover every situation, but it gives me a firmer basis for shipping than trying the change a couple of times. Over time, that becomes part of maintaining the workflow: keeping evidence that it still does the work I expect of it.

Nick Van Exan

Software & Privacy

Hi, I'm Nick Van Exan, a software developer from Toronto. Currently, I work on product and technology at CodeLantern, a software development consultancy that helps teams adopt human-centred agentic workflows for safe and maintainable software.

Nick Van Exan

Writing

  1. How Do We Know the Agent Got Better? 2026.09.30
  2. Thinking in Loops 2026.09.15
  3. What an Agent Should Remember 2026.06.01
  4. The End of Software Development. And Its Beginning. 2026.04.12

About

Hi, I'm Nick Van Exan, a software developer and privacy engineer based in Toronto.

I've been building things on the web for about twenty years, long enough to have been through a few complete turns of the wheel and to have strong opinions about which turns were worth making.

These days I'm focused on agentic development. I build tooling for agentic development workflows, and I help teams onboard those workflows so they can deliver safe and maintainable software. I'm a Co-Founder at CodeLantern, a small consultancy built around that work.

My career in software hasn't always followed a straight line. I started making websites in the early 2000s, and won a SXSW Web Award in 2003. I loved designing and coding applications, but after watching a bit too much Law & Order, I felt the pull towards law school (nobody's perfect). I continued working as a web developer to pay for law school, before going on to practice litigation at a big Toronto firm for a few years. In 2015, I found my way back to software development and later into the field of privacy engineering, which has now led to over a decade of consulting with startups, enterprises, and governments on how to build things that don't hurt the people using them.

When I'm not at a keyboard I'm usually out for a run, listening to jazz, obsessing over stationary, or walking my dog.

Experience

  1. Co-founder & Chief Product Officer
    CodeLantern
    2026–present
  2. Co-founder & Principal Consultant
    Fieldwork Inc.
    2015–present
  3. Product & Privacy Counsel
    Hootsuite
    2017–2018
  4. Litigation Associate
    Davies Ward Phillips & Vineberg LLP
    2011–2015
  5. Software Development Consultant
    ObjectSharp
    2002–2008

Certifications

  1. Certified Information Privacy Professional/Canada (CIPP/C)
    International Association of Privacy Professionals
    2017–present
  2. Certified Information Privacy Professional/Europe (CIPP/E)
    International Association of Privacy Professionals
    2020–present

Open Source

  1. Markdoc
    Contributor to Stripe's markdown authoring framework
    2022–present

Work together

I take on a small number of engagements each year. Areas I'm currently working on include agentic development methodology and tooling, AI governance, and privacy engineering. If you think we might have something worth working on together, send me a note: nick@vanexan.ca. I'd love to hear from you.

Colophon

Designed by
Nick Van Exan
Built with
Astro
Typeset in
Söhne · Söhne Mono, by Klim Type Foundry
Photography
Carmen Cheung
Subscribe
RSS
Tracking
None. No analytics, no cookies.

Designed mostly by not designing much.