Pattern #21 · Tools and integrations

Evals are the new usability test: design teams define what good looks like, then test for it

The agent feels worse is a feeling, not a finding.

Track Bevalsai-agentstestingdesign-opsdesign-tools

Verified as of 2026-08-27

Do

Start with 20 to 50 tasks from real failures and the checks you already run by hand. Grade the output, not the path.

Don't

Don't ship a prompt redesign on vibes and let the support queue run the study. Your angriest users are not a research panel.

The rule. Treat evaluations the way design teams treat usability tests: encode what good AI behavior looks like as graded test cases, run them before shipping, and rerun them at every prompt or model change.

Why. AI features change under their users: models learn, adapt, and get swapped out, which is why the 18 guidelines for human-AI interaction give "over time" its own phase of seven guidelines, including "update and adapt cautiously" and "notify users about changes" (Amershi et al., 2019). Without evals, teams learn about behavior drift when users report the agent feels worse, then debug reactively with no way to verify fixes except guess and check (Anthropic, 2026). The practice is not engineering-only: Anthropic's guidance states that the people closest to product requirements and users are best positioned to define success, and that product managers can contribute eval tasks directly (Anthropic, 2026).

Seen in the wild. Descript built evals for its video-editing agent around three product dimensions (don't break things, do what I asked, do it well), with grading criteria defined by the product team and periodic human calibration (Anthropic, 2026).

Verified as of 2026-08-27 against current Anthropic documentation and engineering guidance.

References

  1. 01

    Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1-13. https://doi.org/10.1145/3290605.3300233

    https://doi.org/10.1145/3290605.3300233
  2. 02

    Anthropic. (2026, January 9). Demystifying evals for AI agents. Anthropic Engineering. Retrieved August 27, 2026, from https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents

    https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
  3. 03

    Anthropic. (n.d.). Define success criteria and build evaluations. Claude Platform documentation. Retrieved August 27, 2026, from https://platform.claude.com/docs/en/test-and-evaluate/develop-tests

    https://platform.claude.com/docs/en/test-and-evaluate/develop-tests