Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.
| Type | Repository |
| Section | GitHub projects |
| Pricing | open source |
| GitHub | hidai25/eval-view |