Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.
| Тип | Репозиторий |
| Категория | GitHub-проекты / Из awesome-списка |
| Цена | открытый код |
| GitHub | hidai25/eval-view |