Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Change your AI with confidence, not crossed fingers. · Prompt evaluation for AI apps. Run prompts against real inputs, score outputs 1-5 with an LLM judge, ship only what improved. Free on Cloud, or self-host. · You changed the prompt. It feels better. You need to know. · Four moves. Then you have evidence. · Point your coding agent at CompletionKit. Walk away. · What you actually get, line by line.
| Type | MCP server |
| Section | MCP servers |
| Pricing | free |
| Platform | Command line |
| Systems | cli, api |
| Hosting | cloud |
| Install | mcp |
| Protocols | mcp |
| Site language | en |
| GitHub | homemade-software-inc/completion-kit |
| Launched | 2026-07-18 |