A Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.
| Type | Repository |
| Section | GitHub projects |
| Pricing | open source |
| GitHub | kaito-project/kaito |