Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.
| Type | Repository |
| Section | GitHub projects |
| Pricing | open source |
| GitHub | defilantech/LLMKube |