Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.
| Тип | Репозиторий |
| Категория | GitHub-проекты / Из awesome-списка |
| Цена | открытый код |
| GitHub | defilantech/LLMKube |