Advanced search

GitHub projects

4 835 entries

GPTCache

github.com

Creating semantic cache to store responses from LLM queries.

Repository GitHub projects open source

Glide

github.com

Cloud-Native LLM Routing Engine. Improve LLM app resilience and speed.

Repository GitHub projects open source

Dstack

github.com

Cost-effective LLM development in any cloud (AWS, GCP, Azure, Lambda, etc).

Repository GitHub projects open source

deeplake

github.com

Stream large multimodal datasets to achieve near 100% GPU utilization. Query, visualize, & version control data. Access data w/o the need to recompute the embeddings for the model finetuning.

Repository GitHub projects open source

Contexto

github.com

Self-hosted context engine for AI agents with persistent conversation memory and recall. Works as a drop-in OpenAI-compatible proxy, OpenClaw plugin, or memory SDK — no code changes required.

Repository GitHub projects open source

Cheshire Cat AI

github.com

Web framework to create vertical AI agents. FastAPI based, plugin system inspired to WordPress, admin panel, vector DB included

Repository GitHub projects open source

BudgetML

github.com

Deploy a ML inference service on a budget in less than 10 lines of code.

Repository GitHub projects open source

AI studio

github.com

A Reliable Open Source AI studio to build core infrastructure stack for your LLM Applications. It allows you to gain visibility, make your application reliable, and prepare it for production with features such as caching, rate limiting, exponential retry, model fallback, and more.

Repository GitHub projects open source

AgentMark

github.com

Type-Safe Markdown-based Agents

Repository GitHub projects open source

semantic-coverage

github.com

Visualizes RAG knowledge gaps and "blind spots" using 2D UMAP clustering and density detection.

Repository GitHub projects open source

Future AGI

github.com

Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.

Repository GitHub projects open source

traceAI

github.com

Open-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.

Repository GitHub projects open source

RagTune

github.com

CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.

Repository GitHub projects open source

onWatch

github.com

Lightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, <50MB RAM, zero telemetry, SQLite storage.

Repository GitHub projects open source

whylogs

github.com

The open standard for data logging

Repository GitHub projects open source

OpenTelemetry-based observability and monitoring for LLM and agents workflows.

Repository GitHub projects open source

Helicone

github.com

Open source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.

Repository GitHub projects open source

Great Expectations

github.com

Always know what to expect from your data.

Repository GitHub projects open source

QWED

github.com

Deterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods.

Repository GitHub projects open source

Fiddler AI

github.com

Evaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity.

Repository GitHub projects open source

EvalView

github.com

Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.

Repository GitHub projects open source

Azure OpenAI Logger

github.com

"Batteries included" logging solution for your Azure OpenAI instance.

Repository GitHub projects open source

Plexiglass

github.com

A Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms.

Repository GitHub projects open source

dstack

github.com

Open-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing.

Repository GitHub projects open source

brood-box

github.com

CLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization.

Repository GitHub projects open source

Kaito

github.com

A Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.

Repository GitHub projects open source

KubeAI

github.com

Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.

Repository GitHub projects open source

Xinference

github.com

Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the clou

Repository GitHub projects open source

ray-llm

github.com

LLMs on Ray - RayLLM (Archived)

Repository GitHub projects open source

lanarky

github.com

FastAPI framework to build production-grade LLM applications

Repository GitHub projects open source

langchain-serve

github.com

Serverless LLM apps on Production with Jina AI Cloud (Archived)

Repository GitHub projects open source

The Triton Inference Server provides an optimized cloud and edge inferencing solution.

Repository GitHub projects open source

Torchserve

github.com

Serve, optimize and scale PyTorch models in production (Archived)

Repository GitHub projects open source

TFServing

github.com

A flexible, high-performance serving system for machine learning models.

Repository GitHub projects open source

Mosec

github.com

A machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.

Repository GitHub projects open source

Jina

github.com

Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native

Repository GitHub projects open source

x-stable-diffusion

github.com

Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived)

Repository GitHub projects open source

whisper.cpp

github.com

Port of OpenAI's Whisper model in C/C++

Repository GitHub projects open source

whisper-ctranslate2

github.com

is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filter

Repository GitHub projects open source

Large Language Model Text Generation Inference

Repository GitHub projects open source

Rapid-MLX

github.com

OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.

Repository GitHub projects open source

Off Grid

github.com

Open-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline.

Repository GitHub projects open source

Modelz-LLM

github.com

OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)

Repository GitHub projects open source

LLMKube

github.com

Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.

Repository GitHub projects open source

FlexGen

github.com

Running large language models on a single GPU for throughput-oriented scenarios. (Archived)

Repository GitHub projects open source

Faster Whisper

github.com

fast inference engine for whisper in C++ using CTranslate2.

Repository GitHub projects open source

Clip-as-a-service

github.com

serving the OpenAI CLIP model

Repository GitHub projects open source

CTranslate2

github.com

fast inference engine for Transformer models in C++

Repository GitHub projects open source

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.