Advanced search

GitHub projects

4 835 entries

Git

github.com

Tools to read, search, and manipulate Git repositories

Tool GitHub projects

Filesystem

github.com

Secure file operations with configurable access controls

Tool GitHub projects

Fetch

github.com

Web content fetching and conversion for efficient LLM usage

Tool GitHub projects

TPOT

automl.info

one of the very first AutoML methods and open-source software packages.

Tool GitHub projects

KubeStellar Console

console.kubestellar.io

Open source AI-powered multi-cluster Kubernetes dashboard for managing LLM workloads across hybrid edge and cloud environments. GPU monitoring, benchmark streaming, real-time observability with 20+ CNCF integrations, and AI-guided cluster operations. CNCF Sandbox project.

Tool GitHub projects

Fiddler AI

github.com

Rich dashboards, reports, and UMAP to perform root cause analysis, pinpoint problem areas, like correctness, safety, and privacy issues, and improve LLM outcomes.

Tool GitHub projects

Flyflow

github.com

Open source, high performance fine tuning as a service for GPT4 quality models with 5x lower latency and 3x lower cost

Tool GitHub projects

Rivestack

rivestack.io

Managed PostgreSQL with pgvector for AI workloads. Built-in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage.

Tool GitHub projects

A suite of LLMOps tools within the developer-first W&B MLOps platform. Utilize W&B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations.

Tool GitHub projects

TreeScale

treescale.com

All In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which

Tool GitHub projects

TeamoRouter

router.teamolab.com

LLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md.

Tool GitHub projects

Puzzlet AI

puzzlet.ai

The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in.

Tool GitHub projects

Prompteams

prompteams.com

Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).

Tool GitHub projects

PromptFoundry

promptfoundry.ai

The simple prompt engineering and evaluation tool designed for developers building AI applications.

Tool GitHub projects

PromptHub

prompthub.us

Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.

Tool GitHub projects

Parea AI

parea.ai

Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground.

Tool GitHub projects

Manag.ai

manag.ai

Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.

Tool GitHub projects

Literal AI

literalai.com

Multi-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback.

Tool GitHub projects

MLflow

github.com

An open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing.

Tool GitHub projects

Izlo

getizlo.com

Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.

Tool GitHub projects

gotoHuman

gotohuman.com

Bring a human into the loop in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results.

Tool GitHub projects

Dataoorts

dataoorts.com

Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.

Tool GitHub projects

ClevAgent

clevagent.io

Runtime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP API.

Tool GitHub projects

Open Responses

docs.julep.ai

Serverless open-source platform for building long-running LLM agents with tool use.

Tool GitHub projects

SmolVLA

huggingface.co

A compact ~450M parameter VLA by Hugging Face, designed to be computationally efficient and accessible, running on consumer GPUs or CPUs. Part of the LeRobot ecosystem.

Tool GitHub projects

Gemma

kaggle.com

Gemma is a family of lightweight, open models built from the research and technology that Google used to create the Gemini models.

Tool GitHub projects

Blueprint

github.com

An open-source set of focused agent skills for designing, implementing, testing, reviewing, and shipping software changes.

Tool GitHub projects

We-Math

we-math.github.io

a benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.

Tool GitHub projects

SuperBench

fm.ai.tsinghua.edu.cn

a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on their performance in different aspects such as natural language understanding, reasoning, and generalization.

Tool GitHub projects

SciBench

scibench-ucla.github.io

benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from domains like chemistry, physics, and mathematics.

Tool GitHub projects

MixEval

mixeval.github.io

a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).

Tool GitHub projects

DreamBench++

dreambenchplus.github.io

a benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.

Tool GitHub projects

BeHonest

gair-nlp.github.io

A pioneering benchmark specifically designed to assess honesty in LLMs comprehensively.

Tool GitHub projects

LiveBench

livebench.ai

A Challenging, Contamination-Free LLM Benchmark.

Tool GitHub projects

Journal of Data Science

jds-online.org

an international journal devoted to applications of statistical methods at large

Tool GitHub projects

Linear Algebra

ocw.mit.edu

Linear Algebra course by Gilbert Strang

Tool GitHub projects

MLSys-NYU-2022

github.com

Slides, scripts and materials for the Machine Learning in Finance course at NYU Tandon, 2022.

Tool GitHub projects

StepByStepML

stepbystepml.com

Interactive calculator that visualizes the step-by-step manual math behind machine learning algorithms for exam prep.

Tool GitHub projects

Tool use

github.com

Learn how to integrate Claude with external tools and functions to extend its capabilities.

Tool GitHub projects

Summarization

github.com

Discover techniques for effective text summarization with Claude.

Tool GitHub projects

Classification

github.com

Explore techniques for text and data classification using Claude.

Tool GitHub projects

MLTables

github.com

Concise cheat sheets containing machine learning best practices.

Tool GitHub projects

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.