Git
Tools to read, search, and manipulate Git repositories
Tools to read, search, and manipulate Git repositories
Secure file operations with configurable access controls
Web content fetching and conversion for efficient LLM usage
WhatsMCP
one of the very first AutoML methods and open-source software packages.
Open source AI-powered multi-cluster Kubernetes dashboard for managing LLM workloads across hybrid edge and cloud environments. GPU monitoring, benchmark streaming, real-time observability with 20+ CNCF integrations, and AI-guided cluster operations. CNCF Sandbox project.
Rich dashboards, reports, and UMAP to perform root cause analysis, pinpoint problem areas, like correctness, safety, and privacy issues, and improve LLM outcomes.
Open source, high performance fine tuning as a service for GPT4 quality models with 5x lower latency and 3x lower cost
Managed PostgreSQL with pgvector for AI workloads. Built-in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage.
A suite of LLMOps tools within the developer-first W&B MLOps platform. Utilize W&B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations.
All In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which
LLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md.
The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in.
Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).
The simple prompt engineering and evaluation tool designed for developers building AI applications.
Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.
Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground.
Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.
Multi-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback.
An open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing.
Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.
Bring a human into the loop in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results.
Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.
Runtime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP API.
Serverless open-source platform for building long-running LLM agents with tool use.
A compact ~450M parameter VLA by Hugging Face, designed to be computationally efficient and accessible, running on consumer GPUs or CPUs. Part of the LeRobot ecosystem.
Gemma is a family of lightweight, open models built from the research and technology that Google used to create the Gemini models.
An open-source set of focused agent skills for designing, implementing, testing, reviewing, and shipping software changes.
a benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on their performance in different aspects such as natural language understanding, reasoning, and generalization.
benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from domains like chemistry, physics, and mathematics.
a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).
a benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.
A pioneering benchmark specifically designed to assess honesty in LLMs comprehensively.
A Challenging, Contamination-Free LLM Benchmark.
Data Science related publications on medium
an international journal devoted to applications of statistical methods at large
Free Download
Linear Algebra course by Gilbert Strang
Slides, scripts and materials for the Machine Learning in Finance course at NYU Tandon, 2022.
Interactive calculator that visualizes the step-by-step manual math behind machine learning algorithms for exam prep.
you can run on the browser with IPython.
Supplement Claude's knowledge with external data sources.
Learn how to integrate Claude with external tools and functions to extend its capabilities.
Discover techniques for effective text summarization with Claude.
Learn how to enhance Claude's responses with external knowledge.
Explore techniques for text and data classification using Claude.
Concise cheat sheets containing machine learning best practices.