distilabel
??️ distilabel is a framework for synthetic data and AI feedback for AI engineers that require high-quality outputs, full data ownership, and overall efficiency.
??️ distilabel is a framework for synthetic data and AI feedback for AI engineers that require high-quality outputs, full data ownership, and overall efficiency.
An open source feature store for machine learning.
The Virtual Feature Store. Turn your existing data infrastructure into a feature store.
A Python library that facilitates fast and easy data exploration by automating the visualization and data analysis process.
A CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.
Modern columnar data format for ML implemented in Rust.
Git-like capabilities for your object storage.
A distributed POSIX file system built on top of Redis and S3.
A self-organizing data hub for S3.
Pachyderm is a version control system for data.
Storage layer that brings scalable, ACID transactions to Apache Spark and other engines.
Git for Data.
A version control system to manage large files. Lake is a dataset format with a simple API for creating, storing, and collaborating on AI datasets of any size.
FastEdit aims to assist developers with injecting fresh and customized knowledge into large language models efficiently using one single command.
AI evaluation platform for interactively exploring data and model outputs.
ML models and internal tensors 3D visualizer.
A python library for decision tree visualization and model interpretation.
Neural network 3D visualization framework, build interactive and intuitive model in browsers, support pre-trained deep learning models from TensorFlow, Keras, TensorFlow.js.
TensorFlow's Visualization Toolkit.
Bring multiple data streams into one dashboard.
A model-agnostic visual debugging tool for machine learning.
A developer first, lightweight, user-friendly experiment tracking and visualization tool for machine learning projects, streamlining collaboration and simplifying MLOps. W&B excels at tracking LLM-powered applications, featuring W&B Prompts for LLM execution flow visualization, input and output mon
LabNotebook is a tool that allows you to flexibly monitor, record, save, and query all your machine learning experiments.
Kedro-Viz is an interactive development tool for building data science pipelines with Kedro. Kedro-Viz also allows users to view and compare different runs in the Kedro project.
Machine Learning automation and tracking.
Experiment tracking, ML developer tools.
Auto-Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-Management
A minimalist neural network library optimized for sparse data and single machine environments.
Machine Learning in Python.
Deep learning framework to train, deploy, and ship AI products Lightning fast.
Machine Learning Framework from Industrial Practice.
OneFlow is a performance-centered and open-source deep learning framework.
MindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.
Metric Learning Algorithms in Python.
MegEngine is a fast, scalable and easy-to-use deep learning framework, with auto-differentiation.
Kedro is an open-source Python framework for creating reproducible, maintainable and modular data science code.
Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
A tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.
Lightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler.
🚀 A simple way to train and use PyTorch models with multi-GPU, TPU, mixed-precision.
Train transformer language models with reinforcement learning.
Efficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance.
An optimized prompt tuning strategy achieving comparable performance to fine-tuning on small/medium-sized models and sequence tagging challenges. (ACL 2022)
Using Low-rank adaptation to quickly fine-tune diffusion models.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
A PyTorch Lightning extension that accelerates and enhances foundation model experimentation with flexible fine-tuning schedules.
Instruct-tune LLaMA on consumer hardware
A build, packaging, and run system for ephemeral multi-container environments.