To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
| Type | Repository |
| Section | GitHub projects |
| Pricing | open source |
| Platform | Self-hosted |
| Systems | исходный код |
| Site language | en |
| GitHub | microsoft/MInference |