[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
-
Updated
Jan 17, 2026 - Cuda
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
🌱 A tiny, readable LLM serving engine with vLLM/SGLang-style features.
Deterministic intermediate representation for AI agents — compile, verify, execute, and replay structured intent.
A lightweight Bun + Express template that connects to the Testune AI API and streams chat responses in real time using Server-Sent Events (SSE)
Anthropic API 兼容层,支持凭据轮换、负载均衡和 Prompt Caching
Anllm is an llm inference engine from scratch powered by mlx on mac devices that supports the Qwen3 and Llama3.x family Int 4bit quantized models
AI gateway and observability infrastructure for production LLM systems
To associate your repository with the llm-infra topic, visit your repo's landing page and select "manage topics."