Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 91.6k 22.1k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 6.8k 1.7k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 415

  4. llm-compressor llm-compressor Public

    Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

    Python 3.8k 661

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 827 218

  6. semantic-router semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    Go 5.8k 937

Repositories

Showing 10 of 49 repositories
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 91,629 Apache-2.0 22,137 2,394 (32 issues need help) 5,000+ Updated Sep 13, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 431 Apache-2.0 311 93 (2 issues need help) 397 Updated Sep 13, 2026
  • ci-infra Public

    This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

    vllm-project/ci-infra's past year of commit activity
    Python 55 Apache-2.0 83 0 70 Updated Sep 13, 2026
  • humming Public

    Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

    vllm-project/humming's past year of commit activity
    Python 230 Apache-2.0 39 7 10 Updated Sep 13, 2026
  • semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    vllm-project/semantic-router's past year of commit activity
    Go 5,779 Apache-2.0 937 373 (3 issues need help) 170 Updated Sep 13, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Ascend

    vllm-project/vllm-ascend's past year of commit activity
    C++ 2,811 Apache-2.0 2,239 1,435 (77 issues need help) 1,782 Updated Sep 13, 2026
  • vllm-omni Public

    A framework for efficient model inference with omni-modality models

    vllm-project/vllm-omni's past year of commit activity
    Python 6,782 Apache-2.0 1,715 826 (152 issues need help) 1,125 Updated Sep 13, 2026
  • vllm-metal Public

    Community maintained hardware plugin for vLLM on Apple Silicon

    vllm-project/vllm-metal's past year of commit activity
    Python 1,734 Apache-2.0 252 12 12 Updated Sep 13, 2026
  • aibrix Public

    Cost-efficient and pluggable Infrastructure components for GenAI inference

    vllm-project/aibrix's past year of commit activity
    Go 5,086 Apache-2.0 689 337 (22 issues need help) 42 Updated Sep 13, 2026
  • vllm-project/vllm-dashboard's past year of commit activity
    TypeScript 13 13 2 10 Updated Sep 13, 2026