中文
欢迎来到 vLLM 中文博客。这里会收录已发布的中文文章,并展示仍待翻译的英文博客,方便持续跟踪翻译进度。
最新中文文章
-
vLLM × Mooncake:用分布式 KV Cache Pool 服务大规模 Agentic 推理
-
vLLM 原生 Hidden States 提取:投机解码训练数据不再绕路
-
Model Runner V2:更模块化、更快速的 vLLM 执行核心
-
P-EAGLE:在 vLLM 中实现并行投机解码加速 LLM 推理
-
vLLM Triton Attention 后端深度解析
-
不止于能跑:vLLM 如何在 AMD ROCm 上做原生推理优化
-
DeepSeek-V3.2 on GB300:性能表现与部署实践
-
在 Blackwell 上推动 vLLM Wide-EP 与大规模推理走向成熟(Part I)
-
GPT-OSS 在 NVIDIA Blackwell 上的性能优化:推动 Pareto 前沿
-
vLLM 流式输入与 Realtime API
-
深入解析 vLLM 新推出的 KV Offloading Connector:智能内存传输助力推理吞吐量最大化
-
vLLM Playground:管理和交互 vLLM 服务的现代化 Web 界面
-
vLLM-Omni 扩散缓存加速:TeaCache 与 Cache-DiT
待翻译文章
-
Run Highly Efficient Multimodal Agentic AI with NVIDIA Nemotron 3 Nano Omni Using vLLM
-
DeepSeek V4 in vLLM: Efficient Long-context Attention
-
The State of FP8 KV-Cache and Attention Quantization in vLLM
-
Disaggregated Serving for Hybrid SSM Models in vLLM
-
vLLM Korea Meetup 2026 Wrap-Up
-
Next-Level Inference: Why Your Single-Node vLLM Setup Needs Prefill-Decode Disaggregation
-
Announcing Gemma 4 on vLLM: Byte for byte, the most capable open models
-
Run Highly Efficient and Accurate Multi-Agent AI with NVIDIA Nemotron 3 Super Using vLLM
-
vLLM Semantic Router v0.2 Athena: ClawOS, Model Refresh, and the System Brain
-
Efficiently serve dozens of fine-tuned models with vLLM on Amazon SageMaker AI and Amazon Bedrock
-
Building Mixture-of-Models on AMD GPUs with vLLM-SR
-
vLLM Semantic Router v0.1 Iris: The First Major Release
-
Announcing vllm.ai Website and Some Community Updates
-
vLLM Large Scale Serving: DeepSeek @ 2.2k tok/s/H200 with Wide-EP
-
AMD × vLLM Semantic Router: Building the System Intelligence Together
-
Run Highly Efficient and Accurate AI Agents with NVIDIA Nemotron 3 Nano on vLLM
-
Encoder Disaggregation for Scalable Multimodal Model Serving
-
Token-Level Truth: Real-Time Hallucination Detection for Production LLMs
-
Diving into speculative decoding training support for vLLM with Speculators v0.3.0
-
vLLM Router: A High-Performance and Prefill/Decode Aware Load Balancer for Large-scale Serving
-
Advancing Low‑Bit Quantization for LLMs: AutoRound x LLM Compressor
-
Tracing Hanging and Complicated GPU Kernels Down To The Source Code
-
Announcing vLLM-Omni: Easy, Fast, and Cheap Omni-Modality Model Serving
-
Streamlined multi-node serving with Ray symmetric-run
-
Building Clean, Maintainable vLLM Modifications Using the Plugin System
-
Docker Model Runner Integrates vLLM for High-Throughput Inferencing
-
Signal-Decision Driven Architecture: Reshaping Semantic Routing at Scale
-
Shared Memory IPC Caching: Accelerating Data Transfer in LLM Inference Systems
-
Fast and Affordable LLMs serving on Intel Arc Pro B-Series GPUs with vLLM
-
No More Train-Inference Mismatch: Bitwise Consistent On-Policy Reinforcement Learning with vLLM and TorchTitan
-
Run Multimodal Reasoning Agents with NVIDIA Nemotron on vLLM
-
Chasing 100% Accuracy: A Deep Dive into Debugging Kimi K2's Tool-Calling on vLLM
-
From Monolithic to Modular: Scaling Semantic Routing with Extensible LoRA
-
Zero-Reload Model Switching with vLLM Sleep Mode
-
Now Serving NVIDIA Nemotron with vLLM
-
No More Retokenization Drift: Returning Token IDs via the OpenAI Compatible API Matters in Agent RL
-
vLLM TPU: A New Unified Backend Supporting PyTorch and JAX on TPU
-
SemiAnalysis InferenceMAX: vLLM and NVIDIA Accelerate Blackwell Inference
-
DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action
-
The First vLLM Meetup in Korea
-
vLLM Now Supports Qwen3-Next: Hybrid Architecture with Extreme Efficiency
-
vLLM Semantic Router: Next Phase in LLM inference
-
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
-
Serving Geospatial, Vision, and Beyond: Enabling Multimodal Output Processing in vLLM
-
Introduction to torch.compile and How It Works with vLLM
-
GLM-4.5 Meets vLLM: Built for Intelligent Agents
-
CUDA Core Dump: An Effective Tool to Debug Memory Access Issues and Beyond
-
vLLM Now Supports gpt-oss
-
MiniMax-M1 Hybrid Architecture Meets vLLM: Long Context, Fast Inference
-
Introducing vLLM Hardware Plugin, Best Practice from Ascend NPU
-
Accelerating RLHF with vLLM, Best Practice from OpenRLHF
-
Transformers modeling backend integration in vLLM
-
Llama 4 in vLLM
-
PTPC-FP8: Boosting vLLM Performance on AMD ROCm
-
Introducing AIBrix: A Scalable, Cost-Effective Control Plane for vLLM
-
Distributed Inference with vLLM
-
Introducing vLLM Inference Provider in Llama Stack
-
vLLM V1: A Major Upgrade to vLLM's Core Architecture
-
High Performance and Easy Deployment of vLLM in K8S with vLLM production-stack
-
Structured Decoding in vLLM: a gentle introduction
-
Installing and Developing vLLM with Ease
-
vLLM 2024 Retrospective and 2025 Vision
-
Serving LLMs on AMD MI300X: Best Practices
-
How Speculative Decoding Boosts vLLM Performance by up to 2.8x
-
vLLM v0.6.0: 2.7x Throughput Improvement and 5x Latency Reduction
-
vLLM’s Open Governance and Performance Roadmap
-
Announcing Llama 3.1 Support in vLLM
-
Notes on vLLM v.s. DeepSpeed-FastGen
-
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention