Open Source Hacker News (LLM)

vLLM v0.28.0

vLLMLLM inferenceKimi-K3DeepSeek V4

vLLM v0.28.0 is a major release with 584 commits from 270 contributors (76 new). A key focus is a performance push for Kimi-K3 across the stack: decode context parallel (DCP) support, fused FlashKDA decode and prefill kernels, SiTU activation support for MegaMoE, GEMM-RS for sequence parallelism, combined all-gathers delivering 1.5–3x kernel-level speedups, an adaptive speculative token budget improving DSpark TTFT by ~60%, and optional shared-expert sharding saving ~17 GiB of memory per GPU. Kimi-K3 also now runs on ROCm via the V2 model runner.

DeepSeek V4 support reaches end-to-end sparse MLA for plain decode, MTP, and DSpark speculative decoding, with AMD Quark NVFP4 support, reasoning-effort prompts and mappings, sparse top-k metadata kernel optimizations, narrowed eager CUDA graph regions, and ROCm enablement on gfx11 and gfx950. Speculative decoding gets DFlash2 with local convolution and a candidate selector, DSpark confidence-scheduled verification, and auto-enabled async scheduling for draft models.

Model Runner V2 matures with E/P/D disaggregation, weight offloading, multi-layer MTP KV cache support, encoder CUDA graphs, decoder token-wise pooling plus Transformers pooling models, attention-free models, and thinking_token_budget support. Tiered KV cache offloading adds disk offloading, out-of-tree secondary tier managers via module_path, partial secondary-tier load results, tiering metrics, and a canonical CPU layout for parallelism-agnostic offload.

The Rust frontend and gRPC path gains a standalone renderer, multimodal image inference over gRPC, explicit data-parallel rank routing, and RL lifecycle control, with protobuf schemas published to Buf. New defaults include max_num_batched_tokens raised from 8192 to 16384, prefix caching on by default for Mamba models, and Blackwell CUDA graph capture default raised to 1024. Breaking changes: bitsandbytes support moved to an out-of-tree plugin, Transformers bumped to 5.15.0, and deprecated calculate_kv_scales and override_attention_dtype removed. Release artifacts include wheels for CUDA 13.0 and ROCm, and Docker images for CUDA 13.0 and CUDA 12.9.

Read original →

← Back to home