vLLM 官方博客T1·68 pts
vLLM Optimizes DeepSeek-V4.1-Flash: 5.3x Agentic Throughput via New Kernels and CUDA Graphs
Original: DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
- DeepSeek-V4.1-Flash achieves 5.3x throughput gain on vLLM, with 1.9x speedup at low concurrency.
- Integration of new kernels like MegaAttention and SWA bounded replay significantly reduces TTFT.
- Leveraging the causal encoder-decoder architecture with FP4 KV Cache drastically compresses memory footprint.