Today for AI

IT之家 · 智能时代 · 10/9/2026, 08:49:58

ByteDance Seed Reveals DeepSeek Long-Context Drift: Chunked KV Cache Compression Causes Phase Sensitivity with 40% Accuracy Fluctuation

By IT之家Original title: 字节 Seed 团队发现 DeepSeek“抽风”原因,长上下文可能性能漂移
78AI Score
Executive Summary

ByteDance Seed team published a paper identifying that DeepSeek's chunked KV cache compression introduces 'phase sensitivity,' causing periodic performance drift in long-context retrieval. The study reveals systematic asymmetry where identical information is easily retrieved at one phase but difficult at another, with accuracy fluctuations reaching up to 40 percentage points in large open-weight models, exposing weaknesses masked by average benchmark scores.

SOURCE COVERAGEOriginal coverage

On October 9, it was reported that the ByteDance Seed team submitted a paper to the preprint platform arXiv in late September. The paper discusses phase sensitivity caused by chunked KV cache compression, directly pointing to the cause of DeepSeek’s erratic behavior.

ByteDance Seed Team Discovers Cause of DeepSeek's Erratic Behavior; Long-Context Performance May Drift

The study evaluated the base and post-trained versions of DeepSeek-V4-Flash and DeepSeek-V4-Pro, as well as the post-trained DeepSeek-V4.1-Flash.

Models utilize chunked KV cache compression to compress windows of consecutive tokens into fewer cache entries with a fixed stride, thereby reducing memory and attention costs for long-context inference.

However, this compression introduces a new positional coordinate: the token’s phase, or its position relative to the compression window boundaries.

ByteDance Seed Team Discovers Cause of DeepSeek's Erratic Behavior; Long-Context Performance May Drift

The team discovered systematic asymmetries in models using this compression: identical information is easily retrieved at one phase but difficult to retrieve at another. The team termed this periodic variation in retrieval performance "phase sensitivity."

In large open-weight models employing such compression, long-context retrieval accuracy can vary by up to 40 percentage points across different phases, revealing periodic weaknesses that average benchmark scores may obscure.

ByteDance Seed Team Discovers Cause of DeepSeek's Erratic Behavior; Long-Context Performance May Drift

IT Home provides the paper link below: