Why does DeepSeek's long-context processing suddenly "glitch"? ByteDance Seed team unravels the mystery of AI performance fluctuations
AIbase Base
PublishedAI News · 1 minute read · Oct 9, 20262In its latest research paper, ByteDance's Seed team has revealed the technical reasons behind the performance fluctuations that large language models exhibit when processing ultra-long texts. The research team points out that this phenomenon mainly stems from the "phase sensitivity" introduced by chunked KV cache compression.
Key mechanism causes retrieval bias
To reduce memory consumption during long-context inference, this technique compresses windows of continuous tokens into fewer entries at a fixed stride. However, this compression mechanism introduces an entirely new position coordinate: the "phase" of a Token relative to the boundary of the compressed window.
Experiments show that when models face exactly the same information, different phases lead to drastic changes in retrieval difficulty. In some large open-source models, the difference in long-text retrieval accuracy caused by this phase gap can reach up to 40 percentage points.
Periodic weaknesses urgently need optimization
This finding reveals the limitations of traditional averaged benchmark testing: beneath some models' seemingly excellent average high scores may lie serious periodic weaknesses. As long-text large model applications become increasingly widespread, this research provides an important theoretical basis for future model architecture optimization and improvements in long-context inference stability.
This article comes from the AIbase Daily
Welcome to the [AI Daily] column! This is your guide to exploring the world of artificial intelligence every day. Each day we present you with hot topics in the AI field, focusing on developers and helping you stay on top of technology trends and learn about innovative AI products and applications.
—— Created by the AIbase Daily team © Copyright AIbase Base 2024, click to view the source -