雷峰网 AI

ByteDance discovers hidden DeepSeek bug: add a few spaces and the model falls apart

DeepSeek's money-saving killer feature — why did it become a fatal blind spot? Author丨Gao Yunyi Editor丨Cen Feng Ask DeepSeek the same question twice, just adding a few extra spaces, and the answers it gives are like one…

DeepSeek's money-saving killer feature — why did it become a fatal blind spot?

    Author丨Gao Yunyi

Editor丨Cen Feng

                                                                                                       

Ask DeepSeek the same question twice, just adding a few extra spaces, and the answers it gives are like one from a "genius" and one from an "idiot." In late September (9月底), the ByteDance team's latest paper"Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression", directly called out the DeepSeek V4 series models'"divine-demonic duality" problem. Caption: paper link https://arxiv.org/pdf/2609.36322 This is essentially "periodic amnesia" caused by corner-cutting in the underlying algorithm.Chunked KV Cache compression technologycuts long contexts into blocks and compresses them, which brings about a problem:"position determines destiny". If your core words happen to fall into the "information blind zone" at a chunk boundary, the model can only spout nonsense. Imagine having a large model review dozens of pages of a contract, and a key clause happens to land in the "information blind zone" — the model simply ignores that information and gives the opposite conclusion. And this error may be caused by you accidentally adding some whitespace characters; the model falls apart outright, and you have no idea where the problem lies. And this inherent "divine-demonic duality" shortcoming is by no means unique to DeepSeek. In this era of shouting "bring down the cost of long texts," any model that adopts this kind of chunked compression or sparsification technology has similar problems to a greater or lesser degree, such as the open-source Llama 3 series.

01


ByteDance's paper catches DeepSeek's weak spot

ByteDance's initial discovery came from having DeepSeek-V4-Flash-Base complete a piece of code. The task was to complete the last Token of an FP8 quantization function. According to the code logic, the correct result should have been 8, but because the researchers added a string of purely decorative equal signs at the very beginning of the code, the model output 32. Even stranger, this error was not random but directly tied to the number of equal signs:When the number of equal signs is divided by 4, if the remainder is 0 or 1, the model most likely gets it wrong, with an error rate as high as 71.3%; if the remainder is 2 or 3, the model gets it right again, with an accuracy as high as 91.5%.Following the thread, the ByteDance team discovered this is an inherent shortcoming left behind after the large model adoptedchunked KV Cache compression technology. Chunked KV Cache compression is DeepSeek's killer feature for solving the memory pressure problem of long contexts. It cuts consecutive Tokens into fixed-length segments, and the Tokens in each segment can be compressed into a "small summary" through a learnable gating layer, which instantly relieves the memory pressure of long contexts. This solves the memory problem but introduces a new one. In a standard Transformer, attention depends only on the relative distance between Tokens and doesn't care about absolute position. No matter where a sentence is placed, as long as the word order is unchanged, its meaning is unchanged.But after DeepSeek adopted chunked KV Cache compression, each Token's position within the compressed segment became important; this position is called the "phase." The gating layer assigns different weights to Tokens at different positions, and this weight is called the slot gain factor.Tokens slotted into strong slots are retained with higher weight, and the model can answer accurately; Tokens slotted into blind-spot slots are almost ignored, and the model gets it wrong.This ranking-sensitivity phenomenon is called phase sensitivity.Digging further into the model's internals, the researchers also found that attention heads formed a natural division of labor: some specialize in handling positions near the front of a segment, others in positions near the back. This phase specialization turns certain positions into weak spots.To verify this was not a coincidence, the ByteDance team also conducted a "needle in a haystack" test. The researchers constructed a context of up to 128K Tokens containing about 16,000 key-value pairs, with each Key corresponding to a Value. With the key-value relationships unchanged, the question unchanged, and the total length of the context kept consistent, if one key-value pair was placed at different positions, could the model still correctly answer the corresponding Value? The results showed that DeepSeek-V4-Flash-Base had a maximum accuracy gap of 40.2 % between different positions. With DeepSeek-V4-Pro-Base, the gap reached 34.8 %, and even with post-training optimization the problem persisted, with gaps of 19.1% and 14.8%. Facing this periodic blind-spot problem, DeepSeek tried DeepSeek-V4.1-Flash, simply halving the compression step. When the compression step is set to 4, each round of compression divides the data into 4 slots. The information weight allocated to different slots varies greatly, and Tokens at edge positions are easily dropped in two consecutive rounds of window compression. After the step was lowered to 2, compression became finer, merging only 2 Tokens at a time. Within one compression cycle, a Token has only two possible relative positions. That way, even if key information falls on a weaker odd position, since the adjacent position differs by only 1 Token, the context is very hard to lose entirely. DeepSeek-V4.1-Flash indeed reduced the accuracy gap further to 6.1 %, but the problem was not completely eliminated. The ByteDance team then trained a batch of models from scratch using the Qwen3-0.6B architecture, with all other configurations identical, changing only the compression mechanism. The results showed:
  • All models using chunked compression exhibited periodic accuracy fluctuations; the full-attention baseline model without compression showed no such phenomenon;
  • The fluctuation period strictly equaled the compression step;
  • Removing the RoPE positional encoding, the fluctuations persisted — it was not a bias introduced by positional encoding;
  • Replacing the learnable gating weights with the simplest uniform allocation still produced periodic fluctuations — it was not poorly tuned parameters.
Caption: with the compression step set to 4, 6, 8, and 12 respectively, the period of the accuracy fluctuation is also 4, 6, 8, and 12. This shows thatas long as chunked compression is performed at fixed lengths, this problem is unavoidable.This is a structural defect of the approach itself, one that cannot be eradicated by tuning parameters, swapping positional encodings, or post-training. Some netizens have pointed out that to truly solve this problem, you have to make the chunking itself "come alive", dynamically deciding boundaries according to the importance of the content. When Tokens have no fixed phase, phase differences become meaningless. There is indeed related research in the AI field, called"dynamic chunking"。

02


Is "dynamic chunking" technology the antidote?

The core of dynamic chunking is abandoning the one-size-fits-all mode of a fixed Token count, flexibly adjusting based on semantics and importance.This technical idea performs a top-down reconstruction of the entire technology stack, from algorithmic mathematics and memory management down to the underlying hardware.

▎Layer One: Rebuilding the Mathematical Foundations of Attention

How exactly should dynamic chunking cut? The core is "size follows content". Low-information-density filler gets merged into large blocks directly; when encountering key turning points and high-density information, it is split finely enough for precise computation. This way, no matter where key information is hidden, the model can keenly capture it, mathematically reconstructing the translation invariance broken by fixed chunking. Going further, this endows the model with a kind of "anticipatory ability": first scan the information density, then decide how large to cut. The inference process no longer reads through at a single pace, but reads quickly where it should and savors slowly where it should, growing an instinct for "fast-slow thinking" at the architectural level.

▎Layer Two: Forcing the Evolution of Memory Management

Previously, to make matrix computation fast, the KV Cache was managed in fixed-size blocks. The advantage of this approach is that it is simple, tidy, and easy to operate, but memory utilization could not improve, and compute was easily wasted. After dynamic chunking, the underlying traditional "static contiguous matrices" can no longer cope, and must inevitably evolve toward "topologically sparse matrices", relying on a layer of dynamic indexing to perform addressing. Even better, research has found that the dynamic chunking boundaries between different Transformer layers are highly similar and can be reused directly. This greatly amortizes the overhead of dynamic addressing. In the end, the upper layer cuts smartly, the lower layer stores more economically, and the addressing in between is not slow.

▎Layer Three: Cracking the Hardware Alignment Problem

This is also the hardest of the three layers to break through. GPUs naturally favor regular matrix operations whose edge lengths are divisible by 8, 16, 32, which is the most realistic hardware reason the whole industry had long insisted on fixed chunking. If chunk sizes are uneven, the acceleration units built for matrix computation (such as Tensor Cores) will idle heavily. For this problem, the industry has already worked out viable solutions. For example, zero-padding algorithms can, at the physical memory level, tightly pack dynamic semantic blocks of varying lengths together without wasting space, while logically preserving the boundaries of each semantic block. Currently,the world's top AI labs are advancing dynamic chunking technology along two directions: "semantic continuity" and "sparse computation"and this technology has become the core breakthrough point for the new generation of long-context large models to gain speed and improve quality.The most representative achievement is ChunkKV, proposed by a team at the Hong Kong University of Science and Technology (Guangzhou), which takes the "semantic continuity" route. In the past, mainstream KV compression methods, whether H2O, SnapKV, or PyramidKV, all shared the core idea of scoring individual Tokens, keeping the important and discarding the unimportant. Such filtering easily shatters a complete sentence of subject, verb, and object into fragments. ChunkKV's core idea is to use contiguous semantic blocks as the basic unit of compression: either keep the whole block or discard the whole block. What remains are complete phrases and sentences. In the NIAH (needle in a haystack) benchmark, which most tests long-text retrieval ability, when the KV cache size is limited to 128, ChunkKV based on LLaMA-3 achieves an accuracy of 73.8%, while the traditional method SnapKV only reaches 58.9%. This technology directly remedies large models' weakness of "not understanding, not finding accurately" in long-text scenarios, greatly improving the model's practical usability.
Besides ChunkKV, for example, Dr. Han Xiao's team at Jina AI, focusing on retrieval-scenario optimization, proposed a late chunking scheme that solves the fatal problem of traditional intelligent retrieval of "first splitting the text, then parsing the content": first let the model read the entire text completely, fully absorbing the global content and contextual logic, and only before finally outputting the result, split the text along natural semantic boundaries, preserving complete information to the greatest extent. On the "sparse computation" side,Microsoft Research Asiadeveloped the MInference framework, specifically designed for ultra-long text inference scenarios at the million-word scale. The team found that when large models read long texts, not every part of the content requires fine-grained computation; the attention matrix contains exploitable fixed sparse patterns. Based on this characteristic, the framework matches the most suitable sparse computation method to each attention head. Without any reduction in retrieval accuracy, it achieves up to a 10x speedup in the prefill stage for million-word long contexts. In the first half of the long-context era, KV compression technology deserves great credit in the race over "who can process longer". But the essence of long context was never length for its own sake. When users start truly using long contexts for real tasks, being usable, useful, and lossless becomes the new KPI. References: https://arxiv.org/pdf/2609.36322https://www.zhihu.com/question/2091513666864682278

Hop on board — Leiphone (WeChat account: Leiphone) takes you through the highlights of top global AI conferences

Exclusive access to:

Expert presentation slides

Full conference reports

Interpretations of popular papers

Interviews with rising academic stars

Scan the QR code above

or click "Read Original" to follow the special section.

This is an original Leiphone article; unauthorized reproduction is prohibited. See thereprint notice for details。

Original source

雷峰网 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original