量子位

The reason why the bytes are strong sometimes and weak sometimes when finding DeepSeek is

Whether to answer correctly depends on the Token's position

Image source · 量子位

Wen Le, from Ao Fei Si

Qbit | Public Account QbitAI

The dual nature of gods and spirits in DeepSeek has been captured by the ByteSeed team.

The same problem, nothing is changed, just a few irrelevant characters added at the beginning, why does the model suddenly stop working?????

And it’s not just occasional outbursts.

Researchers at Seed found that DeepSeek-V4 performance actually changes periodically every 4 Tokens, depending on the position of the information input.

What does that mean? If the model doesn’t remember something, why do we still need to see where that thing appears in the input??

There are two more Tokens in front. The answer might change from wrong to correct, and if there are two more, it will be changed back again.

Okay, okay, the model’s answering starts to take into account its position.

In a 128K-length context retrieval test, researchers found that even when the same piece of information was simply moved to a different position, the retrieval accuracy of the DeepSeek-V4 series models could at most gain 40.2 percentage points.

The team further found that this issue is related to a long context optimization technique used by DeepSeek-V4.

Block KV Cache Compression.

This technology was originally designed to make it possible for models to process long texts more efficiently and with less memory usage.

After the result compression, it makes the model look like a ghost sometimes and a demon other times (doge).

Two more Tokens, and DeepSeek suddenly gets it right

The ByteSeed researchers initially conducted an experiment using DeepSeek’s own code.

The test subject is DeepSeek-V4-Flash-Base.

They extracted a FP8 quantization function from the official reasoning code of DeepSeek-V4, allowing the model to complete the last Token.

The correct answer to the task should be 8, because this code needs to complete the type conversion related to FP8.

But sometimes the model thinks it should be made up to 32.

To figure out what was going on, the researchers added a purely decorative document string at the beginning of the code, which contained some repeated equal signs.

Then, they began to adjust the number of equal signs.

The code itself remains unchanged, the positions that need to be filled are also unchanged, and the correct answers are of course unchanged too... The only change is that several meaningless Tokens have been added at the beginning.

As a result, DeepSeek’s answer began to jump back and forth repeatedly.

When the filling length falls at certain positions, the model tends to answer incorrectly 32.

Move one or two Tokens forward, and it starts to tend towards the correct 8.

Keep moving again, and the wrong answer appears once more—

The entire process repeats every 4 Tokens.

More specifically, among the 16 filling lengths tested in the paper, when the remainder of the length modulo 4 is 0 or 1, the model tends to give the wrong answer 32; when the remainder is 2 or 3, it tends to give the correct answer 8.

The researchers also counted the probabilities that the model assigned to the two candidate answers.

At one set of positions, the average probability for the wrong answer 32 reached 71.3%, while the correct answer 8 had only 26.4%.

When moved to another group position, the situation is directly reversed:

The average probability for the correct answer 8 rose to 91.5%, while the probability for the wrong answer 32 was only 7.2%.

That is to say, the length changes in the previous unrelated characters are sufficient for the model to make completely different judgments on the same question.

This is a bit difficult to evaluate.

Programmers debug code by first checking the logic, variables, and dependencies.

It seems that we might still need to check carefully to see if two extra equal signs were added earlier.

However, a single code completion case is not enough to show how common this issue is.

So, the Seed team continued to expand the testing.

They turned their attention to a classic task within the large model’s long context capabilities, finding a needle in a haystack.

Researchers created a context with 128K tokens, which contains approximately 16,000 key-value pairs.

For example, K1 corresponds to V1, K2 corresponds to V2...

Then let the model find the Value corresponding to the specified Key.

During the testing process, the key-value relationship remains unchanged, the issue persists, and the total length of the context also stays consistent.

Researchers focused on adjusting the position of the target information relative to the compression window boundaries.

It was found that the accuracy curve of the DeepSeek-V4 series shows very obvious periodic fluctuations.

The maximum accuracy gap between different positions for DeepSeek-V4-Flash-Base reaches 40.2 percentage points

DeepSeek-V4-Pro-Base also reached 34.8 percentage points.

After post-training, the situation improved.

The gap for DeepSeek-V4-Flash-0731 reduced to 19.1 percentage points, while that for DeepSeek-V4-Pro-0813 decreased to 14.8 percentage points.

And the updated DeepSeek-V4.1-Flash-0910, the gap further reduced to 6.1 percentage points.

But the periodic differences still exist.

How does this difference arise?

Looking closer, the fluctuation cycle of DeepSeek-V4 is 4 Tokens, while DeepSeek-V4.1 becomes 2 Tokens.

Researchers found that this exactly corresponds to the KV Cache compression step size used by each generation model.

Okay, even the fluctuations in answering performance have matched the underlying compression settings.

Is the problem with KV Cache compression?

So let’s talk about KV Cache compression.

When large models process long contexts, they need to store a large amount of Key and Value information corresponding to historical Tokens for subsequent attention calculations.

The longer the context, the more memory and computational overhead this part of the cache consumes.

In particular, for long tasks that involve hundreds of thousands or even millions of Tokens, the KV Cache easily becomes a bottleneck in reasoning efficiency.

So, DeepSeek-V4 uses chunk KV Cache compression.

The idea is to divide continuous Tokens into individual windows, and compress the information within the windows into fewer cache entries.

Thus, the model does not need to maintain a cache of the same size for each historical token.

It saves memory and reduces the cost of calculating attention for long contexts.

But the Seed team found that the reason why DeepSeek exhibits both supernatural and spiritual characteristics may lie in this chunking process.

Assume that every 4 Tokens constitute a compression step size.

Then, if the same piece of information appears in position 1, 2, 3, or 4 in the window, the conditions under which it is compressed may be different.

The paper refers to this position relative to the compression window boundaries as Phase.

Researchers found that the model has systematic differences in its retrieval ability for information from different phases.

They named this phenomenon Phase Sensitivity, Phase Sensitivity.

For example, the same set of data is given to the model:

The first type of layout ensures that key numbers are placed in positions where the model can easily retain them;

Second layout: just a few more words added at the beginning, and the position of the key numbers relative to the compression window has changed.

It’s still the same data, but whether the numbers can be identified from the model later may differ significantly.

Moreover, this problem cannot be simply attributed to “the information being cut between two windows”.

Researchers found that even when the Key and Value fall within the same compression window, the retrieval accuracy at different positions can vary significantly.

This indicates that the issue also involves how the model writes information into the compression cache, and how it reads it back from the cache later.

For further confirmation, the Seed team simply trained a set of models from scratch by themselves.

They based their work on the Qwen3-0.6B architecture, developed various KV Cache compression schemes, and used a full-attention model without chunking compression as a control group.

The focus is only on changing the compression mechanism, and we will see if periodic fluctuations appear accordingly.

As a result, all block compression models that passed the test showed periodic changes corresponding to the compression step size.

On the contrary, the full attention baseline model did not exhibit the same level of periodicity.

The researchers also adjusted the window size and compression step size separately, and found that the period mainly changed with the compression step size.

The step size is 4, so the performance fluctuates around a cycle of approximately 4 Tokens;

The step length is 6, and the cycle also becomes approximately 6 Tokens;

The step length is 8, and the same applies;

……

Even without using RoPE position encoding, or replacing the learnable compression weights with simple averaging, this phenomenon still exists.

In other words, the problem is not caused by a specific position encoding or a special module alone.

This design of block compression itself may introduce periodic retrieval weaknesses.

The Seed team further intervened in the attention head, observing how the retrieval ability of the model changed across different phases after removing various components.

It was found that the contributions of different attention heads to different phases are not the same.

Some heads are better at processing information in certain positions, while other heads play a greater role in other locations.

Researchers call it Phase Specialization.

That is to say, it seems that a division of labor has formed within the model: different attention components have developed preferences for different positions in the compression window.

This division of labor can help the model complete information retrieval, but it may also make certain locations relatively weak.

The paper also analyzes the training process through a simplified theoretical model and finds that gradient flow may drive the compression module to form stable position preferences.

This also explains why this periodicity is not simply random noise. It may be a natural result of the model learning how to compress information.

But such issues may not be directly visible from a regular Benchmark, as standard evaluations usually aggregate a large number of test results into a average score.

Assume that the model performs well in some positions, but falls behind significantly in others, yet the overall average score may still be good.

So the Seed team suggested that when evaluating models using block KV Cache compression, we should not only focus on the overall retrieval accuracy, but also place the same piece of information in different compression phases and measure them separately.

Even if there are performance deviations caused by such positions, and considering the memory and computational costs of long context inference, compression still has significant practical value.

But after saving cache, the model may no longer treat information at different positions equally.

Of course, based on the results of this experiment, post-training and architecture iteration do indeed significantly reduce this gap.

Paper address: https://arxiv.org/pdf/2609.36322

Original source

量子位

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original