An essay at insufferable.dev argues that the usual story about Chinese AI labs copying Western models is becoming harder to sustain. Its counterexample is model efficiency: Chinese labs are openly publishing techniques that may make long-running AI sessions dramatically cheaper, while OpenAI and Anthropic are cutting prices for cached input.
The technical center of the argument is the KV cache — the working memory a language model keeps so it does not have to recompute every earlier token each time it generates the next one.
- DeepSeek’s earlier MLA design compressed that cache by roughly 15×, according to the essay.
- Later techniques called Compressed Sparse Attention and Heavily Compressed Attention reduced it further.
- The essay says DeepSeek V4.1 Flash gets the global cache down to 890 bytes per token: roughly 437× smaller than DeepSeek V1 for the long-context use case being compared.
- That matters because this cache occupies expensive GPU memory. Shrinking it can lower the cost of serving coding agents and other long sessions.
- The author points to recent cache-read price cuts — 60% for the Anthropic model compared and 80% for the OpenAI model — as evidence that Western labs adopted similar optimizations.
That last step is also the weak point. Prices show that serving economics changed; they do not reveal a proprietary model’s architecture. The essay’s larger contribution is therefore not proof of copying, but a useful inversion: resource constraints pushed Chinese labs toward efficiency, and openly released results can pressure closed labs even while helping them reduce costs.
The 294-comment thread on Hacker News adds a necessary challenge to the essay’s strongest claims and several plausible answers to its unanswered question.
What the thread adds
- user43928 — “That inference wasn’t profitable is a widespread myth… I have no reason to believe that the leading US labs don’t have their own optimizations, or that they learned of this particular optimization from DeepSeek.”
- samuelknight — “The article claims sparse attention was copied from open weight but we can’t know that… sparse attention is an old area of active research.”
- the_origami_fox — “I found the main idea of MLA as very innovative but it was documented in a strange paper with other ideas I couldn’t comprehend as being useful, and they hadn’t properly reported the positives or negatives of MLA.”
- Reptur — “Open releases are just the obvious move when you’re not the incumbent. You commoditize the thing your competitors charge for and get distribution you could never buy.”
- carbonguy — “Making open-weights models even cheaper and easier to run expands that ‘market’ and increases competitive pressure on the Big Two, who still have to charge money.”
The thread’s best correction is epistemic: public prices cannot establish who invented or adopted a private architecture. Its best answer to “why give this away?” is competitive rather than charitable. Open research makes the product sold by dominant labs cheaper and more interchangeable, while expanding the reach of the challenger.
HN handles are pseudonymous and the site publishes no per-comment scores. The ordering is HN’s own ranking, so this is a slice of the thread and not a consensus.