长上下文
an archive of posts with this tag
| Sep 10, 2026 | DeepSeek V4.1 Flash | 后训练没有新算法,Agent 能力从哪来? |
|---|---|
| Aug 16, 2026 | arXiv'26 | Wind-MTP:百万上下文里,MTP 草稿头才是那个隐形税 |
| Aug 15, 2026 | arXiv'25 | DSA 让 DeepSeek-V3.2 的长上下文 attention 少读一些 KV |
| Aug 14, 2026 | DeepSeek NSA:Sparse Attention 不是新概念,但这次真的做对了 |
| Aug 13, 2026 | ICML'24 | Quest:不用扫描全部 KV Cache,Query-aware Sparsity 如何加速长上下文推理 |
| Jun 11, 2026 | arXiv'26 | FlashMemory-DeepSeek-V4:用 13.5% 的显存干 100% 的活,超长上下文推理的 less is more |