Jul 15, 2026 arXiv'26 | StepAudio 2.5:一个底座三种人格,ASR 模式把 RTF 干到 0.0053 Jul 15, 2026 Self-Speculative Decoding 简史:不引入额外模型,用自己给自己打草稿这件事到底能走多远 Jul 15, 2026 arXiv'24 | SenseVoice:10 秒音频 70ms 转完,阿里在 LLM 化浪潮里逆行的 non-autoregressive 路线 Jul 08, 2026 GPU 算力没在算模型,在等 CPU 发号施令——CUDA Graph 是怎么解决这件事的 Jul 07, 2026 推理模型做 decode,GPU 只有 3% 在干活——SparseSpec 把稀疏 Attention 变成了 2.1× 加速 Jun 30, 2026 ICLR'25 | SWIFT:不训练不搜索,Self-Speculative Decoding 的即插即用版 Jun 30, 2026 2025 | SSD for dLLMs:扩散语言模型也能用推测解码,最高 3.46× Jun 30, 2026 arXiv'26 | SSD for ASR:用 CTC 编码器打草稿,LLM 验证,语音识别也能推测解码 Jun 30, 2026 WWW'26 | SS-MoE:显存墙下的 MoE 推理,Self-Speculative Decoding 怎么救场? Jun 30, 2026 EMNLP'23 | 不用额外模型,LLM 自己给自己加速——Self-Speculative Decoding 原理详解 Jun 30, 2026 ACL'24 | LayerSkip:Meta 把 Self-Speculative Decoding 从头训到尾,直接干到 2.16× Jun 25, 2026 SpecASR:ASR 专属 Speculative Decoding,让 LLM 语音识别快 3.79 倍 Apr 26, 2026 KVCOMM:让多 Agent 系统的 KV Cache 真正“通起来”,TTFT 直接砍掉 7.8 倍