📋 待办 & 收藏

觉得有价值的论文,持续更新

arXiv: 2602.22603 JP Morgan Chase

用 LLM 本身做 KV eviction decision,parallel auxiliary thread 判断 cursor expiration,Peak token ↓65%,Throughput ↑83.9%。核心洞察:past importance ≠ future relevance。

KV Cache Agent Context Compression
arXiv: 2505.19433 HKUST

首个评估压缩对 agentic 能力影响的 benchmark。4-bit 量化保留 workflow/tool-use(仅降 1-3%),但 real-world 应用降 10-15%,DeepSeek-R1-Distill 降 32%。JSON structured output 比 string 降级更严重。

LLM Compression Agent Benchmark