ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
Researchers introduced ResKV, a KV cache compression method that reconstructs omitted attention contributions to improve long-context LLM inference efficiency.