RESEARCHInvestigateNEXT 12 MONTHS
CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research introduces CoSA, a new sparse attention mechanism designed to accelerate long-context inference in large language models by co-designing proxy and kernel.
Open source