OneBench
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss | OneBench: AI Insights