What matters in AI.

Subscribe

RaReCache lets a small model prefill the cache for a large model

The authors write that the large model keeps 95 to 99% of its precision when it recomputes only 30% of positions.

Claimed, not confirmed

This is a brief. We point to the report and do not rewrite it. Read it at the source below.

Recomputing 30% of positions keeps 95-99% of accuracy A ten by ten grid of positions: 30 are recomputed (red), 70 are reused (grey). On a 23x parameter gap, Qwen3-0.6B to 14B, this keeps 95-99% of target accuracy. Llama3-8B to 70B: 40% recomputed gives 96.5%. Qwen3-0.6B to 14B 23x parameter gap 30% recomputed 70% reused 95-99% of target accuracy retained Llama3-8B to 70B: 40% → 96.5%
On a 23x size gap, recomputing 30% of positions retains 95-99% of the large model's accuracy, with the rest reused from the small model's cache.

Sources

  1. RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputationarxiv.org
AI MATTER · NEWS · AI MATTER · NEWS ·9 OCT2026

Posted

Tags

More in Compute

All Compute news