Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

Hugging Face Daily Papers Papers

Summary

This paper introduces byte-exact KV-cache grafting, a technique that makes frozen small language models both more capable and cheaper by depositing verified knowledge as a byte-exact state artifact and restoring it during inference, achieving dramatic token and energy reductions without weight changes.

We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh inference context. The restore is bit-exact: under a pinned deterministic configuration, the grafted logits are byte-for-byte identical to a fresh computation (SHA-256 equality), with zero KL divergence and 100% argmax agreement over fifty samples. We show that own-position graft is the unique numerically exact operating point on a model with floating-point rotary encoding, and we verify byte-exactness on two model scales (12B, 31B) and two GPU targets, one through a pre-registered replay. On AIME 2025, a frozen Gemma-4-12B moves from 80.0% to 93.3% once a verified solution library is grafted, above its own 77.5% and its 31B sibling's 89.2% published anchors. On the recurring case, eight problems the base model never solves within a 401,026-token budget are answered from cached verified solutions in 61 total decode tokens, a factor of 6,574 fewer tokens and about 8,700x less energy; the capability claim proper rests on held-out transfer (7 of 7 at 31B). The same byte-exact store widens usable context from 32,768 to 2,854,766 tokens at zero extra accelerator memory, and moves byte-identical between machines of the same architecture. We describe the system at the behavior level; the engine is proprietary, and every reported number is backed by committed input and output hashes so the scoring can be re-checked without it.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:43 AM

Paper page - Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

Source: https://huggingface.co/papers/2607.14431

Abstract

Wereportawaytomakeafrozensmalllanguagemodelbothmorecapableanddramaticallycheaperatonce,withoutchanginganyweights.Verifiedknowledgeisdepositedonceasabyte-exactkey-value(KV)stateartifactandlaterrestored,bygraft,intoafreshinferencecontext.Therestoreisbit-exact:underapinneddeterministicconfiguration,thegraftedlogitsarebyte-for-byteidenticaltoafreshcomputation(SHA-256equality),withzeroKLdivergenceand100%argmaxagreementoverfiftysamples.Weshowthatown-positiongraftistheuniquenumericallyexactoperatingpointonamodelwithfloating-pointrotaryencoding,andweverifybyte-exactnessontwomodelscales(12B,31B)andtwoGPUtargets,onethroughapre-registeredreplay.OnAIME2025,afrozenGemma-4-12Bmovesfrom80.0%to93.3%onceaverifiedsolutionlibraryisgrafted,aboveitsown77.5%andits31Bsibling’s89.2%publishedanchors.Ontherecurringcase,eightproblemsthebasemodelneversolveswithina401,026-tokenbudgetareansweredfromcachedverifiedsolutionsin61totaldecodetokens,afactorof6,574fewertokensandabout8,700xlessenergy;thecapabilityclaimproperrestsonheld-outtransfer(7of7at31B).Thesamebyte-exactstorewidensusablecontextfrom32,768to2,854,766tokensatzeroextraacceleratormemory,andmovesbyte-identicalbetweenmachinesofthesamearchitecture.Wedescribethesystematthebehaviorlevel;theengineisproprietary,andeveryreportednumberisbackedbycommittedinputandoutputhashessothescoringcanbere-checkedwithoutit.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2607\.14431

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.14431 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.14431 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.14431 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Models Take Notes at Prefill: KV Cache Can Be Editable and Composable

arXiv cs.LG

This paper proposes that the KV cache in transformers acts as a notebook of memoized conclusions, enabling surgical editing and composition without full recomputation. The method achieves significant latency reductions while preserving decision equivalence across model scales.

Enabling KV Caching of Shared Prefix for Diffusion Language Models

arXiv cs.LG

This paper proposes BiCache, a novel KV caching technique for shared prefixes in diffusion language models, which avoids accuracy collapse by dynamically reusing cached keys and values in shallow layers and achieves 36.3%–98.3% throughput improvement.

KV Cache Compression 900000x Beyond TurboQuant and Per-Vector Shannon Limit

Hacker News Top

A new paper proposes sequential KV cache compression using probabilistic language tries and predictive delta coding, achieving theoretical compression ratios of ~914,000× beyond TurboQuant by exploiting the sequential structure of language model tokens rather than treating vectors independently.