Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
Summary
This paper introduces byte-exact KV-cache grafting, a technique that makes frozen small language models both more capable and cheaper by depositing verified knowledge as a byte-exact state artifact and restoring it during inference, achieving dramatic token and energy reductions without weight changes.
View Cached Full Text
Cached at: 07/20/26, 09:43 AM
Paper page - Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
Source: https://huggingface.co/papers/2607.14431
Abstract
Wereportawaytomakeafrozensmalllanguagemodelbothmorecapableanddramaticallycheaperatonce,withoutchanginganyweights.Verifiedknowledgeisdepositedonceasabyte-exactkey-value(KV)stateartifactandlaterrestored,bygraft,intoafreshinferencecontext.Therestoreisbit-exact:underapinneddeterministicconfiguration,thegraftedlogitsarebyte-for-byteidenticaltoafreshcomputation(SHA-256equality),withzeroKLdivergenceand100%argmaxagreementoverfiftysamples.Weshowthatown-positiongraftistheuniquenumericallyexactoperatingpointonamodelwithfloating-pointrotaryencoding,andweverifybyte-exactnessontwomodelscales(12B,31B)andtwoGPUtargets,onethroughapre-registeredreplay.OnAIME2025,afrozenGemma-4-12Bmovesfrom80.0%to93.3%onceaverifiedsolutionlibraryisgrafted,aboveitsown77.5%andits31Bsibling’s89.2%publishedanchors.Ontherecurringcase,eightproblemsthebasemodelneversolveswithina401,026-tokenbudgetareansweredfromcachedverifiedsolutionsin61totaldecodetokens,afactorof6,574fewertokensandabout8,700xlessenergy;thecapabilityclaimproperrestsonheld-outtransfer(7of7at31B).Thesamebyte-exactstorewidensusablecontextfrom32,768to2,854,766tokensatzeroextraacceleratormemory,andmovesbyte-identicalbetweenmachinesofthesamearchitecture.Wedescribethesystematthebehaviorlevel;theengineisproprietary,andeveryreportednumberisbackedbycommittedinputandoutputhashessothescoringcanbere-checkedwithoutit.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.14431
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.14431 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.14431 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.14431 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference
KVBoost is a chunk-level key-value cache reuse system for efficient large language model inference that achieves high cache hit rates and significant speedup in time-to-first-token without quality loss, using dual-hash keying and deviation-guided recomputation.
Models Take Notes at Prefill: KV Cache Can Be Editable and Composable
This paper proposes that the KV cache in transformers acts as a notebook of memoized conclusions, enabling surgical editing and composition without full recomputation. The method achieves significant latency reductions while preserving decision equivalence across model scales.
Enabling KV Caching of Shared Prefix for Diffusion Language Models
This paper proposes BiCache, a novel KV caching technique for shared prefixes in diffusion language models, which avoids accuracy collapse by dynamically reusing cached keys and values in shallow layers and achieves 36.3%–98.3% throughput improvement.
A big chunk of AI cost is just the model re-reading the same text over and over. Interesting attempt to fix it, with public proofs
Corbenic AI claims to offer lossless KV cache reuse for LLMs, allowing stored model memory to be restored bit-for-bit across machines and GPU generations, verified via public checksums. The project includes an open-sourced small model trained for ~600 EUR to make the full pipeline inspectable.
KV Cache Compression 900000x Beyond TurboQuant and Per-Vector Shannon Limit
A new paper proposes sequential KV cache compression using probabilistic language tries and predictive delta coding, achieving theoretical compression ratios of ~914,000× beyond TurboQuant by exploiting the sequential structure of language model tokens rather than treating vectors independently.