Tag
NVIDIA's paper introduces a method to transfer KV caches between LLM models, enabling target models to skip prefill and achieve 2.7 to 25x faster conversion than reprocessing the context.
This arXiv paper investigates whether simple linear transformations can translate representations across nine heterogeneous text embedding models, finding that shared structure and transferability depend jointly on architecture, training objective, pooling, and data distribution, challenging the notion of universal latent compatibility.