Tag
Samsung published a paper on offloading KV-cache over a CXL memory pool, demonstrating that even with early-gen CXL hardware, GPUs can be fed as effectively as with real DRAM, making the approach easily replicable.