model-loading

Tag

Cards List
#model-loading

I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5)

Reddit r/LocalLLaMA · 2d ago

A llama.cpp/GGML PR (26291) speeds up RPC-based 300GB model loads by ~300%, cutting load time from ~5 minutes to ~1.5 minutes on consumer hardware with the GGML_RPC_LOAD_THREADS setting.

0 favorites 0 likes
#model-loading

@ariG23498: I have always admired @stevhliu's work. I consider his technical writeups to be among the best there is. In the latest …

X AI KOLs Timeline · 2026-07-08 Cached

A thread highlighting a technical blog series on how Hugging Face's transformers library loads models efficiently, covering meta device, safetensors, CUDA caching, and more.

0 favorites 0 likes
← Back to home

Submit Feedback