llm-runtime

Tag

Cards List
#llm-runtime

I somehow got GPT-OSS 120B running locally at 21 tok/s on a 4070 Ti with 32gb ram lol🏗😤🤣

Reddit r/ArtificialInteligence · 5h ago

A developer got GPT-OSS 120B running locally on a 4070 Ti with 32GB RAM by exploiting its MoE architecture, streaming cold experts from NVMe and caching hot experts on GPU, reaching 21 tok/s with a top-1 approximation.

0 favorites 0 likes
← Back to home

Submit Feedback