ktransformers

Tag

Cards List
#ktransformers

@techNmak: The smartest way to run a giant MoE model is not to add more GPUs. It is to stop treating every expert as GPU-worthy. L…

X AI KOLs Timeline · 2026-07-12 Cached

KTransformers is a framework that optimizes inference and fine-tuning of large Mixture-of-Experts models by dynamically placing only active experts on the GPU while keeping the rest in CPU memory, enabling large models like DeepSeek-V3 to run on limited consumer GPU memory.

0 favorites 0 likes
#ktransformers

@TheAhmadOsman: Yannick is criminally underfollowed in the Local AI space for the depth of his work

X AI KOLs Timeline · 2026-06-26 Cached

Yannick Nick demonstrates running DeepSeek V4 Flash with native FP4+FP8 precision on 2x RTX Pro 6000 GPUs using KTransformers, enabling efficient inference on resource-constrained systems.

0 favorites 0 likes
#ktransformers

I have (even faster) DeepSeek V4 Pro at home

Reddit r/LocalLLaMA · 2026-05-15

A user reports successfully running the DeepSeek V4 Pro model locally using ktransformers and sharing detailed benchmark results across various context depths, demonstrating improved inference speeds.

0 favorites 0 likes
#ktransformers

kvcache-ai/ktransformers

GitHub Trending (daily) · 2026-07-20 Cached

KTransformers is a flexible research framework for cutting-edge LLM inference and fine-tuning, leveraging CPU-GPU heterogeneous computing and supporting many recent models.

0 favorites 0 likes
← Back to home

Submit Feedback