Tag
Kimi K3 weights are being released today. The model has 2.8T parameters, MoE with 896 experts, 1M context, vision, and MXFP4 quantization. Deployment requires multiple nodes for A100s and H200s, but fits in single B300 node. Benchmarks for tok/s, ttft, and cost per M token across GPU configs are expected by end of week.
An open dataset on GitHub maps which local LLMs fit various RAM tiers (8GB to 128GB), providing memory sizing rules, per-tier model lists, and Ollama commands, with a JSON API for programmatic access.