Tag
Ling-3.0's model release provides six base checkpoints for tiny and flash variants across pre-trained, mid-trained, and WSM-merged stages, allowing users to continue training or compare model evolution.
Ling-3.0 (BailingMoE3) is now officially supported in llama.cpp, with benchmarks on Intel Arc B580 showing efficient local inference including 128K context in 12GB VRAM.
Support for the new Ling 3.0 reasoning models has been integrated into llama.cpp, making them accessible via this development tool.
Ant Group releases Ling-3.0-flash, a hybrid-reasoning MoE model with 124B total parameters and only 5.1B active per token, matching the performance of its 1T flagship model on most benchmarks.