Tag
The Marin team is training the largest fully open model with 535B parameters, providing unprecedented transparency with live tracking and open data logs.
Tencent has released Hy4 Preview, an open-weight text-only large language model with 770B total parameters, 49B active parameters, and a 1M token context window, featuring reasoning capabilities with 'high' and 'no_think' effort levels.
Recommends installing Office CLI; combined with large language models and PPT/data-analysis Skills, it can significantly boost office productivity.
Alatkins announces a last-minute talk at AI4 Conference in Vegas, covering an updated version of the Trinity Large talk with teasers about training a 400B MoE model to 17T tokens without loss spikes.
Alibaba unveils Qwen3.8-Max, its largest model at 2.4T parameters, showing a 2% higher Terminal-Bench result than Fable 5, with open weights to be released next week.
Moonshot has released Kimi K3, an extremely large AI model that is virtually impossible to run on local workstations even with multiple RTX 6000 Blackwell GPUs, prompting the author to seek ways to run it via hacking or distillation.
Fields Medal winner Jacob Tsimerman announces his departure from academia to join OpenAI's safety team, while NVIDIA explores massive financing for a data center and the open-source Kimi K3 model (2.8T parameters) is released.
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.
Inkling is an open-weights 975B multimodal AI model designed for fine-tuning.
Intern Tian Keyu, who was previously fired by ByteDance for maliciously attacking a large model and compensated 8 million yuan, has now obtained hundreds of millions in financing, marking a dramatic turn of events.
Tsinghua University NLP lab's graduate destination statistics show that PhDs in large model direction can earn an annual salary of over 6 million yuan, and master's degree holders over 1 million yuan, sparking criticism of self-media's distorted reporting.
MiMo-V2.5-Pro-UltraSpeed is a tool for visualizing differences in attention mechanisms of large models, boasting ultra-fast speed.
MiMo-V2.5-Pro-UltraSpeed is an ultra-fast large model training pipeline.
Meta is about to launch a new model, Muse Spark. According to Alexandr Wang, the model has major improvements in coding and agentic capabilities, aiming to compete with other leading models, and the model is very large in scale.
Announces a server configuration with 4 Nvidia V100 GPUs and 128GB Tesla memory, targeting AI large model workloads.
Ant Group recently released a large number of recruitment positions involving Large Models and health fields, including campus and social recruitment, providing opportunities for job seekers.
Summarizes several main implementation directions for current multimodal large model startups, including game AI NPC, enterprise-level multimodal Agent, content generation, embodied intelligence, and visual code assistants.
Garry Tan tweets that he fine-tuned a Qwen3.5-397B model in a couple hours using Thinking Machines, praising its speed and usability for multimodal personal AI.
An observation on talent flow in the AI large model industry: many former CTOs and early Kuaishou employees join relevant companies, and the phenomenon of founders selling their companies, then interning, becoming regular employees, and starting to build teams.