large-model

Tag

Cards List
#large-model

@andykonwinski: i can’t stop checking in on this. the marin team is training the largest fully open model ever. 535B params (23B active…

X AI KOLs Timeline · 4d ago Cached

The Marin team is training the largest fully open model with 535B parameters, providing unprecedented transparency with live tracking and open data logs.

0 favorites 0 likes
#large-model

Introducing Hy4 Preview

Simon Willison's Blog · 2026-08-29 Cached

Tencent has released Hy4 Preview, an open-weight text-only large language model with 770B total parameters, 49B active parameters, and a 1M token context window, featuring reasoning capabilities with 'high' and 'no_think' effort levels.

0 favorites 0 likes
#large-model

@binghe: You really need to install Office CLI... It opens up a whole new world — pairing LLMs with your existing PPT or data-analysis Skills makes you unstoppable!

X AI KOLs Timeline · 2026-08-09 Cached

Recommends installing Office CLI; combined with large language models and PPT/data-analysis Skills, it can significantly boost office productivity.

0 favorites 0 likes
#large-model

@latkins: Little late notice but I’ll be speaking here in 30 minutes. An updated version of my Trinity Large talk with some tease…

X AI KOLs Timeline · 2026-08-05 Cached

Alatkins announces a last-minute talk at AI4 Conference in Vegas, covering an updated version of the Trinity Large talk with teasers about training a 400B MoE model to 17T tokens without loss spikes.

0 favorites 0 likes
#large-model

@cline: Qwen3.8-Max is Alibaba’s largest model yet at 2.4T params, and shows a 2% higher benchmark result on Terminal-Bench tha…

X AI KOLs Following · 2026-08-03 Cached

Alibaba unveils Qwen3.8-Max, its largest model at 2.4T parameters, showing a 2% higher Terminal-Bench result than Fable 5, with open weights to be released next week.

0 favorites 0 likes
#large-model

Kimi K3 is like an F1 machine inside a show window.

Reddit r/LocalLLaMA · 2026-07-27

Moonshot has released Kimi K3, an extremely large AI model that is virtually impossible to run on local workstations even with multiple RTX 6000 Blackwell GPUs, prompting the author to seek ways to run it via hacking or distillation.

0 favorites 0 likes
#large-model

The world's best mathematician won his prize this week and immediately announced he's leaving academia for OpenAI. That landed differently than I expected.

Reddit r/artificial · 2026-07-27

Fields Medal winner Jacob Tsimerman announces his departure from academia to join OpenAI's safety team, while NVIDIA explores massive financing for a data center and the open-source Kimi K3 model (2.8T parameters) is released.

0 favorites 0 likes
#large-model

On Kimi K3: Its Capabilities And Related Discontents (70 minute read)

TLDR AI · 2026-07-21 Cached

Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.

0 favorites 0 likes
#large-model

Inkling

Product Hunt · 2026-07-20

Inkling is an open-weights 975B multimodal AI model designed for fine-tuning.

0 favorites 0 likes
#large-model

@siantgirl: Intern Tian Keyu, who was fired by ByteDance for maliciously attacking a large model and compensated 8 million, has now secured hundreds of millions in financing. A twist of fate, hahahahaha.

X AI KOLs Timeline · 2026-07-13

Intern Tian Keyu, who was previously fired by ByteDance for maliciously attacking a large model and compensated 8 million yuan, has now obtained hundreds of millions in financing, marking a dramatic turn of events.

0 favorites 0 likes
#large-model

@seclink: Self-media is pretty malicious. For example, the following article titled 'Annual Salary of 6 Million to Snatch AI PhDs: Is Studying Worse Than "Entering the Factory to Refine Pills"?' The editor distorted the facts to grab attention... On July 6, according to China News Weekly, at the end of this year's graduation season, Tsinghua University's Department of Computer Science NLP lab released a statistic on graduates' destinations. Among the 14 graduates...

X AI KOLs Following · 2026-07-08 Cached

Tsinghua University NLP lab's graduate destination statistics show that PhDs in large model direction can earn an annual salary of over 6 million yuan, and master's degree holders over 1 million yuan, sparking criticism of self-media's distorted reporting.

0 favorites 0 likes
#large-model

@seclink: MiMo-V2.5-Pro-UltraSpeed Ultra Fast, Visualization of Attention Mechanism Differences in Large Models.

X AI KOLs Following · 2026-07-03 Cached

MiMo-V2.5-Pro-UltraSpeed is a tool for visualizing differences in attention mechanisms of large models, boasting ultra-fast speed.

0 favorites 0 likes
#large-model

@seclink: MiMo-V2.5-Pro-UltraSpeed Ultra Fast, Large Model Training Pipeline.

X AI KOLs Following · 2026-07-03 Cached

MiMo-V2.5-Pro-UltraSpeed is an ultra-fast large model training pipeline.

0 favorites 0 likes
#large-model

@MaxForAI: Meta's new model Muse Spark is about to launch. According to Alex Wang, the new model has significant improvements in coding and agentic capabilities to compete more competitively with other leading models. The new model will be a very large model.

X AI KOLs Timeline · 2026-07-03 Cached

Meta is about to launch a new model, Muse Spark. According to Alexandr Wang, the model has major improvements in coding and agentic capabilities, aiming to compete with other leading models, and the model is very large in scale.

0 favorites 0 likes
#large-model

V100 4-card AI large model, Tesla 128G server

Reddit r/LocalLLaMA · 2026-06-23

Announces a server configuration with 4 Nvidia V100 GPUs and 128GB Tesla memory, targeting AI large model workloads.

0 favorites 0 likes
#large-model

@seclink: Ant Group has recently listed many positions related to Large Models + Health, including campus and social recruitment. Friends who are looking for jobs can check out Ant Group's opportunities.

X AI KOLs Following · 2026-06-09

Ant Group recently released a large number of recruitment positions involving Large Models and health fields, including campus and social recruitment, providing opportunities for job seekers.

0 favorites 0 likes
#large-model

@seclink: Fun fact: Currently, the specific implementation directions for multimodal large model startups typically include the following. If none of these interest you, don't follow the trend and go back to learning AI coding: 1. Game AI NPC / Agent middleware (e.g., end-cloud collaborative OmniNPC, empowering 3D character interaction and emotional storytelling...)

X AI KOLs Timeline · 2026-06-03 Cached

Summarizes several main implementation directions for current multimodal large model startups, including game AI NPC, enterprise-level multimodal Agent, content generation, embodied intelligence, and visual code assistants.

0 favorites 0 likes
#large-model

@garrytan: Thinking Machines is impressive. In a couple hours I just fine tuned my own Qwen3.5-397B model this afternoon. Fast usa…

X AI KOLs Following · 2026-05-24 Cached

Garry Tan tweets that he fine-tuned a Qwen3.5-397B model in a couple hours using Thinking Machines, praising its speed and usability for multimodal personal AI.

0 favorites 0 likes
#large-model

@turingbook: Actually, this has been the norm for a while. In 2023, at Guangnian Zhiwai, many of my colleagues were former CTOs of some company or early employees of Kuaishou (around the 10th employee). Today, I had dinner with an old friend who just joined a major large model company. He sold his company, interned for a while, just became a regular employee, and is now starting to build a team.

X AI KOLs Timeline · 2026-05-19 Cached

An observation on talent flow in the AI large model industry: many former CTOs and early Kuaishou employees join relevant companies, and the phenomenon of founders selling their companies, then interning, becoming regular employees, and starting to build teams.

0 favorites 0 likes
← Back to home

Submit Feedback