large-language-model

Tag

Cards List
#large-language-model

@rohanpaul_ai: Alibaba released Qwen3.8-Max, a 2.4 trillion parameter model that activates only about 95 bn parameters per token. A th…

X AI KOLs Timeline · yesterday Cached

Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter sparse MoE model with 95B active parameters per token, 1M token context, and strong agentic and benchmark results, including autonomously coding for days, circuit design, and outperforming rivals on Terminal Bench and PaperBench.

0 favorites 0 likes
#large-language-model

ByteDance is at an early stage of training a model with as many as 10 trillion parameters

Reddit r/singularity · 2d ago

ByteDance is in the early stages of training a large language model with up to 10 trillion parameters, signaling a massive scale-up in AI development.

0 favorites 0 likes
#large-language-model

Gemini

Reddit r/singularity · 2d ago

Google's Gemini AI model represents a significant advancement in multimodal AI capabilities.

0 favorites 0 likes
#large-language-model

K-EXAONE 2.0 Technical Report

arXiv cs.CL · 3d ago Cached

LG AI Research presents K-EXAONE 2.0, a 750B-parameter MoE foundation model upcycled from K-EXAONE, supporting 256K context and six languages, with self-speculative decoding for efficient inference.

0 favorites 0 likes
#large-language-model

Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model

arXiv cs.CL · 4d ago Cached

This paper introduces an activation-guided neuron intervention framework using Qwen3-8B to induce Alzheimer's-related computational language phenotypes, demonstrating that amplifying AD-associated neurons produces graded impairments in multiple cognitive domains.

0 favorites 0 likes
#large-language-model

@ChiragAsarpota: Bro this so HUGE for open models Even Qwen finally folded and is open-weighting a Max model. They never opened any of t…

X AI KOLs Timeline · 6d ago Cached

Qwen announces Qwen3.8-Max, a 2.4T parameter model, will be released with open weights next week, along with the Qwen3.8-27B, marking a major push for open-source AI and potentially competitive with top proprietary models.

0 favorites 0 likes
#large-language-model

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

arXiv cs.CL · 2026-07-31 Cached

The paper proposes SentiLLM, a framework that uses semantic-aligned structural abstraction to distill non-verbal modalities into text-like tokens for multimodal sentiment analysis with LLMs. It introduces a dual-stream salience-context calibration mechanism and achieves superior performance on four datasets.

0 favorites 0 likes
#large-language-model

Anyone tried the Q1 Kimi K3 yet? (555GB)

Reddit r/LocalLLaMA · 2026-07-29 Cached

Kimi K3 is a massive 2.9 trillion parameter mixture-of-experts model with 104B active parameters, 1M context length, and native MXFP4 training, now available in GGUF quantizations ranging from 540GB to smaller sizes, though requiring substantial hardware to run.

0 favorites 0 likes
#large-language-model

@_akhaliq: A.X K2 just dropped on Hugging Face Large-Scale Sparse MoE (688B / 33B Active) https://huggingface.co/skt/A.X-K2

X AI KOLs Following · 2026-07-29 Cached

SKT released A.X K2, a 688B-parameter sparse MoE language model with 33B active parameters, natively trained in FP8 and featuring Think/Non-Think reasoning modes, on Hugging Face.

0 favorites 0 likes
#large-language-model

@interjc: 期待 Grok 4.6,卷起来才能有更多重置

X AI KOLs Following · 2026-07-29 Cached

埃隆·马斯克宣布Grok 4.6将于8月7日左右发布,这是一个1.5T参数的模型,经过改进的SFT和RL;随后几周将发布更好的Grok 4.7(2.1T参数)。

0 favorites 0 likes
#large-language-model

@elonmusk: Interesting. Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL. Grok …

X AI KOLs Timeline · 2026-07-28 Cached

Elon Musk announces that Grok 4.6 will be released around August 7, featuring a 1.5T parameter model with improved SFT and RL, followed by Grok 4.7 with 2.1T parameters.

0 favorites 0 likes
#large-language-model

@seclink: Official page: https://huggingface.co/moonshotai/Kimi-K3… This is Moonshot AI's open-weight model with 2.8 trillion parameters (MoE architecture, ~104B active), supporting native multimodal (text+image…)

X AI KOLs Timeline · 2026-07-28 Cached

Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight MoE model with native multimodal capabilities and a 1 million token context window, claiming it as the world's first open 3T-class model.

0 favorites 0 likes
#large-language-model

Releasing the model weights and technical report of Kimi K3 (2 minute read)

TLDR AI · 2026-07-28 Cached

Kimi Moonshot released Kimi K3, a 2.8-trillion-parameter multimodal model with a 1M context window and architectural innovations like Kimi Delta Attention and Attention Residuals, claiming significant efficiency gains and outperforming Claude Opus 4.8 and GPT-5.5 on internal benchmarks.

0 favorites 0 likes
#large-language-model

moonshotai/Kimi-K3

Simon Willison's Blog · 2026-07-27 Cached

Moonshot released the weights for their 2.8 trillion parameter Kimi K3 model under a modified license requiring a separate agreement for large commercial Model-as-a-Service businesses.

0 favorites 0 likes
#large-language-model

Kimi K3 Now Available via Telnyx Inference API

Hacker News Top · 2026-07-27 Cached

Kimi K3, a 2.8-trillion-parameter open-source model from Moonshot AI with 1M token context and native vision, is now available on the Telnyx Inference API for tasks like coding, reasoning, and multimodality.

0 favorites 0 likes
#large-language-model

Kimi-K3 Technical Report [pdf]

Hacker News Top · 2026-07-27 Cached

MoonshotAI releases Kimi-K3, a 2.8T-parameter open-weight multimodal agentic model with a 1M-token context window, built on new Kimi Delta Attention and Attention Residuals architecture, achieving significant scaling improvements.

0 favorites 0 likes
#large-language-model

@levie: The k3 weights have arrived

X AI KOLs Timeline · 2026-07-27 Cached

Kimi.ai released the model weights and technical report for Kimi K3, a 2.8T parameter MoE model with native visual understanding and a 1M-token context window, claiming 2.5x intelligence per unit of compute.

0 favorites 0 likes
#large-language-model

@AdinaYakup: Macaron V1 from @Macaron0fficial drops in two variants: Venti: 748B flagship (744B base + 4×1B LoRA specialists). Post …

X AI KOLs Following · 2026-07-27 Cached

Macaron V1 is released in two variants: Venti (748B flagship with 4×1B LoRA specialists) and Tall (50B for local deployment with 4×3.7B LoRA specialists), post-trained on GLM-5.2 and Qwen 3.6 respectively.

0 favorites 0 likes
#large-language-model

Kimi K3 is the largest open-weight model ever released. You still can't run it.

Reddit r/AI_Agents · 2026-07-27

Moonshot released the largest open-weight model, Kimi K3 (2.8T parameters), boasting strong benchmarks but requiring massive hardware for self-hosting, making API access the practical choice for most users.

0 favorites 0 likes
#large-language-model

Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode

Reddit r/LocalLLaMA · 2026-07-27

Krasis, a MoE-focused runtime, enables running the 397B-parameter Ornith model on a single RTX PRO 6000 Blackwell 96GB GPU with ~20-24 tok/s decode by dynamically managing expert residency in VRAM.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback