large-language-model

Tag

Cards List
#large-language-model

Kimi K3: Open Frontier Intelligence

Hugging Face Daily Papers · 2026-07-27 Cached

Kimi K3 is a 2.8 trillion parameter Mixture-of-Experts model with 104 billion active parameters, native vision, and a 1-million-token context window, achieving frontier-level performance across multiple domains and released as open weights.

0 favorites 0 likes
#large-language-model

Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 (9 minute read)

TLDR AI · 2026-07-27 Cached

Compares two new AI models for agentic workloads: the compact Nanbeige4.2-3B with a looped transformer architecture and the large Mixture-of-Experts Laguna S2.1, both released on Hugging Face.

0 favorites 0 likes
#large-language-model

GLM-5.5 Leak: It Beats Fable 5

Reddit r/LocalLLaMA · 2026-07-25 Cached

A leaked report claims that Z.ai's upcoming GLM-5.5 model, with over 1 trillion parameters and 1M-token context, will compete directly with Fable 5 and Mythos, targeting an August release and remaining open-weight.

0 favorites 0 likes
#large-language-model

Kimi Linear 48B A3B?

Reddit r/LocalLLaMA · 2026-07-25

Kimi Linear 48B A3B appears to be a new large language model with 48 billion parameters, likely from Moonshot AI.

0 favorites 0 likes
#large-language-model

Claude Opus 5

Hacker News Top · 2026-07-24 Cached

Anthropic releases Claude Opus 5, a new large language model with enhanced capabilities and safety features, as detailed in its system card.

0 favorites 0 likes
#large-language-model

CLAUDE 5 OPUS - System Card

Reddit r/singularity · 2026-07-24 Cached

Anthropic publishes the system card for Claude 5 Opus, detailing its capabilities, safety evaluations, and deployment details.

0 favorites 0 likes
#large-language-model

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

arXiv cs.CL · 2026-07-24 Cached

This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.

0 favorites 0 likes
#large-language-model

@Chinazhidx: Ant Group just released Ling-3.0-flash • 124B MoE • 5.1B active params/token • 256K context, expandable to 1M Just 1/8 …

X AI KOLs Timeline · 2026-07-24 Cached

Ant Group released Ling-3.0-flash, a 124B MoE model with 5.1B active parameters per token and 256K context expandable to 1M, matching or outperforming their 1T flagship model on most benchmarks.

0 favorites 0 likes
#large-language-model

Genesis-Science-1 (GS1), 1T open-weight model later this year from Arcee AI

Reddit r/LocalLLaMA · 2026-07-22

Arcee AI announces Genesis-Science-1 (GS1), a 1 trillion parameter open-weight model to be released later this year.

0 favorites 0 likes
#large-language-model

Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro

Reddit r/LocalLLaMA · 2026-07-21 Cached

Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.

0 favorites 0 likes
#large-language-model

@NousResearch: The new Laguna S 2.1 model by @poolsideai is now free for 2 weeks on Nous Portal. At 118B total parameters with 8B acti…

X AI KOLs Following · 2026-07-21 Cached

The new Laguna S 2.1 model by poolsideai, with 118B total parameters and 8B active, is now free for two weeks on Nous Portal.

0 favorites 0 likes
#large-language-model

poolside/Laguna-S-2.1 released! Finally an interesting 120B contender!

Reddit r/LocalLLaMA · 2026-07-21

poolside released Laguna-S-2.1, a new 120 billion parameter language model that emerges as a strong contender in the LLM landscape.

0 favorites 0 likes
#large-language-model

@seclink: This fastjson 0day can be used to test the vulnerability discovery capabilities of LLMs. Without providing explicit information, only setting the goal (achieving a gadget-free RCE vulnerability), see how many steps the LLM takes to discover it and verify itself. Practice has shown: Claude indeed…

X AI KOLs Following · 2026-07-21 Cached

Using the fastjson 0day to test LLM vulnerability discovery capabilities, found Claude and Doubao performed outstandingly.

0 favorites 0 likes
#large-language-model

Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph

arXiv cs.CL · 2026-07-21 Cached

Debate-on-Graph (DoG) is a framework that enhances LLM reasoning by leveraging uncertain knowledge graphs (UKGs) with confidence scores, using a heuristic search and multi-agent debate mechanism to produce reliable answers. It achieves state-of-the-art performance on four QA benchmarks.

0 favorites 0 likes
#large-language-model

Sparse By Design (5 minute read)

TLDR AI · 2026-07-21 Cached

Moonshot's Kimi K3, a 2.8 trillion parameter open weights model with 896 experts (16 active per token), exemplifies the trend of scaling total parameters while holding active compute constant, and uses attention compression to reduce KV cache size, making frontier inference more accessible but with high storage costs.

0 favorites 0 likes
#large-language-model

Motif-Technologies/Motif-3-Beta

Hugging Face Models Trending · 2026-07-20 Cached

Motif Technologies releases an intermediate beta checkpoint of Motif-3, a large-scale Mixture-of-Experts language model with ~314B total parameters (~13B active), 256K context length, and custom architectures like Grouped Differential Latent Attention (GDLA), openly available for non-commercial research.

0 favorites 0 likes
#large-language-model

@savipww: a 744B parameter model just booted on a laptop with 25 gigs of ram and i read the repo twice before i believed it repo …

X AI KOLs Timeline · 2026-07-20 Cached

A 744B parameter mixture-of-experts model boots on a laptop with 25GB RAM by storing expert weights on SSD and only loading the active ~40B parameters per token, enabling local inference despite the model's size.

0 favorites 0 likes
#large-language-model

Knowledge-Centric Agents for Workflow Generation

arXiv cs.AI · 2026-07-20 Cached

The paper introduces a knowledge-centric framework for generating ComfyUI workflows by distilling hierarchical knowledge (pseudo-codes, skeletons, strategies) from real workflows and using LLMs to perform reasoning from task descriptions to executable structures, achieving higher node diversity and execution success rates.

0 favorites 0 likes
#large-language-model

Qwen3.8 Is Going Open-Weight (1 minute read)

TLDR AI · 2026-07-20 Cached

Qwen3.8, a 2.4-trillion-parameter language model, is being released as open-weight by Alibaba, with a preview available on Alibaba's Token Plan, Qoder, and QoderWork.

0 favorites 0 likes
#large-language-model

Tested the new Qwen 3.8 model (2.4T parameters)

Reddit r/LocalLLaMA · 2026-07-19

Tested the new Qwen 3.8 model, a large language model with 2.4 trillion parameters.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback