parameters

Tag

Cards List
#parameters

Deepseek v4.1 flash finally has engrams, what do you expect from 4.1 pro?

Reddit r/LocalLLaMA · 2026-09-10

Speculation on Deepseek v4.1 pro's specifications and future developments, including parameter counts and engram technology.

0 favorites 0 likes
#parameters

@ShreyashJokare: 200 billion parameters

X AI KOLs Timeline · 2026-08-29 Cached

The tweet highlights an AI model with 200 billion parameters, suggesting a significant scale in machine learning development.

0 favorites 0 likes
#parameters

@rohanpaul_ai: Brilliant piece by Zhipu Founder Tang Jie. AI scaling is moving past parameter growth. “How many parameters?” is becomi…

X AI KOLs Following · 2026-08-20 Cached

Zhipu Founder Tang Jie discusses how AI scaling is evolving beyond parameter count to include factors like training data, compute per forward pass, and post-training, with GLM-5.3 as an example.

0 favorites 0 likes
#parameters

The Most Attractive Quadrant is already occupied ! It's happening !

Reddit r/singularity · 2026-08-17

An analysis on Artificial Analysis shows that the most attractive quadrant for AI models is already occupied, indicating significant advancements in open-source models.

0 favorites 0 likes
#parameters

Weirdly, no one talks about Temperature setting for the Qwen3.8 27b

Reddit r/LocalLLaMA · 2026-08-17

The post discusses the default temperature setting of 1.0 for Qwen3.8 27b and suggests setting it to 0.7 to reduce unnecessary reasoning, questioning the impact on capabilities and optimal settings for various tasks.

0 favorites 0 likes
#parameters

@Dinosaur_liu: A horror story for all Chinese and American model makers: DeepSeek V4 Flash has only 284b parameters, with a mere 13b active parameters

X AI KOLs Timeline · 2026-08-03 Cached

It is claimed that DeepSeek V4 Flash has only 284 billion parameters and only 13 billion active parameters, posing an efficiency shock to model manufacturers in China and the US.

0 favorites 0 likes
#parameters

Deepseek, please explain to me how you make a 300B parameter model that is cheaper than a 9B parameter model by SO MUCH.

Reddit r/singularity · 2026-07-31

A Reddit user asks how DeepSeek can make a 300B parameter model cheaper than a 9B parameter model, sparking discussion on efficiency and cost.

0 favorites 0 likes
#parameters

Gemini last models: temperature, top_p, and top_k are deprecated and ignored

Hacker News Top · 2026-07-21

Google's Gemini models are deprecating and ignoring the temperature, top_p, and top_k parameters, likely simplifying inference configuration.

0 favorites 0 likes
#parameters

The untuned 27B beat the tuned 75B as an agent

Reddit r/LocalLLaMA · 2026-07-10

A comparison showing that an untuned 27B parameter model outperforms a tuned 75B parameter model in agent tasks, highlighting potential inefficiencies in scaling and fine-tuning.

0 favorites 0 likes
#parameters

@witcheer: Gemma 4 dropped a 12B. I put it on RTX 5090 against its 31B sibling. when you cut a model from 31B to 12B, what do you …

X AI KOLs Timeline · 2026-06-03 Cached

A comparison of Gemma 4 12B and 31B models shows that the smaller model retains reasoning capabilities nearly intact but suffers significant knowledge loss, making it ideal for reasoning tasks while the larger model is better for broad knowledge Q&A.

0 favorites 0 likes
#parameters

@NeoAIForecast: https://x.com/NeoAIForecast/status/2058479806048792583

X AI KOLs Timeline · 2026-05-24 Cached

A full educational series on local LLMs, covering inference, tokens, weights, and system-level understanding for beginners and reference.

0 favorites 0 likes
← Back to home

Submit Feedback