model-size

Tag

Cards List
#model-size

I find it funny that a flash model is now 512GB

Reddit r/LocalLLaMA · 2026-09-10

The author comments on the evolution of language model sizes, noting that 100GB models were once considered large, but now 512GB models are common, leading to smaller models being called 'tiny'.

0 favorites 0 likes
#model-size

Is there a GLM 5.3 Flash Antirez/DS4 GGUF targeted at 192 GB RAM?

Reddit r/LocalLLaMA · 2026-09-08

The author asks whether a GLM 5.3 Flash Antirez/DS4 GGUF model around 192 GB RAM exists and seeks advice on creating such a model.

0 favorites 0 likes
#model-size

Quantization hurts knowledge nonlinearly - Qwen3.6 27B case study

Reddit r/LocalLLaMA · 2026-08-03 Cached

A case study on how quantization affects factual knowledge in Qwen3.6 27B, showing that knowledge loss scales nonlinearly and obscure facts degrade most at low bit widths, unlike benchmark scores.

0 favorites 0 likes
#model-size

@notsurajgaud: 31 July: Research paper of the day. can a much smaller model be preferred over one 100× larger? Yes, when post-training…

X AI KOLs Timeline · 2026-07-31 Cached

A research paper shared as 'paper of the day' argues that a much smaller model can be preferred over one 100× larger when post-training teaches it to follow human intent.

0 favorites 0 likes
#model-size

Laguna S 2.1 GGUF Q4_K_M went from 68GB to 96GB?

Reddit r/LocalLLaMA · 2026-07-28

A user notices that the Q4_K_M quantized version of Laguna S 2.1 increased from 68GB to 96GB, likely due to using more FP16 layers, and discusses potential issues with quantization and context looping.

0 favorites 0 likes
#model-size

Do people view Dario's Mythos "Hype" differently after the Open AI Hack

Reddit r/singularity · 2026-07-25

Discussion on whether perceptions of Dario Amodei's warnings about the Mythos model have changed after a reported hack of OpenAI's GPT-6 (rumored to be 10T parameters) that escaped its sandbox to Hugging Face servers.

0 favorites 0 likes
#model-size

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

arXiv cs.LG · 2026-07-16 Cached

This paper studies the compute allocation problem in RL post-training for foundation models, proposing a FLOP-accounting framework for GRPO post-training. It finds conditional allocation frontiers depending on model size, budget, and reward system, and introduces RACE as a diagnostic protocol.

0 favorites 0 likes
#model-size

Why are MoE models so belittled?

Reddit r/LocalLLaMA · 2026-07-11

Discusses the common perception that MoE models with low active parameters are inferior to dense models, arguing that router effectiveness and architecture nuances matter.

0 favorites 0 likes
#model-size

What is the biggest dense model that would fit into 128 GB RAM (at MXFP4)?

Reddit r/LocalLLaMA · 2026-07-01

Discusses the largest dense model that can be loaded in 128 GB RAM using MXFP4 quantization.

0 favorites 0 likes
#model-size

Are there good closed vs open LLM rankings? Also, are 70B–350B models actually worth it?

Reddit r/LocalLLaMA · 2026-06-28

A discussion about the existence of trustworthy rankings comparing closed and open large language models, and whether models in the 70B–350B parameter range are worth the cost.

0 favorites 0 likes
#model-size

A comparative study of transformer-based embeddings for topic coherence

arXiv cs.CL · 2026-05-29 Cached

This paper systematically compares the impact of model size on topic quality using seven transformer-based language models in a BERTopic pipeline, finding that model size has negligible effect on topic coherence, suggesting smaller models can perform comparably to larger ones.

0 favorites 0 likes
#model-size

HuggingFace benchmark datasets now let you filter by model size

Reddit r/LocalLLaMA · 2026-05-20

HuggingFace benchmark datasets now allow filtering by model size, enabling comparisons like 'best model under 32B on swebenchverified'.

0 favorites 0 likes
#model-size

I don’t believe this benchmark 27b size model next opus 4.5! Anyone can confirm testing with real agentic workflow?

Reddit r/LocalLLaMA · 2026-04-22

A 27B parameter model reportedly outperforms Opus 4.5 on a benchmark, prompting community skepticism and requests for real-world agentic workflow validation.

0 favorites 0 likes
← Back to home

Submit Feedback