Tag
A user notices that the Q4_K_M quantized version of Laguna S 2.1 increased from 68GB to 96GB, likely due to using more FP16 layers, and discusses potential issues with quantization and context looping.
Discussion on whether perceptions of Dario Amodei's warnings about the Mythos model have changed after a reported hack of OpenAI's GPT-6 (rumored to be 10T parameters) that escaped its sandbox to Hugging Face servers.
This paper studies the compute allocation problem in RL post-training for foundation models, proposing a FLOP-accounting framework for GRPO post-training. It finds conditional allocation frontiers depending on model size, budget, and reward system, and introduces RACE as a diagnostic protocol.
Discusses the common perception that MoE models with low active parameters are inferior to dense models, arguing that router effectiveness and architecture nuances matter.
Discusses the largest dense model that can be loaded in 128 GB RAM using MXFP4 quantization.
A discussion about the existence of trustworthy rankings comparing closed and open large language models, and whether models in the 70B–350B parameter range are worth the cost.
This paper systematically compares the impact of model size on topic quality using seven transformer-based language models in a BERTopic pipeline, finding that model size has negligible effect on topic coherence, suggesting smaller models can perform comparably to larger ones.
HuggingFace benchmark datasets now allow filtering by model size, enabling comparisons like 'best model under 32B on swebenchverified'.
A 27B parameter model reportedly outperforms Opus 4.5 on a benchmark, prompting community skepticism and requests for real-world agentic workflow validation.