mmlu

Tag

Cards List
#mmlu

@IndieDevHailey: Guys, local players can now experiment without boundaries. Alibaba's Qwen3.8-27B has been completely obliterated (safety restrictions removed), with the V2 version achieving zero rejection. Even better, MMLU (Massive Multitask Language Understanding) score jumped from 85.3% to 86.3%.

X AI KOLs Timeline · 2026-08-21 Cached

Alibaba's Qwen3.8-27B model's safety restrictions have been completely removed, with the V2 version achieving zero rejection, and the MMLU score improved to 86.3%, using complementary ablation blending to achieve uncensored and stronger performance.

0 favorites 0 likes
#mmlu

Understanding Tone-Dependent Inference Cost in Large Language Models

arXiv cs.CL · 2026-07-28 Cached

This paper investigates how the tone of prompts affects both the accuracy and the inference cost (output token consumption) of large language models, finding that output token length can vary by up to 44.3% across tones while accuracy changes are smaller, and identifies optimal tones for different models.

0 favorites 0 likes
#mmlu

Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network

arXiv cs.CL · 2026-07-22 Cached

This paper reports on a collaboration between the European Commission's Directorate-General for Translation and the European Master's in Translation network to localize the MMLU dataset into 11 European languages, creating a more inclusive benchmark for LLM evaluation while providing authentic training for translation students.

0 favorites 0 likes
#mmlu

Quantifying Ranking Uncertainty in LLM Benchmarks

arXiv cs.LG · 2026-07-21 Cached

This paper analyzes sources of ranking uncertainty in LLM benchmarks like MMLU, proposing modifications to hypothesis tests for constructing rank confidence intervals, and shows that variability across subjects is substantial.

0 favorites 0 likes
#mmlu

@LiorOnAI: A ~10K parameter router can beat every individual open model on MMLU by learning which model should answer which questi…

X AI KOLs Following · 2026-07-05 Cached

A tiny ~10K parameter router called tinyrouter learns which open model to use per question on MMLU, outperforming individual models by optimizing allocation.

0 favorites 0 likes
#mmlu

A "93% answer-flip" headline in our prompt-formatting study was mostly a parsing artifact. Here's the correction, and why I think parse-failure handling is an under-reported analyst degree of freedom.

Reddit r/ArtificialInteligence · 2026-06-23

A study on MMLU accuracy shifts across prompt formats found that the headline "93% answer-flip" rate was partly due to parsing artifacts. The author recommends pre-registering parse-failure handling and reporting parse-corrected metrics to separate genuine model sensitivity from parser brittleness.

0 favorites 0 likes
#mmlu

Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

Hugging Face Daily Papers · 2026-06-14 Cached

This paper introduces a controlled protocol to evaluate answer stability in large language models by challenging correct answers with plausible counterarguments, revealing large variation in flip rates across models that accuracy metrics alone do not capture. The authors release the protocol, challenge records, and a curated MaxFlip challenge set to support stability evaluation.

0 favorites 0 likes
#mmlu

HRM-Text: Efficient Pretraining Beyond Scaling

arXiv cs.CL · 2026-05-21 Cached

HRM-Text introduces a Hierarchical Recurrent Model that decouples computation into slow and fast layers, enabling efficient pretraining from scratch on only 40 billion tokens and a $1,500 budget, achieving competitive performance with larger models.

0 favorites 0 likes
#mmlu

@akshay_pachaar: The Operating System for Al Research Labs. TransformerLab orchestrates GPUs across any cloud and runs any training or e…

X AI KOLs Following · 2026-05-20 Cached

TransformerLab is an open-source platform that orchestrates GPUs across clouds and provides pre-built templates for AI training and evaluation workflows like LoRA, DPO, and MMLU.

0 favorites 0 likes
#mmlu

Capability Conditioned Scaffolding for Professional Human LLM Collaboration

arXiv cs.CL · 2026-05-18 Cached

Introduces Capability Conditioned Scaffolding, a framework for LLM collaboration that adapts intervention based on user expertise domains to prevent Professional Domain Drift, with pilot evaluation on MMLU subsets.

0 favorites 0 likes
#mmlu

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

arXiv cs.CL · 2026-05-11 Cached

This study presents a 33-model atlas analyzing domain-level metacognitive monitoring in frontier LLMs using MMLU benchmarks, revealing significant variations in confidence calibration across different knowledge domains that are obscured by aggregate metrics.

0 favorites 0 likes
← Back to home

Submit Feedback