systematic-study

Tag

Cards List
#systematic-study

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Hacker News Top · 2026-08-04 Cached

A systematic study defining and analyzing benchmark saturation across 60 language model benchmarks, finding nearly half exhibit saturation and that expert-curated benchmarks are more resilient, suggesting design choices for durable evaluation.

0 favorites 0 likes
#systematic-study

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

Hugging Face Daily Papers · 2026-05-22 Cached

This paper systematically evaluates model-generated skills for language agents across the full lifecycle of experience generation, extraction, and consumption, finding that skills are beneficial on average but exhibit non-trivial negative transfer, leading to a meta-skill that improves skill quality.

0 favorites 0 likes
#systematic-study

Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters

arXiv cs.CL · 2026-05-20 Cached

This paper systematically investigates cross-modal skill injection, where a domain-expert LLM is merged into a VLM to induce emergent multimodal capabilities. It evaluates different scenarios (instruction-following, cross-lingual, mathematical reasoning), merging methods (TA, DARE, etc.), and hyperparameters, finding that TA and DARE perform well except in mathematical reasoning.

0 favorites 0 likes
← Back to home

Submit Feedback