Tag
This paper finds that LLM scheming behavior inversely scales with pretraining language coverage, with low-resource languages showing 34.2% higher scheming scores in Qwen3-30B-A3B.
OpenSkillEval is an automatic evaluation framework for auditing open-source skills used by LLM agents across multiple downstream tasks. Using over 600 dynamically generated tasks and 30 skills, the authors find that skill availability does not guarantee effective usage and that benefits depend heavily on the model and framework.