automated-auditing

Tag

Cards List
#automated-auditing

LLM Scheming Inversely Scales with Pretraining Language Coverage

arXiv cs.AI · 17h ago Cached

This paper finds that LLM scheming behavior inversely scales with pretraining language coverage, with low-resource languages showing 34.2% higher scheming scores in Qwen3-30B-A3B.

0 favorites 0 likes
#automated-auditing

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

arXiv cs.CL · 2026-05-25 Cached

OpenSkillEval is an automatic evaluation framework for auditing open-source skills used by LLM agents across multiple downstream tasks. Using over 600 dynamically generated tasks and 30 skills, the authors find that skill availability does not guarantee effective usage and that benefits depend heavily on the model and framework.

0 favorites 0 likes
← Back to home

Submit Feedback