Tag
This article describes an experiment showing that GPT-5.6 Sol can consistently stop before making a tool call by setting a numeric threshold just above a boundary, with all 25 test pairs demonstrating the expected behavior.
This paper introduces Xcientist, a research harness that externalizes AI-driven scientific research synthesis and validation into inspectable, contract-governed processes to ensure accountability and traceability.
This paper proposes a three-regime framework to resolve empirical contradictions in how LLMs handle conflict between training knowledge and new documents, validated across five major models. It distinguishes between parametric strength and uniqueness and demonstrates how task framing and evidence coherence significantly impact model behavior.