Tag
Humanly is a configurable writing platform that records the writing process to provide traceable evidence of human-AI collaboration, with features like sealed certificates and anomaly detection.
This paper presents the first rigorous study of how LLM watermarking schemes affect medical performance, evaluating five watermarks across multiple LLMs and VLMs on clinical reasoning tasks. The authors find that watermarks can cause degradation in medical text quality, including hallucinations and lexical corruption, which are masked by general-domain benchmarks.
ArcKit is an open-source enterprise architecture governance harness that works with AI coding assistants like Claude Code and GitHub Copilot to manage governance workflows including requirements, design reviews, and traceability.
This paper proposes a multi-level validation and traceability framework for AI-generated telescope scheduling decisions, integrating data reference validation, logical consistency checks, and observational constraints to improve executability and reliability in high-reliability astronomical observation tasks.
ModSleuth is a new tool that traces the dependencies of modern LLMs, revealing that models like OLMo 3 and Nemotron 3 rely on hundreds of other models and datasets, highlighting the shift from human-only to AI-generated training data.
The article critiques current AI memory systems as mere write-only logs that lack the ability to be corrected, updated, or traced to their source, arguing that true memory requires a governance layer.
A developer built a fully traceable and forkable research agent using Active Graph and monid_ai, ensuring every claim is natively traced to its source, and got it working in about 30 minutes.
The author argues that current agent observability provides a trace of actions but lacks runtime justification for why actions were permitted, which is critical for production deployments involving money, data, or communications.