Tag
The article advises keeping if clauses side-effect free to improve code readability, using examples to illustrate common pitfalls and clearer alternatives.
The article discusses where to enforce approval for write actions in MCP tools, considering client-side and server-side approaches, and what to log for checking against approved actions.
The author argues that AI agent reliability in production should focus not just on observability but also on ensuring actions with real side effects produce the expected outcomes.
The article explores an early concept called agentuptime, which addresses verifying AI agent actions by independently checking outcomes to ensure that an agent's completion claim matches the actual state of external systems.
This paper investigates whether side effects of activation steering in language models can be predicted before intervention, constructing a cross-effect matrix across 67 behaviors and finding that side effects are systematic and forecastable from unsteered representations.
This paper introduces a new problem setting called side-effect introspection, where the goal is to detect alignment degradation in fine-tuned LLMs that occurs as an unintended side effect rather than explicitly implanted behavior. The authors propose a novel Delta-Aware Introspection Adapter (DAIA) that outperforms existing introspection adapters in generalizing to unseen models and safety categories.
This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.