Tag
A group of 25 Fields Medal-winning mathematicians has signed an open letter accusing AI labs, particularly OpenAI, of threatening mathematical research through unverified proofs and poor attribution practices, escalating tensions and raising concerns about the integrity of open science.
The paper reveals that only a small fraction of transformer components are necessary for token predictions and introduces direct methods to read and write to these components with minimal changes.
This paper investigates the role of citations in human and LLM preferences for scientific question answering, finding that humans prefer diverse citations but fewer overall, while LLMs exhibit stronger citation-related preferences despite lacking source access.
A user on X seeks assistance from X and Space xAI to address content theft by another user who fails to provide attribution.
The article discusses a demo of an AI agent autonomously executing git workflows, highlighting critical issues with attribution, identity, and compliance, and proposes cryptographic agent identities while questioning accountability models.
LongRCA Bench introduces a benchmark for diagnosing failures in long-horizon agent trajectories, and the RCTA method improves responsible role and root-cause step attribution.
This paper proposes using signed, fusion-aware Integrated Gradients for attributing predictions in feature-tokenized transformers like BiomeGPT, overcoming limitations of CLS attention weights and revealing disease-supporting versus protective microbial signals.
This paper proposes a model-agnostic, post-hoc interpretation framework for black-box LLMs using sentence-level energy landscapes. A surrogate Energy-Based Model simulates the target LLM, and a lightweight interpreter network identifies which prompt sentences most influence a given output, without requiring further API calls.
This paper introduces an Integrated Gradients-based token attribution method to diagnose sycophancy in LLMs at the token level, and proposes attribution-guided contrastive activation steering to reduce sycophantic behavior during inference without retraining.
LlamaParse now supports multi-layered bounding boxes (region, line, word) for granular document text attribution, improving auditability in invoices, research reports, and more.
LAPO proposes a leave-one-turn attribution method for self-generated process rewards in multi-turn search reasoning, enabling fine-grained credit assignment without external reward models. It achieves state-of-the-art results across seven datasets.
This paper introduces a training-free attribution method to identify sparse inter-layer dependencies in Transformer FFN neurons, showing that small subsets of preceding activations suffice to preserve neuron activations with high fidelity.
This paper introduces a class of nonlinear axiomatic attribution methods for cooperative games to overcome the limitations of the linear Shapley value, which has an excessively large null space. Experimental results demonstrate the potential effectiveness of these methods in terms of inclusion AUC metric compared to Shapley value variants.
The article discusses the challenge of attribution in AI agent monetization, where determining credit for conversions becomes complex when agents recommend products and influence users before clicks.
Introduces MultAttnAttrib, a training-free method for multimodal attribution in long document QA, along with the MultAttrEval benchmark. It outperforms prompting-based methods and matches frontier models like GPT-5.4.
This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.
Proposes Gradient-Based Connections (GBC), a method that models multi-agent LLM systems as computational graphs and uses gradient signals to attribute errors to specific agents, enabling better system-level optimization.
Introduces the Member vs Generated Inference (MGI) task to distinguish training members from generated outputs in generative models, and proposes Data Circuit Breaker (DCB), a three-stage method combining autoencoder and latent generator signals, which outperforms existing methods across autoregressive and diffusion models.
This paper introduces CAMS, a modular multi-document summarization framework that extracts atomic claims with token-level provenance, clusters equivalent claims, and rewrites them into summaries with fine-grained, multi-source traceability, significantly improving faithfulness and citation precision.
Proposes an anytime-valid attribution method that uses a human-labeled anchor set and a betting e-process to distinguish whether score drift in LLM evaluation pipelines comes from the system or the judge, resolving the ambiguity caused by silent judge changes.