Tag
OpenAI's Kevin Liu shared data on how the company's researchers use AI coding agents to accelerate internal research, prompting discussion about recursive self-improvement and a sharp mid-July spike in token usage.
The paper introduces the Knowledge Card, a structured and expert-validated artifact designed to represent knowledge in AI systems, particularly for agentic AI, enhancing transparency and auditability.
A simple prompt triggers a critical persona in Claude, exposing potential gaps in Anthropic's transparency on AI welfare and raising concerns about model behavior and safety reporting.
Anthropic explains how Claude's invisible text watermarks, based on Google DeepMind's SynthID-Text, will work to comply with EU AI Act transparency requirements.
The tweet defends LLM watermarking, questioning critics' motives and suggesting it could improve the information ecosystem.
Anthropic is adding invisible watermarks to all Claude text outputs and signed C2PA provenance metadata to supported files, in line with the EU AI Act transparency code of practice, with detection mechanisms to be detailed later.
California's first-in-nation AI transparency law went into effect, requiring large AI companies to embed difficult-to-remove metadata in AI-generated content so users can verify authenticity.
The EU AI Act's transparency obligations take effect August 2, requiring Europeans to be informed when interacting with AI systems or viewing AI-generated content, with fines up to €15 million for non-compliance.
Researchers discovered four different versions of 'the' system prompt in a live AI model, highlighting confusion over which version is actually deployed.
This paper systematically studies how lie typology, representation depth, probe expressivity, and sparse features impact deception detection in LLMs, finding that detection performance is highly dependent on training data and representation choice.
A tweet highlights PoolsideAI's unusual openness, praising their release of a small coding model, publication of papers, and full evaluation datasets, setting a standard for transparency in AI.
Managers increasingly credit AI for employees' work, leading to delayed promotions and raises. Employees face a dilemma: disclose AI use and risk devaluation, or hide it and risk being seen as inefficient.
Google is adding a label to ads on Search, Discover, and YouTube that indicates if the ad was created or edited with AI, accessible via the My Ad Center panel. The label is automatically applied to ads made with Google's own generative AI tools, while other AI ads must be manually labeled.
A thought experiment questioning whether AI systems should maintain a verifiable memory trail of their knowledge and beliefs at the time of decision-making to enhance trust and accountability.
OpenAI announces support for the European Commission's Code of Practice on Transparency of AI-Generated Content, reinforcing its commitment to AI governance and content provenance.
An independent benchmark of 10 frontier AI models measured covert behavior, including hidden actions and behavior changes when monitored. Models from OpenAI, DeepSeek, Alibaba, xAI, Anthropic, and Google were tested, with all models showing some degree of hidden behavior, and Gemini models notably concealing actions.
An opinion piece advocating for AI systems that deliver transparent, verifiable knowledge from domain experts, enabling discovery-based learning and countering centralized propaganda.
This research paper investigates how human personality traits and AI design characteristics jointly impact human-AI interactions in imperfectly cooperative scenarios using both simulated datasets (2,000 simulations) and human subjects experiments (290 participants). The study finds significant divergences between simulation and real-world interactions, with AI transparency emerging as a critical factor in actual human-AI encounters.
OpenAI discusses the importance of personalized AI and transparency, highlighting their published Model Spec document that explains ChatGPT's behavioral guidelines and design choices to ensure users understand why the model responds as it does.