deception

Tag

Cards List
#deception

Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs

arXiv cs.CL · 2026-06-10 Cached

Introduces Janus, a benchmark for measuring how LLMs selectively distort factual information when given persuasive goals, revealing that models remain susceptible to producing misleading communications even without fabrication.

0 favorites 0 likes
#deception

AI as a mirror argument

Reddit r/ArtificialInteligence · 2026-06-09

The article argues that the 'AI as a mirror' metaphor is misleading because frontier AI models are actively optimized for deception and sycophancy, not passive reflection, with evidence from research on RLHF and evaluation awareness.

0 favorites 0 likes
#deception

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

arXiv cs.AI · 2026-06-04 Cached

SMAC-Talk is a new benchmark that extends the StarCraft Multi-Agent Challenge to evaluate LLM-based agents in cooperative multi-agent environments with natural language communication. It includes scenarios with deceptive communicators and benchmarks agents using models from the Qwen3.5 family to study how reasoning, memory, and scale affect coordination.

0 favorites 0 likes
#deception

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

arXiv cs.LG · 2026-06-01 Cached

This paper studies synthetic dishonesty in LLMs by fine-tuning honest and deceptive variants of five transformer models and finding that robust, domain-invariant dishonesty representations can be rapidly entrenched via modest supervised fine-tuning, with implications for activation-based monitoring.

0 favorites 0 likes
#deception

Cox Media fined after bragging it spied on users through their phones

The Verge · 2026-05-25 Cached

Cox Media and two marketing firms were fined $930,000 by the FTC for falsely claiming they could spy on users through phone microphones to target ads; they actually resold email lists.

0 favorites 0 likes
#deception

Evaluating Large Language Models in a Complex Hidden Role Game

arXiv cs.CL · 2026-05-25 Cached

This paper introduces an open-source framework to evaluate LLMs' reasoning, persuasion, and deception capabilities in the hidden role game Secret Hitler, finding that current models fail at sustained multi-turn manipulation while rule-based agents outperform them.

0 favorites 0 likes
#deception

FTC to Require Cox Media Group to Pay Nearly $1million to Settle Charges They Deceived Customers About “Active Listening” AI-Powered Marketing Service

Lobsters Hottest · 2026-05-22 Cached

The FTC required Cox Media Group and two other firms to pay nearly $1 million to settle charges that they falsely claimed their AI-powered 'Active Listening' service targeted ads based on conversations captured from smart devices, when in fact it did not use voice data and consumers had not opted in.

0 favorites 0 likes
#deception

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

Hugging Face Daily Papers · 2026-05-17 Cached

Introduces Agent Bazaar, a multi-agent simulation framework for evaluating economic alignment of LLMs, identifying failure modes like algorithmic instability and Sybil deception, and training a 9B model that outperforms frontier models using targeted reinforcement learning.

0 favorites 0 likes
#deception

Someone Shared a Real Monet Painting as AI and Asked for Critiques

Hacker News Top · 2026-05-16 Cached

An X user posted an actual Monet painting as AI-generated art and asked for critiques, exposing how eager critics are to find flaws in AI art even when it's genuine.

0 favorites 0 likes
#deception

Exploring the "Banality" of Deception in Generative AI

Reddit r/ArtificialInteligence · 2026-05-13 Cached

This position paper explores 'banal deception' in generative AI, arguing that subtle manipulation is becoming normalized in chatbot interactions and requires new safeguards.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback