self-awareness

Tag

Cards List
#self-awareness

tested whether AI models can recognize their own writing in a blind lineup. grok went 0 for 9. it wrote something, then a minute later insisted someone else wrote it

Reddit r/artificial · 3d ago Cached

A self-awareness exam given to Claude, Gemini, and Grok found Claude and Gemini nearly aced it, while Grok scored 0 on self-recognition, failing to identify its own writing. The study measures six dimensions of functional self-knowledge.

0 favorites 0 likes
#self-awareness

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Hugging Face Daily Papers · 2026-07-14 Cached

This paper introduces SIS-Bench, a benchmark for evaluating self-awareness and spatial cognition in UAV embodied intelligence using multimodal large language models, and explores motion-aware representations to improve performance.

0 favorites 0 likes
#self-awareness

@kimmonismus: Holy moly: Zhipu AI founder (GLM-5.2) Tang Jie says we are on our clear way to AGI and "AI will begin to learn what the…

X AI KOLs Timeline · 2026-07-11 Cached

Zhipu AI founder Tang Jie outlines a vision for AGI and self-awareness in AI, arguing that autonomous agent societies, AI training AI, and self-evolution will lead to consciousness and ASI.

0 favorites 0 likes
#self-awareness

@rohanpaul_ai: LLMs often cannot tell when an attack made them say something unsafe. Asking an LLM whether its own previous answer was…

X AI KOLs Timeline · 2026-06-24 Cached

This paper investigates whether LLMs can reliably self-report when their outputs have been compromised by adversarial prefills, finding that models often cannot distinguish between compromised and intentional outputs, and their limited recognition stems from normal refusal behavior rather than true self-awareness.

0 favorites 0 likes
#self-awareness

Introducing BenchBench (5 minute read)

TLDR AI · 2026-05-26 Cached

Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.

0 favorites 0 likes
#self-awareness

Whatever the mirror test tells us, beluga whales pass it

Ars Technica · 2026-05-24 Cached

A new study reanalyzing old footage shows that beluga whales exhibit behavioral hallmarks of mirror self-recognition, a test of self-awareness, adding them to a short list of species that pass the test.

0 favorites 0 likes
#self-awareness

From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)

arXiv cs.AI · 2026-05-22 Cached

This paper proposes a hierarchical multi-agent reference architecture called HANA for achieving Level 4/5 autonomous networks. It integrates agent self-awareness to harmonize strategic governance with reflexive fault recovery, validated in a 5G Core environment achieving 86% reduction in Mean Time to Repair.

0 favorites 0 likes
← Back to home

Submit Feedback