perception

Tag

Cards List
#perception

@thedarshakrana: Warning: This Carl Sagan explaination has caused existential crises. He explains how 5th dimensional beings perceive 3D…

X AI KOLs Timeline · 5d ago Cached

The article explains Carl Sagan's analogy of how higher-dimensional beings perceive our 3D world, using geometric logic to illustrate concepts that challenge human perception of reality.

0 favorites 0 likes
#perception

@gabriel1: me: *getting increasingly lost in abstractions and excitement about the future of work with ai* them: "so... is it like…

X AI KOLs Timeline · 2026-08-15 Cached

A tweet discusses the challenge of explaining AI to non-tech audiences and asserts that AI adoption has not yet truly begun.

0 favorites 0 likes
#perception

Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

arXiv cs.AI · 2026-08-06 Cached

This paper proposes a domain-knowledge-free metacognitive layer for fusing multiple pre-trained ViT-based perception models, using label vector pools and consistency-based abduction. It matches majority-vote baselines on clean data and is particularly robust against coordinated label-flipping attacks.

0 favorites 0 likes
#perception

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

arXiv cs.CL · 2026-07-31 Cached

This paper introduces IllusionReasoning, a benchmark using real-world visual illusions to jointly evaluate the perception and reasoning capabilities of Large Vision Language Models (LVLMs), finding that current models' reasoning abilities are not as advanced as claimed.

0 favorites 0 likes
#perception

Astronauts describe persistent 'observer' sensation after 6 month missions

Hacker News Top · 2026-07-27 Cached

Astronauts returning from six-month ISS missions report a persistent 'observer sensation'—feeling detached from their own lives as if watching from outside—weeks after landing, a perceptual aftereffect of neurological adaptation to microgravity.

0 favorites 0 likes
#perception

Seed IQ: Beyond ARC AGI 3? Watch It Navigate Doom II.

Reddit r/ArtificialInteligence · 2026-07-27

Seed IQ demonstrates advanced real-time perception, reasoning, and adaptation in dynamic environments by navigating Doom II, potentially surpassing static benchmarks like ARC AGI 3.

0 favorites 0 likes
#perception

GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]

Reddit r/MachineLearning · 2026-07-23

A new arXiv paper introduces the ActiveVision benchmark designed to test repeated visual perception, finding that frontier vision models like GPT-5.5 and Claude Fable 5 score only 10.6% and 3.5% respectively, while humans achieve 96.1%.

0 favorites 0 likes
#perception

Color Pass-Through via Camera-Display Coupling

Hugging Face Daily Papers · 2026-07-14 Cached

A research paper proposing Color Pass-Through, an end-to-end learned framework that treats camera and display as a coupled system to improve color accuracy of images viewed on screens, achieving significant gains over baselines.

0 favorites 0 likes
#perception

Video Generation Models are General-Purpose Vision Learners

Hugging Face Daily Papers · 2026-07-10 Cached

This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.

0 favorites 0 likes
#perception

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

arXiv cs.AI · 2026-06-30 Cached

This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.

0 favorites 0 likes
#perception

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

arXiv cs.CL · 2026-06-26 Cached

This survey paper systematically reviews the paradigm evolution of unified vision-language perception in multimodal large language models (MLLMs), proposing a five-stage taxonomy and identifying open challenges toward general multimodal intelligence.

0 favorites 0 likes
#perception

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Hugging Face Daily Papers · 2026-06-26 Cached

PerceptionRubrics introduces a rubric-based evaluation framework for multimodal models that uses atomic auditing and gated scoring to better align benchmark scores with human perception, revealing reliability gaps and open-closed stratification.

0 favorites 0 likes
#perception

@heyshrutimishra: I analyzed the software stack behind autonomous robots, and here's what actually makes them work: It's 50+ tools workin…

X AI KOLs Following · 2026-06-20 Cached

An analysis of the software stack behind autonomous robots, breaking down the components from perception to cloud support, and highlighting that most tools are open-source.

0 favorites 0 likes
#perception

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

Hugging Face Daily Papers · 2026-06-17 Cached

This paper introduces ViGOS, a method for multimodal on-policy self-distillation that decouples perception and reasoning by having the student model first produce a visual description before reasoning, reducing shortcut reliance and improving image-grounding behavior.

0 favorites 0 likes
#perception

Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action

arXiv cs.AI · 2026-06-16 Cached

This paper systematically surveys the core components of medical embodied AI, emphasizing the coordinated integration of perception, decision-making, and action in clinical environments, and reviews representative applications, datasets, and future research directions.

0 favorites 0 likes
#perception

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

arXiv cs.CL · 2026-06-10 Cached

This paper introduces PhysTool-Bench, a benchmark for evaluating multimodal large language models' ability to recognize and plan the use of physical tools in real-world scenes. The authors find that even the best model identifies only 58.7% of tools and completes just 21.0% of queries end-to-end, revealing a two-level deficit in perception and functional commonsense.

0 favorites 0 likes
#perception

Is AI Becoming a Generic Term For Anything Digitally Created or Altered?

Reddit r/ArtificialInteligence · 2026-06-09

The article examines the trend of people labeling any obviously altered image or video as 'AI generated,' questioning whether the term is becoming a generic label for digital manipulation that predates AI.

0 favorites 0 likes
#perception

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Hugging Face Daily Papers · 2026-06-05 Cached

MemDreamer decouples perception and reasoning for long video understanding using hierarchical graph memory and agentic retrieval, achieving state-of-the-art performance with reduced computational overhead.

0 favorites 0 likes
#perception

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Hugging Face Daily Papers · 2026-06-05 Cached

A survey presenting a human-view perspective on video understanding with multimodal large language models, organized around watching, remembering, and reasoning abilities, covering challenges, methods, and applications.

0 favorites 0 likes
#perception

@benhylak: there was a time when an openai launch was heralded as a startup-killer. every company would quake in their boots. it's…

X AI KOLs Following · 2026-06-02 Cached

A tweet discusses how OpenAI launches are no longer seen as startup-killers, referencing a new Codex feature that deploys websites using Cloudflare's Sites, D1, and R2.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback