Tag
The article explains Carl Sagan's analogy of how higher-dimensional beings perceive our 3D world, using geometric logic to illustrate concepts that challenge human perception of reality.
A tweet discusses the challenge of explaining AI to non-tech audiences and asserts that AI adoption has not yet truly begun.
This paper proposes a domain-knowledge-free metacognitive layer for fusing multiple pre-trained ViT-based perception models, using label vector pools and consistency-based abduction. It matches majority-vote baselines on clean data and is particularly robust against coordinated label-flipping attacks.
This paper introduces IllusionReasoning, a benchmark using real-world visual illusions to jointly evaluate the perception and reasoning capabilities of Large Vision Language Models (LVLMs), finding that current models' reasoning abilities are not as advanced as claimed.
Astronauts returning from six-month ISS missions report a persistent 'observer sensation'—feeling detached from their own lives as if watching from outside—weeks after landing, a perceptual aftereffect of neurological adaptation to microgravity.
Seed IQ demonstrates advanced real-time perception, reasoning, and adaptation in dynamic environments by navigating Doom II, potentially surpassing static benchmarks like ARC AGI 3.
A new arXiv paper introduces the ActiveVision benchmark designed to test repeated visual perception, finding that frontier vision models like GPT-5.5 and Claude Fable 5 score only 10.6% and 3.5% respectively, while humans achieve 96.1%.
A research paper proposing Color Pass-Through, an end-to-end learned framework that treats camera and display as a coupled system to improve color accuracy of images viewed on screens, achieving significant gains over baselines.
This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
This survey paper systematically reviews the paradigm evolution of unified vision-language perception in multimodal large language models (MLLMs), proposing a five-stage taxonomy and identifying open challenges toward general multimodal intelligence.
PerceptionRubrics introduces a rubric-based evaluation framework for multimodal models that uses atomic auditing and gated scoring to better align benchmark scores with human perception, revealing reliability gaps and open-closed stratification.
An analysis of the software stack behind autonomous robots, breaking down the components from perception to cloud support, and highlighting that most tools are open-source.
This paper introduces ViGOS, a method for multimodal on-policy self-distillation that decouples perception and reasoning by having the student model first produce a visual description before reasoning, reducing shortcut reliance and improving image-grounding behavior.
This paper systematically surveys the core components of medical embodied AI, emphasizing the coordinated integration of perception, decision-making, and action in clinical environments, and reviews representative applications, datasets, and future research directions.
This paper introduces PhysTool-Bench, a benchmark for evaluating multimodal large language models' ability to recognize and plan the use of physical tools in real-world scenes. The authors find that even the best model identifies only 58.7% of tools and completes just 21.0% of queries end-to-end, revealing a two-level deficit in perception and functional commonsense.
The article examines the trend of people labeling any obviously altered image or video as 'AI generated,' questioning whether the term is becoming a generic label for digital manipulation that predates AI.
MemDreamer decouples perception and reasoning for long video understanding using hierarchical graph memory and agentic retrieval, achieving state-of-the-art performance with reduced computational overhead.
A survey presenting a human-view perspective on video understanding with multimodal large language models, organized around watching, remembering, and reasoning abilities, covering challenges, methods, and applications.
A tweet discusses how OpenAI launches are no longer seen as startup-killers, referencing a new Codex feature that deploys websites using Cloudflare's Sites, D1, and R2.