Tag
The article reviews Space Bunny, an AI model with strong visual understanding and coding capabilities, excelling in tasks like image-to-code, 3D modeling, and creative prototyping, available in OpenCode.
Researchers from MIT Senseable City Lab discuss the use of visual AI to analyze urban environments, highlighting its potential for urban planning while addressing concerns about privacy and fairness in a new book.
The AI City Challenge at ECCV 2026 showcased winners developing robust visual AI systems, with top solutions in video forecasting using NVIDIA Cosmos world foundation models for tasks like scene understanding and prediction.
The article highlights the most impressive application of 'Image to Video' technology seen so far this year.
Perceptron, a startup founded by ex-Meta scientists, has launched Isaac 0.5, an open-weight visual AI model designed to help robots perceive, reason, and act in industrial environments like warehouses and factories.
NewEyes AI by Collov Labs is a visual assistant that understands camera input contextually—styling outfits, counting calories, and reminding about plant care—rather than just capturing pixels.
ComfyUI is an open-source, node-based AI creation engine that lets builders and visual professionals design complex generation workflows for image, video, audio, and 3D without coding, with partial re-execution and broad model support.
Former DeepMind researcher Andrew Dai raised a $55 million seed round at a $300 million valuation for his visual AI startup Elorian, discussing fundraising strategy and the importance of visual understanding in AI.
A comment expressing nostalgia for the time when AI-generated images were easily distinguishable from real ones, highlighting the increasing sophistication of visual AI.
A major shift has occurred in the visual AI field: top tools no longer directly generate final outputs, but instead generate the source code behind them. a16z partner Yoko Li has provided an in-depth analysis of this.
The article argues that the next frontier of visual AI is generating code (e.g., SVG, HTML/CSS, React components) instead of raw pixels, enabling editability, iteration, and integration into professional design and development workflows.
Martin Scorsese joins Black Forest Labs as an advisor and uses FLUX for storyboarding, showcasing AI's role in visual creativity.
EyeBench-V3 visual benchmark evaluates Claude Opus 4.8, finding it still fails basic vision tasks, similar to IBench. The benchmark is introduced via a Twitter thread by Adonis Singh.
OpenAI released ChatGPT Images 2.0, claiming a GPT-3-to-GPT-5 leap; Simon Willison benchmarks it with a "Where's Waldo"-style raccoon-and-ham-radio prompt against gpt-image-1, Google Nano Banana 2 and Pro, showing mixed hide-and-seek success.
ChatGPT Images 2.0 in "Thinking" mode can turn 1,000-word prompts or 70-page PDFs into ready-to-use infographics, slide decks, and academic posters without manual editing.