Tag
Flexport CEO Ryan Petersen and Y Combinator CEO Garry Tan discuss how startup founders' initial visions often evolve and expand as they learn from customers, arguing that too much certainty about the vision is probably wrong.
The author wrote a pure-MLX runtime for the vision and audio towers of Nvidia's Nemotron Omni model, enabling the full multimodal model to run locally on Apple Silicon Macs. All 23 tests pass against the PyTorch reference, and it achieves 67-152 tok/s on an M5 Max.
Kimi K3, an open model that leads agentic coding benchmarks with native vision and a 1M-token context window, is now available to run in Codex on Modal via its Shared Endpoint.
Paul Graham explains why founders shouldn't try to be visionaries, arguing that big things start small and that having a blurry vision like Columbus beats a precise one.
The Gemini Live API enables low-latency, real-time voice and vision interactions with Gemini, supporting continuous streams of audio, images, and text for building natural conversational agents.
Baseten released GLM 5.2 Vision on Hugging Face, integrating a vision encoder from Kimi k2.6 into the GLM 5.2 model, addressing the lack of vision capabilities.
Kimi K3 is a 2.8 trillion parameter Mixture-of-Experts model with 104 billion active parameters, native vision, and a 1-million-token context window, achieving frontier-level performance across multiple domains and released as open weights.
This article revisits Apple's 1987 concept video "Knowledge Navigator," showcasing its early vision of AI agents and praising Apple's foresight.
A developer demonstrates adding vision capabilities to the GLM language model, showcasing a significant multimodal extension.
InvincibleHunter claims GPT-5.6 Sol is the most impressive model ever, with huge improvements in mathematics and vision capabilities, surpassing other models.
SenseTime releases SenseNova-Vision-7B-MoT, a fully open-sourced unified multimodal model that handles multiple vision tasks using natural language instructions, supporting detection, OCR, depth, segmentation, and more.
Empero releases Qwythos-27B-v1, an open-weight reasoning model based on Qwen3.5-27B that preserves native multi-token prediction, the full vision tower, and a 1,048,576-token context window. It demonstrates strong agentic terminal performance and improved closed-book reasoning over its 9B sibling.
The article discusses how AI agents have been given capabilities for physical interaction (hands) and voice communication, but most still lack visual perception (eyes), highlighting a key limitation.
This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.
Vantage is an open-source Windows app that uses an LLM to automate desktop tasks via natural language, enabling users to drive any desktop app with plain English commands.
LingBot Vision is a new visual foundation model released under Apache 2.0, available in four sizes (small, base, large, giant) and pretrained with masked boundary modeling. It powers a depth estimation system that tops multiple benchmarks.
LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.
A new study identifies a neural circuit that integrates visual and cognitive processing in the brain.
A new open benchmark called VCCB tests how accurately multimodal LLMs can extract calendar events from screenshots. Early results show humans near 99%, frontier models around 80-85%, and local models significantly lower, prompting a call for community submissions.
A tweet by Ahmad Osman expressing a goal to make local AI the default.