Tag
A demonstration shows that Opus 5.5, an AI model, can generate impressive SVG animations in a zero-shot setting, highlighting its advanced generative capabilities.
GPT-6 Astra is evolving robotic hand dexterity through simulation training, achieving impressive zero-shot fine manipulation capabilities like twirling a pen and solving a Rubik's Cube, with the goal of transferring these skills to the real world.
MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.
This paper introduces SVEET, a framework for high-quality streaming video editing that leverages a pretrained video diffusion model to enable auto-regressive editing with real-time performance on a single GPU.
Figure released Helix 2.5, a humanoid control model that performed household tasks zero-shot across 30 homes using pretrained data.
Brett Adcock posted a video of a robot performing zero-shot tasks across 30 rental homes, showcasing its real-world application capabilities.
This paper explores building self-adaptive physical AI agents using LLMs to manage long-horizon tasks in a zero-shot manner, showing they can adapt effectively to environmental changes compared to reinforcement learning agents.
HarnessVLN is a zero-shot, training-free framework for embodied navigation that unifies perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface, achieving state-of-the-art results on benchmarks like R2R and RxR.
IBM released the Granite Time Series PatchTST-FM-r2 model, a 385M-parameter foundation model for zero-shot time-series forecasting with top performance on the GIFT-Eval benchmark and a commercial-friendly Apache 2.0 license.
Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.
ZipDepth is a lightweight zero-shot monocular depth estimation model that achieves the best accuracy-efficiency trade-off, running in real time on any device from mobile phones to server GPUs, and has been accepted at ECCV 2026.
This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.
TC-Next is a multimodal deep learning model that leverages foundation model forecasts and satellite imagery for zero-shot tropical cyclone track and intensity forecasting, showing significant error reduction over conventional trackers.
SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.
Google introduces TimesFM-3, a state-of-the-art zero-shot foundation model for multivariate time series forecasting, capable of handling multiple targets and covariates in a single forward pass without fine-tuning.
TontaubeV1 is an open-weight text-to-speech model released for local long-form generation, supporting English and German with zero-shot voice cloning and low-latency inference on GPUs.
J-Zero is a unified framework for co-evolving Challenger, Solver, and Judge models from zero data, enabling self-improvement in language models across both verifiable and unverifiable domains with performance surpassing baselines.
The paper introduces VietAIDetector, an open-source zero-shot tool for detecting Vietnamese AI-generated text, featuring a Gradio web interface and superior performance over existing methods.
This paper proposes a zero-shot time-series anomaly detection framework that enhances LLM inputs with frequency-domain evidence from FFT to better capture structural anomalies.
A technique is shared for using the Segment Anything Model (SAM) to segment images without prompting by exploiting similarities in microscopy images, enabling zero-shot mask propagation similar to video.