Tag
Audio8 TTS Preview 0.1b is a compact zero-shot text-to-speech model with approximately 170M parameters for the main model, supporting voice cloning and multiple languages.
Roboflow Playground is a new developer tool that enables side-by-side comparison of over 30 computer vision models across tasks like object detection and classification, streamlining the evaluation process.
TinyCast is a compact zero-shot time series foundation model with computed periodicity, enabling efficient probabilistic forecasting on edge devices like Cortex-M7.
This paper compares natively multimodal embedding models (Gemini Embedding 2, Amazon Nova 2) against frontier LLMs (GPT-4.1, Claude Sonnet 4.6) for hard-negative text-to-image retrieval, finding comparable accuracy but much lower latency for embedding-based ranking.
This paper introduces a continuous metric field framework trained by a single causal contrastive loss that unifies geometric structure discovery from robot navigation to black hole emergence, demonstrating zero-shot generalization across dimensions.
This paper proposes EchoPrompt, a training-free detector for LLM-generated text that restores a latent prompt dependency by prepending a generic prefix and measuring likelihood gain differences between instruction-tuned and base models, achieving state-of-the-art zero-shot detection performance.
This paper systematically studies perturbation-based continued pre-training (CPT) for improving zero-shot dialect robustness in multilingual LLMs, comparing six training conditions across German, Italian, and Arabic. It finds that character-noised CPT is the most effective general strategy and reveals that different perturbation methods induce distinct robustness mechanisms.
Introduces HyperODE, a zero-shot surrogate that maps ODE structures to hypergraphs, enabling simulation and parameter inference across entire families of dynamical systems without retraining.
Introduces DE-NER, a dialogue elicitation framework for zero-shot named entity recognition that uses self-play between questioner and roleplayer LLMs to clarify entity boundaries, achieving an average 3.75% F1 improvement over baselines.
A study evaluating zero-shot GPT-4o-mini for predicting child stunting from Bangladesh Demographic and Health Survey data, comparing against a random forest baseline and assessing fairness across demographic groups and temporal robustness. Results show comparable balanced accuracy but notable fairness disparities across residence and wealth categories.
This paper introduces SwanTale, a unified multi-speaker speech and audio generation model supporting both zero-shot and instruct tasks, along with SwanData-Caption for data annotation and SwanVAE for high-quality multi-audio-modality generation.
This paper systematically evaluates 41 open-weight language models (135M–9B) for zero-shot intent classification across 8 datasets, analyzing accuracy, calibration, robustness, and deployment efficiency. It finds instruction-tuned 3B models can beat 7B base models and that some benchmarks like SNIPS are saturated.
This paper introduces EC-Reason-Bench, a training-free diagnostic benchmark to analyze why general LLMs fail on enzyme EC number prediction. It finds that external knowledge is decisive and must precede reasoning, and that reasoning over evidence acts as an arbiter of conflicting nearest neighbors rather than a source of new knowledge.
ReMem introduces a dual-level memory-augmented keyframe selection framework for training-free long video understanding, achieving state-of-the-art zero-shot performance on multiple benchmarks.
Introduces OVEarth-Bench, a benchmark for open-vocabulary Earth observation that broadens category coverage and query diversity, revealing that current methods remain limited and MLLM-based approaches perform best.
MissionBench is a new benchmark for evaluating multimodal large language models (MLLMs) on long-horizon embodied tasks in aerial 3D environments, revealing that even the best models succeed on fewer than 35% of missions compared to 84.4% human performance.
Introduces DWT-Fusion, a training-free framework using discrete wavelet analysis of token log-probabilities for detecting LLM-generated text, achieving strong AUROC results on multiple datasets.
Presents a lightweight knowledge-injection framework for zero-shot ICU delirium prediction that augments structured EHR data summaries with external clinical knowledge at inference time, improving AUROC by up to 8.57 percentage points on LLaMA models without fine-tuning.
This paper presents a novel framework for zero-shot Digital Twins that integrates real-time visual perception with a geometry-agnostic, physics-informed Graph Neural Network. The approach uses a Thermodynamics-Informed GNN to enforce energy conservation and entropy production, achieving physically accurate simulations on unseen geometries without retraining.
This paper presents a comprehensive evaluation of five large language models for citation function classification, achieving new state-of-the-art results on the ACL-ARC dataset with a fine-tuned Falcon 7B model. It also introduces the AC3 dataset, which includes a seven-category annotation scheme distinguishing neutral acknowledgments from evaluative stances.