Tag
The article speculates on when LLMs might begin bootstrapping themselves, potentially leading to an AI singularity with exponential progress. It invites discussion and resources on this topic.
The tweet announces the start of Interspeech 2026 in Sydney, highlighting the presentation of 2 tutorials and 19 papers on spoken language models and conversational speech recognition.
C5R has built a research facility that is entirely run by the AI model GPT-6 Astra, enabling autonomous design, execution, and observation of experiments across multiple scientific disciplines.
NVIDIA and Stanford developed a Contrastive Language Model (CLM) that frames decision-making in AI as a retrieval problem, achieving up to 9x lower latency compared to existing systems like Jev by caching action embeddings.
FB-GDM introduces a fully-Bayesian guided diffusion method for high-dimensional linear inverse problems, eliminating per-task hyperparameter tuning and demonstrating robust performance improvements over existing techniques.
PixelJev is introduced as a native-image decision interface using small open multimodal models to map images, instructions, and candidate sets to structured choices. The study demonstrates adaptation improves accuracy on benchmarks like Pets and highlights challenges in calibration and generalization.
This paper proposes an LLM-driven unified conversion framework that enables automatic and semantically consistent transformation between behavior trees and finite state machines in autonomous intelligent systems, improving scalability and maintainability.
This paper introduces SkillPivot, a deviation-guided framework for skill self-evolution in LLM agents that improves skills by contrasting failed and successful trajectories from a common prefix.
This paper introduces MI-SARSA, a reinforcement learning algorithm that models bounded rationality by incorporating mutual-information regularization to balance policy complexity and reaction time under cognitive constraints.
The paper proposes the Slicing-Graphing-Alignment (SGA) method to quantify uncertainty in multi-step forecasting for time series foundation models, demonstrating superior performance and revealing an empirical scaling law.
The paper proposes VINTAGE-TS, a revision-aware adaptation of a time-series foundation model that distinguishes observation time from information-availability time, and provides a framework for evaluating forecasts with data revisions.
The paper proposes DEEPO, a dual-stage reinforcement learning optimization method to reduce hallucination in multimodal large language models by addressing weaknesses in the correction chain from reward to parameter update.
The paper introduces CARE, a condition-aware representation regularization framework for diffusion models that improves sample quality and training efficiency by dynamically modulating feature distributions based on condition similarity. Empirically, it achieves significant reductions in FID and faster convergence for both class-to-image and text-to-image tasks.
The paper presents TW3Cast, a time-series forecasting system that uses a frozen router of lightly fine-tuned foundation models to achieve top performance on the GIFT-Eval benchmark without agents or language models.
ModularSQL introduces a runtime guardrail to address the multiplicity blind spot in Text-to-SQL systems, detecting and correcting multiplicity errors with low overhead to improve execution safety in production.
This paper evaluates the causal link between explanations and model predictions in vision-language reasoning through generation order interventions, finding that larger models are required for rationale-first reasoning and that answer-first generation reduces format-related errors.
This paper introduces a lightweight backchannel head for full-duplex spoken dialogue models to predict and control the timing of backchannels, improving natural conversation dynamics.
The paper introduces Active Taskless Distillation (ATD), a method that transfers capabilities from a teacher model to a student model using only single-word responses on task-unrelated prompts, probing the behavioral shadows of post-training.
A researcher rants about problems with AAAI peer review, including unblinded papers, AI-like reviews, and poor organizer communication.
A user excitedly announces that a preprint will be released soon, related to the NeurIPS conference.