Tag
The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.
The paper presents a method for one-step code generation by making language continuous, using diffusion models, and distilling the trajectory, with accompanying code release.
This paper proposes distilling deep optical flow stereo methods into a single-satellite model for efficient retrieval of dense three-dimensional wind fields, improving accuracy over traditional atmospheric motion vectors in certain spectral bands.
This paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level to improve on-policy distillation, leading to better performance and increased GPU utilization in asynchronous setups.
The paper introduces PersonaLink, a training-free method that distills user history into a bounded persona, matching retrieval on classification tasks but not on regression, highlighting a task-type asymmetry.
This paper proposes a method to close the MLP reachability gap in Low-Rank Clone distillation by training the full deployed matrix, resulting in significant improvements in token efficiency and model performance at no additional inference cost.
FlashRender is a few-step generative rendering framework that accelerates video synthesis by aligning representations and using mean-flow objectives, achieving comparable quality to multi-step methods with significantly reduced sampling cost.
DISTAL is a dual-prior framework that combines self-supervised compositional pretraining and structure-aware knowledge distillation to achieve robust structure-agnostic materials property prediction in low-data settings, improving performance across multiple benchmark tasks.
The author describes an experiment where AI agents like Claude Code and Codex are given a workbook to document their decisions and tradeoffs during tasks, exploring the potential use of these reasoning traces as a signal for distillation.
FastVideo has released an open-source, post-trained version of Minimax H3 that generates video 50x faster, producing 5 seconds of video from 3 seconds of compute for text-to-audio-video generation.
This paper demonstrates that scaling up off-task model-generated distillation data can amplify latent teacher traits in students, even when the data appears benign, suggesting the need for trait-aware curation in AI training.
EditaLive is a novel framework for real-time human-centric live-stream video editing that adapts image animation models using causal generation and distillation for efficient streaming inference.
This paper proposes a Sparse-Activation-ReLU (SAR) layer for low-latency, energy-efficient virtual sensing, achieving significant improvements in latency-error-energy metrics and reducing errors through synthetic knowledge distillation.
PROOF-Gen presents a per-scenario reflective optimization method to recover failed trajectories for distilling tool-calling capabilities, enhancing data quality and model performance in distillation pipelines.
STAR-OPD is a novel on-policy distillation method for ABSA quadruple extraction that uses set-structured rewards to correct structural errors in student models, improving performance and narrowing the gap with teacher models.
The paper identifies factual access failures in large language models after supervised fine-tuning and introduces Recall-Anchored Distillation (RAD) to preserve out-of-distribution factual recall without labeled data.
This paper proposes Intention Distillation (INDI) to distill behavior intent into the action decoder of Vision-Language-Action models, improving performance on benchmarks like SimplerEnv-Bridge and real-world tasks.
A talk introducing young economists to post-training in AI, discussing methods like SFT, DPO, and future challenges in world adaptation.
A software developer discusses how OpenAI and Anthropic's strategy of keeping models internal may influence open weight releases and the replication of models by Chinese companies.
The paper proposes TUP, a method for BoN-style distillation via rank-based classification that truncates low-ranked completions and upweights high-ranked ones to improve alignment efficiency and performance.