distillation

Tag

Cards List
#distillation

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Hugging Face Daily Papers · yesterday Cached

The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.

0 favorites 0 likes
#distillation

continuous diffusion for code generation in one step

Reddit r/artificial · yesterday

The paper presents a method for one-step code generation by making language continuous, using diffusion models, and distilling the trajectory, with accompanying code release.

0 favorites 0 likes
#distillation

Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields

arXiv cs.LG · 5d ago Cached

This paper proposes distilling deep optical flow stereo methods into a single-satellite model for efficient retrieval of dense three-dimensional wind fields, improving accuracy over traditional atmospheric motion vectors in certain spectral bands.

0 favorites 0 likes
#distillation

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

arXiv cs.LG · 5d ago Cached

This paper introduces Teacher-Gated On-Policy Distillation (TGOPD), a method that verifies teacher reliability at the prompt level to improve on-policy distillation, leading to better performance and increased GPU utilization in asynchronous setups.

0 favorites 0 likes
#distillation

Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

arXiv cs.CL · 5d ago Cached

The paper introduces PersonaLink, a training-free method that distills user history into a bounded persona, matching retrieval on classification tasks but not on regression, highlighting a task-type asymmetry.

0 favorites 0 likes
#distillation

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

arXiv cs.LG · 6d ago Cached

This paper proposes a method to close the MLP reachability gap in Low-Rank Clone distillation by training the full deployed matrix, resulting in significant improvements in token efficiency and model performance at no additional inference cost.

0 favorites 0 likes
#distillation

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

Hugging Face Daily Papers · 6d ago Cached

FlashRender is a few-step generative rendering framework that accelerates video synthesis by aligning representations and using mean-flow objectives, achieving comparable quality to multi-step methods with significantly reduced sampling cost.

0 favorites 0 likes
#distillation

DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction

arXiv cs.LG · 2026-09-02 Cached

DISTAL is a dual-prior framework that combines self-supervised compositional pretraining and structure-aware knowledge distillation to achieve robust structure-agnostic materials property prediction in low-data settings, improving performance across multiple benchmark tasks.

0 favorites 0 likes
#distillation

Give your agent somewhere to think loud watch its decisions unfold live

Reddit r/AI_Agents · 2026-09-01

The author describes an experiment where AI agents like Claude Code and Codex are given a workbook to document their decisions and tradeoffs during tasks, exploring the potential use of these reasoning traces as a signal for distillation.

0 favorites 0 likes
#distillation

@DeRonin_: these guys built an infinite movie generation machine... @fal's post-trained Minimax H3 Max is 50x faster than the orig…

X AI KOLs Timeline · 2026-08-29 Cached

FastVideo has released an open-source, post-trained version of Minimax H3 that generates video 50x faster, producing 5 seconds of video from 3 seconds of compute for text-to-audio-video generation.

0 favorites 0 likes
#distillation

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

arXiv cs.LG · 2026-08-28 Cached

This paper demonstrates that scaling up off-task model-generated distillation data can amplify latent teacher traits in students, even when the data appears benign, suggesting the need for trait-aware curation in AI training.

0 favorites 0 likes
#distillation

EditaLive! Unified Character Video Editing for Live Streaming

Hugging Face Daily Papers · 2026-08-27 Cached

EditaLive is a novel framework for real-time human-centric live-stream video editing that adapts image animation models using causal generation and distillation for efficient streaming inference.

0 favorites 0 likes
#distillation

Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Edge-Deployable Virtual Sensing

arXiv cs.LG · 2026-08-26 Cached

This paper proposes a Sparse-Activation-ReLU (SAR) layer for low-latency, energy-efficient virtual sensing, achieving significant improvements in latency-error-energy metrics and reducing errors through synthetic knowledge distillation.

0 favorites 0 likes
#distillation

PROOF-Gen: From Optimized Data to Better Distillation

arXiv cs.AI · 2026-08-26 Cached

PROOF-Gen presents a per-scenario reflective optimization method to recover failed trajectories for distilling tool-calling capabilities, enhancing data quality and model performance in distillation pipelines.

0 favorites 0 likes
#distillation

STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction

arXiv cs.CL · 2026-08-24 Cached

STAR-OPD is a novel on-policy distillation method for ABSA quadruple extraction that uses set-structured rewards to correct structural errors in student models, improving performance and narrowing the gap with teacher models.

0 favorites 0 likes
#distillation

Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation

arXiv cs.AI · 2026-08-24 Cached

The paper identifies factual access failures in large language models after supervised fine-tuning and introduces Recall-Anchored Distillation (RAD) to preserve out-of-distribution factual recall without labeled data.

0 favorites 0 likes
#distillation

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Hugging Face Daily Papers · 2026-08-24 Cached

This paper proposes Intention Distillation (INDI) to distill behavior intent into the action decoder of Vision-Language-Action models, improving performance on benchmarks like SimplerEnv-Bridge and real-world tasks.

0 favorites 0 likes
#distillation

@ethayarajh: I recently gave a talk introducing young economists to post-training. The slides are now up! https://kawine.github.io/a…

X AI KOLs Timeline · 2026-08-21 Cached

A talk introducing young economists to post-training in AI, discussing methods like SFT, DPO, and future challenges in world adaptation.

0 favorites 0 likes
#distillation

Open weight progression with no frontier release

Reddit r/singularity · 2026-08-21

A software developer discusses how OpenAI and Anthropic's strategy of keeping models internal may influence open weight releases and the replication of models by Chinese companies.

0 favorites 0 likes
#distillation

Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification

arXiv cs.LG · 2026-08-21 Cached

The paper proposes TUP, a method for BoN-style distillation via rank-based classification that truncates low-ranked completions and upweights high-ranked ones to improve alignment efficiency and performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback