bilevel-optimization

Tag

Cards List
#bilevel-optimization

To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity

arXiv cs.LG ↗ · 7h ago Cached

This paper proposes PROBE, a perturbed gradient algorithm for bilevel optimization with nonconvex lower levels, using second-order stationarity to achieve finite-time convergence and outperform state-of-the-art methods in experiments on LLM-based tasks and meta-learning.

0 favorites 0 likes
#bilevel-optimization

Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

arXiv cs.LG ↗ · 2026-09-21 Cached

The paper proposes a single-level optimization formulation that incorporates KKT conditions into decision-focused learning for mean-variance portfolio optimization, improving performance in experiments on real-world ETF data.

0 favorites 0 likes
#bilevel-optimization

Federated stochastic bilevel optimization with fully first-order gradients

arXiv cs.LG ↗ · 2026-09-16 Cached

The paper introduces a federated stochastic bilevel optimization algorithm that uses only first-order gradients to avoid second-order matrix computations, reducing running time, and includes a novel learning rate mechanism with experimental confirmation.

0 favorites 0 likes
#bilevel-optimization

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper proposes Global-ImpactCache (GCache), a bilevel optimization framework that learns cache reuse policies for diffusion models by aligning error weighting with final generation quality, instead of relying on local similarity heuristics. It achieves significant speedups and quality improvements on image and video generation tasks, including a 2.17x speedup on Wan2.1 with lower LPIPS.

0 favorites 0 likes
#bilevel-optimization

Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization

arXiv cs.LG ↗ · 2026-08-05 Cached

This paper presents a bilevel reformulation for grey-box Bayesian optimization that separates black-box and white-box variables, reducing surrogate dimensionality and improving regret and wall-clock time on benchmark problems.

0 favorites 0 likes
#bilevel-optimization

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

arXiv cs.LG ↗ · 2026-07-14 Cached

This paper proposes a bilevel optimization framework for Direct Preference Optimization under noisy preference labels, introducing a metadata-free meta-reweighting method that uses central-difference approximation and LoRA fine-tuning to improve alignment performance.

0 favorites 0 likes
#bilevel-optimization

Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

arXiv cs.AI ↗ · 2026-06-24 Cached

This paper presents Agentic-LTPO, a nested bilevel optimization framework that uses agentic AI to adapt physical layer configurations under dynamic operator policies, achieving 57.2% long-term performance improvement in cell-free MIMO beamforming.

0 favorites 0 likes
#bilevel-optimization

FastMix: Fast Data Mixture Optimization via Gradient Descent

arXiv cs.LG ↗ · 2026-06-16 Cached

FastMix is a novel framework that automates data mixture discovery for training large models using a single proxy model and bilevel optimization, achieving state-of-the-art performance with significant efficiency gains.

0 favorites 0 likes
#bilevel-optimization

Pseudospectral Bounds for Transient Amplification in Coupled Gradient Descent

arXiv cs.LG ↗ · 2026-06-04 Cached

This paper develops a sharp pseudospectral theory for block-triangular Jacobians in coupled gradient descent, proving Kreiss-constant bounds and establishing iteration complexity results. The work exposes non-asymptotic, instance-dependent transient amplification phenomena relevant to bilevel optimization, two-time-scale stochastic approximation, and GAN training.

0 favorites 0 likes
#bilevel-optimization

IGT-OMD: Implicit Gradient Transport for Decision-Focused Learning under Delayed Feedback

arXiv cs.LG ↗ · 2026-05-14 Cached

This paper identifies 'staleness amplification' in bilevel optimization under delayed feedback and proposes IGT-OMD, which uses Implicit Gradient Transport to achieve sublinear regret and improve decision loss on benchmarks like Warcraft shortest-path and LQR.

0 favorites 0 likes
#bilevel-optimization

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

Hugging Face Daily Papers ↗ · 2026-05-11 Cached

The paper introduces FocuSFT, a bilevel optimization framework that enhances long-context language model performance by addressing attention dilution through parametric memory. It demonstrates significant improvements in accuracy and context engagement on benchmarks like BABILong and RULER.

0 favorites 0 likes
← Back to home

Submit Feedback