contextual-bandits

Tag

Cards List
#contextual-bandits

When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence

arXiv cs.LG · yesterday Cached

This paper studies KL-regularized contextual bandits and shows that greedy sampling can achieve logarithmic regret without explicit eluder-dimension dependence for both reward and preference feedback.

0 favorites 0 likes
#contextual-bandits

COBRA-Skills: Contextual Bandits for Efficient Agent Skill Optimization (Open Source)

Reddit r/artificial · 2d ago

The article presents COBRA-Skills, a method using contextual bandits to efficiently optimize agent skills, reducing costs by avoiding repeated LLM-based trajectory analysis and skill rewriting.

0 favorites 0 likes
#contextual-bandits

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Hugging Face Daily Papers · 6d ago Cached

COBRA-Skills introduces a method for optimizing agent skills using contextual bandits and evolutionary operators, achieving 55-58% lower optimization costs across multiple AI agent benchmarks.

0 favorites 0 likes
#contextual-bandits

Online Learning with LLM Experts from Limited Feedback

Hugging Face Daily Papers · 2026-09-05 Cached

This paper formulates the adaptive routing of prompts to large language model experts as a contextual bandit problem with limited feedback, proposing algorithms that achieve sublinear regret and demonstrate efficient learning of high-quality routing strategies.

0 favorites 0 likes
#contextual-bandits

Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions

arXiv cs.LG · 2026-08-28 Cached

This paper proposes High-Probability Constrained UCB for contextual bandits with continuous actions, emphasizing realized-cost constraints over expected-cost to improve safety, and provides theoretical regret bounds and experimental validation.

0 favorites 0 likes
#contextual-bandits

Learning What to Fail On: Failure-Mode Contextual Bandits for Adversarial Data Curation

arXiv cs.CL · 2026-08-20 Cached

The paper introduces a failure-aware adversarial retrieval-augmented framework using contextual bandits to improve robustness in natural language understanding, with significant improvements on benchmarks like SNLI, ANLI, and MultiNLI.

0 favorites 0 likes
#contextual-bandits

Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing

arXiv cs.LG · 2026-08-14 Cached

Introduces Tree-Coupled A/B Testing (TCAB), an exact feedback-sharing design for comparing multiple adaptive policies with fewer reward queries while preserving each policy's trajectory law.

0 favorites 0 likes
#contextual-bandits

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv cs.LG · 2026-08-13 Cached

This paper presents a diagnostic protocol for selecting rewards and policies in delayed-feedback contextual bandits, arguing that standard offline evaluation can mislead and validating the approach on benchmarks and a deployed push system.

0 favorites 0 likes
#contextual-bandits

Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints

arXiv cs.LG · 2026-08-13 Cached

This paper proposes new reoptimization algorithms for contextual bandits with knapsack constraints, achieving an average regret bound of O((ln T)^3 / T) and improving existing results.

0 favorites 0 likes
#contextual-bandits

Bootstrap-Conditioned Action Selection with Tabular Foundation Models

arXiv cs.LG · 2026-08-10 Cached

The paper proposes BC-ICL, a bootstrap-conditioned action selection method that leverages pretrained tabular foundation models with in-context learning for contextual bandits, improving exploration and regret performance under strict online protocols.

0 favorites 0 likes
#contextual-bandits

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

arXiv cs.LG · 2026-07-27 Cached

This paper introduces cross-domain off-policy evaluation and learning (OPE/L) for contextual bandits, allowing the use of logged data from multiple source domains to improve policy evaluation and learning in target domains with challenging conditions like few-shot data, deterministic logging policies, and new actions.

0 favorites 0 likes
#contextual-bandits

Information-Directed Sampling for Causal Bandits

arXiv cs.LG · 2026-07-20 Cached

This paper studies contextual causal bandits with non-manipulable variables, proposing causal variants of Thompson Sampling and Information-Directed Sampling (IDS) that exploit shared causal mechanisms to accelerate decision-making. Theoretical regret bounds and experiments on synthetic tasks show that the proposed methods outperform causal and non-causal baselines.

0 favorites 0 likes
#contextual-bandits

Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing

arXiv cs.LG · 2026-07-13 Cached

This paper proposes correlation-aware contextual bandit algorithms that leverage surrogate reward signals from machine learning models for LLM routing, achieving improved accuracy-cost trade-offs and sample efficiency compared to standard baselines.

0 favorites 0 likes
#contextual-bandits

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

arXiv cs.AI · 2026-07-02 Cached

This paper introduces a contextual-bandit team game with two-sided informational asymmetry for runtime human oversight of AI agents, characterizing gaps between team-optimal and myopic human oversight strategies.

0 favorites 0 likes
#contextual-bandits

Contextual Slate GLM Bandits with Limited Adaptivity

arXiv cs.LG · 2026-07-01 Cached

Proposes algorithms for contextual slate bandits with generalized linear rewards under limited adaptivity, achieving regret bounds independent of the non-linearity parameter. The batched and rarely-switching algorithms are computationally efficient and empirically outperform baselines, including in a language model example selection task.

0 favorites 0 likes
#contextual-bandits

Diagnosing and Repairing Factual Errors in RAG under Budget Constraints

arXiv cs.AI · 2026-06-30 Cached

This paper proposes D2R-RAG, a model-agnostic and resource-aware framework that diagnoses and repairs factual errors in RAG systems under latency and VRAM constraints, achieving better accuracy-efficiency trade-offs on FEVER and HotpotQA.

0 favorites 0 likes
#contextual-bandits

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces

arXiv cs.LG · 2026-06-29 Cached

Proposes GraphDR-LinUCB, a method for contextual bandits with graph-structured arms that projects features onto the graph's low-frequency spectral subspace. Achieves the first regret bound for spectral-projection-based contextual bandits and demonstrates 15x regret reduction on real datasets over full-dimensional LinUCB.

0 favorites 0 likes
#contextual-bandits

Contextual Bandits for Maximizing Stimulated Word-of-Mouth Rewards

arXiv cs.LG · 2026-06-16 Cached

This paper presents a contextual multi-armed bandit framework that learns individual spillover probabilities in social networks to optimize stimulated word-of-mouth marketing, achieving higher rewards by targeting connected users.

0 favorites 0 likes
#contextual-bandits

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

arXiv cs.LG · 2026-06-16 Cached

This paper formalizes embedding model routing as an adversarial contextual linear bandit with low-rank experts, proposing the Hypentropy Policy Gradient (HPG) algorithm that achieves O~(s√(MT)) policy regret, avoiding the curse of dimensionality.

0 favorites 0 likes
#contextual-bandits

Online Pandora's Box for Contextual LLM Cascading

arXiv cs.AI · 2026-06-08 Cached

This paper introduces an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs, proposing a learning approach that combines GMM estimation with UCB-style confidence bounds and proving dimension-dependent regret bounds.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback