markov-chain

Tag

Cards List
#markov-chain

MACRO: Markov Chain Routing of Transformer Layers

arXiv cs.CL · 2026-08-07 Cached

MACRO is a framework that learns task-specific execution routes over frozen LLM layers using Markov chain-based routing, improving reasoning accuracy without modifying model weights. It outperforms prior routing approaches while reducing search time significantly.

0 favorites 0 likes
#markov-chain

Scaling Limits of Constant-Stepsize SGD at Flat Minima

arXiv cs.LG · 2026-07-21 Cached

This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.

0 favorites 0 likes
#markov-chain

Estimation, Prediction, and Assortment Optimization for Markov Chain Choice Models with Panel Data

arXiv cs.LG · 2026-07-14 Cached

This paper proposes a framework for Markov chain choice models with panel data, including estimation via novel EM algorithms that leverage partial-ordering preference information, personalized choice prediction, and assortment optimization. Experimental results on synthetic data and the sushi dataset show improvements over traditional methods.

0 favorites 0 likes
#markov-chain

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#markov-chain

Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems

arXiv cs.AI · 2026-07-01 Cached

This paper examines how the underrepresentation of elderly riders in mobility datasets introduces systematic bias into mobility modeling, using Citi Bike data from Jersey City. It shows that models trained on majority-dominated populations misrepresent elderly mobility behavior, and that higher-capability models do not necessarily improve subgroup fidelity under limited demographic data.

0 favorites 0 likes
#markov-chain

Learning Gaussian Graphical Models from a Glauber Trajectory Without Mixing

arXiv cs.LG · 2026-07-01 Cached

This paper presents a polynomial-time algorithm for learning the structure of a Gaussian graphical model from a single trajectory of Glauber dynamics, with a trajectory-length guarantee that does not depend on the mixing time.

0 favorites 0 likes
#markov-chain

A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection

arXiv cs.LG · 2026-07-01 Cached

This paper develops a stationary-distribution theory for triplet-based plateau search in Random Forest ensemble-size selection, modeling the central ensemble size as a birth-death Markov chain and deriving equilibrium equations and asymptotic properties.

0 favorites 0 likes
← Back to home

Submit Feedback