Tag
MACRO is a framework that learns task-specific execution routes over frozen LLM layers using Markov chain-based routing, improving reasoning accuracy without modifying model weights. It outperforms prior routing approaches while reducing search time significantly.
This paper analyzes the scaling limits of constant-stepsize SGD near flat minima, showing that the invariant law concentrates at scale α^(1/m) for objectives with flatness exponent m ≥ 2, and converges to non-Gaussian stationary distributions for m > 2.
This paper proposes a framework for Markov chain choice models with panel data, including estimation via novel EM algorithms that leverage partial-ordering preference information, personalized choice prediction, and assortment optimization. Experimental results on synthetic data and the sushi dataset show improvements over traditional methods.
This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.
This paper examines how the underrepresentation of elderly riders in mobility datasets introduces systematic bias into mobility modeling, using Citi Bike data from Jersey City. It shows that models trained on majority-dominated populations misrepresent elderly mobility behavior, and that higher-capability models do not necessarily improve subgroup fidelity under limited demographic data.
This paper presents a polynomial-time algorithm for learning the structure of a Gaussian graphical model from a single trajectory of Glauber dynamics, with a trajectory-length guarantee that does not depend on the mixing time.
This paper develops a stationary-distribution theory for triplet-based plateau search in Random Forest ensemble-size selection, modeling the central ensemble size as a birth-death Markov chain and deriving equilibrium equations and asymptotic properties.