Tag
A HEC Montréal/MILA paper proposes fair policy optimization for major-minor weakly coupled MDPs, replacing the utilitarian objective with monotone concave fairness functions and introducing a count-proportion-based deep RL approach with a priority-based sampler, validated on machine replacement and NYC taxi pricing/relocation tasks.
ACV-Gate is an adaptive candidate evaluation framework that enhances reconstruction quality in generative image communication by selectively applying counterfactual reasoning to informative visual tokens, thereby reducing computation costs.
The author suggests that heavy AI users may inefficiently allocate model capacity by not optimizing model selection, proposing that work should be routed to the least expensive capable model to save costs and improve efficiency.
Introduces PrimeScientist, a method for strategic allocation of research effort in autonomous research agents using an adaptive MCTS-based policy to improve research quality and sample efficiency under resource constraints.
This paper proposes a measurement-driven multi-layer digital twin framework for terahertz wireless data centers, using tri-band channel measurements to enable AI-based channel reconstruction and system optimization for future AI computing demands.
FL-MAESTRO introduces a multi-agent LLM orchestrator to jointly optimize communication topology, resource allocation, and aggregation rules in federated learning, significantly reducing wasted energy and improving efficiency on resource-constrained edge devices.
FleetSieve introduces a decision-critical profiling method for SLO-aware LLM fleet configuration that optimizes resource allocation by reducing unnecessary measurements, achieving efficiency gains over uniform profiling.
This paper proposes a cross-layer polar code based federated learning scheme to address communication bottlenecks and channel impairments, providing convergence analysis and resource optimization that demonstrates performance gains over uncoded and LDPC-based benchmarks.
A simulation study on GPT-4o-mini finds that distributing triage decisions across a multi-agent pipeline with an audit stage does not reduce biased outcomes, but audit capacity significantly affects whether bias is caught. Reordering audits by estimated risk recovers most lost coverage under load.
Introduces FairFund-Bench, a benchmark for evaluating distributive bias in LLM resource allocation, showing that audit format changes the direction and magnitude of bias, and that causal framing effects dominate demographic effects.
This paper proposes formulating the bridge between planned tasks and performed actions as a quasi-linear Fisher market, allowing fractional credit assignment. It introduces instruments for conservation and junk filtering, and extends the model with entropy regularization to handle noise, unifying it with optimal transport.
A novel decision-aware machine learning framework was deployed nationwide in Sierra Leone to allocate essential medicines, achieving a 19% increase in consumption and covering 2 million women and children under five.
Proposes a two-timescale multi-layer deep reinforcement learning framework with latent action space for joint service placement, computational delegation, and power control in hierarchical edge-cloud computing, achieving up to 20.8% latency reduction and 13% resource utilization improvement.
This paper presents Quota Marketplace, a market-based dynamic pricing mechanism deployed at Google for efficient allocation of ML training accelerators across business units, achieving Pareto efficiency and max-min fairness under heterogeneous workload values.
A discussion on strategies for managing token budgets when deploying multiple AI agents in production, covering cost and efficiency considerations.
Jeff Bezos argues that AI data centers should be prioritized for water and energy resources over human consumption to enable the development of superintelligence, sparking controversy over AI's environmental impact.
This paper introduces CEO-Bench, a multi-agent benchmark for evaluating LLMs on CEO-level strategic resource reallocation, revealing systematic failure modes and a structural integration–boldness tradeoff.
STARIXNet is a lightweight neural network that improves cloud resource allocation by capturing multivariate spatio-temporal relationships among system metrics, prioritizing service stability over forecast accuracy. Deployed at Walmart, it achieved 10-50% cost savings while maintaining service reliability.
This paper analyzes tradeoffs between latency, reliability, and cost in LLM-enabled agentic workflows, introducing performance models and deriving optimal resource allocation policies like water-filling token allocation.
This paper introduces Computable Fair Division (CFD), a framework using Boltzmann-Softmax control to balance efficiency and fairness in AI resource allocation, with real-time adaptation via AHC++.