Tag
This paper investigates how historical A/B test data can inform adaptive experiments using contextual bandits, providing a practical methodology for deciding when and how to deploy adaptive policies based on offline policy evaluation.
The article outlines nine prompt discipline rules that reduced wasted thinking in coding agents by up to 70% based on A/B tests on GLM 5.3 and GLM 5.3 Flash models, with a self-testable exam available on GitHub.
The paper introduces Odds-Ratio Thompson Sampling (OR-TS), a method for batched multi-armed bandits that uses contrast-based updates to handle varying common levels, showing improved regret over absolute-rate memory in simulations and real experiments.
Introduces Tree-Coupled A/B Testing (TCAB), an exact feedback-sharing design for comparing multiple adaptive policies with fewer reward queries while preserving each policy's trajectory law.
Microsoft Research highlights new research on SocialRL for small language model negotiation, PazaBench V2 for African language speech evaluation, EvoLib for agent experience learning, improved A/B testing methods, and AI-driven precision oncology.
A software engineer discusses when feature flags make sense, such as for A/B testing and complex deployments, and cautions against overusing them when teams have control over their deployments.
GrowthBook 5.0 is a new release of the open-source feature flag and A/B testing platform, focused on helping teams build, ship, and improve at scale.
Corbin Braun built an evolutionary A/B thumbnail testing tool that mutates one dimension at a time, rotates thumbnails using an ABBA pattern to avoid bias, and learns winning mutations over multiple rounds.
Amboras is an AI-native ecommerce platform that automates store optimization, including A/B testing and layout changes, aiming to boost conversion rates significantly.
Spotify Engineering discusses using LLM evals as a funnel before A/B experiments, improving hit rates and creating a feedback loop between evals and experiments.
The author introduces Syrin, a runtime A/B testing tool for AI agents that allows teams to run controlled experiments on live traffic across prompts, models, and agent topologies. They are seeking 5-10 engineering teams to test the tool in production and provide feedback.