ab-testing

Tag

Cards List
#ab-testing

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

arXiv cs.LG ↗ · 15h ago Cached

This paper investigates how historical A/B test data can inform adaptive experiments using contextual bandits, providing a practical methodology for deciding when and how to deploy adaptive policies based on offline policy evaluation.

0 favorites 0 likes
#ab-testing

9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)

Reddit r/LocalLLaMA ↗ · yesterday

The article outlines nine prompt discipline rules that reduced wasted thinking in coding agents by up to 70% based on A/B tests on GLM 5.3 and GLM 5.3 Flash models, with a self-testable exam available on GitHub.

0 favorites 0 likes
#ab-testing

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

arXiv cs.LG ↗ · 2026-09-18 Cached

The paper introduces Odds-Ratio Thompson Sampling (OR-TS), a method for batched multi-armed bandits that uses contrast-based updates to handle varying common levels, showing improved regret over absolute-rate memory in simulations and real experiments.

0 favorites 0 likes
#ab-testing

Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing

arXiv cs.LG ↗ · 2026-08-14 Cached

Introduces Tree-Coupled A/B Testing (TCAB), an exact feedback-sharing design for comparing multiple adaptive policies with fewer reward queries while preserving each policy's trajectory law.

0 favorites 0 likes
#ab-testing

@MSFTResearch: Small language models learn to negotiate with SocialRL, PazaBench V2 expands speech AI evaluation across African langua…

X AI KOLs Following ↗ · 2026-08-03 Cached

Microsoft Research highlights new research on SocialRL for small language model negotiation, PazaBench V2 for African language speech evaluation, EvoLib for agent experience learning, improved A/B testing methods, and AI-driven precision oncology.

0 favorites 0 likes
#ab-testing

When Feature Flags Do and Don't Make Sense

Hacker News Top ↗ · 2026-08-03 Cached

A software engineer discusses when feature flags make sense, such as for A/B testing and complex deployments, and cautions against overusing them when teams have control over their deployments.

0 favorites 0 likes
#ab-testing

GrowthBook 5.0

Product Hunt ↗ · 2026-07-29

GrowthBook 5.0 is a new release of the open-source feature flag and A/B testing platform, focused on helping teams build, ship, and improve at scale.

0 favorites 0 likes
#ab-testing

@corbin_braun: https://x.com/corbin_braun/status/2077244527988113420

X AI KOLs Following ↗ · 2026-07-15 Cached

Corbin Braun built an evolutionary A/B thumbnail testing tool that mutates one dimension at a time, rotates thumbnails using an ABBA pattern to avoid bias, and learns winning mutations over multiple rounds.

0 favorites 0 likes
#ab-testing

@ycombinator: Amboras (@amboras_inc) puts your entire ecommerce stack on autopilot. AI runs, optimizes, and A/B tests your store end …

X AI KOLs Following ↗ · 2026-05-22 Cached

Amboras is an AI-native ecommerce platform that automates store optimization, including A/B testing and layout changes, aiming to boost conversion rates significantly.

0 favorites 0 likes
#ab-testing

Better Experiments with LLM Evals — A funnel, not a fork (6 minute read)

TLDR AI ↗ · 2026-05-21 Cached

Spotify Engineering discusses using LLM evals as a funnel before A/B experiments, improving hit rates and creating a feedback loop between evals and experiments.

0 favorites 0 likes
#ab-testing

Built a runtime A/B testing layer for AI agents in production/dev - looking for 5-10 teams to break it

Reddit r/AI_Agents ↗ · 2026-05-13

The author introduces Syrin, a runtime A/B testing tool for AI agents that allows teams to run controlled experiments on live traffic across prompts, models, and agent topologies. They are seeking 5-10 engineering teams to test the tool in production and provide feedback.

0 favorites 0 likes
← Back to home

Submit Feedback