Tag
This academic paper identifies and characterizes Simpson's paradox in behavioral curve modeling, demonstrating how aggregation systematically distorts parametric estimates of user dynamics due to survival bias. The authors validate this distortion across datasets like Goodreads and Amazon Electronics and propose hierarchical peak estimation methods to mitigate the issue.
This paper identifies and addresses aggregation bias in GRPO-style reinforcement learning for LLMs, proposing Balanced Aggregation (BA) which improves training stability and final performance by computing token-level means separately for positive and negative subsets.