Tag
This tweet discusses replicating a multi-agent debate system from US stocks to Chinese A-shares, with modifications for local market features like T+1 trading and price limits, emphasizing it's for research purposes only.
Inherent Labs introduces Faraday, a 27B-parameter AI Scientist agent trained via long-horizon reinforcement learning to replicate scientific research, outperforming Claude Opus 4.8 and GPT-5.5 on paper replication tasks.
Mattshumer reports that first replications of an AI model are coming in, with a user noting impressive game quality generated from a single prompt.
OpenAI introduces PaperBench, a benchmark evaluating AI agents' ability to replicate state-of-the-art AI research by replicating 20 ICML 2024 papers with 8,316 gradable tasks. The best-performing model (Claude 3.5 Sonnet) achieves only 21% replication score, below human PhD-level performance, highlighting current limitations in autonomous research capabilities.