Tag
CRAFT converts rubric-based evaluation into hierarchical capability diagnosis for LLMs, identifying specific weaknesses and generating targeted fine-tuning data, achieving stronger results on finance and legal benchmarks across four open-source models.
GergelyOrosz comments on the rapid progress in agentic capabilities, especially cloud coding agents. Alexandr Wang responds that Scale AI's next Muse Spark update will bring major improvements in coding and agentic abilities.
After Meta acquired Manus, Scale AI founder Alexandr Wang unfollowed ManusAI, sparking discussion.
This article explores the challenge of applying reinforcement learning to tasks that lack clear verifiability, citing Dario Amodei's prediction about achieving a 'country of geniuses in a data center' and discussing techniques such as RLVR, RLHF, Constitutional AI, and rubric-based rewards from Scale AI.
Gergely Orosz highlights that software engineers at Scale AI and Meta are being assigned manual data labeling tasks, a practice that new leadership at Scale AI stopped after finding it concerning.
A detailed report on Meta's efforts to catch up in AI, including the hiring of Alexandr Wang and the release of the Muse Spark model, with mixed opinions on progress.
Scale AI CEO Alexandr Wang shares how Paul Graham's 'Schlep Blindness' essay inspired the company's focus on solving the unglamorous but critical problem of building high-quality data sets for machine learning.