@matanSF: ProgramBench is a benchmark where agents must reproduce the observable behavior of real software, completely from scrat…
Summary
ProgramBench is a benchmark for agents to reproduce real software behavior from scratch. Factory is announced as the most advanced reverse-engineering system with breakthroughs in model-agnostic long-horizon agents.
View Cached Full Text
Cached at: 08/29/26, 07:59 PM
ProgramBench is a benchmark where agents must reproduce the observable behavior of real software, completely from scratch.
With our latest breakthrough in model-agnostic long-horizon agents, Factory is now the most advanced reverse-engineering system in the world.
Mark Attar (@mark_attar): Theo is the absolute man, he has incredible intuition for model behavior and how the harness influences it
Similar Articles
ProgramBench (5 minute read)
ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.
ProgramBench Vetted: Reverse Engineering from a Runnable Binary
ProgramBench Vetted is a benchmark with 50 tasks that evaluate AI agents on reconstructing programs from runnable binaries, designed to study long-context coordination and enable dense reward for reinforcement learning.
@KLieret: You can evaluate on ProgramBench yourself: https://github.com/facebookresearch/ProgramBench/… We will open the leaderbo…
ProgramBench is a new benchmark that tests AI agents' ability to reconstruct a complete codebase from a compiled binary and its documentation. The leaderboard will open for submissions soon.
Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
Introduces ToolBench-X, a benchmark for evaluating large language model agents under various tool-environment reliability hazards, revealing a substantial gap in performance compared to clean environments.
Benchmark Everything Everywhere All at Once
Introduces Benchmark Agent, a fully autonomous system for creating diverse benchmarks with minimal human intervention, enabling continuous model assessment across domains.