Tag
Epoch AI created the Furniture Assembly Benchmark (FAB) to test AI models' visual reasoning in spotting mistakes in IKEA assembly photos. OpenAI's GPT-6 Astra leads with 80% accuracy, showing major improvements over previous models.
Introduces Flat-Pack Bench, a benchmark for evaluating fine-grained spatio-temporal reasoning in large vision-language models using furniture assembly tasks. Experiments show current LVLMs struggle with tracking and spatial interactions.