Tag
JarvisBench introduces a benchmark for evaluating the coordination between humans and AI agents, focusing on attention allocation in long-horizon tasks. It provides a reference implementation with a full-duplex speech interface.