Tag
SenseTime releases SenseNova-Vision-7B-MoT, a fully open-sourced unified multimodal model that handles multiple vision tasks using natural language instructions, supporting detection, OCR, depth, segmentation, and more.
The video demonstrates the ability of multiple robots under the Gemini Robotics 2 system to collaborate on cleaning and organizing tasks through natural language instructions, showcasing task allocation and coordination between robots.