Tag
The article discusses structuring AI video prompts and introduces a 'white model' that helps reduce randomness in spatial relationships, recommending its use for users with cinematographic sense to enhance control over generation.
GST-Bench is a new VQA benchmark for evaluating global spatial awareness in video understanding, testing whether VLMs can build coherent global scene representations from long-horizon egocentric video. Evaluation of 22 state-of-the-art VLMs shows a large gap versus humans, with the best model scoring 42.68 vs 79.08.