Tag
This paper empirically characterizes uncertainty in thinking-mode visual language models, demonstrating that the thinking chain entropy is a more reliable signal for hallucination detection than conventional answer token distribution, which collapses in these models.
The Gemma 4 Technical Report introduces a new generation of open-weight, natively multimodal language models with diverse architectures, enhanced reasoning capabilities, and improved performance across tasks. The models range from 2.3B to 31B parameters and feature a thinking mode for generating reasoning traces.
Xiaokang Chen shares two prompts, 'Think with Grounding' and 'Think with Pointing', to improve model performance in domains like counting in Thinking mode. These prompts use bounding boxes and points to make the MLLM's reasoning more human-like.
ChatGPT Images 2.0 in "Thinking" mode can turn 1,000-word prompts or 70-page PDFs into ready-to-use infographics, slide decks, and academic posters without manual editing.