ShutterMuse: Capture-Time Photography Guidance with MLLMs

Hugging Face Daily Papers Papers

Summary

Researchers introduce CaptureGuide-Bench, a benchmark for capture-time photography guidance, and ShutterMuse, a unified multimodal LLM trained to provide composition and pose recommendations, demonstrating improved performance over general-purpose models.

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction and overlook subject-side recommendations, leaving the capture-time guidance capabilities of multimodal large language models (MLLMs) underexplored. To address this gap, we introduce CaptureGuide-Bench, a benchmark with two complementary tasks: photographer-side composition decision and refinement, and subject-side scene-conditioned pose recommendation. Our evaluation reveals limitations: general-purpose MLLMs can make composition decisions but lack precise refinement localization, while specialized aesthetic cropping models localize crops effectively but are limited to refinement; neither provides actionable pose guidance. To support model development, we further construct CaptureGuide-Dataset, comprising 130K samples with textual rationales and structured visual annotations, and develop ShutterMuse, a unified MLLM trained with supervised and reinforcement fine-tuning. Experiments on CaptureGuide-Bench show that ShutterMuse achieves the best overall photographer-side performance among evaluated baselines and competitive subject-side pose recommendation with substantially lower inference cost, demonstrating the potential of MLLMs as interactive assistants for photography during image capture.
Original Article
View Cached Full Text

Cached at: 06/25/26, 09:10 AM

Paper page - ShutterMuse: Capture-Time Photography Guidance with MLLMs

Source: https://huggingface.co/papers/2606.25763

Abstract

Researchers developed a new benchmark and dataset for photography assistance, along with a unified multimodal model that provides both composition guidance and pose recommendations during image capture.

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existingaesthetic croppingbenchmarks mainly evaluate post-hoc crop prediction and overlook subject-side recommendations, leaving the capture-time guidance capabilities ofmultimodal large language models(MLLMs) underexplored. To address this gap, we introduce CaptureGuide-Bench, a benchmark with two complementary tasks:photographer-side compositiondecision and refinement, and subject-side scene-conditioned pose recommendation. Our evaluation reveals limitations: general-purpose MLLMs can make composition decisions but lack precise refinement localization, while specializedaesthetic croppingmodels localize crops effectively but are limited to refinement; neither provides actionable pose guidance. To support model development, we further construct CaptureGuide-Dataset, comprising 130K samples with textual rationales and structuredvisual annotations, and develop ShutterMuse, a unified MLLM trained with supervised andreinforcement fine-tuning. Experiments on CaptureGuide-Bench show that ShutterMuse achieves the best overall photographer-side performance among evaluated baselines and competitivesubject-side pose recommendationwith substantially lower inference cost, demonstrating the potential of MLLMs as interactive assistants for photography during image capture.

View arXiv pageView PDFProject pageGitHub11Add to collection

Get this paper in your agent:

hf papers read 2606\.25763

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### ShutterMuse/ShutterMuse Image-Text-to-Text• 9B• Updatedabout 3 hours ago

Datasets citing this paper1

#### ShutterMuse/CaptureGuide-Bench Updatedabout 3 hours ago • 5

Spaces citing this paper1

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles