image-captioning

Tag

Cards List
#image-captioning

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

SimLoss uses embedding-space contrastive supervision to enable single-pass fine-grained image captioning that matches multi-stage quality at much lower latency, with variants like SimLoss FFT achieving high precision.

0 favorites 0 likes
#image-captioning

@HuggingModels: Want to build an AI that can see images and describe them in natural language? This new model does just that. It's a vi…

X AI KOLs Following ↗ · 2026-07-14

A new vision-encoder-decoder model is introduced that can process both images and text to generate human-like responses, suitable for tasks like image captioning and visual question answering.

0 favorites 0 likes
#image-captioning

PRX Part 4: Our Data Strategy

Hugging Face Blog ↗ · 2026-07-06 Cached

Photoroom details their data strategy for training PRX, including assembling diverse datasets, re-captioning with a VLM, and using Mosaic Data Shards for efficient training.

0 favorites 0 likes
#image-captioning

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

arXiv cs.LG ↗ · 2026-05-21 Cached

Introduces ClaimDiff-RL, a reinforcement learning framework for long-form image captioning that uses typed, verifiable claim differences as reward units to separately measure and balance hallucination and missing facts, improving faithfulness and coverage.

0 favorites 0 likes
#image-captioning

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

Hugging Face Daily Papers ↗ · 2026-05-08 Cached

The paper introduces BalCapRL, a balanced reinforcement learning framework for multimodal large language models that jointly optimizes correctness, coverage, and linguistic quality in image captioning. It demonstrates improved performance over existing methods by addressing trade-offs between utility and fluency through reward decoupling and length-conditional masking.

0 favorites 0 likes
#image-captioning

lucataco/florence-2-large

Replicate Explore ↗ · 2026-09-09 Cached

Florence-2 is an advanced vision foundation model from Microsoft that uses a prompt-based approach to handle various vision and vision-language tasks like captioning and object detection, trained on a large annotated dataset.

0 favorites 0 likes
← Back to home

Submit Feedback