lucataco/florence-2-large

Replicate Explore 模型

摘要

Florence-2 是微软推出的一款先进的视觉基础模型,它采用基于提示的方法来处理各种视觉和视觉语言任务,如图像描述和目标检测,训练于大规模标注数据集。

lucataco / florence-2-large
查看原文
查看缓存全文

缓存时间: 2026/09/09 04:00

# lucataco/florence-2-large – Replicate 源码:https://replicate.com/lucataco/florence-2-large ## Florence\-2:推进多种视觉任务的统一表示 ## 模型摘要 本Hub仓库包含微软Florence\-2模型的HuggingFace `transformers`实现。Florence\-2是一种先进的视觉基础模型,采用基于提示的方法来处理广泛的视觉和视觉\-语言任务。Florence\-2可以解释简单的文本提示以执行图像描述、目标检测和分割等任务。它利用我们的FLD\-5B数据集(包含1.26亿图像上的54亿标注)来掌握多任务学习。该模型的序列到序列架构使其在零样本和微调设置中都表现出色,证明了其作为具有竞争力的视觉基础模型的能力。资源与技术文档: +Florence\-2技术报告(https://arxiv.org/abs/2311.06242)。 +用于Florence\-2\-large推理和可视化的Jupyter Notebook(https://huggingface.co/microsoft/Florence-2-large/blob/main/sample_inference.ipynb) 模型 模型大小 模型描述 Florence\-2\-base [\[HF\]](https://huggingface.co/microsoft/Florence-2-base) 0\.23B 使用FLD\-5B的预训练模型 Florence\-2\-large [\[HF\]](https://huggingface.co/microsoft/Florence-2-large) 0\.77B 使用FLD\-5B的预训练模型 Florence\-2\-base\-ft [\[HF\]](https://huggingface.co/microsoft/Florence-2-base-ft) 0\.23B 在下游任务集合上微调的模型 Florence\-2\-large\-ft [\[HF\]](https://huggingface.co/microsoft/Florence-2-large-ft) 0\.77B 在下游任务集合上微调的模型 ## 任务 此模型能够通过更改提示来执行不同的任务。 ### 图像描述 `` prompt = "" run_example(prompt) `` ### 详细描述 `` prompt = "" run_example(prompt) `` ### 更详细描述 `` prompt = "" run_example(prompt) `` ### 描述到短语定位 描述到短语定位任务需要额外的文本输入,即描述。描述到短语定位结果格式: \{'<\|phrase\|>': \{'bboxes': \[\[x1, y1, x2, y2\], ...\], 'labels': \['<\|text\|>', '<\|text\|>', ...\]\}\} `` task_prompt = "<\|phrase\|>" results = run_example(task_prompt, text_input="A green car parked in front of a yellow building.") `` ### 目标检测 目标检测结果格式: \{'<\|loc\|>': \{'bboxes': \[\[x1, y1, x2, y2\], ...\], 'labels': \['<\|label\|>', '<\|label\|>', ...\]\} \} `` prompt = "<\|loc\|>" run_example(prompt) `` ### 密集区域描述 密集区域描述结果格式: \{'<\|loc\|>' : \{'bboxes': \[\[x1, y1, x2, y2\], ...\], 'labels': \['<\|label\|>', '<\|label\|>', ...\]\} \} `` prompt = "<\|loc\|>" run_example(prompt) `` ### 区域提议 密集区域描述结果格式: \{'<\|proposal\|>': \{'bboxes': \[\[x1, y1, x2, y2\], ...\], 'labels': \['<\|text\|>', '<\|text\|>', ...\]\}\} `` prompt = "<\|proposal\|>" run_example(prompt) `` ### OCR `` prompt = "<\|ocr\|>" run_example(prompt) `` ### 带区域的OCR 带区域的OCR输出格式: \{'<\|ocr\|>': \{'quad\_boxes': \[\[x1, y1, x2, y2, x3, y3, x4, y4\], ...\], 'labels': \['<\|text\|>', ...\]\}\} `` prompt = "<\|ocr\|>" run_example(prompt) `` 更多详细示例,请参考notebook(https://huggingface.co/microsoft/Florence-2-large/blob/main/sample_inference.ipynb)。 ## 基准测试 ## Florence\-2 零样本性能 下表展示了通用视觉基础模型在图像描述和目标检测评估任务上的零样本性能。这些模型在其训练阶段未接触过评估任务的训练数据。 方法 #参数 COCO Cap\. test CIDEr NoCaps val CIDEr TextCaps val CIDEr COCO Det\. val2017 mAP Flamingo 80B 84\.3 \- \- \- Florence\-2\-base 0\.23B 133\.0 118\.7 70\.1 34\.7 Florence\-2\-large 0\.77B 135\.6 120\.8 72\.8 37\.5 下表继续比较了其他视觉\-语言评估任务上的性能。 方法 Flickr30k test R@1 Refcoco val Accuracy Refcoco test\-A Accuracy Refcoco test\-B Accuracy Refcoco\+ val Accuracy Refcoco\+ test\-A Accuracy Refcoco\+ test\-B Accuracy Refcocog val Accuracy Refcocog test Accuracy Refcoco RES val mIoU Kosmos\-2 78\.7 52\.3 57\.4 47\.3 45\.5 50\.7 42\.2 60\.6 61\.7 \- Florence\-2\-base 83\.6 53\.9 58\.4 49\.7 51\.5 56\.4 47\.9 66\.3 65\.1 34\.6 Florence\-2\-large 84\.4 56\.3 61\.6 51\.4 53\.6 57\.9 49\.9 68\.0 67\.0 35\.8 ## Florence\-2 微调性能 我们使用一组下游任务对Florence\-2模型进行微调,得到两个通用模型*Florence\-2\-base\-ft*和*Florence\-2\-large\-ft*,它们可以执行广泛的下游任务。下表比较了专用模型和通用模型在各种描述和视觉问答(VQA)任务上的性能。专用模型专门针对每个任务进行微调,而通用模型则以与任务无关的方式在所有任务上进行微调。符号“▲”表示使用外部OCR作为输入。 方法 #参数 COCO Caption Karpathy test CIDEr NoCaps val CIDEr TextCaps val CIDEr VQAv2 test\-dev Acc TextVQA test\-dev Acc VizWiz VQA test\-dev Acc **专用模型** CoCa 2\.1B 143\.6 122\.4 \- 82\.3 \- \- BLIP\-2 7\.8B 144\.5 121\.6 \- 82\.2 \- \- GIT 25\.1B 145\.0 126\.9 148\.6 81\.7 67\.3 71\.0 Flamingo 80B 138\.1 \- \- 82\.0 54\.1 65\.7 PaLI 17B 149\.1 127\.0 160\.0 ▲84\.3 58\.8 / 73\.1 ▲71\.6 / 74\.4 ▲ PaLI\-X 55B 149\.2 126\.3 147\.0 / 163\.7 ▲86\.0 71\.4 / 80\.8 ▲70\.9 / 74\.6 ▲ **通用模型** Unified\-IO 2\.9B \- 100\.0 \- 77\.9 \- 57\.4 Florence\-2\-base\-ft 0\.23B 140\.0 116\.7 143\.9 79\.7 63\.6 63\.6 Florence\-2\-large\-ft 0\.77B 143\.3 124\.9 151\.1 81\.7 73\.5 72\.6 方法 #参数 COCO Det\. val2017 mAP Flickr30k test R@1 RefCOCO val Accuracy RefCOCO test\-A Accuracy RefCOCO test\-B Accuracy RefCOCO\+ val Accuracy RefCOCO\+ test\-A Accuracy RefCOCO\+ test\-B Accuracy RefCOCOg val Accuracy RefCOCOg test Accuracy RefCOCO RES val mIoU **专用模型** SeqTR \- \- \- 83\.7 86\.5 81\.2 71\.5 76\.3 64\.9 74\.9 74\.2 \- PolyFormer \- \- \- 90\.4 92\.9 87\.2 85\.0 89\.8 78\.0 85\.8 85\.9 76\.9 UNINEXT 0\.74B 60\.6 \- 92\.6 94\.3 91\.5 85\.2 89\.6 79\.8 88\.7 89\.4 \- Ferret 13B \- \- 89\.5 92\.4 84\.4 82\.8 88\.1 75\.2 85\.8 86\.3 \- **通用模型** UniTAB \- \- \- 88\.6 91\.1 83\.8 81\.0 85\.4 71\.6 84\.6 84\.7 \- Florence\-2\-base\-ft 0\.23B 41\.4 84\.0 92\.6 94\.8 91\.5 86\.8 91\.7 82\.2 89\.8 82\.2 78\.0 Florence\-2\-large\-ft 0\.77B 43\.4 85\.2 93\.4 95\.3 92\.0 88\.3 92\.9 83\.6 91\.2 91\.7 80\.5 ## BibTex和引用信息 `` @article{xiao2023florence, title={Florence-2: Advancing a unified representation for a variety of vision tasks}, author={Xiao, Bin and Wu, Haiping and Xu, Weijian and Dai, Xiyang and Hu, Houdong and Lu, Yumao and Zeng, Michael and Liu, Ce and Yuan, Lu}, journal={arXiv preprint arXiv:2311.06242}, year={2023} } `` 模型创建于1年多前

相似文章

lucataco/moondream2

Replicate Explore

moondream2是一个紧凑的视觉语言模型,专为高效的边缘设备推理而设计,提供了基准测试结果和使用说明。

microsoft/Lens-Turbo

Hugging Face Models Trending

微软发布了Lens,一个拥有38亿参数的基础文本到图像模型,具备高效的训练和快速的高分辨率生成能力,采用密集字幕预训练和混合分辨率学习。

LiquidAI/LFM2.5-VL-3B · Hugging Face

Reddit r/LocalLLaMA

LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.

Lens:重新思考基础文本到图像模型的训练效率

Hugging Face Daily Papers

Lens是微软推出的一款紧凑型38亿参数文本到图像模型,在训练计算量显著降低的同时,通过密集描述、多分辨率批处理和高效架构,达到了与更大模型竞争甚至超越的性能。