Florence-2

Small Microsoft vision-language model that switches between captioning, detection and OCR by prompt.

Open SourceSelf HostedOffline CapableGPU Required (6GB+ VRAM)
0.0 (0)

About

Instead of loading a separate network for every vision job, Florence-2 uses one sequence-to-sequence model and picks the task from a short token at the start of the prompt. Microsoft published a 0.23 billion parameter base and a 0.77 billion large version, each with a fine-tuned twin, trained on FLD-5B, roughly 5.4 billion machine-made labels spanning 126 million photos. Output arrives as text that a bundled helper converts into boxes, polygons, labels or prose, covering detection, region captions, phrase grounding, referring segmentation, OCR with optional region boxes, and captions at three levels of detail. The original checkpoints load in Transformers with trust_remote_code, and converted copies under florence-community use the native Florence2 classes now in the library. Being this small, they run on a laptop GPU or even a CPU, which is why ComfyUI nodes and dataset captioning scripts lean on them. Code and weights are MIT licensed, so commercial use is fine. Open-ended visual questions are better left to large multimodal LLMs; this model shines at structured, repeatable labeling.

Should you use Florence-2?

Pick it when

Pick Florence-2 when one small MIT model should cover captioning, grounding, detection, OCR and segmentation through text prompts, for example auto-labeling datasets or adding vision features to an app on a 6 GB GPU.

Look elsewhere when

Skip it if users need open-ended visual chat or reasoning, since it works through fixed task prompts; SmolVLM answers free-form questions. For a purpose-built open-set detector, Grounding DINO is the dedicated tool.

Alternatives to Florence-2

  • SmolVLM

    Generates free-form answers about images and video under Apache 2.0, even on CPU, but is not built to return boxes or masks.

  • Grounding DINO

    Dedicated open-set detector for free-text prompts with Apache 2.0 code and a 4 GB floor, but it only outputs boxes, no captions or OCR.

  • Grounded SAM 2

    Pairs a text-prompted detector, which can be Florence-2, with SAM 2 for precise masks and video tracking, but needs about 8 GB of VRAM.

  • PaddleOCR

    Specialized OCR stack for documents and multilingual text, the better pick when reading text is the main job rather than one task of many.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Minimum VRAM
6 GB
Added
Apr 3, 2026

Tags