dots.ocr
Compact vision-language model that does document layout analysis and text recognition in one pass.
About
A single compact vision-language model replaces the usual OCR pipeline in dots.ocr, the document parser from Xiaohongshu's rednote-hilab. Instead of chaining a layout detector, a text recognizer, and a reading-order module, the model, built on a 1.7B LLM, treats all of it as one task: switching between layout extraction, plain text, or grounded output is just a prompt change. Results interleave bounding boxes with Markdown text, HTML tables, and LaTeX formulas across 11 layout categories, and the team's dots.ocr-bench stresses 1,493 PDF pages spanning about 100 languages, where the model leads open alternatives on low-resource scripts; it also posts top scores among compact models on OmniDocBench and olmOCR-bench. Weights are MIT licensed on Hugging Face, inference targets a CUDA GPU, and vLLM has carried official support since version 0.11, with Docker images and a plain transformers path for smaller setups. The same lab has since extended the line with dots.mocr, a 3B successor released with its own paper and weights in March 2026.
Should you use dots.ocr?
Pick it when
Pick dots.ocr when you need layout boxes, Markdown text, HTML tables, and LaTeX formulas from one compact model, MIT weights for commercial use, and coverage of low-resource scripts across roughly 100 languages.
Look elsewhere when
Skip it without a CUDA GPU; a CPU path like PaddleOCR or Docling is more practical. If you want a full ingestion pipeline with chunking and framework integrations rather than a raw model, Docling or Unstructured fit better.
Alternatives to dots.ocr
- PaddleOCR
Runs on CPU with mature detection, table, and layout modules, trading one-model simplicity for a multi-stage pipeline that is easier to put on edge devices.
- Dolphin
Similar layout-first parsing with 21 element types and attribute extraction, but its weights are research-only where dots.ocr's are MIT.
- DeepSeek-OCR
Also MIT and vLLM-ready, built around compressing pages into few visual tokens; pick it when GPU throughput per page matters more than layout output.
- Chandra
Stronger stated focus on handwriting and forms, but its weights are free only for organizations under 2 million dollars in revenue or funding.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- OCR & Document Processing
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MIT
- Added
- Aug 24, 2026