DINOv3
Meta's self-supervised vision backbones trained on 1.7B images for dense visual features.
About
No labels went into DINOv3: Meta trained its third-generation vision foundation models with self-supervised learning on LVD-1689M, a curated set of 1.7 billion web images, plus a 493-million-image satellite corpus for geospatial variants. The result is a family of backbones from a 21M-parameter ViT-S up to a 6.7B ViT-7B, with ConvNeXt distillations from 29M to 198M for tighter deployment budgets, all producing high-resolution dense features that transfer to detection, segmentation, depth estimation, 3D correspondence, and video tracking without fine-tuning the backbone. A Gram-matrix anchoring technique keeps patch-level features clean during long training runs, the key fix over DINOv2, and frozen-backbone probes beat weakly supervised rivals on many dense prediction benchmarks. Checkpoints are distributed under the custom DINOv3 License, which allows commercial use with conditions rather than following a standard OSI license, through a request flow on Meta's site and Hugging Face. The repo provides PyTorch code (2.7.1 or newer), starter notebooks, and a text-alignment recipe, with a GPU expected for practical feature extraction.
Should you use DINOv3?
Pick it when
Pick DINOv3 when dense features drive the job, such as segmentation, depth, correspondence or tracking on a frozen backbone, or you need satellite-trained variants, and you can accept Meta's custom license and request flow.
Look elsewhere when
Skip it when you need a standard OSI license or ungated weights; DINOv2 is Apache 2.0 and simpler to obtain. It also expects PyTorch 2.7.1 or newer, and the 7B model is far beyond small GPU budgets.
Alternatives to DINOv3
- DINOv2
Apache 2.0, ungated and well integrated with ready heads, at the cost of less clean patch features and a 1.1B size ceiling.
- timm
Ships DINOv3 backbones behind create_model next to hundreds of other architectures, easier for fine-tuning comparisons, though the DINOv3 license still applies.
- SigLIP
Text-aligned image embeddings for zero-shot labels and retrieval under Apache 2.0, aimed at image-level matching rather than dense per-patch features.
- SAM 3
Gives segmentation masks straight from text or box prompts, so you skip training your own heads on frozen DINOv3 features, but under the SAM License.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- DINOv3 License
- Added
- Aug 24, 2026