NeMo Curator

GPU accelerated toolkit for filtering, deduplicating, and transforming multimodal AI training data.

Open SourceSelf HostedOffline CapableGPU Required (16GB+ VRAM)
0.0 (0)

About

NVIDIA's answer to data curation bottlenecks moves the whole pipeline onto GPUs: NeMo Curator uses RAPIDS libraries such as cuDF, cuML, and cuGraph, orchestrated over Ray based executors, to filter, classify, deduplicate, and transform training data across text, image, video, and audio modalities. The team reports fuzzy deduplication of a RedPajama v2 subset dropping from 10.7 hours to 0.65 hours, roughly a 16x speedup, on three H100 nodes compared with CPU alternatives. Text tooling covers language identification, quality classifiers, and exact, fuzzy, and semantic dedup; images get aesthetic and NSFW filtering plus embedding generation; video supports scene detection and clip extraction; audio adds transcription and quality filtering. The same stack powers the data pipelines behind NVIDIA's Nemotron model family. The GPU path expects Linux x86_64, CUDA 12, and about 16 GB of GPU memory, with installation through uv packages or Docker containers. The code is Apache 2.0 licensed with about 1,700 GitHub stars, aimed squarely at teams curating pretraining scale corpora on GPU clusters.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
Apache-2.0
Minimum VRAM
16 GB
Added
Jul 29, 2026

Related Tools

Featured

Fine-tune LLMs 2x faster with 80% less memory

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Featured

Tool for fine-tuning LLMs with various configurations

Open SourceSelf HostedOfflineGPU 16GB+
Advanced
0.0 (0)

Open source data labeling platform for ML projects

Open SourceSelf HostedOffline
Easy
0.0 (0)
Featured

Library for accessing and sharing ML datasets

Open SourceSelf HostedOffline
Easy
0.0 (0)

Hugging Face library for parameter-efficient fine-tuning

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Python library for synthetic data generation and curation pipelines with batch LLM inference.

Open SourceSelf HostedOffline
Easy
0.0 (0)
Browse all Datasets & Training tools