NVIDIA shares streaming robotics pipeline for Cosmos3-DROID dataset
NVIDIA’s tutorial demonstrates how to build a streaming robotics learning pipeline for the Cosmos3-DROID dataset—707 GB in size—without downloading it locally. The approach uses metadata graphs, HTTP byte-range access with PyArrow, and selective Parquet row-group reading to extract state-action trajectories. It also decodes only required AV1 video windows via seek-based PyAV/FFmpeg access,…
Key points
- Pipeline streams 707 GB Cosmos3-DROID dataset without local download using metadata graphs and PyArrow byte-range access
- Decodes only required AV1 video windows via seek-based PyAV/FFmpeg, reducing storage needs to ~hundreds of MB
- Trains a multimodal behavior-cloning policy in PyTorch with optional visual conditioning and temporal ensembling
The workflow avoids full dataset materialization by leveraging Hugging Face’s filesystem and dataset statistics. It supports optional visual conditioning, chunked training, and temporal ensembling for smoother action predictions. Evaluation metrics include per-joint MSE and R² against a mean-action baseline. The tutorial includes code snippets for setup, data loading, model training, and rollout evaluation, emphasizing scalability across shards, failure demonstrations, and alternative action representations.
Building a Streaming Robotics Learning Pipeline Using NVIDIA Cosmos3-DROID
MarkTechPost · 5 October 2026
Loading the full article…
This text was published by MarkTechPost and written by Sana Hassan. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Google launches prototype satellite for Project Suncatcher to test AI chips in space · 2 src
- AI agents identify two room-temperature magnetic semiconductor candidates · 1 src
- OpenAI claims to solve Navier–Stokes problem, sparking math community debate · 2 src
- Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model · 1 src
- Microsoft Research unveils Quine, a biological world model for life sciences · 1 src
Comments
via GitHub Discussions