Note
This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.
This is a FiftyOne dataset with 68 samples.
Installation#
If you haven’t already, install FiftyOne:
pip install -U fiftyone
Usage#
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for droid_3d (68-episode FiftyOne RGB subset)#

A 68-episode, RGB-only subset of droid_3d, a large-scale robot manipulation dataset built on the DROID data collection platform, reformatted into LeRobotDataset v3.0 with added depth video and point cloud modalities. The full repo (58,201 episodes, ~1.3 TB) is published at ZibinDong/droid_3d; this repo holds 68 of those episodes’ RGB streams only (depth and point clouds excluded — see “Parsing decisions”), re-packaged as a self-contained LeRobotDataset v3.0 export and loaded into FiftyOne for exploration.
Dataset Details#
Dataset Description#
Curated by: Zibin Dong, Fei Ni, Yifu Yuan, Yinchuan Li, Jianye Hao (EmbodiedMAE paper authors;
droid_3dreformatting)Shared by: ZibinDong (original
droid_3dLeRobot v3 reformatting); this FiftyOne subset shared by the FiftyOne communityLanguage(s): English (task/language annotations)
License: MIT
Dataset Sources#
Repository: ZibinDong/droid_3d
Paper: EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation (arXiv:2505.10105)
Demo: ZibinDong/lerobotdataset3d (reader package for the full RGB+depth+pointcloud dataset) · Visualize on the LeRobot dataset visualizer
Uses#
Direct Use#
Exploring and visualizing robot manipulation episodes’ RGB camera views (wrist +
2 external cameras) alongside 8-dim actions and per-episode natural-language task
descriptions in the FiftyOne App; prototyping data loaders and filters before working
with the full 58,201-episode droid_3d repo (which additionally includes depth and
point-cloud modalities not present here).
Out-of-Scope Use#
This 68-episode subset is not a statistically representative sample of the full dataset
(it is simply the first contiguous block of episodes whose data and all 3 RGB video
streams share the first storage shard) and should not be used to draw conclusions about
task or scenario distributions across the full droid_3d dataset. It contains no
depth or point-cloud data — it is not suitable for any 3D-aware or depth-conditioned
policy work; use the full source repo (with the lerobotdataset3d reader) for that.
Dataset Structure#
This is a multimodal FiftyOne dataset (dataset.media_type == "multimodal") with
68 samples, one sample per episode. Each sample’s media (3 RGB video streams) is not
copied into per-sample files; instead it is resolved through a media_reference that
points into the exported LeRobotDataset v3.0 source (data/, videos/, meta/ in this
repo) at import time — this is how FiftyOne represents LeRobot episodes natively.
Fields#
Field |
FiftyOne type |
Description |
|---|---|---|
|
|
FiftyOne sample id |
|
|
Pointer into the LeRobot source’s |
|
|
FiftyOne sample tags (empty by default) |
|
|
Standard FiftyOne sample metadata |
|
|
FiftyOne bookkeeping timestamps |
|
|
Episode index within this subset ( |
|
|
Primary task caption for the episode, equal to source |
|
|
Full task list for the episode (length 1 for every episode here) |
|
|
Number of frames in the episode (verbatim from source) |
|
|
Episode duration in seconds ( |
|
|
Empty — source |
|
|
Recording frame rate, 15.0 (verbatim from source) |
The per-frame numeric/string features and the 3 per-frame RGB video streams are not
flattened into sample fields — they remain in the LeRobot data/*.parquet and
videos/*/*.mp4 files referenced by media_reference, and are surfaced by the
FiftyOne App’s State & Action, Streams, and Statistics viewer tabs rather than as
queryable sample-level fields. Per source meta/info.json, the per-frame features
imported into this subset are:
feature |
shape |
description |
|---|---|---|
|
|
robot action vector |
|
string |
first natural-language task annotation (this is the string surfaced as sample-level |
|
string |
second, independently-worded natural-language task annotation for the same episode |
|
string |
third natural-language task annotation |
|
|
3 RGB camera views: wrist-mounted + 2 external cameras |
Not imported (present in the source, excluded from this subset):
feature |
dtype (source) |
why excluded |
|---|---|---|
|
|
Not a standard LeRobot v3 |
|
|
Same reason — not a |
Both excluded modalities can be read with the community reader package
lerobotdataset3d
(LeRobotDatasetDepthPointcloud) against the full source repo, but bringing them into
FiftyOne would require custom per-frame decode + a manual 3D scene/field construction
step — out of scope for this subset.
Label types and why#
There are no traditional detection/classification/segmentation labels. task is stored
as a plain StringField rather than fo.Classification because it is a free-form
natural-language instruction (23,858 distinct across the full source dataset, 44 in this
subset), not a fixed closed-set taxonomy.
dataset.info contents#
{
"lerobot": {
"format": "LeRobotDataset",
"format_major": 3,
"episode_count": 58201, # total episodes in the full source dataset
"imported_episode_count": 68, # episodes actually imported into this subset
"skipped_episodes": [],
}
}
Parsing decisions#
RGB only — depth and point clouds deliberately excluded. See the “Not imported” table above. This was a scoped decision, not a data-quality issue: the source dataset is specifically designed around RGB + depth + point cloud + action + language for 3D representation learning (EmbodiedMAE), and this subset only covers the RGB + action + language slice of that.
Which episodes, and why: episodes
0–67were selected because they are the largest contiguous, zero-gap block of episodes whosedata/chunk-000/file-000.parquetshard and every one of the 3 RGBvideos/<key>/chunk-000/file-000.mp4shards are shared — i.e. the smallest set of source files that had to be downloaded (~700 MB) to get a complete, non-truncated set of episodes, given a limited local disk budget (ignoring the excluded depth/pointcloud shard boundaries entirely). It is not a curated or stratified sample.taskuseslanguage_1, not a merged/combined string. Each episode in the source has up to 3 independent natural-language annotations (language_1/_2/_3) that are not paraphrases of each other so much as differently-focused descriptions (e.g. one episode’s three annotations: “Pick up the peach towel and fold it.”, “Fold the sweatshirt”, “Fold the orange cloth” — describing the same folding action with different object framing/emphasis).meta/tasks.parquet’stask_index(which drives this subset’stask/tasksfields) is keyed tolanguage_1specifically;language_2/language_3are not surfaced as separate FiftyOne fields since they live in the frame-level parquet data, not per-episode metadata.Re-export, not a thin reference to the original repo: this repo is a self-contained LeRobotDataset v3.0 export (via FiftyOne’s
LeRobotDatasetExporter), not a pointer back toZibinDong/droid_3d. Task indices were remapped to only the tasks actually present in this subset (44 of the source’s 23,858). Per-episode and global statistics were recomputed from the exported rows, not carried over from the source’s global stats.robot_typeis empty, mirroring the source’s"robot_type": null— DROID data was collected across multiple robot platforms without a single embodiment tag recorded at this level.
Dataset Creation#
Curation Rationale#
droid_3d was built to support training 3D-aware, vision-language-action robot
policies and unified multi-modal representations (per the EmbodiedMAE paper) by pairing
DROID’s large-scale manipulation demonstrations with depth and point-cloud
reconstructions alongside the existing RGB + action + language data. This subset exists
purely as a lightweight, disk-budget-friendly RGB slice for exploration and tooling in
FiftyOne; it was not re-curated for scenario content, and intentionally omits the 3D
modalities that motivate the parent dataset.
Source Data#
Data Collection and Processing#
Platform: collected with the DROID (Distributed Robot Interaction Dataset) data-collection platform — a standardized manipulation data-collection rig used across many research labs; this reformatting does not record a single
robot_typeper episode.Cameras: 3 fixed viewpoints per episode — a wrist-mounted camera and two external cameras — each recording 224×398 RGB video at 15 fps, AV1-encoded in this repo. In the full source repo, each RGB stream has a matched depth stream (H.265,
uint12_mmscale, 2000 mm depth range) and a matched point-cloud stream (quantized XYZ, up to 2048 points/frame, ranges x∈[−1,1] y∈[−1,1] z∈[0,1.6] m) — not included here. Downstream, EmbodiedMAE fuses RGB with these depth/point-cloud streams into a unified 3D multi-modal representation for manipulation.Actions: 8-dimensional per-frame action vectors (
action).Language: up to 3 independent natural-language task annotations per episode (
language_1/_2/_3), differing in phrasing/emphasis rather than being verbatim duplicates.
Who are the source data producers?#
Collected under the DROID data-collection effort (multi-institution robot manipulation
data collection); the droid_3d depth/point-cloud reformatting and this LeRobot v3
packaging were produced by Zibin Dong and collaborators for the EmbodiedMAE project.
Annotations#
Annotation process#
Task/language annotations (language_1/_2/_3) accompany each episode from the
underlying DROID collection process; this reformatting does not describe an additional
annotation pass beyond carrying these forward and deriving the LeRobot task_index from
language_1.
Personal and Sensitive Information#
Episodes are RGB video of tabletop robot manipulation (folding cloth, picking up objects, etc.); no depth/point-cloud (excluded here) or audio is present in this subset. No specific personal or sensitive information is called out by the source.
Citation#
BibTeX:
@article{dong2025embodiedmae,
title = {EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation},
author = {Dong, Zibin and Ni, Fei and Yuan, Yifu and Li, Yinchuan and Hao, Jianye},
journal = {arXiv preprint arXiv:2505.10105},
year = {2025}
}
APA:
Dong, Z., Ni, F., Yuan, Y., Li, Y., & Hao, J. (2025). EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation. arXiv:2505.10105.
More Information#
This is a 68-episode, RGB-only subset of ZibinDong/droid_3d
(58,201 episodes, ~1.3 TB including depth and point clouds), produced for local
exploration under a limited disk budget. To work with the full RGB + depth + point-cloud
dataset, use lerobotdataset3d
(LeRobotDatasetDepthPointcloud), which is a standalone reader with no LeRobot fork
required.