Note

This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.

Hugging Face

This is a FiftyOne dataset with 68 samples.

Installation#

If you haven’t already, install FiftyOne:

pip install -U fiftyone

Usage#

import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")

# Launch the App
session = fo.launch_app(dataset)

Dataset Card for droid_3d (68-episode FiftyOne RGB subset)#

droid_3d preview

A 68-episode, RGB-only subset of droid_3d, a large-scale robot manipulation dataset built on the DROID data collection platform, reformatted into LeRobotDataset v3.0 with added depth video and point cloud modalities. The full repo (58,201 episodes, ~1.3 TB) is published at ZibinDong/droid_3d; this repo holds 68 of those episodes’ RGB streams only (depth and point clouds excluded — see “Parsing decisions”), re-packaged as a self-contained LeRobotDataset v3.0 export and loaded into FiftyOne for exploration.

Dataset Details#

Dataset Description#

  • Curated by: Zibin Dong, Fei Ni, Yifu Yuan, Yinchuan Li, Jianye Hao (EmbodiedMAE paper authors; droid_3d reformatting)

  • Shared by: ZibinDong (original droid_3d LeRobot v3 reformatting); this FiftyOne subset shared by the FiftyOne community

  • Language(s): English (task/language annotations)

  • License: MIT

Dataset Sources#

Uses#

Direct Use#

Exploring and visualizing robot manipulation episodes’ RGB camera views (wrist + 2 external cameras) alongside 8-dim actions and per-episode natural-language task descriptions in the FiftyOne App; prototyping data loaders and filters before working with the full 58,201-episode droid_3d repo (which additionally includes depth and point-cloud modalities not present here).

Out-of-Scope Use#

This 68-episode subset is not a statistically representative sample of the full dataset (it is simply the first contiguous block of episodes whose data and all 3 RGB video streams share the first storage shard) and should not be used to draw conclusions about task or scenario distributions across the full droid_3d dataset. It contains no depth or point-cloud data — it is not suitable for any 3D-aware or depth-conditioned policy work; use the full source repo (with the lerobotdataset3d reader) for that.

Dataset Structure#

This is a multimodal FiftyOne dataset (dataset.media_type == "multimodal") with 68 samples, one sample per episode. Each sample’s media (3 RGB video streams) is not copied into per-sample files; instead it is resolved through a media_reference that points into the exported LeRobotDataset v3.0 source (data/, videos/, meta/ in this repo) at import time — this is how FiftyOne represents LeRobot episodes natively.

Fields#

Field

FiftyOne type

Description

id

ObjectIdField

FiftyOne sample id

media_reference

MediaReferenceField

Pointer into the LeRobot source’s data/, videos/*, and meta/ files for this episode (chunk/file indexes, frame range, per-video timestamp ranges) — resolved on demand, not duplicated per sample

tags

ListField(StringField)

FiftyOne sample tags (empty by default)

metadata

Metadata (size_bytes, mime_type)

Standard FiftyOne sample metadata

created_at / last_modified_at

DateTimeField

FiftyOne bookkeeping timestamps

episode_index

IntField

Episode index within this subset (0–67, contiguous with the source dataset’s original indices since this block happened to start at 0)

task

StringField

Primary task caption for the episode, equal to source language_1 for that episode (verbatim from source meta/tasks.parquet); may be an empty string for a small number of episodes where the source task label itself is blank

tasks

ListField(StringField)

Full task list for the episode (length 1 for every episode here)

length

IntField

Number of frames in the episode (verbatim from source)

duration

FloatField

Episode duration in seconds (length / fps)

robot_type

StringField

Empty — source meta/info.json has "robot_type": null (DROID spans multiple robot platforms/embodiments and does not record a single type per this reformatting)

fps

FloatField

Recording frame rate, 15.0 (verbatim from source)

The per-frame numeric/string features and the 3 per-frame RGB video streams are not flattened into sample fields — they remain in the LeRobot data/*.parquet and videos/*/*.mp4 files referenced by media_reference, and are surfaced by the FiftyOne App’s State & Action, Streams, and Statistics viewer tabs rather than as queryable sample-level fields. Per source meta/info.json, the per-frame features imported into this subset are:

feature

shape

description

action

(8,) float32

robot action vector

language_1

string

first natural-language task annotation (this is the string surfaced as sample-level task)

language_2

string

second, independently-worded natural-language task annotation for the same episode

language_3

string

third natural-language task annotation

observation.images.{wrist, external_0, external_1}

(224, 398, 3) video, AV1

3 RGB camera views: wrist-mounted + 2 external cameras

Not imported (present in the source, excluded from this subset):

feature

dtype (source)

why excluded

observation.depth.{wrist, external_0, external_1}

depth_video (H.265, uint12_mm scale, 2000 mm range)

Not a standard LeRobot v3 video/image dtype; FiftyOne’s LeRobot v3 importer only recognizes video and image features. The source README itself notes standard lerobot also can’t read this dtype without a custom reader.

observation.pointcloud.{wrist, external_0, external_1}

pointcloud (Parquet, quantized XYZ, max 2048 pts/frame)

Same reason — not a video/image dtype, no built-in FiftyOne import path.

Both excluded modalities can be read with the community reader package lerobotdataset3d (LeRobotDatasetDepthPointcloud) against the full source repo, but bringing them into FiftyOne would require custom per-frame decode + a manual 3D scene/field construction step — out of scope for this subset.

Label types and why#

There are no traditional detection/classification/segmentation labels. task is stored as a plain StringField rather than fo.Classification because it is a free-form natural-language instruction (23,858 distinct across the full source dataset, 44 in this subset), not a fixed closed-set taxonomy.

dataset.info contents#

{
    "lerobot": {
        "format": "LeRobotDataset",
        "format_major": 3,
        "episode_count": 58201,         # total episodes in the full source dataset
        "imported_episode_count": 68,   # episodes actually imported into this subset
        "skipped_episodes": [],
    }
}

Parsing decisions#

  • RGB only — depth and point clouds deliberately excluded. See the “Not imported” table above. This was a scoped decision, not a data-quality issue: the source dataset is specifically designed around RGB + depth + point cloud + action + language for 3D representation learning (EmbodiedMAE), and this subset only covers the RGB + action + language slice of that.

  • Which episodes, and why: episodes 0–67 were selected because they are the largest contiguous, zero-gap block of episodes whose data/chunk-000/file-000.parquet shard and every one of the 3 RGB videos/<key>/chunk-000/file-000.mp4 shards are shared — i.e. the smallest set of source files that had to be downloaded (~700 MB) to get a complete, non-truncated set of episodes, given a limited local disk budget (ignoring the excluded depth/pointcloud shard boundaries entirely). It is not a curated or stratified sample.

  • task uses language_1, not a merged/combined string. Each episode in the source has up to 3 independent natural-language annotations (language_1/_2/_3) that are not paraphrases of each other so much as differently-focused descriptions (e.g. one episode’s three annotations: “Pick up the peach towel and fold it.”, “Fold the sweatshirt”, “Fold the orange cloth” — describing the same folding action with different object framing/emphasis). meta/tasks.parquet’s task_index (which drives this subset’s task/tasks fields) is keyed to language_1 specifically; language_2/language_3 are not surfaced as separate FiftyOne fields since they live in the frame-level parquet data, not per-episode metadata.

  • Re-export, not a thin reference to the original repo: this repo is a self-contained LeRobotDataset v3.0 export (via FiftyOne’s LeRobotDatasetExporter), not a pointer back to ZibinDong/droid_3d. Task indices were remapped to only the tasks actually present in this subset (44 of the source’s 23,858). Per-episode and global statistics were recomputed from the exported rows, not carried over from the source’s global stats.

  • robot_type is empty, mirroring the source’s "robot_type": null — DROID data was collected across multiple robot platforms without a single embodiment tag recorded at this level.

Dataset Creation#

Curation Rationale#

droid_3d was built to support training 3D-aware, vision-language-action robot policies and unified multi-modal representations (per the EmbodiedMAE paper) by pairing DROID’s large-scale manipulation demonstrations with depth and point-cloud reconstructions alongside the existing RGB + action + language data. This subset exists purely as a lightweight, disk-budget-friendly RGB slice for exploration and tooling in FiftyOne; it was not re-curated for scenario content, and intentionally omits the 3D modalities that motivate the parent dataset.

Source Data#

Data Collection and Processing#

  • Platform: collected with the DROID (Distributed Robot Interaction Dataset) data-collection platform — a standardized manipulation data-collection rig used across many research labs; this reformatting does not record a single robot_type per episode.

  • Cameras: 3 fixed viewpoints per episode — a wrist-mounted camera and two external cameras — each recording 224×398 RGB video at 15 fps, AV1-encoded in this repo. In the full source repo, each RGB stream has a matched depth stream (H.265, uint12_mm scale, 2000 mm depth range) and a matched point-cloud stream (quantized XYZ, up to 2048 points/frame, ranges x∈[−1,1] y∈[−1,1] z∈[0,1.6] m) — not included here. Downstream, EmbodiedMAE fuses RGB with these depth/point-cloud streams into a unified 3D multi-modal representation for manipulation.

  • Actions: 8-dimensional per-frame action vectors (action).

  • Language: up to 3 independent natural-language task annotations per episode (language_1/_2/_3), differing in phrasing/emphasis rather than being verbatim duplicates.

Who are the source data producers?#

Collected under the DROID data-collection effort (multi-institution robot manipulation data collection); the droid_3d depth/point-cloud reformatting and this LeRobot v3 packaging were produced by Zibin Dong and collaborators for the EmbodiedMAE project.

Annotations#

Annotation process#

Task/language annotations (language_1/_2/_3) accompany each episode from the underlying DROID collection process; this reformatting does not describe an additional annotation pass beyond carrying these forward and deriving the LeRobot task_index from language_1.

Personal and Sensitive Information#

Episodes are RGB video of tabletop robot manipulation (folding cloth, picking up objects, etc.); no depth/point-cloud (excluded here) or audio is present in this subset. No specific personal or sensitive information is called out by the source.

Citation#

BibTeX:

@article{dong2025embodiedmae,
  title   = {EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation},
  author  = {Dong, Zibin and Ni, Fei and Yuan, Yifu and Li, Yinchuan and Hao, Jianye},
  journal = {arXiv preprint arXiv:2505.10105},
  year    = {2025}
}

APA:

Dong, Z., Ni, F., Yuan, Y., Li, Y., & Hao, J. (2025). EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation. arXiv:2505.10105.

More Information#

This is a 68-episode, RGB-only subset of ZibinDong/droid_3d (58,201 episodes, ~1.3 TB including depth and point clouds), produced for local exploration under a limited disk budget. To work with the full RGB + depth + point-cloud dataset, use lerobotdataset3d (LeRobotDatasetDepthPointcloud), which is a standalone reader with no LeRobot fork required.

Dataset Card Authors#

Harpreet Sahota

Dataset Card Contact#

Harpreet Sahota