#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

This is a [FiftyOne](https://github.com/voxel51/fiftyone) dataset with 68 samples.

# Installation

If you haven’t already, install FiftyOne:

```bash
pip install -U fiftyone
```

# Usage

```python
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")

# Launch the App
session = fo.launch_app(dataset)
```

# Dataset Card for droid_3d (68-episode FiftyOne RGB subset)

![droid_3d preview](https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep/resolve/main/droid-3d-rgb.gif)

A 68-episode, **RGB-only** subset of **droid_3d**, a large-scale robot manipulation
dataset built on the DROID data collection platform, reformatted into LeRobotDataset
v3.0 with added depth video and point cloud modalities. The full repo (58,201 episodes,
~1.3 TB) is published at [ZibinDong/droid_3d](https://huggingface.co/datasets/ZibinDong/droid_3d);
this repo holds 68 of those episodes’ **RGB streams only** (depth and point clouds
excluded — see “Parsing decisions”), re-packaged as a self-contained LeRobotDataset
v3.0 export and loaded into FiftyOne for exploration.

## Dataset Details

### Dataset Description

- **Curated by:** Zibin Dong, Fei Ni, Yifu Yuan, Yinchuan Li, Jianye Hao (EmbodiedMAE paper authors; `droid_3d` reformatting)
- **Shared by:** ZibinDong (original `droid_3d` LeRobot v3 reformatting); this FiftyOne subset shared by the FiftyOne community
- **Language(s):** English (task/language annotations)
- **License:** MIT

### Dataset Sources

- **Repository:** [ZibinDong/droid_3d](https://huggingface.co/datasets/ZibinDong/droid_3d)
- **Paper:** [EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation (arXiv:2505.10105)](https://arxiv.org/abs/2505.10105)
- **Demo:** [ZibinDong/lerobotdataset3d](https://github.com/ZibinDong/lerobotdataset3d) (reader package for the full RGB+depth+pointcloud dataset) · [Visualize on the LeRobot dataset visualizer](https://huggingface.co/spaces/lerobot/visualize_dataset?path=ZibinDong/droid_3d)

## Uses

### Direct Use

Exploring and visualizing robot manipulation episodes’ RGB camera views (wrist +
2 external cameras) alongside 8-dim actions and per-episode natural-language task
descriptions in the FiftyOne App; prototyping data loaders and filters before working
with the full 58,201-episode `droid_3d` repo (which additionally includes depth and
point-cloud modalities not present here).

### Out-of-Scope Use

This 68-episode subset is not a statistically representative sample of the full dataset
(it is simply the first contiguous block of episodes whose data and all 3 RGB video
streams share the first storage shard) and should not be used to draw conclusions about
task or scenario distributions across the full `droid_3d` dataset. **It contains no
depth or point-cloud data** — it is not suitable for any 3D-aware or depth-conditioned
policy work; use the full source repo (with the `lerobotdataset3d` reader) for that.

## Dataset Structure

This is a **multimodal** FiftyOne dataset (`dataset.media_type == "multimodal"`) with
**68 samples**, one sample per episode. Each sample’s media (3 RGB video streams) is not
copied into per-sample files; instead it is resolved through a `media_reference` that
points into the exported LeRobotDataset v3.0 source (`data/`, `videos/`, `meta/` in this
repo) at import time — this is how FiftyOne represents LeRobot episodes natively.

### Fields

| Field                             | FiftyOne type                          | Description                                                                                                                                                                                                                        |
|-----------------------------------|----------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `id`                              | `ObjectIdField`                        | FiftyOne sample id                                                                                                                                                                                                                 |
| `media_reference`                 | `MediaReferenceField`                  | Pointer into the LeRobot source’s `data/`, `videos/*`, and `meta/` files for this episode (chunk/file indexes, frame range, per-video timestamp ranges) — resolved on demand, not duplicated per sample                            |
| `tags`                            | `ListField(StringField)`               | FiftyOne sample tags (empty by default)                                                                                                                                                                                            |
| `metadata`                        | `Metadata` (`size_bytes`, `mime_type`) | Standard FiftyOne sample metadata                                                                                                                                                                                                  |
| `created_at` / `last_modified_at` | `DateTimeField`                        | FiftyOne bookkeeping timestamps                                                                                                                                                                                                    |
| `episode_index`                   | `IntField`                             | Episode index within this subset (`0`–`67`, contiguous with the source dataset’s original indices since this block happened to start at 0)                                                                                         |
| `task`                            | `StringField`                          | Primary task caption for the episode, equal to source `language_1` for that episode (verbatim from source `meta/tasks.parquet`); may be an empty string for a small number of episodes where the source task label itself is blank |
| `tasks`                           | `ListField(StringField)`               | Full task list for the episode (length 1 for every episode here)                                                                                                                                                                   |
| `length`                          | `IntField`                             | Number of frames in the episode (verbatim from source)                                                                                                                                                                             |
| `duration`                        | `FloatField`                           | Episode duration in seconds (`length / fps`)                                                                                                                                                                                       |
| `robot_type`                      | `StringField`                          | Empty — source `meta/info.json` has `"robot_type": null` (DROID spans multiple robot platforms/embodiments and does not record a single type per this reformatting)                                                                |
| `fps`                             | `FloatField`                           | Recording frame rate, 15.0 (verbatim from source)                                                                                                                                                                                  |

The per-frame numeric/string features and the 3 per-frame RGB video streams are **not**
flattened into sample fields — they remain in the LeRobot `data/*.parquet` and
`videos/*/*.mp4` files referenced by `media_reference`, and are surfaced by the
FiftyOne App’s State & Action, Streams, and Statistics viewer tabs rather than as
queryable sample-level fields. Per source `meta/info.json`, the per-frame features
**imported into this subset** are:

| feature                                              | shape                      | description                                                                                 |
|------------------------------------------------------|----------------------------|---------------------------------------------------------------------------------------------|
| `action`                                             | `(8,)` float32             | robot action vector                                                                         |
| `language_1`                                         | string                     | first natural-language task annotation (this is the string surfaced as sample-level `task`) |
| `language_2`                                         | string                     | second, independently-worded natural-language task annotation for the same episode          |
| `language_3`                                         | string                     | third natural-language task annotation                                                      |
| `observation.images.{wrist, external_0, external_1}` | `(224, 398, 3)` video, AV1 | 3 RGB camera views: wrist-mounted + 2 external cameras                                      |

**Not imported (present in the source, excluded from this subset):**

| feature                                                  | dtype (source)                                            | why excluded                                                                                                                                                                                                                        |
|----------------------------------------------------------|-----------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `observation.depth.{wrist, external_0, external_1}`      | `depth_video` (H.265, `uint12_mm` scale, 2000 mm range)   | Not a standard LeRobot v3 `video`/`image` dtype; FiftyOne’s LeRobot v3 importer only recognizes `video` and `image` features. The source README itself notes standard `lerobot` also can’t read this dtype without a custom reader. |
| `observation.pointcloud.{wrist, external_0, external_1}` | `pointcloud` (Parquet, quantized XYZ, max 2048 pts/frame) | Same reason — not a `video`/`image` dtype, no built-in FiftyOne import path.                                                                                                                                                        |

Both excluded modalities can be read with the community reader package
[`lerobotdataset3d`](https://github.com/ZibinDong/lerobotdataset3d)
(`LeRobotDatasetDepthPointcloud`) against the full source repo, but bringing them into
FiftyOne would require custom per-frame decode + a manual 3D scene/field construction
step — out of scope for this subset.

### Label types and why

There are no traditional detection/classification/segmentation labels. `task` is stored
as a plain `StringField` rather than `fo.Classification` because it is a free-form
natural-language instruction (23,858 distinct across the full source dataset, 44 in this
subset), not a fixed closed-set taxonomy.

### `dataset.info` contents

```python
{
    "lerobot": {
        "format": "LeRobotDataset",
        "format_major": 3,
        "episode_count": 58201,         # total episodes in the full source dataset
        "imported_episode_count": 68,   # episodes actually imported into this subset
        "skipped_episodes": [],
    }
}
```

### Parsing decisions

- **RGB only — depth and point clouds deliberately excluded.** See the “Not imported”
  table above. This was a scoped decision, not a data-quality issue: the source dataset
  is specifically designed around RGB + depth + point cloud + action + language for 3D
  representation learning (EmbodiedMAE), and this subset only covers the RGB + action +
  language slice of that.
- **Which episodes, and why:** episodes `0`–`67` were selected because they are the
  largest contiguous, zero-gap block of episodes whose `data/chunk-000/file-000.parquet`
  shard **and** every one of the 3 RGB `videos/<key>/chunk-000/file-000.mp4` shards are
  shared — i.e. the smallest set of source files that had to be downloaded (~700 MB) to
  get a complete, non-truncated set of episodes, given a limited local disk budget
  (ignoring the excluded depth/pointcloud shard boundaries entirely). It is not a
  curated or stratified sample.
- **`task` uses `language_1`, not a merged/combined string.** Each episode in the source
  has up to 3 independent natural-language annotations (`language_1`/`_2`/`_3`) that are
  not paraphrases of each other so much as differently-focused descriptions (e.g. one
  episode’s three annotations: “Pick up the peach towel and fold it.”, “Fold the
  sweatshirt”, “Fold the orange cloth” — describing the same folding action with
  different object framing/emphasis). `meta/tasks.parquet`’s `task_index` (which drives
  this subset’s `task`/`tasks` fields) is keyed to `language_1` specifically;
  `language_2`/`language_3` are not surfaced as separate FiftyOne fields since they live
  in the frame-level parquet data, not per-episode metadata.
- **Re-export, not a thin reference to the original repo:** this repo is a
  self-contained LeRobotDataset v3.0 export (via FiftyOne’s `LeRobotDatasetExporter`),
  not a pointer back to `ZibinDong/droid_3d`. Task indices were remapped to only the
  tasks actually present in this subset (44 of the source’s 23,858). Per-episode and
  global statistics were recomputed from the exported rows, not carried over from the
  source’s global stats.
- **`robot_type` is empty**, mirroring the source’s `"robot_type": null` — DROID data
  was collected across multiple robot platforms without a single embodiment tag recorded
  at this level.

## Dataset Creation

### Curation Rationale

`droid_3d` was built to support training 3D-aware, vision-language-action robot
policies and unified multi-modal representations (per the EmbodiedMAE paper) by pairing
DROID’s large-scale manipulation demonstrations with depth and point-cloud
reconstructions alongside the existing RGB + action + language data. This subset exists
purely as a lightweight, disk-budget-friendly RGB slice for exploration and tooling in
FiftyOne; it was not re-curated for scenario content, and intentionally omits the 3D
modalities that motivate the parent dataset.

### Source Data

#### Data Collection and Processing

- **Platform:** collected with the DROID (Distributed Robot Interaction Dataset)
  data-collection platform — a standardized manipulation data-collection rig used across
  many research labs; this reformatting does not record a single `robot_type` per
  episode.
- **Cameras:** 3 fixed viewpoints per episode — a wrist-mounted camera and two external
  cameras — each recording 224×398 RGB video at 15 fps, AV1-encoded in this repo. In the
  full source repo, each RGB stream has a matched depth stream (H.265, `uint12_mm`
  scale, 2000 mm depth range) and a matched point-cloud stream (quantized XYZ, up to
  2048 points/frame, ranges x∈[−1,1] y∈[−1,1] z∈[0,1.6] m) — not included here.
  Downstream, EmbodiedMAE fuses RGB with these depth/point-cloud streams into a unified
  3D multi-modal representation for manipulation.
- **Actions:** 8-dimensional per-frame action vectors (`action`).
- **Language:** up to 3 independent natural-language task annotations per episode
  (`language_1`/`_2`/`_3`), differing in phrasing/emphasis rather than being verbatim
  duplicates.

#### Who are the source data producers?

Collected under the DROID data-collection effort (multi-institution robot manipulation
data collection); the `droid_3d` depth/point-cloud reformatting and this LeRobot v3
packaging were produced by Zibin Dong and collaborators for the EmbodiedMAE project.

### Annotations

#### Annotation process

Task/language annotations (`language_1`/`_2`/`_3`) accompany each episode from the
underlying DROID collection process; this reformatting does not describe an additional
annotation pass beyond carrying these forward and deriving the LeRobot `task_index` from
`language_1`.

#### Personal and Sensitive Information

Episodes are RGB video of tabletop robot manipulation (folding cloth, picking up
objects, etc.); no depth/point-cloud (excluded here) or audio is present in this subset.
No specific personal or sensitive information is called out by the source.

## Citation

**BibTeX:**

```bibtex
@article{dong2025embodiedmae,
  title   = {EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation},
  author  = {Dong, Zibin and Ni, Fei and Yuan, Yifu and Li, Yinchuan and Hao, Jianye},
  journal = {arXiv preprint arXiv:2505.10105},
  year    = {2025}
}
```

**APA:**

Dong, Z., Ni, F., Yuan, Y., Li, Y., & Hao, J. (2025). *EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation*. arXiv:2505.10105.

## More Information

This is a 68-episode, **RGB-only** subset of [ZibinDong/droid_3d](https://huggingface.co/datasets/ZibinDong/droid_3d)
(58,201 episodes, ~1.3 TB including depth and point clouds), produced for local
exploration under a limited disk budget. To work with the full RGB + depth + point-cloud
dataset, use [`lerobotdataset3d`](https://github.com/ZibinDong/lerobotdataset3d)
(`LeRobotDatasetDepthPointcloud`), which is a standalone reader with no LeRobot fork
required.

## Dataset Card Authors

[Harpreet Sahota](https://huggingface.co/harpreetsahota)

## Dataset Card Contact

[Harpreet Sahota](https://huggingface.co/harpreetsahota)
