#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/rh20t-cfg1-28ep" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

This is a [FiftyOne](https://github.com/voxel51/fiftyone) dataset with 28 samples.

# Installation

If you haven’t already, install FiftyOne:

```bash
pip install -U fiftyone
```

# Usage

```python
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/rh20t-cfg1-28ep")

# Launch the App
session = fo.launch_app(dataset)
```

# Dataset Card for RH20T cfg1 (28-episode FiftyOne subset)

![rh20t_cfg1 preview](https://huggingface.co/datasets/Voxel51/rh20t-cfg1-28ep/resolve/main/rh20t-cfg.gif)

A 28-episode subset of **rh20t_cfg1**, an unofficial community LeRobotDataset v3
reformatting (config 1, `flexiv` robot) of **RH20T**, a large-scale multi-modal robotic
manipulation dataset. The full `rh20t_cfg1` repo (4,258 episodes) is published at
[robot-lev/rh20t_cfg1](https://huggingface.co/datasets/robot-lev/rh20t_cfg1); this repo
holds episodes `0`–`27` re-packaged as a self-contained LeRobotDataset v3.0 export and
loaded into FiftyOne for exploration.

## Dataset Details

### Dataset Description

- **Curated by:** Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, Cewu Lu (original RH20T dataset); LeRobot v3 reformatting by the community (`lvjonok/rh20t-lerobot-port`)
- **Shared by:** robot-lev (LeRobot v3 reformatting); this FiftyOne subset shared by the FiftyOne community
- **Language(s):** English (task metadata)
- **License:** Dual-licensed **per scene**, mirrored from the original RH20T license split (see “Parsing decisions” below): scenes 1–5 → CC BY-SA 4.0; scenes 6–10 → CC BY-NC 4.0. **This 28-episode subset mixes both**: 16 episodes are CC BY-SA 4.0, 12 are CC BY-NC 4.0 (non-commercial only) — see the per-episode table below before any commercial use.

### Dataset Sources

- **Repository:** [robot-lev/rh20t_cfg1](https://huggingface.co/datasets/robot-lev/rh20t_cfg1) · reformatting code: `lvjonok/rh20t-lerobot-port`
- **Paper:** [RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot (arXiv:2307.00595)](https://arxiv.org/abs/2307.00595)
- **Demo:** [RH20T project page](https://rh20t.github.io/) · [RH20T API](https://github.com/rh20t/rh20t_api)

## Uses

### Direct Use

Exploring and visualizing multi-camera robotic manipulation episodes in the FiftyOne
App — inspecting 10 synchronized RGB camera views alongside end-effector/joint state,
force/torque and 6-axis fingertip wrench readings, and 8-dim actions (next end-effector
pose + gripper); prototyping data loaders and filters before working with the full
4,258-episode `rh20t_cfg1` repo or the broader RH20T dataset (110,000+ sequences across
many robot configs).

### Out-of-Scope Use

This 28-episode subset is not a statistically representative sample of the full dataset
(it is simply the first contiguous block of episodes whose data and all 10 video streams
share the first storage shard) and should not be used to draw conclusions about task,
scene, or robot-configuration distributions across the full RH20T dataset. **Do not use
the CC BY-NC-licensed episodes (or models trained on them) for commercial purposes** —
see the per-episode license table below. This subset also excludes audio and depth
entirely (see Parsing decisions); it is not suitable for tasks requiring those
modalities.

## Dataset Structure

This is a **multimodal** FiftyOne dataset (`dataset.media_type == "multimodal"`) with
**28 samples**, one sample per episode. Each sample’s media (10 video streams) is not
copied into per-sample files; instead it is resolved through a `media_reference` that
points into the exported LeRobotDataset v3.0 source (`data/`, `videos/`, `meta/` in this
repo) at import time — this is how FiftyOne represents LeRobot episodes natively.

### Fields

| Field                             | FiftyOne type                          | Description                                                                                                                                                                                             |
|-----------------------------------|----------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `id`                              | `ObjectIdField`                        | FiftyOne sample id                                                                                                                                                                                      |
| `media_reference`                 | `MediaReferenceField`                  | Pointer into the LeRobot source’s `data/`, `videos/*`, and `meta/` files for this episode (chunk/file indexes, frame range, per-video timestamp ranges) — resolved on demand, not duplicated per sample |
| `tags`                            | `ListField(StringField)`               | FiftyOne sample tags (empty by default)                                                                                                                                                                 |
| `metadata`                        | `Metadata` (`size_bytes`, `mime_type`) | Standard FiftyOne sample metadata                                                                                                                                                                       |
| `created_at` / `last_modified_at` | `DateTimeField`                        | FiftyOne bookkeeping timestamps                                                                                                                                                                         |
| `episode_index`                   | `IntField`                             | Episode index within this subset (`0`–`27`, remapped on export; contiguous with the source dataset’s original indices since this block happened to start at 0)                                          |
| `task`                            | `StringField`                          | Task label for the episode, one of `"task 1"`–`"task 6"` (verbatim placeholder task strings from source `meta/tasks.parquet` — see note below)                                                          |
| `tasks`                           | `ListField(StringField)`               | Full task list for the episode (length 1 for every episode here)                                                                                                                                        |
| `length`                          | `IntField`                             | Number of frames in the episode (verbatim from source)                                                                                                                                                  |
| `duration`                        | `FloatField`                           | Episode duration in seconds (`length / fps`)                                                                                                                                                            |
| `robot_type`                      | `StringField`                          | `flexiv` (verbatim from source `meta/info.json`)                                                                                                                                                        |
| `fps`                             | `FloatField`                           | Recording frame rate, 10.0 (verbatim from source)                                                                                                                                                       |

The per-frame numeric features and the 10 per-frame video streams are **not** flattened
into sample fields — they remain in the LeRobot `data/*.parquet` and `videos/*/*.mp4`
files referenced by `media_reference`, and are surfaced by the FiftyOne App’s State &
Action, Streams, and Statistics viewer tabs rather than as queryable sample-level fields.
Per source `meta/info.json`, these per-frame features are:

| feature                                | shape                      | description                                                                                                        |
|----------------------------------------|----------------------------|--------------------------------------------------------------------------------------------------------------------|
| `observation.state`                    | `(15,)`                    | concatenation of `ee_pose` (7) + `joint` (7) + `gripper` (1)                                                       |
| `observation.state.ee_pose`            | `(7,)`                     | end-effector pose (position + quaternion)                                                                          |
| `observation.state.joint`              | `(7,)`                     | joint positions (all-zero for episodes where the source `has_joint` flag is `False` — see below)                   |
| `observation.state.gripper`            | `(1,)`                     | gripper opening, mm (0–95 for the `dahuan_ag95` gripper)                                                           |
| `observation.force`                    | `(3,)`                     | end-effector force (Fx, Fy, Fz)                                                                                    |
| `observation.torque`                   | `(3,)`                     | end-effector torque (Mx, My, Mz)                                                                                   |
| `observation.robot_ft`                 | `(6,)`                     | robot-mounted force/torque sensor reading                                                                          |
| `action`                               | `(8,)`                     | next end-effector pose (7) + gripper target (1)                                                                    |
| `meta.rating`                          | `(1,)` int64               | per-episode demonstration quality rating (source RH20T annotation)                                                 |
| `observation.images.cam_<serial>` × 10 | `(360, 640, 3)` video, AV1 | one video stream per camera serial number; `has_audio: false` at the container level for every stream in this repo |

**Task labels are placeholders, not descriptions.** `meta/tasks.parquet` in the source
maps `task_index` to plain strings `"task 1"`, `"task 2"`, … rather than natural-language
instructions — the LeRobot v3 port did not carry over RH20T’s original task
descriptions/language annotations, only the numeric `task_id` (visible in
`meta/rh20t_episodes.json`, not surfaced as a FiftyOne field). Treat `task`/`tasks` here
as an opaque task-cluster id, not a human-readable instruction.

### Label types and why

There are no traditional detection/classification/segmentation labels. `task` is a plain
`StringField` (not `fo.Classification`) because, as noted above, it is a placeholder
identifier rather than a meaningful closed-set category with semantic content worth
modeling as a classification label in this port.

### `dataset.info` contents

```python
{
    "lerobot": {
        "format": "LeRobotDataset",
        "format_major": 3,
        "episode_count": 4258,          # total episodes in the full source dataset
        "imported_episode_count": 28,   # episodes actually imported into this subset
        "skipped_episodes": [],
    }
}
```

### Per-episode license and scene mapping

Source scene/task metadata for each imported episode (`new_episode_index` → source
`episode_index`, `meta/rh20t_episodes.json`); episode indices are identical here since
this contiguous block starts at 0:

|   episode_index |   scene_id |   task_id | license      |
|-----------------|------------|-----------|--------------|
|               0 |          1 |         1 | CC BY-SA 4.0 |
|               1 |          7 |         1 | CC BY-NC 4.0 |
|               2 |          3 |         1 | CC BY-SA 4.0 |
|               3 |          9 |         1 | CC BY-NC 4.0 |
|               4 |          5 |         1 | CC BY-SA 4.0 |
|               5 |          1 |         2 | CC BY-SA 4.0 |
|               6 |          7 |         2 | CC BY-NC 4.0 |
|               7 |          3 |         2 | CC BY-SA 4.0 |
|               8 |          9 |         2 | CC BY-NC 4.0 |
|               9 |          5 |         2 | CC BY-SA 4.0 |
|              10 |          1 |         2 | CC BY-SA 4.0 |
|              11 |          7 |         3 | CC BY-NC 4.0 |
|              12 |          3 |         3 | CC BY-SA 4.0 |
|              13 |          9 |         3 | CC BY-NC 4.0 |
|              14 |          5 |         3 | CC BY-SA 4.0 |
|              15 |          1 |         3 | CC BY-SA 4.0 |
|              16 |          7 |         4 | CC BY-NC 4.0 |
|              17 |          4 |         4 | CC BY-SA 4.0 |
|              18 |         10 |         4 | CC BY-NC 4.0 |
|              19 |          6 |         4 | CC BY-NC 4.0 |
|              20 |          2 |         4 | CC BY-SA 4.0 |
|              21 |          8 |         5 | CC BY-NC 4.0 |
|              22 |          4 |         5 | CC BY-SA 4.0 |
|              23 |         10 |         5 | CC BY-NC 4.0 |
|              24 |          6 |         5 | CC BY-NC 4.0 |
|              25 |          2 |         6 | CC BY-SA 4.0 |
|              26 |          8 |         6 | CC BY-NC 4.0 |
|              27 |          4 |         6 | CC BY-SA 4.0 |

### Parsing decisions

- **Which episodes, and why:** episodes `0`–`27` were selected because they are the
  largest contiguous, zero-gap block of episodes whose `data/chunk-000/file-000.parquet`
  shard **and** every one of the 10 `videos/<key>/chunk-000/file-000.mp4` shards are
  shared — i.e. the smallest set of source files that had to be downloaded to get a
  complete, non-truncated set of episodes, given a limited local disk budget (~2 GB for
  all 28). It is not a curated or stratified sample.
- **Audio deliberately excluded.** The source `rh20t_cfg1` repo ships a per-episode
  audio sidecar (`audio/episode_NNNNNN.wav`, one file per episode across the full 4,258
  episodes) that is **not** declared as a feature in `meta/info.json` — it sits outside
  the LeRobot v3 schema entirely, and no video stream in this repo has an embedded audio
  track (`has_audio: false` for all 10 cameras). FiftyOne’s LeRobot importer only
  imports `dtype: "video"`/`"image"` features declared in `meta/info.json`, so this
  audio is not part of this FiftyOne dataset. This was also a deliberate choice given
  the source’s stated sensitivity note (see below) — audio may contain
  volunteer/operator voices.
- **Sensitive content (per source RH20T/port README):** the source data is
  volunteer-recorded human-robot interaction that may include faces in video and voices
  in the (excluded) audio sidecars. Handle this dataset for model-training purposes
  only; avoid casually browsing or redistributing episodes beyond that use.
- **Dual license mirrors the source, not simplified.** RH20T licenses its data
  per-scene (scenes 1–5 permissive CC BY-SA, scenes 6–10 non-commercial CC BY-NC); the
  `rh20t_cfg1` LeRobot port preserves this per-episode via `meta/rh20t_episodes.json`
  (`folder` field encodes `scene_NNNN`). This card’s frontmatter uses `license: other`
  with a link to the source, and the per-episode table above is the actual authority —
  do not assume a single blanket license for this repo.
- **Re-export, not a thin reference to the original repo:** this repo is a
  self-contained LeRobotDataset v3.0 export (via FiftyOne’s `LeRobotDatasetExporter`),
  not a pointer back to `robot-lev/rh20t_cfg1`. Task indices were remapped to only the
  tasks actually present in this subset (6 of the source’s 124). Per-episode and global
  statistics were recomputed from the exported rows, not carried over from the source’s
  global stats.
- **`observation.state.joint` may be all-zero.** Per source `meta/rh20t_episodes.json`,
  some episodes have `has_joint: false` (joint encoders were not recorded/valid for that
  session); those episodes’ `observation.state.joint` (and the corresponding slice of
  `observation.state`) is filled with zeros rather than omitted.
- **RGB only, no depth.** RH20T’s original release includes depth for some
  configurations; this LeRobot v3 port is RGB-only (`video.is_depth_map: false` for all
  10 streams).

## Dataset Creation

### Curation Rationale

RH20T was collected to give the robot-learning community a large-scale,
contact-rich, multi-modal manipulation dataset spanning hundreds of skills, robots, and
camera viewpoints, specifically to support one-shot imitation learning research where
a single demonstration should transfer to new tasks. This subset exists purely as a
lightweight, disk-budget-friendly slice of one robot configuration (`cfg1`, `flexiv`)
for exploration and tooling in FiftyOne; it was not re-curated for scenario content.

### Source Data

#### Data Collection and Processing

- **Robot/gripper:** Flexiv arm (7 DOF) with a `dahuan_ag95` parallel gripper
  (0–95 mm), config 1 (`cfg_0001`). Force/torque sensing includes both an end-effector
  force/torque estimate (`observation.force`/`observation.torque`) and a dedicated
  robot-mounted 6-axis sensor (`observation.robot_ft`).
- **Cameras:** 10 fixed RealSense-style camera serials per config, each recording
  360×640 RGB video at 10 fps (down-sampled from higher native capture rates —
  `native_hz` in `meta/rh20t_episodes.json` varies per episode, e.g. 5.85–10.81 Hz for
  the episodes in this subset); one camera is designated `master_camera` per episode for
  temporal alignment, and not every camera serial is present in every episode
  (`cameras_present` varies).
- **Original RH20T collection (per the paper):** over 110,000 contact-rich manipulation
  sequences across diverse skills, contexts, robots, and camera viewpoints, each with
  visual, force, audio, and action information, a corresponding human demonstration
  video, and a language description — collected across many robot platforms and
  configurations (`cfg1`–`cfg7`+), of which this repo covers only `cfg1` (Flexiv).
- **Quality rating:** each episode carries a `meta.rating` (per-frame, constant within
  an episode) reflecting demonstration quality, e.g. rating 8–9 for the episodes in this
  subset.

#### Who are the source data producers?

Collected by the RH20T authors’ team and volunteer human operators demonstrating and
teleoperating the robots across the dataset’s task/scene sessions.

### Annotations

#### Annotation process

Per-episode task grouping (`task_id`) and demonstration quality (`meta.rating`) come
from the original RH20T collection protocol. The LeRobot v3 port
(`lvjonok/rh20t-lerobot-port`) does not carry forward RH20T’s original natural-language
task descriptions into `meta/tasks.parquet` — only placeholder `"task N"` strings keyed
by `task_id` are present (see “Task labels are placeholders” above).

#### Who are the annotators?

Original RH20T data collectors/reviewers for `meta.rating`; not documented in this port’s metadata.

#### Personal and Sensitive Information

Yes — per the source README, this dataset contains volunteer-recorded human-robot
interaction that may include **faces in video** and **voices in the (excluded) audio
sidecars**. Exercise care to avoid inspecting or sharing sensitive content; use this
dataset for model-training purposes only.

## Citation

**BibTeX:**

```bibtex
@article{fang2023rh20t,
  title={RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot},
  author={Fang, Hao-Shu and Fang, Hongjie and Tang, Zhenyu and Liu, Jirong and Wang, Chenxi and Wang, Junbo and Zhu, Haoyi and Lu, Cewu},
  journal={arXiv preprint arXiv:2307.00595},
  year={2023}
}
```

**APA:**

Fang, H.-S., Fang, H., Tang, Z., Liu, J., Wang, C., Wang, J., Zhu, H., & Lu, C. (2023). *RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot*. arXiv:2307.00595.

## More Information

This is a 28-episode subset of one robot configuration (`cfg1`, Flexiv) of the
unofficial LeRobot v3 port [robot-lev/rh20t_cfg1](https://huggingface.co/datasets/robot-lev/rh20t_cfg1)
(4,258 episodes), itself one config of the full RH20T dataset (110,000+ sequences across
many robot platforms). See the [RH20T project page](https://rh20t.github.io/) and
[API](https://github.com/rh20t/rh20t_api) for the full dataset and other configs.

## Dataset Card Authors

[Harpreet Sahota](https://huggingface.co/harpreetsahota)

## Dataset Card Contact

[Harpreet Sahota](https://huggingface.co/harpreetsahota)
