Note
This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.
This is a FiftyOne dataset with 28 samples.
Installation#
If you haven’t already, install FiftyOne:
pip install -U fiftyone
Usage#
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/rh20t-cfg1-28ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for RH20T cfg1 (28-episode FiftyOne subset)#

A 28-episode subset of rh20t_cfg1, an unofficial community LeRobotDataset v3
reformatting (config 1, flexiv robot) of RH20T, a large-scale multi-modal robotic
manipulation dataset. The full rh20t_cfg1 repo (4,258 episodes) is published at
robot-lev/rh20t_cfg1; this repo
holds episodes 0–27 re-packaged as a self-contained LeRobotDataset v3.0 export and
loaded into FiftyOne for exploration.
Dataset Details#
Dataset Description#
Curated by: Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, Cewu Lu (original RH20T dataset); LeRobot v3 reformatting by the community (
lvjonok/rh20t-lerobot-port)Shared by: robot-lev (LeRobot v3 reformatting); this FiftyOne subset shared by the FiftyOne community
Language(s): English (task metadata)
License: Dual-licensed per scene, mirrored from the original RH20T license split (see “Parsing decisions” below): scenes 1–5 → CC BY-SA 4.0; scenes 6–10 → CC BY-NC 4.0. This 28-episode subset mixes both: 16 episodes are CC BY-SA 4.0, 12 are CC BY-NC 4.0 (non-commercial only) — see the per-episode table below before any commercial use.
Dataset Sources#
Repository: robot-lev/rh20t_cfg1 · reformatting code:
lvjonok/rh20t-lerobot-portPaper: RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot (arXiv:2307.00595)
Demo: RH20T project page · RH20T API
Uses#
Direct Use#
Exploring and visualizing multi-camera robotic manipulation episodes in the FiftyOne
App — inspecting 10 synchronized RGB camera views alongside end-effector/joint state,
force/torque and 6-axis fingertip wrench readings, and 8-dim actions (next end-effector
pose + gripper); prototyping data loaders and filters before working with the full
4,258-episode rh20t_cfg1 repo or the broader RH20T dataset (110,000+ sequences across
many robot configs).
Out-of-Scope Use#
This 28-episode subset is not a statistically representative sample of the full dataset (it is simply the first contiguous block of episodes whose data and all 10 video streams share the first storage shard) and should not be used to draw conclusions about task, scene, or robot-configuration distributions across the full RH20T dataset. Do not use the CC BY-NC-licensed episodes (or models trained on them) for commercial purposes — see the per-episode license table below. This subset also excludes audio and depth entirely (see Parsing decisions); it is not suitable for tasks requiring those modalities.
Dataset Structure#
This is a multimodal FiftyOne dataset (dataset.media_type == "multimodal") with
28 samples, one sample per episode. Each sample’s media (10 video streams) is not
copied into per-sample files; instead it is resolved through a media_reference that
points into the exported LeRobotDataset v3.0 source (data/, videos/, meta/ in this
repo) at import time — this is how FiftyOne represents LeRobot episodes natively.
Fields#
Field |
FiftyOne type |
Description |
|---|---|---|
|
|
FiftyOne sample id |
|
|
Pointer into the LeRobot source’s |
|
|
FiftyOne sample tags (empty by default) |
|
|
Standard FiftyOne sample metadata |
|
|
FiftyOne bookkeeping timestamps |
|
|
Episode index within this subset ( |
|
|
Task label for the episode, one of |
|
|
Full task list for the episode (length 1 for every episode here) |
|
|
Number of frames in the episode (verbatim from source) |
|
|
Episode duration in seconds ( |
|
|
|
|
|
Recording frame rate, 10.0 (verbatim from source) |
The per-frame numeric features and the 10 per-frame video streams are not flattened
into sample fields — they remain in the LeRobot data/*.parquet and videos/*/*.mp4
files referenced by media_reference, and are surfaced by the FiftyOne App’s State &
Action, Streams, and Statistics viewer tabs rather than as queryable sample-level fields.
Per source meta/info.json, these per-frame features are:
feature |
shape |
description |
|---|---|---|
|
|
concatenation of |
|
|
end-effector pose (position + quaternion) |
|
|
joint positions (all-zero for episodes where the source |
|
|
gripper opening, mm (0–95 for the |
|
|
end-effector force (Fx, Fy, Fz) |
|
|
end-effector torque (Mx, My, Mz) |
|
|
robot-mounted force/torque sensor reading |
|
|
next end-effector pose (7) + gripper target (1) |
|
|
per-episode demonstration quality rating (source RH20T annotation) |
|
|
one video stream per camera serial number; |
Task labels are placeholders, not descriptions. meta/tasks.parquet in the source
maps task_index to plain strings "task 1", "task 2", … rather than natural-language
instructions — the LeRobot v3 port did not carry over RH20T’s original task
descriptions/language annotations, only the numeric task_id (visible in
meta/rh20t_episodes.json, not surfaced as a FiftyOne field). Treat task/tasks here
as an opaque task-cluster id, not a human-readable instruction.
Label types and why#
There are no traditional detection/classification/segmentation labels. task is a plain
StringField (not fo.Classification) because, as noted above, it is a placeholder
identifier rather than a meaningful closed-set category with semantic content worth
modeling as a classification label in this port.
dataset.info contents#
{
"lerobot": {
"format": "LeRobotDataset",
"format_major": 3,
"episode_count": 4258, # total episodes in the full source dataset
"imported_episode_count": 28, # episodes actually imported into this subset
"skipped_episodes": [],
}
}
Per-episode license and scene mapping#
Source scene/task metadata for each imported episode (new_episode_index → source
episode_index, meta/rh20t_episodes.json); episode indices are identical here since
this contiguous block starts at 0:
episode_index |
scene_id |
task_id |
license |
|---|---|---|---|
0 |
1 |
1 |
CC BY-SA 4.0 |
1 |
7 |
1 |
CC BY-NC 4.0 |
2 |
3 |
1 |
CC BY-SA 4.0 |
3 |
9 |
1 |
CC BY-NC 4.0 |
4 |
5 |
1 |
CC BY-SA 4.0 |
5 |
1 |
2 |
CC BY-SA 4.0 |
6 |
7 |
2 |
CC BY-NC 4.0 |
7 |
3 |
2 |
CC BY-SA 4.0 |
8 |
9 |
2 |
CC BY-NC 4.0 |
9 |
5 |
2 |
CC BY-SA 4.0 |
10 |
1 |
2 |
CC BY-SA 4.0 |
11 |
7 |
3 |
CC BY-NC 4.0 |
12 |
3 |
3 |
CC BY-SA 4.0 |
13 |
9 |
3 |
CC BY-NC 4.0 |
14 |
5 |
3 |
CC BY-SA 4.0 |
15 |
1 |
3 |
CC BY-SA 4.0 |
16 |
7 |
4 |
CC BY-NC 4.0 |
17 |
4 |
4 |
CC BY-SA 4.0 |
18 |
10 |
4 |
CC BY-NC 4.0 |
19 |
6 |
4 |
CC BY-NC 4.0 |
20 |
2 |
4 |
CC BY-SA 4.0 |
21 |
8 |
5 |
CC BY-NC 4.0 |
22 |
4 |
5 |
CC BY-SA 4.0 |
23 |
10 |
5 |
CC BY-NC 4.0 |
24 |
6 |
5 |
CC BY-NC 4.0 |
25 |
2 |
6 |
CC BY-SA 4.0 |
26 |
8 |
6 |
CC BY-NC 4.0 |
27 |
4 |
6 |
CC BY-SA 4.0 |
Parsing decisions#
Which episodes, and why: episodes
0–27were selected because they are the largest contiguous, zero-gap block of episodes whosedata/chunk-000/file-000.parquetshard and every one of the 10videos/<key>/chunk-000/file-000.mp4shards are shared — i.e. the smallest set of source files that had to be downloaded to get a complete, non-truncated set of episodes, given a limited local disk budget (~2 GB for all 28). It is not a curated or stratified sample.Audio deliberately excluded. The source
rh20t_cfg1repo ships a per-episode audio sidecar (audio/episode_NNNNNN.wav, one file per episode across the full 4,258 episodes) that is not declared as a feature inmeta/info.json— it sits outside the LeRobot v3 schema entirely, and no video stream in this repo has an embedded audio track (has_audio: falsefor all 10 cameras). FiftyOne’s LeRobot importer only importsdtype: "video"/"image"features declared inmeta/info.json, so this audio is not part of this FiftyOne dataset. This was also a deliberate choice given the source’s stated sensitivity note (see below) — audio may contain volunteer/operator voices.Sensitive content (per source RH20T/port README): the source data is volunteer-recorded human-robot interaction that may include faces in video and voices in the (excluded) audio sidecars. Handle this dataset for model-training purposes only; avoid casually browsing or redistributing episodes beyond that use.
Dual license mirrors the source, not simplified. RH20T licenses its data per-scene (scenes 1–5 permissive CC BY-SA, scenes 6–10 non-commercial CC BY-NC); the
rh20t_cfg1LeRobot port preserves this per-episode viameta/rh20t_episodes.json(folderfield encodesscene_NNNN). This card’s frontmatter useslicense: otherwith a link to the source, and the per-episode table above is the actual authority — do not assume a single blanket license for this repo.Re-export, not a thin reference to the original repo: this repo is a self-contained LeRobotDataset v3.0 export (via FiftyOne’s
LeRobotDatasetExporter), not a pointer back torobot-lev/rh20t_cfg1. Task indices were remapped to only the tasks actually present in this subset (6 of the source’s 124). Per-episode and global statistics were recomputed from the exported rows, not carried over from the source’s global stats.observation.state.jointmay be all-zero. Per sourcemeta/rh20t_episodes.json, some episodes havehas_joint: false(joint encoders were not recorded/valid for that session); those episodes’observation.state.joint(and the corresponding slice ofobservation.state) is filled with zeros rather than omitted.RGB only, no depth. RH20T’s original release includes depth for some configurations; this LeRobot v3 port is RGB-only (
video.is_depth_map: falsefor all 10 streams).
Dataset Creation#
Curation Rationale#
RH20T was collected to give the robot-learning community a large-scale,
contact-rich, multi-modal manipulation dataset spanning hundreds of skills, robots, and
camera viewpoints, specifically to support one-shot imitation learning research where
a single demonstration should transfer to new tasks. This subset exists purely as a
lightweight, disk-budget-friendly slice of one robot configuration (cfg1, flexiv)
for exploration and tooling in FiftyOne; it was not re-curated for scenario content.
Source Data#
Data Collection and Processing#
Robot/gripper: Flexiv arm (7 DOF) with a
dahuan_ag95parallel gripper (0–95 mm), config 1 (cfg_0001). Force/torque sensing includes both an end-effector force/torque estimate (observation.force/observation.torque) and a dedicated robot-mounted 6-axis sensor (observation.robot_ft).Cameras: 10 fixed RealSense-style camera serials per config, each recording 360×640 RGB video at 10 fps (down-sampled from higher native capture rates —
native_hzinmeta/rh20t_episodes.jsonvaries per episode, e.g. 5.85–10.81 Hz for the episodes in this subset); one camera is designatedmaster_cameraper episode for temporal alignment, and not every camera serial is present in every episode (cameras_presentvaries).Original RH20T collection (per the paper): over 110,000 contact-rich manipulation sequences across diverse skills, contexts, robots, and camera viewpoints, each with visual, force, audio, and action information, a corresponding human demonstration video, and a language description — collected across many robot platforms and configurations (
cfg1–cfg7+), of which this repo covers onlycfg1(Flexiv).Quality rating: each episode carries a
meta.rating(per-frame, constant within an episode) reflecting demonstration quality, e.g. rating 8–9 for the episodes in this subset.
Who are the source data producers?#
Collected by the RH20T authors’ team and volunteer human operators demonstrating and teleoperating the robots across the dataset’s task/scene sessions.
Annotations#
Annotation process#
Per-episode task grouping (task_id) and demonstration quality (meta.rating) come
from the original RH20T collection protocol. The LeRobot v3 port
(lvjonok/rh20t-lerobot-port) does not carry forward RH20T’s original natural-language
task descriptions into meta/tasks.parquet — only placeholder "task N" strings keyed
by task_id are present (see “Task labels are placeholders” above).
Who are the annotators?#
Original RH20T data collectors/reviewers for meta.rating; not documented in this port’s metadata.
Personal and Sensitive Information#
Yes — per the source README, this dataset contains volunteer-recorded human-robot interaction that may include faces in video and voices in the (excluded) audio sidecars. Exercise care to avoid inspecting or sharing sensitive content; use this dataset for model-training purposes only.
Citation#
BibTeX:
@article{fang2023rh20t,
title={RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot},
author={Fang, Hao-Shu and Fang, Hongjie and Tang, Zhenyu and Liu, Jirong and Wang, Chenxi and Wang, Junbo and Zhu, Haoyi and Lu, Cewu},
journal={arXiv preprint arXiv:2307.00595},
year={2023}
}
APA:
Fang, H.-S., Fang, H., Tang, Z., Liu, J., Wang, C., Wang, J., Zhu, H., & Lu, C. (2023). RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot. arXiv:2307.00595.
More Information#
This is a 28-episode subset of one robot configuration (cfg1, Flexiv) of the
unofficial LeRobot v3 port robot-lev/rh20t_cfg1
(4,258 episodes), itself one config of the full RH20T dataset (110,000+ sequences across
many robot platforms). See the RH20T project page and
API for the full dataset and other configs.