Note
This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.
This dataset was created using LeRobot and is presented here as a FiftyOne dataset.
Installation#
pip install -U fiftyone
Usage#
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
dataset = load_from_hub("Voxel51/tavis-head-gr1t2-800ep")
session = fo.launch_app(dataset)
Dataset Card for TAVIS Head GR1T2 (800-episode FiftyOne export)#

Dataset Details#
Dataset Description#
TAVIS Head GR1T2 is part of the (TAVIS) Torso Active-Vision Imitation Suite, a benchmark of teleoperated bimanual manipulation episodes collected on a humanoid platform. Each episode pairs four synchronized camera streams (a head-mounted active vision camera, a fixed scene camera, and two wrist cameras) with 60 Hz proprioceptive state, dual end-effector pose, and 19-dimensional action vectors, spanning 5 task families: cluttered pick-and-place, conditional picking, multi-shelf scanning, and wait-then-act behaviors. This card covers a full FiftyOne re-export of the source LeRobot v3.0 repository: all 800 episodes / 267,046 frames.
Curated by: tavis-benchmark (Hugging Face organization; individual authors not disclosed in the source repository)
Shared by: tavis-benchmark
License: cc-by-4.0
Dataset Sources#
Repository: https://huggingface.co/datasets/tavis-benchmark/tavis-head-gr1t2
Uses#
Direct Use#
Training and evaluating imitation-learning / vision-language-action policies for
bimanual humanoid manipulation, benchmarking active-vision (moving head camera)
strategies against fixed-camera baselines, and studying multi-task generalization
across pick, conditional-pick, multi-shelf, and wait-then-act behaviors. Two
pi0-based policies fine-tuned on this dataset are published by the same organization
(tavis-benchmark/pi0-tavis-head-gr1t2-fixedcam, tavis-benchmark/pi0-tavis-head-gr1t2-headcam).
Out-of-Scope Use#
Not intended for deployment on physical robots without further safety validation. Not annotated for object detection, segmentation, or any non-robotics computer-vision task.
Dataset Structure#
Fields#
Field |
Type |
Description |
|---|---|---|
|
|
Pointer to the episodeβs synchronized video/data streams |
|
|
Episode index (re-indexed |
|
|
Free-text task name for the episode (frame-0 label; see below) |
|
|
All task labels associated with the episode |
|
|
Number of frames in the episode (range 173β605) |
|
|
Episode duration in seconds (range 2.88β10.08) |
|
|
Robot identifier; not populated in the source metadata |
|
|
Frame rate, 60.0 for every episode |
Label types and why#
task is a StringField, not a Classification, because the source stores it as
free-text per episode (one task name per episode in this repo) rather than a
per-frame categorical label. Use tasks if you need the list form, and
count_values("task") to tabulate task frequency (see below).
Per-frame streams (from meta/info.json, not FiftyOne sample fields)#
Shown in the FiftyOne Appβs State & Action / Streams tabs, not as top-level sample fields:
Feature |
dtype |
Shape |
Notes |
|---|---|---|---|
|
video |
[480, 640, 3] |
AV1, yuv420p, 60 fps, head-mounted active-vision camera |
|
video |
[480, 640, 3] |
AV1, yuv420p, 60 fps, fixed scene camera |
|
video |
[480, 640, 3] |
AV1, yuv420p, 60 fps |
|
video |
[480, 640, 3] |
AV1, yuv420p, 60 fps |
|
float32 |
[19] |
Environment/joint actions |
|
float32 |
[44] |
Full proprioceptive state |
|
float32 |
[3] |
Left end-effector position |
|
float32 |
[4] |
Left end-effector orientation (quaternion) |
|
float32 |
[3] |
Right end-effector position |
|
float32 |
[4] |
Right end-effector orientation (quaternion) |
|
string |
[1] |
Natural-language instruction for the episode |
|
float32/int64 |
[1] |
Standard LeRobot bookkeeping columns |
dataset.info contents#
ds.info["lerobot"] records the original LeRobot metadata (codebase version,
fps, feature declarations) and skipped_episodes (empty β all 800 episodes imported
cleanly).
Parsing decisions#
Full dataset, not a subset. All 800 episodes from the source repository were imported; this export is a re-encoding of the complete dataset rather than a sampled subset, so episode indices are unchanged in content (0-based, contiguous) and no episodes were excluded.
meta/tasks.parquetrepair. The source file stored task names as the pandas DataFrame index (columns['task_index', '__index_level_0__']) instead of thetask_index+taskcolumns the LeRobot v3 reader and FiftyOneβs export path require. Repaired locally by resetting the index into ataskcolumn before ingest; no task content was changed, only its column layout.meta/episodesdata/file_indexrepair. The source episode metadata declareddata/file_indexvalues0β3(implying 4 data shard files), but the repo only ships a singledata/chunk-000/file-000.parquet, which in fact contains all 267,046 rows for all 800 episodes. The stale file_index values caused export to look for nonexistentfile-001.parquet/file-002.parquet/file-003.parquet. Repaired locally by settingdata/chunk_index/data/file_indexto0for every episode; no row data was moved or altered, only the shard pointer.No modalities excluded. All 4 declared video streams and all declared low-dimensional features (
dtype: "video"/ numeric) are supported by FiftyOneβs LeRobot importer and are all present in this export.Re-export. As with any FiftyOne β LeRobot push, episode indices were re-verified against
0..799and per-episode/global stats were recomputed by the exporter; only the shard files referenced by the imported episodes are included (here: all shards, since all episodes were imported).robot_typeisnullin the sourcemeta/info.jsonand is therefore empty for every sample; the humanoid platform is implied by the eef/state field layout (dual arm, 44-dim state, 19-dim action) but not named in the source metadata.
Dataset Creation#
Curation Rationale#
Collected as part of the TAVIS benchmark to study the effect of active (head-mounted, moving) vision versus fixed cameras on bimanual manipulation policies, across a mixture of pick-and-place, conditional, multi-object, and waiting behaviors.
Source Data#
Data Collection and Processing#
Episodes were recorded via teleoperation on a bimanual humanoid platform at 60 fps,
with 4 synchronized AV1-encoded camera streams and paired proprioceptive/action logs,
stored in LeRobot v3.0 format (single Parquet data shard, video streams sharded per
camera). Task distribution: ClutterPickLiftTask and MultiShelfScanTask each 250
episodes; ClutterPickCubeTask, ConditionalPickTask, and WaitThenActTask each 100
episodes (800 total).
Annotations#
Annotation process#
No manual annotations beyond the task label and language instruction recorded at collection time.
Personal and Sensitive Information#
No personal or sensitive information is expected; episodes depict a robot manipulating objects on a workbench, not people.
More Information#
Codec note: all video streams use AV1 (yuv420p), which decodes reliably in
Chromium-based browsers; playback support may vary in other browsers.