Note

This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.

Hugging Face

This dataset was created using LeRobot and is presented here as a FiftyOne dataset.

Installation#

pip install -U fiftyone

Usage#

import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

dataset = load_from_hub("Voxel51/tavis-head-gr1t2-800ep")
session = fo.launch_app(dataset)

Dataset Card for TAVIS Head GR1T2 (800-episode FiftyOne export)#

TAVIS Head GR1T2 preview

Dataset Details#

Dataset Description#

TAVIS Head GR1T2 is part of the (TAVIS) Torso Active-Vision Imitation Suite, a benchmark of teleoperated bimanual manipulation episodes collected on a humanoid platform. Each episode pairs four synchronized camera streams (a head-mounted active vision camera, a fixed scene camera, and two wrist cameras) with 60 Hz proprioceptive state, dual end-effector pose, and 19-dimensional action vectors, spanning 5 task families: cluttered pick-and-place, conditional picking, multi-shelf scanning, and wait-then-act behaviors. This card covers a full FiftyOne re-export of the source LeRobot v3.0 repository: all 800 episodes / 267,046 frames.

  • Curated by: tavis-benchmark (Hugging Face organization; individual authors not disclosed in the source repository)

  • Shared by: tavis-benchmark

  • License: cc-by-4.0

Dataset Sources#

  • Repository: https://huggingface.co/datasets/tavis-benchmark/tavis-head-gr1t2

Uses#

Direct Use#

Training and evaluating imitation-learning / vision-language-action policies for bimanual humanoid manipulation, benchmarking active-vision (moving head camera) strategies against fixed-camera baselines, and studying multi-task generalization across pick, conditional-pick, multi-shelf, and wait-then-act behaviors. Two pi0-based policies fine-tuned on this dataset are published by the same organization (tavis-benchmark/pi0-tavis-head-gr1t2-fixedcam, tavis-benchmark/pi0-tavis-head-gr1t2-headcam).

Out-of-Scope Use#

Not intended for deployment on physical robots without further safety validation. Not annotated for object detection, segmentation, or any non-robotics computer-vision task.

Dataset Structure#

Fields#

Field

Type

Description

media_reference

MediaReferenceField

Pointer to the episode’s synchronized video/data streams

episode_index

IntField

Episode index (re-indexed 0..799 on export)

task

StringField

Free-text task name for the episode (frame-0 label; see below)

tasks

ListField(StringField)

All task labels associated with the episode

length

IntField

Number of frames in the episode (range 173–605)

duration

FloatField

Episode duration in seconds (range 2.88–10.08)

robot_type

StringField

Robot identifier; not populated in the source metadata

fps

FloatField

Frame rate, 60.0 for every episode

Label types and why#

task is a StringField, not a Classification, because the source stores it as free-text per episode (one task name per episode in this repo) rather than a per-frame categorical label. Use tasks if you need the list form, and count_values("task") to tabulate task frequency (see below).

Per-frame streams (from meta/info.json, not FiftyOne sample fields)#

Shown in the FiftyOne App’s State & Action / Streams tabs, not as top-level sample fields:

Feature

dtype

Shape

Notes

observation.images.OBS_HEAD

video

[480, 640, 3]

AV1, yuv420p, 60 fps, head-mounted active-vision camera

observation.images.OBS_FIXED

video

[480, 640, 3]

AV1, yuv420p, 60 fps, fixed scene camera

observation.images.OBS_WRIST_LEFT

video

[480, 640, 3]

AV1, yuv420p, 60 fps

observation.images.OBS_WRIST_RIGHT

video

[480, 640, 3]

AV1, yuv420p, 60 fps

action

float32

[19]

Environment/joint actions

observation.state

float32

[44]

Full proprioceptive state

observation.left_eef_pos

float32

[3]

Left end-effector position

observation.left_eef_quat

float32

[4]

Left end-effector orientation (quaternion)

observation.right_eef_pos

float32

[3]

Right end-effector position

observation.right_eef_quat

float32

[4]

Right end-effector orientation (quaternion)

language_instruction

string

[1]

Natural-language instruction for the episode

timestamp, frame_index, episode_index, index, task_index

float32/int64

[1]

Standard LeRobot bookkeeping columns

dataset.info contents#

ds.info["lerobot"] records the original LeRobot metadata (codebase version, fps, feature declarations) and skipped_episodes (empty β€” all 800 episodes imported cleanly).

Parsing decisions#

  • Full dataset, not a subset. All 800 episodes from the source repository were imported; this export is a re-encoding of the complete dataset rather than a sampled subset, so episode indices are unchanged in content (0-based, contiguous) and no episodes were excluded.

  • meta/tasks.parquet repair. The source file stored task names as the pandas DataFrame index (columns ['task_index', '__index_level_0__']) instead of the task_index + task columns the LeRobot v3 reader and FiftyOne’s export path require. Repaired locally by resetting the index into a task column before ingest; no task content was changed, only its column layout.

  • meta/episodes data/file_index repair. The source episode metadata declared data/file_index values 0–3 (implying 4 data shard files), but the repo only ships a single data/chunk-000/file-000.parquet, which in fact contains all 267,046 rows for all 800 episodes. The stale file_index values caused export to look for nonexistent file-001.parquet/file-002.parquet/file-003.parquet. Repaired locally by setting data/chunk_index / data/file_index to 0 for every episode; no row data was moved or altered, only the shard pointer.

  • No modalities excluded. All 4 declared video streams and all declared low-dimensional features (dtype: "video" / numeric) are supported by FiftyOne’s LeRobot importer and are all present in this export.

  • Re-export. As with any FiftyOne β†’ LeRobot push, episode indices were re-verified against 0..799 and per-episode/global stats were recomputed by the exporter; only the shard files referenced by the imported episodes are included (here: all shards, since all episodes were imported).

  • robot_type is null in the source meta/info.json and is therefore empty for every sample; the humanoid platform is implied by the eef/state field layout (dual arm, 44-dim state, 19-dim action) but not named in the source metadata.

Dataset Creation#

Curation Rationale#

Collected as part of the TAVIS benchmark to study the effect of active (head-mounted, moving) vision versus fixed cameras on bimanual manipulation policies, across a mixture of pick-and-place, conditional, multi-object, and waiting behaviors.

Source Data#

Data Collection and Processing#

Episodes were recorded via teleoperation on a bimanual humanoid platform at 60 fps, with 4 synchronized AV1-encoded camera streams and paired proprioceptive/action logs, stored in LeRobot v3.0 format (single Parquet data shard, video streams sharded per camera). Task distribution: ClutterPickLiftTask and MultiShelfScanTask each 250 episodes; ClutterPickCubeTask, ConditionalPickTask, and WaitThenActTask each 100 episodes (800 total).

Annotations#

Annotation process#

No manual annotations beyond the task label and language instruction recorded at collection time.

Personal and Sensitive Information#

No personal or sensitive information is expected; episodes depict a robot manipulating objects on a workbench, not people.

More Information#

Codec note: all video streams use AV1 (yuv420p), which decodes reliably in Chromium-based browsers; playback support may vary in other browsers.

Dataset Card Authors#

Harpreet Sahota

Dataset Card Contact#

Harpreet Sahota