#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/tavis-head-gr1t2-800ep" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

This dataset was created using [LeRobot](https://github.com/huggingface/lerobot) and
is presented here as a [FiftyOne](https://github.com/voxel51/fiftyone) dataset.

# Installation

```bash
pip install -U fiftyone
```

# Usage

```python
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

dataset = load_from_hub("Voxel51/tavis-head-gr1t2-800ep")
session = fo.launch_app(dataset)
```

# Dataset Card for TAVIS Head GR1T2 (800-episode FiftyOne export)

![TAVIS Head GR1T2 preview](https://huggingface.co/datasets/Voxel51/tavis-head-gr1t2-800ep/resolve/main/tavis-head-gr1t2.gif)

## Dataset Details

### Dataset Description

TAVIS Head GR1T2 is part of the **(TAVIS) Torso Active-Vision Imitation Suite**, a
benchmark of teleoperated bimanual manipulation episodes collected on a humanoid
platform. Each episode pairs four synchronized camera streams (a head-mounted active
vision camera, a fixed scene camera, and two wrist cameras) with 60 Hz proprioceptive
state, dual end-effector pose, and 19-dimensional action vectors, spanning 5 task
families: cluttered pick-and-place, conditional picking, multi-shelf scanning, and
wait-then-act behaviors. This card covers a full FiftyOne re-export of the source
LeRobot v3.0 repository: all 800 episodes / 267,046 frames.

- **Curated by:** tavis-benchmark (Hugging Face organization; individual authors not
  disclosed in the source repository)
- **Shared by:** tavis-benchmark
- **License:** cc-by-4.0

### Dataset Sources

- **Repository:** https://huggingface.co/datasets/tavis-benchmark/tavis-head-gr1t2

## Uses

### Direct Use

Training and evaluating imitation-learning / vision-language-action policies for
bimanual humanoid manipulation, benchmarking active-vision (moving head camera)
strategies against fixed-camera baselines, and studying multi-task generalization
across pick, conditional-pick, multi-shelf, and wait-then-act behaviors. Two
pi0-based policies fine-tuned on this dataset are published by the same organization
(`tavis-benchmark/pi0-tavis-head-gr1t2-fixedcam`, `tavis-benchmark/pi0-tavis-head-gr1t2-headcam`).

### Out-of-Scope Use

Not intended for deployment on physical robots without further safety validation.
Not annotated for object detection, segmentation, or any non-robotics computer-vision
task.

## Dataset Structure

### Fields

| Field             | Type                     | Description                                                    |
|-------------------|--------------------------|----------------------------------------------------------------|
| `media_reference` | `MediaReferenceField`    | Pointer to the episode’s synchronized video/data streams       |
| `episode_index`   | `IntField`               | Episode index (re-indexed `0..799` on export)                  |
| `task`            | `StringField`            | Free-text task name for the episode (frame-0 label; see below) |
| `tasks`           | `ListField(StringField)` | All task labels associated with the episode                    |
| `length`          | `IntField`               | Number of frames in the episode (range 173–605)                |
| `duration`        | `FloatField`             | Episode duration in seconds (range 2.88–10.08)                 |
| `robot_type`      | `StringField`            | Robot identifier; not populated in the source metadata         |
| `fps`             | `FloatField`             | Frame rate, 60.0 for every episode                             |

### Label types and why

`task` is a `StringField`, not a `Classification`, because the source stores it as
free-text per episode (one task name per episode in this repo) rather than a
per-frame categorical label. Use `tasks` if you need the list form, and
`count_values("task")` to tabulate task frequency (see below).

### Per-frame streams (from `meta/info.json`, not FiftyOne sample fields)

Shown in the FiftyOne App’s State & Action / Streams tabs, not as top-level sample
fields:

| Feature                                                            | dtype         | Shape         | Notes                                                   |
|--------------------------------------------------------------------|---------------|---------------|---------------------------------------------------------|
| `observation.images.OBS_HEAD`                                      | video         | [480, 640, 3] | AV1, yuv420p, 60 fps, head-mounted active-vision camera |
| `observation.images.OBS_FIXED`                                     | video         | [480, 640, 3] | AV1, yuv420p, 60 fps, fixed scene camera                |
| `observation.images.OBS_WRIST_LEFT`                                | video         | [480, 640, 3] | AV1, yuv420p, 60 fps                                    |
| `observation.images.OBS_WRIST_RIGHT`                               | video         | [480, 640, 3] | AV1, yuv420p, 60 fps                                    |
| `action`                                                           | float32       | [19]          | Environment/joint actions                               |
| `observation.state`                                                | float32       | [44]          | Full proprioceptive state                               |
| `observation.left_eef_pos`                                         | float32       | [3]           | Left end-effector position                              |
| `observation.left_eef_quat`                                        | float32       | [4]           | Left end-effector orientation (quaternion)              |
| `observation.right_eef_pos`                                        | float32       | [3]           | Right end-effector position                             |
| `observation.right_eef_quat`                                       | float32       | [4]           | Right end-effector orientation (quaternion)             |
| `language_instruction`                                             | string        | [1]           | Natural-language instruction for the episode            |
| `timestamp`, `frame_index`, `episode_index`, `index`, `task_index` | float32/int64 | [1]           | Standard LeRobot bookkeeping columns                    |

### `dataset.info` contents

`ds.info["lerobot"]` records the original LeRobot metadata (codebase version,
fps, feature declarations) and `skipped_episodes` (empty — all 800 episodes imported
cleanly).

### Parsing decisions

- **Full dataset, not a subset.** All 800 episodes from the source repository were
  imported; this export is a re-encoding of the complete dataset rather than a
  sampled subset, so episode indices are unchanged in content (0-based, contiguous)
  and no episodes were excluded.
- **`meta/tasks.parquet` repair.** The source file stored task names as the pandas
  DataFrame index (columns `['task_index', '__index_level_0__']`) instead of the
  `task_index` + `task` columns the LeRobot v3 reader and FiftyOne’s export path
  require. Repaired locally by resetting the index into a `task` column before
  ingest; no task content was changed, only its column layout.
- **`meta/episodes` `data/file_index` repair.** The source episode metadata declared
  `data/file_index` values `0`–`3` (implying 4 data shard files), but the repo only
  ships a single `data/chunk-000/file-000.parquet`, which in fact contains all
  267,046 rows for all 800 episodes. The stale file_index values caused export to
  look for nonexistent `file-001.parquet`/`file-002.parquet`/`file-003.parquet`.
  Repaired locally by setting `data/chunk_index` / `data/file_index` to `0` for
  every episode; no row data was moved or altered, only the shard pointer.
- **No modalities excluded.** All 4 declared video streams and all declared
  low-dimensional features (`dtype: "video"` / numeric) are supported by FiftyOne’s
  LeRobot importer and are all present in this export.
- **Re-export.** As with any FiftyOne → LeRobot push, episode indices were
  re-verified against `0..799` and per-episode/global stats were recomputed by the
  exporter; only the shard files referenced by the imported episodes are included
  (here: all shards, since all episodes were imported).
- **`robot_type`** is `null` in the source `meta/info.json` and is therefore empty
  for every sample; the humanoid platform is implied by the eef/state field layout
  (dual arm, 44-dim state, 19-dim action) but not named in the source metadata.

## Dataset Creation

### Curation Rationale

Collected as part of the TAVIS benchmark to study the effect of active (head-mounted,
moving) vision versus fixed cameras on bimanual manipulation policies, across a
mixture of pick-and-place, conditional, multi-object, and waiting behaviors.

### Source Data

#### Data Collection and Processing

Episodes were recorded via teleoperation on a bimanual humanoid platform at 60 fps,
with 4 synchronized AV1-encoded camera streams and paired proprioceptive/action logs,
stored in LeRobot v3.0 format (single Parquet data shard, video streams sharded per
camera). Task distribution: `ClutterPickLiftTask` and `MultiShelfScanTask` each 250
episodes; `ClutterPickCubeTask`, `ConditionalPickTask`, and `WaitThenActTask` each 100
episodes (800 total).

### Annotations

#### Annotation process

No manual annotations beyond the task label and language instruction recorded at
collection time.

#### Personal and Sensitive Information

No personal or sensitive information is expected; episodes depict a robot
manipulating objects on a workbench, not people.

## More Information

Codec note: all video streams use AV1 (`yuv420p`), which decodes reliably in
Chromium-based browsers; playback support may vary in other browsers.

## Dataset Card Authors

[Harpreet Sahota](https://huggingface.co/harpreetsahota)

## Dataset Card Contact

[Harpreet Sahota](https://huggingface.co/harpreetsahota)
