#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/tavis-head-reachy2-800ep" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

This dataset was made available by [FiftyOne](https://github.com/voxel51/fiftyone).

# Installation

If you haven’t already, install FiftyOne:

```bash
pip install -U fiftyone
```

# Usage

```python
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/tavis-head-reachy2-800ep")

# Launch the App
session = fo.launch_app(dataset)
```

# Dataset Card for TAVIS-HEAD Reachy2 (800-episode FiftyOne dataset)

![TAVIS-HEAD Reachy2 preview](https://huggingface.co/datasets/Voxel51/tavis-head-reachy2-800ep/resolve/main/tavis-head-reachy2.gif)

This is a FiftyOne-formatted, self-contained LeRobot v3.0 export of the full
`tavis-benchmark/tavis-head-reachy2` dataset (all 800 episodes) loaded via
`fiftyone.utils.huggingface.load_from_hub`.

## Dataset Details

### Dataset Description

TAVIS-HEAD/Reachy2 is one of four multi-task suites in **TAVIS (A Benchmark for
Egocentric Active Vision and Anticipatory Gaze in Imitation Learning)**. It contains
800 VR-teleoperated demonstrations of a simulated Pollen Reachy 2 humanoid torso
performing five active-vision manipulation tasks (global visual search and clutter
handling) in NVIDIA IsaacLab. Episodes record a 19-D bimanual+neck action space, full
proprioceptive state, end-effector poses, and four synchronized RGB video streams
(head, two wrists, fixed external). This card documents the FiftyOne re-export of the
complete source dataset; no episodes were excluded.

- **Curated by:** the TAVIS benchmark authors (org: `tavis-benchmark` on Hugging Face;
  paper under double-blind review, authors not yet disclosed)
- **Shared by:** `tavis-benchmark`; this FiftyOne re-export shared by the FiftyOne
  community
- **Language(s):** English (33 distinct `language_instruction` values across the
  source dataset; 5 task classes on this Reachy2 shard)
- **License:** CC-BY-4.0 (datasets); code is MIT. Robot models, YCB objects, and task
  assets carry their own upstream licenses (NVIDIA IsaacLab, the YCB project, and the
  original robot-model authors) — see the [tavis-benchmark org card](https://huggingface.co/tavis-benchmark)

### Dataset Sources

- **Repository:** [tavis-benchmark/tavis-head-reachy2](https://huggingface.co/datasets/tavis-benchmark/tavis-head-reachy2)
- **Org / benchmark overview:** [tavis-benchmark](https://huggingface.co/tavis-benchmark)
- **Code (anonymous mirror, review period):** https://anonymous.4open.science/r/tavis-F5D7
- **Paper:** Under double-blind review at time of this export. Cite by name:
  “TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation
  Learning” — BibTeX to be posted once review concludes.
- **Pretrained baselines:** `tavis-benchmark/pi0-tavis-head-reachy2-headcam`,
  `tavis-benchmark/pi0-tavis-head-reachy2-fixedcam` (π₀ policies trained on this suite)

## Uses

### Direct Use

- Training and evaluating imitation-learning policies (diffusion policy, π₀, etc.) on
  active-vision, humanoid manipulation tasks.
- Studying anticipatory gaze behavior via the GALT (Gaze-Action Lead Time) metric,
  computed from the 19-D commanded-action trajectory (neck + arm channels) — see
  `docs/galt.md` in the code mirror.
- Comparing head-mounted vs. fixed-camera policies (paired π₀ checkpoints are
  released for both camera modes).
- Browsing, filtering, and re-exporting episode subsets with FiftyOne before
  training.

### Out-of-Scope Use

- This is simulated data (NVIDIA IsaacLab); it is not validated for direct sim-to-real
  transfer without further work.
- Not intended as a general humanoid-manipulation dataset outside the five TAVIS-HEAD
  task definitions below.

## Dataset Structure

This dataset is a `multimodal` FiftyOne dataset. Each sample is one episode; the
underlying frame-level Parquet rows and MP4 videos are referenced (not duplicated)
through `media_reference` and are exposed in the FiftyOne App’s State & Action /
Streams tabs.

### Fields

| Field                             | FiftyOne type            | Meaning                                                                                                |
|-----------------------------------|--------------------------|--------------------------------------------------------------------------------------------------------|
| `id`                              | `ObjectIdField`          | FiftyOne sample id                                                                                     |
| `media_reference`                 | `MediaReferenceField`    | Pointer to the episode’s Parquet row range and per-camera video byte/time ranges                       |
| `tags`                            | `ListField(StringField)` | FiftyOne tags (empty by default)                                                                       |
| `metadata`                        | `EmbeddedDocumentField`  | Standard FiftyOne sample metadata (unset for multimodal samples)                                       |
| `created_at` / `last_modified_at` | `DateTimeField`          | FiftyOne bookkeeping timestamps                                                                        |
| `episode_index`                   | `IntField`               | Episode index, re-indexed `0..799` on export                                                           |
| `task`                            | `StringField`            | Task class name for the episode’s first (only) task, e.g. `ClutterPickCubeTask`                        |
| `tasks`                           | `ListField(StringField)` | Full list of task labels associated with the episode (length 1 for every episode here)                 |
| `length`                          | `IntField`               | Number of frames in the episode (157–537)                                                              |
| `duration`                        | `FloatField`             | Episode duration in seconds (2.62–8.95)                                                                |
| `robot_type`                      | `StringField`            | Robot type string; `None` for every episode (source `meta/info.json` also declares `robot_type: null`) |
| `fps`                             | `FloatField`             | Recording frame rate, 60.0 for all episodes                                                            |

### Per-frame features (referenced via `media_reference`, from `meta/info.json`)

| Feature                                                            | dtype                | Shape         | Notes                                                                      |
|--------------------------------------------------------------------|----------------------|---------------|----------------------------------------------------------------------------|
| `action`                                                           | float32              | [19]          | Bimanual arm IK targets + 3-DoF neck + 2 gripper scalars                   |
| `observation.images.OBS_HEAD`                                      | video (av1, yuv420p) | [480, 640, 3] | Head-mounted camera, 60 fps                                                |
| `observation.images.OBS_WRIST_LEFT`                                | video (av1, yuv420p) | [480, 640, 3] | Left wrist camera, 60 fps                                                  |
| `observation.images.OBS_WRIST_RIGHT`                               | video (av1, yuv420p) | [480, 640, 3] | Right wrist camera, 60 fps                                                 |
| `observation.images.OBS_FIXED`                                     | video (av1, yuv420p) | [480, 640, 3] | Fixed external camera, 60 fps                                              |
| `observation.state`                                                | float32              | [36]          | Full proprioceptive state                                                  |
| `observation.left_eef_pos` / `.right_eef_pos`                      | float32              | [3]           | End-effector Cartesian position                                            |
| `observation.left_eef_quat` / `.right_eef_quat`                    | float32              | [4]           | End-effector orientation quaternion                                        |
| `language_instruction`                                             | string               | [1]           | Free-text instruction per frame (33 distinct values in the source dataset) |
| `timestamp`, `frame_index`, `episode_index`, `index`, `task_index` | float32 / int64      | [1]           | Standard LeRobot bookkeeping columns                                       |

### Label types and why

`task` is a `StringField`, not a `Classification` — the source stores task identity
as a per-episode string keyed against `meta/tasks.parquet` (5 distinct task classes:
`ClutterPickCubeTask`, `ClutterPickLiftTask`, `ConditionalPickTask`,
`MultiShelfScanTask`, `WaitThenActTask`), and every episode in this shard has exactly
one task, so it is imported as free text on the sample rather than a per-frame
primitive. `language_instruction` (the natural-language variant, e.g. “Find the red
cube and pick it up.”) remains inside the referenced per-frame Parquet data, not on
the sample.

### `dataset.info` contents

`dataset.info["lerobot"]` records the source `codebase_version` (`v3.0`), the
original `chunks_size` / file-size targets, and `skipped_episodes` (`[]` — every
episode from the source imported cleanly).

### Parsing decisions

- **Full dataset, no subsetting.** All 800 episodes and all 4 camera streams were
  imported; the source dataset’s single data shard (`data/chunk-000/file-000.parquet`)
  and 30 video shard files (4 cameras × up to 9 files each) were downloaded in full
  (~4.3 GB), so no shard-0-only subset was needed.
- **`meta/tasks.parquet` repair.** The source file stored task names as the pandas
  index (`__index_level_0__`) rather than a `task` column, which FiftyOne’s importer
  tolerates but the LeRobot export path (`lerobot_export._read_source_tasks`)
  requires as literal `task_index` + `task` columns. Repaired locally with
  `reset_index().rename(columns={"index": "task"})` before ingest; no data values were
  changed.
- **Stale `data/file_index` in `meta/episodes` repair.** The source episode metadata
  table declared `data/file_index` values `0`–`4` (as if the frame data were split
  across five Parquet shards), but the repo ships only one data file
  (`data/chunk-000/file-000.parquet`), which was verified to contain all 231,286
  frames across all 800 episodes. This is a source metadata bug (the data was likely
  consolidated into one file after the per-episode index was written but the index
  wasn’t refreshed). Repaired locally by setting `data/file_index = 0` for every row
  before ingest/export; no frame data was moved or altered. Video shard indices
  (`videos/<key>/file_index`, genuinely split across up to 9 files per camera) were
  left untouched.
- **Re-export.** Pushing via `dataset_type=fo.types.LeRobotDataset` re-indexes
  `episode_index` to a contiguous `0..799`, remaps `task_index`, and recomputes
  per-episode and global stats over the full 800-episode set — all of which match the
  source since no episodes were dropped.
- **`robot_type: null`.** The source `meta/info.json` declares `"robot_type": null`
  despite the dataset name/org page identifying the robot as a simulated Pollen
  Reachy 2; this is a source metadata gap, not a FiftyOne import artifact, and is
  preserved as-is.
- **Codec caveat.** Video is AV1/yuv420p, which decodes fine in Chromium-based
  browsers (the FiftyOne App); other browsers may have weaker AV1 support.
- **Excluded modalities:** none. Every feature declared in `meta/info.json` is either
  a supported `video` stream or a scalar/vector field carried in the referenced
  Parquet data.

## Dataset Creation

### Curation Rationale

TAVIS-HEAD isolates *global* visual-search behavior (as opposed to TAVIS-HANDS, which
tests local/occlusion-driven perception) by requiring the robot to scan clutter,
shelves, or cue cards with its head before acting. The Reachy2 shard pairs with a
matching GR1T2 shard so the benchmark can compare embodiments on an identical task
and action-space definition.

### Source Data

#### Data Collection and Processing

Demonstrations were collected via VR teleoperation (Meta Quest) in NVIDIA IsaacLab
simulation, built on the LeRobot framework and IsaacLab-Arena. Robot models are from
the upstream Pollen Reachy 2 USD distribution; manipulated objects are from the YCB
project. Episodes were exported as LeRobot v3.0 (Parquet action/state tables + AV1
video shards, chunked by file size).

#### Who are the source data producers?

VR teleoperators recorded the demonstrations under the TAVIS benchmark project (org
`tavis-benchmark`); individual operator identities are not disclosed in the source
metadata.

### Annotations

No separate annotation pass — task identity and language instructions are recorded
directly during teleoperation as `task_index`/`task` and per-frame
`language_instruction` fields.

#### Personal and Sensitive Information

None. All data is synthetic (simulated humanoid, simulated objects); no real people,
locations, or personal data are involved.

## Citation

Until the formal citation is posted, cite by name: “TAVIS: A Benchmark for Egocentric
Active Vision and Anticipatory Gaze in Imitation Learning.”

**APA:**

TAVIS Benchmark Authors. (Under review). *TAVIS: A Benchmark for Egocentric Active
Vision and Anticipatory Gaze in Imitation Learning* [Data set]. Hugging Face.
https://huggingface.co/datasets/tavis-benchmark/tavis-head-reachy2

## More Information

Part of the four-suite TAVIS benchmark (`tavis-head-gr1t2`, `tavis-head-reachy2`,
`tavis-hands-gr1t2`, `tavis-hands-reachy2`; 2,200 total demonstrations). See the
[tavis-benchmark org page](https://huggingface.co/tavis-benchmark) for the full task
list, evaluation protocol (`id` / `ood_spatial` / `ood_init_pose`), and pretrained π₀
checkpoints.

## Dataset Card Authors

[Harpreet Sahota](https://huggingface.co/harpreetsahota)

## Dataset Card Contact

[Harpreet Sahota](https://huggingface.co/harpreetsahota)
