#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/AgiBotWorld2026-sample-10ep" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

This is a subset of the [agibot-world/AgiBotWorld2026](https://huggingface.co/datasets/agibot-world/AgiBotWorld2026) dataset created for use with FiftyOne.

# Installation

If you haven’t already, install FiftyOne:

```bash
pip install -U fiftyone
```

# Usage

```python
import fiftyone as fo
import fiftyone.utils.huggingface as fouh

dataset = fouh.load_from_hub("Voxel51/AgiBotWorld2026-sample-10ep")

session = fo.launch_app(dataset)
```

# Dataset Card for AgiBot World 2026 (10-episode FiftyOne subset)

![AgiBot World 2026 preview](https://huggingface.co/datasets/Voxel51/AgiBotWorld2026-sample-10ep/resolve/main/agibotworld.gif)

## Dataset Details

### Dataset Description

AgiBot World 2026 is a large-scale, real-world embodied intelligence dataset collected on the AgiBot G2 dual-arm humanoid platform. It spans three top-level splits on the Hugging Face repo — `ImitationLearning`, `ReinforcementLearning`, and `RichInteraction` (plus a separate `simulation` split from the companion GenieSim project) — covering commercial-space, home, and other general-purpose manipulation scenarios.

This subset contains **10 episodes** taken from a single `ReinforcementLearning` shard (`ReinforcementLearning/Home/task_12192/HG-DAgger/20857245_20857259.tar.gz`) of the source repo, all for the task  *“Insert the loose change into the coin slot of the piggy bank dagger”*. Each episode records the dual-arm robot performing this insertion task under HG-DAgger-style intervention/rollout collection, with 7 synchronized camera streams and full proprioceptive state/action vectors.

- **Curated by:** AgiBot World Team (source data); FiftyOne subset prepared for browsing/exploration in FiftyOne.
- **Funded by:** Not specified in the source dataset.
- **Language(s):** English (task instruction text)
- **License:** CC BY-NC-SA 4.0 (inherited from the source dataset)

### Dataset Sources

- **Repository:** https://huggingface.co/datasets/agibot-world/AgiBotWorld2026

## Uses

### Direct Use

Exploring, filtering, and visually inspecting AgiBot World 2026 robot-manipulation episodes in the FiftyOne App; prototyping data loaders and visualization pipelines for LeRobot-format multimodal robot-learning data before scaling up to the full 12+ TB source dataset; understanding the shape and structure of AgiBot’s HG-DAgger reinforcement-learning collection format (reward/done/intervened/succeed/truncated per-frame signals).

### Out-of-Scope Use

Not intended as a representative sample of the full AgiBot World 2026 dataset — it covers a single task, single collection mode (`HG-DAgger`), and single category (`ReinforcementLearning`) out of the many tasks/scenes/categories in the source. Not suitable for training production policies given its small size (10 episodes). Commercial use is restricted by the CC BY-NC-SA 4.0 license.

## Dataset Structure

This dataset uses FiftyOne’s multimodal / LeRobot representation: one **sample = one episode**. Each sample references the episode’s per-frame Parquet data (state/action vectors) and its 7 per-camera MP4 video streams via a `media_reference`; there are no per-sample image/video fields directly on the sample.

### Fields

| Field             | Type                     | Description                                                                                         |
|-------------------|--------------------------|-----------------------------------------------------------------------------------------------------|
| `id`              | `ObjectIdField`          | FiftyOne sample ID                                                                                  |
| `media_reference` | `MediaReferenceField`    | Pointer to the episode’s Parquet (state/action) and MP4 (camera) shard files                        |
| `episode_index`   | `IntField`               | Episode index within this subset (0-9)                                                              |
| `task`            | `StringField`            | Free-text task instruction for the episode (all episodes share the same instruction in this subset) |
| `tasks`           | `ListField(StringField)` | List form of the task instruction(s) associated with the episode                                    |
| `length`          | `IntField`               | Number of frames in the episode (2176-2582 in this subset)                                          |
| `duration`        | `FloatField`             | Episode duration in seconds (72.5-86.1s in this subset)                                             |
| `robot_type`      | `StringField`            | Robot platform, `g2a` (AgiBot G2) for all episodes                                                  |
| `fps`             | `FloatField`             | Frame rate, 30.0 for all episodes                                                                   |

`task` is a `StringField` rather than a `Classification` because the source stores it as free-text natural-language instructions, not a fixed label taxonomy.

Per-frame data is not materialized as FiftyOne fields — it is read on demand from the referenced Parquet/MP4 shards and surfaced in the FiftyOne App’s State & Action / Streams tabs. The per-frame schema (from `meta/info.json`) is:

| Per-frame feature                                                  | dtype / shape               | Description                                                                                                                                                                                                                                                                                                                                                                     |
|--------------------------------------------------------------------|-----------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `observation.images.top_head`                                      | video, 400x640x3            | Head-mounted top camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                           |
| `observation.images.hand_left`                                     | video, 528x640x3            | Left hand/wrist camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                            |
| `observation.images.hand_right`                                    | video, 528x640x3            | Right hand/wrist camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                           |
| `observation.images.head_left_fisheye`                             | video, 768x960x3            | Head left fisheye camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                          |
| `observation.images.head_right_fisheye`                            | video, 768x960x3            | Head right fisheye camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                         |
| `observation.images.head_back_fisheye`                             | video, 768x960x3            | Head rear-facing fisheye camera (AV1/yuv420p)                                                                                                                                                                                                                                                                                                                                   |
| `observation.images.head_depth`                                    | video, 400x640x3, depth map | Head depth stream, encoded as PNG-codec video with `gray16be` pixel format                                                                                                                                                                                                                                                                                                      |
| `observation.state`                                                | float32[216]                | Full proprioceptive state vector: effector positions, end-effector wrench/velocity/orientation/position, joint position/effort/mode/velocity, head/waist position & mode, robot base position/orientation, plus 6 camera extrinsic rotation matrices + translation vectors (see `meta/info.json` -> `features.observation.state.field_descriptions` for the exact index layout) |
| `action`                                                           | float32[44]                 | Action vector: effector positions, end-effector position/orientation, joint position, head/waist position, robot base velocity (see `field_descriptions` for index layout)                                                                                                                                                                                                      |
| `reward`                                                           | float32[1]                  | Per-frame RL reward signal                                                                                                                                                                                                                                                                                                                                                      |
| `done`                                                             | bool[1]                     | Episode-termination flag                                                                                                                                                                                                                                                                                                                                                        |
| `succeed`                                                          | bool[1]                     | Task-success flag                                                                                                                                                                                                                                                                                                                                                               |
| `intervened`                                                       | bool[1]                     | Whether a human intervened at this frame (HG-DAgger collection signal)                                                                                                                                                                                                                                                                                                          |
| `truncated`                                                        | bool[1]                     | Episode-truncation flag                                                                                                                                                                                                                                                                                                                                                         |
| `action_prob`                                                      | float32[1]                  | Policy action probability/confidence at this frame                                                                                                                                                                                                                                                                                                                              |
| `episode_index`, `frame_index`, `index`, `task_index`, `timestamp` | int64 / float32             | Standard LeRobot bookkeeping columns                                                                                                                                                                                                                                                                                                                                            |

### Label types and why

There are no detection/classification/segmentation label fields in this subset — it is raw multimodal sensorimotor data (video + proprioceptive state/action time series), not annotated with FiftyOne label types.

### `dataset.info` contents

```python
{
    "lerobot": {
        "format": "LeRobotDataset",
        "format_major": 3,
        "episode_count": 10,
        "imported_episode_count": 10,
        "skipped_episodes": [],
    }
}
```

### Parsing decisions

- **Why these episodes:** The source `agibot-world/AgiBotWorld2026` repo (12.3 TB, 420 files) is not laid out as a standard flat LeRobot repo — it has no root-level `meta/info.json`. Instead it ships as per-task `.tar.gz` archives nested under `ImitationLearning/`, `ReinforcementLearning/`, `RichInteraction/`, and `simulation/`, each containing a self-contained LeRobot v2.1 export (`data/`, `meta/`, `videos/`) for one or more episodes of one task. This subset uses all 10 episodes from a single archive, `ReinforcementLearning/Home/task_12192/HG-DAgger/20857245_20857259.tar.gz`, chosen because it was the smallest available `ReinforcementLearning` archive (~3.8 GB) and contains a complete, internally consistent set of episodes for one task.
- **Why only `ReinforcementLearning`, not a mix of categories:** `ImitationLearning`, `ReinforcementLearning`, and `RichInteraction` archives were sampled and inspected, but they declare **incompatible per-frame schemas** — `observation.state` is shape `[193]` for `ImitationLearning`, `[216]` for `ReinforcementLearning` (this subset), and `[169]` for `RichInteraction`, and only `ReinforcementLearning` episodes carry the `reward`/`done`/`succeed`/`intervened`/`truncated`/`action_prob` columns. A single LeRobot v3 export needs one consistent Parquet schema, so mixing categories was not possible; this subset stays within the single-schema `ReinforcementLearning` / `HG-DAgger` slice.
- **Excluded modalities:** none dropped from the declared feature set. `observation.images.head_depth` is declared as `dtype: "video"` (not the FiftyOne-importer-unsupported `depth_video`), so it imported like an RGB stream even though it is a `gray16be` PNG-codec depth map — no pixel values are log/linearly re-encoded and no decode parameters are needed, unlike some other LeRobot depth conventions.
- **v2.1 -> v3.0 conversion:** the source ships LeRobot **v2.1**. It was converted in place with the official `lerobot.scripts.convert_dataset_v21_to_v30` (adding the `total_chunks`/`total_videos`/`total_tasks` fields the converter expects) before import. Extra top-level `info.json` keys not part of the LeRobot v2.1/v3 spec (`camera_parameters`, `data_version`, `h5_path`, `instruction_segments`, `key_frame`, `take_over` — AgiBot’s additional annotation layers for subtask segments, 2D bounding boxes, and instruction-level step labels) were dropped by the converter and are **not** present in this FiftyOne dataset or its re-export.
- **Re-export on push:** FiftyOne’s `push_to_hub` re-exports the 10 selected episodes as a fresh, self-contained LeRobot v3 repo — `episode_index` is remapped to `0..9`, `task_index` is remapped to the single task present, and per-episode/global statistics are recomputed from the 10 episodes actually included (they are not copied from the much larger source dataset’s statistics).
- **Codec caveats:** the 6 RGB/fisheye camera streams use AV1 (`libdav1d`/`yuv420p`), which decodes fine in Chromium-based browsers; the depth stream uses a PNG codec with `gray16be`, an unusual combination that may not render as a normal video in all players/browsers.
- **Per-episode video statistics are zero:** the source `meta/episodes_stats.jsonl` reports `min`/`max`/`mean`/`std` of `0.0` for every video feature — this is a known upstream artifact of AgiBot’s stats pipeline (documented on the source dataset card as not affecting usage), not an error introduced by this subset.

## Dataset Creation

### Curation Rationale

Created to make a small, browsable, schema-consistent slice of the very large (12.3 TB) AgiBot World 2026 dataset available for exploration in FiftyOne, and to document the dataset’s non-standard multi-archive repo layout and per-category schema differences for future subsetting work.

### Source Data

#### Data Collection and Processing

Collected on the AgiBot G2 dual-arm robot platform in 100% real-world environments (commercial spaces, homes, and other general-purpose scenarios) via free-form teleoperation and, for the `ReinforcementLearning` category, an HG-DAgger-style human-in-the-loop intervention/rollout collection process. See the [source dataset card](https://huggingface.co/datasets/agibot-world/AgiBotWorld2026) for full collection methodology.

#### Who are the source data producers?

The AgiBot World Team.

### Annotations

#### Annotation process

No additional annotations were added for this subset. The source dataset’s extended annotation layers (subtask “Task Frame” segments, 2D bounding boxes, instruction-level “skill” segments, stored under `key_frame`/`instruction_segments` in the source `meta/info.json`) are not carried into this FiftyOne subset — see *Parsing decisions*.

#### Personal and Sensitive Information

Episodes are recorded in real-world commercial and home environments and may contain incidental appearances of people, personal spaces, or identifiable objects/branding. No explicit PII redaction is documented by the source dataset.

## Citation

**BibTeX:**

```bibtex
@misc{agibotworld2026,
    title        = {AgiBot World 2026},
    author       = {AgiBot World Team},
    howpublished = {\url{https://huggingface.co/datasets/agibot-world/AgiBotWorld2026}},
    year         = {2026}
}
```

**APA:**

AgiBot World Team. (2026). *AgiBot World 2026* [Data set]. Hugging Face. https://huggingface.co/datasets/agibot-world/AgiBotWorld2026

## Dataset Card Authors

[Harpreet Sahota](https://huggingface.co/harpreetsahota)

## Dataset Card Contact

[Harpreet Sahota](https://huggingface.co/harpreetsahota)
