Note
This is a Hugging Face dataset. For large datasets, ensure huggingface_hub>=1.1.3 to avoid rate limits. Learn more in the Hugging Face integration docs.
This is a subset of the agibot-world/AgiBotWorld2026 dataset created for use with FiftyOne.
Installation#
If you haven’t already, install FiftyOne:
pip install -U fiftyone
Usage#
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
dataset = fouh.load_from_hub("Voxel51/AgiBotWorld2026-sample-10ep")
session = fo.launch_app(dataset)
Dataset Card for AgiBot World 2026 (10-episode FiftyOne subset)#

Dataset Details#
Dataset Description#
AgiBot World 2026 is a large-scale, real-world embodied intelligence dataset collected on the AgiBot G2 dual-arm humanoid platform. It spans three top-level splits on the Hugging Face repo — ImitationLearning, ReinforcementLearning, and RichInteraction (plus a separate simulation split from the companion GenieSim project) — covering commercial-space, home, and other general-purpose manipulation scenarios.
This subset contains 10 episodes taken from a single ReinforcementLearning shard (ReinforcementLearning/Home/task_12192/HG-DAgger/20857245_20857259.tar.gz) of the source repo, all for the task “Insert the loose change into the coin slot of the piggy bank dagger”. Each episode records the dual-arm robot performing this insertion task under HG-DAgger-style intervention/rollout collection, with 7 synchronized camera streams and full proprioceptive state/action vectors.
Curated by: AgiBot World Team (source data); FiftyOne subset prepared for browsing/exploration in FiftyOne.
Funded by: Not specified in the source dataset.
Language(s): English (task instruction text)
License: CC BY-NC-SA 4.0 (inherited from the source dataset)
Dataset Sources#
Repository: https://huggingface.co/datasets/agibot-world/AgiBotWorld2026
Uses#
Direct Use#
Exploring, filtering, and visually inspecting AgiBot World 2026 robot-manipulation episodes in the FiftyOne App; prototyping data loaders and visualization pipelines for LeRobot-format multimodal robot-learning data before scaling up to the full 12+ TB source dataset; understanding the shape and structure of AgiBot’s HG-DAgger reinforcement-learning collection format (reward/done/intervened/succeed/truncated per-frame signals).
Out-of-Scope Use#
Not intended as a representative sample of the full AgiBot World 2026 dataset — it covers a single task, single collection mode (HG-DAgger), and single category (ReinforcementLearning) out of the many tasks/scenes/categories in the source. Not suitable for training production policies given its small size (10 episodes). Commercial use is restricted by the CC BY-NC-SA 4.0 license.
Dataset Structure#
This dataset uses FiftyOne’s multimodal / LeRobot representation: one sample = one episode. Each sample references the episode’s per-frame Parquet data (state/action vectors) and its 7 per-camera MP4 video streams via a media_reference; there are no per-sample image/video fields directly on the sample.
Fields#
Field |
Type |
Description |
|---|---|---|
|
|
FiftyOne sample ID |
|
|
Pointer to the episode’s Parquet (state/action) and MP4 (camera) shard files |
|
|
Episode index within this subset (0-9) |
|
|
Free-text task instruction for the episode (all episodes share the same instruction in this subset) |
|
|
List form of the task instruction(s) associated with the episode |
|
|
Number of frames in the episode (2176-2582 in this subset) |
|
|
Episode duration in seconds (72.5-86.1s in this subset) |
|
|
Robot platform, |
|
|
Frame rate, 30.0 for all episodes |
task is a StringField rather than a Classification because the source stores it as free-text natural-language instructions, not a fixed label taxonomy.
Per-frame data is not materialized as FiftyOne fields — it is read on demand from the referenced Parquet/MP4 shards and surfaced in the FiftyOne App’s State & Action / Streams tabs. The per-frame schema (from meta/info.json) is:
Per-frame feature |
dtype / shape |
Description |
|---|---|---|
|
video, 400x640x3 |
Head-mounted top camera (AV1/yuv420p) |
|
video, 528x640x3 |
Left hand/wrist camera (AV1/yuv420p) |
|
video, 528x640x3 |
Right hand/wrist camera (AV1/yuv420p) |
|
video, 768x960x3 |
Head left fisheye camera (AV1/yuv420p) |
|
video, 768x960x3 |
Head right fisheye camera (AV1/yuv420p) |
|
video, 768x960x3 |
Head rear-facing fisheye camera (AV1/yuv420p) |
|
video, 400x640x3, depth map |
Head depth stream, encoded as PNG-codec video with |
|
float32[216] |
Full proprioceptive state vector: effector positions, end-effector wrench/velocity/orientation/position, joint position/effort/mode/velocity, head/waist position & mode, robot base position/orientation, plus 6 camera extrinsic rotation matrices + translation vectors (see |
|
float32[44] |
Action vector: effector positions, end-effector position/orientation, joint position, head/waist position, robot base velocity (see |
|
float32[1] |
Per-frame RL reward signal |
|
bool[1] |
Episode-termination flag |
|
bool[1] |
Task-success flag |
|
bool[1] |
Whether a human intervened at this frame (HG-DAgger collection signal) |
|
bool[1] |
Episode-truncation flag |
|
float32[1] |
Policy action probability/confidence at this frame |
|
int64 / float32 |
Standard LeRobot bookkeeping columns |
Label types and why#
There are no detection/classification/segmentation label fields in this subset — it is raw multimodal sensorimotor data (video + proprioceptive state/action time series), not annotated with FiftyOne label types.
dataset.info contents#
{
"lerobot": {
"format": "LeRobotDataset",
"format_major": 3,
"episode_count": 10,
"imported_episode_count": 10,
"skipped_episodes": [],
}
}
Parsing decisions#
Why these episodes: The source
agibot-world/AgiBotWorld2026repo (12.3 TB, 420 files) is not laid out as a standard flat LeRobot repo — it has no root-levelmeta/info.json. Instead it ships as per-task.tar.gzarchives nested underImitationLearning/,ReinforcementLearning/,RichInteraction/, andsimulation/, each containing a self-contained LeRobot v2.1 export (data/,meta/,videos/) for one or more episodes of one task. This subset uses all 10 episodes from a single archive,ReinforcementLearning/Home/task_12192/HG-DAgger/20857245_20857259.tar.gz, chosen because it was the smallest availableReinforcementLearningarchive (~3.8 GB) and contains a complete, internally consistent set of episodes for one task.Why only
ReinforcementLearning, not a mix of categories:ImitationLearning,ReinforcementLearning, andRichInteractionarchives were sampled and inspected, but they declare incompatible per-frame schemas —observation.stateis shape[193]forImitationLearning,[216]forReinforcementLearning(this subset), and[169]forRichInteraction, and onlyReinforcementLearningepisodes carry thereward/done/succeed/intervened/truncated/action_probcolumns. A single LeRobot v3 export needs one consistent Parquet schema, so mixing categories was not possible; this subset stays within the single-schemaReinforcementLearning/HG-DAggerslice.Excluded modalities: none dropped from the declared feature set.
observation.images.head_depthis declared asdtype: "video"(not the FiftyOne-importer-unsupporteddepth_video), so it imported like an RGB stream even though it is agray16bePNG-codec depth map — no pixel values are log/linearly re-encoded and no decode parameters are needed, unlike some other LeRobot depth conventions.v2.1 -> v3.0 conversion: the source ships LeRobot v2.1. It was converted in place with the official
lerobot.scripts.convert_dataset_v21_to_v30(adding thetotal_chunks/total_videos/total_tasksfields the converter expects) before import. Extra top-levelinfo.jsonkeys not part of the LeRobot v2.1/v3 spec (camera_parameters,data_version,h5_path,instruction_segments,key_frame,take_over— AgiBot’s additional annotation layers for subtask segments, 2D bounding boxes, and instruction-level step labels) were dropped by the converter and are not present in this FiftyOne dataset or its re-export.Re-export on push: FiftyOne’s
push_to_hubre-exports the 10 selected episodes as a fresh, self-contained LeRobot v3 repo —episode_indexis remapped to0..9,task_indexis remapped to the single task present, and per-episode/global statistics are recomputed from the 10 episodes actually included (they are not copied from the much larger source dataset’s statistics).Codec caveats: the 6 RGB/fisheye camera streams use AV1 (
libdav1d/yuv420p), which decodes fine in Chromium-based browsers; the depth stream uses a PNG codec withgray16be, an unusual combination that may not render as a normal video in all players/browsers.Per-episode video statistics are zero: the source
meta/episodes_stats.jsonlreportsmin/max/mean/stdof0.0for every video feature — this is a known upstream artifact of AgiBot’s stats pipeline (documented on the source dataset card as not affecting usage), not an error introduced by this subset.
Dataset Creation#
Curation Rationale#
Created to make a small, browsable, schema-consistent slice of the very large (12.3 TB) AgiBot World 2026 dataset available for exploration in FiftyOne, and to document the dataset’s non-standard multi-archive repo layout and per-category schema differences for future subsetting work.
Source Data#
Data Collection and Processing#
Collected on the AgiBot G2 dual-arm robot platform in 100% real-world environments (commercial spaces, homes, and other general-purpose scenarios) via free-form teleoperation and, for the ReinforcementLearning category, an HG-DAgger-style human-in-the-loop intervention/rollout collection process. See the source dataset card for full collection methodology.
Who are the source data producers?#
The AgiBot World Team.
Annotations#
Annotation process#
No additional annotations were added for this subset. The source dataset’s extended annotation layers (subtask “Task Frame” segments, 2D bounding boxes, instruction-level “skill” segments, stored under key_frame/instruction_segments in the source meta/info.json) are not carried into this FiftyOne subset — see Parsing decisions.
Personal and Sensitive Information#
Episodes are recorded in real-world commercial and home environments and may contain incidental appearances of people, personal spaces, or identifiable objects/branding. No explicit PII redaction is documented by the source dataset.
Citation#
BibTeX:
@misc{agibotworld2026,
title = {AgiBot World 2026},
author = {AgiBot World Team},
howpublished = {\url{https://huggingface.co/datasets/agibot-world/AgiBotWorld2026}},
year = {2026}
}
APA:
AgiBot World Team. (2026). AgiBot World 2026 [Data set]. Hugging Face. https://huggingface.co/datasets/agibot-world/AgiBotWorld2026