#### NOTE
This is a **Hugging Face dataset**. For large datasets, ensure `huggingface_hub>=1.1.3` to avoid rate limits. Learn more in the <a href="https://docs.voxel51.com/integrations/huggingface.html#loading-datasets-from-the-hub" target="_blank">Hugging Face integration docs</a>.

<a href="https://huggingface.co/datasets/Voxel51/oxford-spires-multimodal" target="_blank">![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-yellow)</a>

# Dataset Card for Oxford Spires Multimodal (MCAP)

![image/png](https://huggingface.co/datasets/Voxel51/oxford-spires-multimodal/resolve/main/oxford_spires.gif)

A FiftyOne build of the **Oxford Spires Dataset**, the large-scale
LiDAR-visual localisation, reconstruction and radiance-field benchmark from
the Oxford Robotics Institute (ORI). This build repackages 6 of the 24
source sequences — one per historic Oxford landmark — as time-synchronised
[MCAP](https://mcap.dev/) recordings for FiftyOne’s native
[multimodal dataset support](https://docs.voxel51.com/user_guide/multimodal.html)
(FiftyOne 1.19+). Each sample is one episode, viewable in FiftyOne’s tiled
multimodal viewer with synchronised three-camera fisheye imagery,
motion-undistorted LiDAR point clouds, IMU telemetry, and a live 6-DoF pose
track from LiDAR-inertial SLAM.

Alongside the raw sensor streams, each episode also carries the three
per-camera products of the source devkit’s own `generate_depth.py`
pipeline — 16-bit depth maps, HSV depth overlays on the camera image, and
surface-normal maps — logged as additional streams so the devkit’s canonical
visualisation is reproducible inside the App.

Oxford Spires is a raw multi-sensor dataset for benchmarking SLAM,
Structure-from-Motion, Multi-View Stereo, NeRF and 3D Gaussian Splatting
methods; it carries **no object-level annotations**. “Ground truth” in the
source dataset means millimetre-accurate Terrestrial LiDAR Scanner (TLS) 3D
models and the centimetre-accurate trajectories registered against them.
This repackaging does not add or alter any ground truth; see
[Dataset Creation](#dataset-creation) for exactly what was kept, converted,
and left out.

This is a [FiftyOne](https://github.com/voxel51/fiftyone) dataset with 6 samples.

## Installation

If you haven’t already, install FiftyOne:

```bash
pip install -U fiftyone
```

## Usage

```python
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub

# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/oxford-spires-multimodal")

# Launch the App
session = fo.launch_app(dataset)
```

## Dataset Details

### Dataset Description

The Oxford Spires Dataset was captured in and around six well-known
historic landmarks in Oxford, UK, using a custom handheld multi-sensor
perception unit called **Frontier**, carried in a backpack at walking pace.
The unit comprises three synchronised global-shutter colour fisheye cameras
(forward-, left- and right-facing), a 64-beam automotive 3D LiDAR, and an
inertial sensor, all precisely calibrated. Each site is additionally
covered by a millimetre-accurate reference 3D model captured with a
Terrestrial LiDAR Scanner, which the authors use both as reconstruction
ground truth and — via ICP registration of the mobile LiDAR scans — as the
source of centimetre-accurate ground-truth trajectories.

In total the source dataset contains 24 sequences across the six sites,
covering more than 125,000 m² (about the size of a small town), with the
average distance travelled per sequence exceeding 400 metres. The three
forward/left/right camera configuration is a distinguishing feature: it
widens the field of view for texture mapping and supplies the extra view
constraints that vision-only methods need to infer 3D structure from a
single linear pass through an environment. The authors establish three
benchmarks on this data — localisation, 3D reconstruction, and novel-view
synthesis — and use them to show that state-of-the-art radiance field
methods overfit to training poses and generalise poorly to out-of-sequence
viewpoints.

This FiftyOne build covers 6 full-length episodes, one per site (see
[Curation Rationale](#curation-rationale)).

- **Curated by:** Oxford Robotics Institute, Department of Engineering
  Science, University of Oxford, in collaboration with the Group of
  Automation, Robotics and Computer Vision (AUROVA), University of Alicante
  — original data collection, sensor calibration, TLS reference models, and
  ground-truth trajectory post-processing. This MCAP/FiftyOne multimodal
  repackaging (episode authoring, dataset card) was prepared independently
  by Harpreet Sahota.
- **Funded by:** Partly funded by the National Research Foundation of Korea
  (NRF) grant funded by the Korea government (MSIT), No. RS-2024-00461409.
  Miguel Ángel Muñoz-Bañón is supported by the Valencian Community
  Government and the European Union through the CIBEST/2023/44 fellowship
  and the PROMETEO/2021/075 project.
- **Shared by:** Harpreet Sahota (this repackaging); the original Oxford
  Spires Dataset is shared by the Oxford Robotics Institute via
  https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/ and the
  [`ori-drs/oxford_spires_dataset`](https://huggingface.co/datasets/ori-drs/oxford_spires_dataset)
  Hugging Face dataset repository.
- **Language(s):** N/A (sensor data — camera, LiDAR, IMU, pose; no text).
- **License:** [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)
  — Copyright (c) 2024, University of Oxford; intended for non-commercial
  academic use. Commercial use requires contacting the original authors at
  oxfordspiresdataset@robots.ox.ac.uk.

### Dataset Sources

- **Repository:** https://github.com/ori-drs/oxford_spires_dataset (devkit,
  `oxspires_tools`); data at
  https://huggingface.co/datasets/ori-drs/oxford_spires_dataset
- **Paper:** Tao, Y., Muñoz-Bañón, M. Á., Zhang, L., Wang, J., Fu, L. F. T.,
  & Fallon, M. (2025). *The Oxford Spires Dataset: Benchmarking Large-Scale
  LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods*. The
  International Journal of Robotics Research.
  arXiv:[2411.10546](https://arxiv.org/abs/2411.10546);
  [SAGE](https://journals.sagepub.com/doi/10.1177/02783649251369905)
- **Demo:** https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/ (official
  project site) and https://www.youtube.com/watch?v=AKZ-YrOob_4 (dataset
  video)

## Uses

### Direct Use

- Exercising/demoing FiftyOne’s multimodal MCAP support: synchronised
  playback of three fisheye camera streams, motion-undistorted LiDAR point
  clouds, IMU telemetry, and a moving 6-DoF sensor pose across
  full-length, real handheld recordings (3.5–13.7 minutes each).
- Inspecting LiDAR-camera calibration quality interactively — the depth
  overlay streams (`/overlay_cam_*`) reproduce the devkit’s own
  LiDAR-on-image visualisation, which the paper uses (Figs. 4 and 5) to
  demonstrate calibration and motion-undistortion accuracy.
- Prototyping analyses that need per-frame LiDAR-to-camera projection:
  16-bit euclidean depth maps and surface-normal maps are logged per camera
  at every image-synchronised keyframe.
- Qualitative review of LiDAR-inertial SLAM trajectory behaviour, with the
  `world` → `base` transform track driving the 3D tile’s camera and
  point-cloud placement.
- Browsing site-to-site variation (limestone vs. Keble’s alternating red and
  white brick; open squares vs. narrow passages; indoor/outdoor lighting
  transitions) via the `location` sample field.

### Out-of-Scope Use

- Reproducing the paper’s localisation, reconstruction, or novel-view
  synthesis benchmark numbers. Those benchmarks require artifacts that are
  **not** in this build: the TLS ground-truth 3D models, the TLS-registered
  ground-truth trajectories (`gt-tum.txt`), the HBA and COLMAP trajectory
  variants, and the COLMAP SfM outputs. See
  [Parsing decisions](#parsing-decisions).
- Any task needing object-level labels (detection, segmentation,
  classification). No such annotations exist anywhere in the source dataset.
- Metric evaluation against the poses in this build. The pose track here is
  the **VILENS-SLAM estimate** shipped in `slam-poses.csv`, which is itself
  one of the systems *evaluated* in the paper’s localisation benchmark
  (ATE 0.03–1.11 m depending on sequence), not the TLS-registered ground
  truth (≈1–2 cm accuracy).
- Photometrically consistent colour reconstruction from the merged
  `/lidar_rgb` clouds. The paper explicitly flags this as an open problem
  for this data: camera auto-exposure was enabled, so the same 3D structure
  observed from different viewpoints has inconsistent pixel intensity, and
  merging colourised LiDAR clouds “would lead to a mixture of different
  colours in the reconstruction.”
- Map-based workflows. There is no GNSS/GPS stream anywhere in the source
  dataset, so the App’s Map tile is empty for every episode by design.

## Dataset Structure

This is a flat (ungrouped) FiftyOne dataset with `media_type: "multimodal"`
and **6 samples**. Each sample is one **episode**, stored as one `.mcap`
file; FiftyOne infers the multimodal media type automatically from the
`.mcap` extension. There are no separate per-frame image or point-cloud
samples — the episode is the sample unit, and every stream inside it
(cameras, LiDAR, IMU, transforms, and the derived depth/overlay/normal
products) is decoded live by FiftyOne’s multimodal viewer.

The dataset carries no per-sample tags, no temporal tags, and `dataset.info`
is empty — there is no extra dataset-level metadata beyond the per-sample
fields below. Sensor calibration is not stored in `dataset.info`; it lives
inside each MCAP as `foxglove.CameraCalibration` and
`foxglove.FrameTransform` messages, so the viewer can use it directly. The
built-in `metadata` field is unpopulated (`None`) because
`compute_metadata()` was not run.

Totals across the 6 episodes: 1,091,709 MCAP messages, 2,481.3 seconds
(41.4 minutes) of recording, 37 GB of MCAP on disk.

### Episodes in this dataset

| `sequence_id`                       | `location`            | `recording_date`   |   `duration_s` |   `message_count` |   `channel_count` |   `/lidar` clouds | `has_rgb_lidar`   |
|-------------------------------------|-----------------------|--------------------|----------------|-------------------|-------------------|-------------------|-------------------|
| `2024-03-12-keble-college-02`       | `keble-college`       | 2024-03-12         |          300.1 |           148,939 |                18 |               805 | False             |
| `2024-03-13-observatory-quarter-02` | `observatory-quarter` | 2024-03-13         |          275.1 |           131,263 |                18 |               361 | False             |
| `2024-03-14-blenheim-palace-05`     | `blenheim-palace`     | 2024-03-14         |          372.9 |           160,664 |                19 |               361 | True              |
| `2024-03-20-christ-church-05`       | `christ-church`       | 2024-03-20         |          822.1 |           336,685 |                19 |               713 | True              |
| `2024-05-20-bodleian-library-02`    | `bodleian-library`    | 2024-05-20         |          503   |           219,393 |                19 |               521 | True              |
| `2024-07-09-new-college-01`         | `new-college`         | 2024-07-09         |          208.1 |            94,765 |                19 |               192 | True              |

The two 18-channel episodes lack `/lidar_rgb` because their source point
clouds already ship with colour baked in — see
[Parsing decisions](#parsing-decisions).

The effective image rate varies markedly between episodes (from ≈4.8 Hz for
`christ-church-05` to ≈20 Hz for `keble-college-02`, computed as image count
divided by episode duration). The paper documents the raw camera streams as
20 Hz, so the `raw/images.zip` archives for some sequences appear to be a
decimated subset rather than the full stream; the counts logged here match
those archives exactly.

### Fields

| Field                    | FiftyOne type            | Description                                                                                                                     |
|--------------------------|--------------------------|---------------------------------------------------------------------------------------------------------------------------------|
| `filepath`               | `StringField`            | Absolute path to the episode’s `.mcap` file — the sample’s multimodal media                                                     |
| `sequence_id`            | `StringField`            | Source sequence name (`YYYY-MM-DD-<site>-<NN>`), verbatim from the source repository’s folder name                              |
| `location`               | `StringField`            | Site slug parsed from `sequence_id` (e.g. `keble-college`); one of 6                                                            |
| `recording_date`         | `StringField`            | Recording date parsed from `sequence_id`, `YYYY-MM-DD`                                                                          |
| `run_number`             | `IntField`               | Run index at that site, parsed from the `sequence_id` suffix                                                                    |
| `duration_s`             | `FloatField`             | Episode duration in seconds, computed from the MCAP’s message-time span                                                         |
| `message_count`          | `IntField`               | Total MCAP message count across all channels                                                                                    |
| `channel_count`          | `IntField`               | Total MCAP channel (topic) count — 18 or 19                                                                                     |
| `topics`                 | `ListField(StringField)` | Every MCAP topic present (see [MCAP topics](#mcap-topics-inside-each-episode))                                                  |
| `schemas`                | `ListField(StringField)` | Every distinct message schema present in the episode                                                                            |
| `has_image`              | `BooleanField`           | Has an Image-tile-decodable stream (`foxglove.CompressedImage` or `foxglove.RawImage`) — `True` for all                         |
| `has_pointcloud`         | `BooleanField`           | Has a 3D-tile-decodable point-cloud stream (`foxglove.PointCloud`) — `True` for all                                             |
| `has_gps`                | `BooleanField`           | Has a Map-tile GPS fix stream (`foxglove.LocationFix`) — `False` for all; no GNSS exists in the source                          |
| `has_imu`                | `BooleanField`           | Has IMU telemetry — `True` for all; hardcoded, because the stream is logged as generic JSON rather than a recognised IMU schema |
| `has_logs`               | `BooleanField`           | Has a Logs-tile stream (`foxglove.Log`) — `False` for all                                                                       |
| `has_trajectory`         | `BooleanField`           | Has a 6-DoF pose track (the `world` → `base` `FrameTransform` series) — `True` for all                                          |
| `has_camera_calibration` | `BooleanField`           | Has `foxglove.CameraCalibration` streams — `True` for all                                                                       |
| `has_depth`              | `BooleanField`           | Has per-camera 16-bit depth maps (`/depth_*`) — `True` for all                                                                  |
| `has_depth_overlay`      | `BooleanField`           | Has per-camera depth overlays on the camera image (`/overlay_*`) — `True` for all                                               |
| `has_surface_normals`    | `BooleanField`           | Has per-camera surface-normal maps (`/normal_*`) — `True` for all                                                               |
| `has_rgb_lidar`          | `BooleanField`           | Has the derived camera-colourised full-sweep cloud (`/lidar_rgb`) — `True` for 4 of 6 episodes                                  |
| `num_cameras`            | `IntField`               | Number of cameras in the rig — 3 for every episode                                                                              |

Standard FiftyOne bookkeeping fields (`id`, `tags`, `metadata`,
`created_at`, `last_modified_at`) are also present but not source-specific.

### MCAP topics (inside each episode)

Cameras use the devkit’s own labels from `configs/sensor.yaml`:
`cam_front`, `cam_left`, `cam_right` (source directories `cam0`, `cam1`,
`cam2` respectively).

| Topic(s)                                                                    | Schema                                                         | Tile           | Notes                                                                                                                                                                                                                                                                   |
|-----------------------------------------------------------------------------|----------------------------------------------------------------|----------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `/cam_front/image_raw`, `/cam_left/image_raw`, `/cam_right/image_raw`       | `foxglove.CompressedImage` (jpeg)                              | Image          | Raw, still-distorted fisheye JPEGs (1440×1080), byte-for-byte from the source `raw/cam{0,1,2}` folders                                                                                                                                                                  |
| `/cam_front/calibration`, `/cam_left/calibration`, `/cam_right/calibration` | `foxglove.CameraCalibration`                                   | 3D (frustums)  | Static, logged once at episode start; `equidistant` distortion model with `K`/`D` from `configs/sensor.yaml`                                                                                                                                                            |
| `/lidar`                                                                    | `foxglove.PointCloud`                                          | 3D             | VILENS-SLAM motion-undistorted keyframe clouds (≈1 Hz, one per SLAM pose-graph node), in frame `base`. Fields `x,y,z,intensity,normal_x,normal_y,normal_z,curvature`, or `x,y,z,red,green,blue,normal_*,curvature` for the two episodes whose source clouds ship colour |
| `/lidar_rgb` (4 of 6 episodes)                                              | `foxglove.PointCloud`                                          | 3D             | **Derived, not a devkit product.** The full sweep with RGB sampled by projecting every point into all three cameras; points no camera sees keep an intensity-derived grey. Fields `x,y,z,red,green,blue`                                                                |
| `/depth_cam_front`, `/depth_cam_left`, `/depth_cam_right`                   | `foxglove.RawImage` (`16UC1`)                                  | Image          | Euclidean depth × 256 as uint16, equivalent to the devkit’s `depths_euc_accum_0` output                                                                                                                                                                                 |
| `/overlay_cam_front`, `/overlay_cam_left`, `/overlay_cam_right`             | `foxglove.CompressedImage` (jpeg)                              | Image          | Depth-coloured LiDAR points drawn over the camera image (HSV colormap, radius-2 circles), equivalent to the devkit’s `*_overlay` output                                                                                                                                 |
| `/normal_cam_front`, `/normal_cam_left`, `/normal_cam_right`                | `foxglove.CompressedImage` (png)                               | Image          | Surface normals from the PCD `normal_*` fields, rotated into the camera frame and encoded `(n+1)/2·255`, equivalent to the devkit’s `normals_euc_accum_0` output                                                                                                        |
| `/tf`                                                                       | `foxglove.FrameTransform`                                      | 3D             | One static `base` → `lidar` transform at episode start, plus per SLAM keyframe: `world` → `base` and `world` → {`cam_front`, `cam_left`, `cam_right`}                                                                                                                   |
| `/imu`                                                                      | generic JSON (auto-named `jsonschema`, e.g. `schema-f5wrpOiP`) | Plot / Message | `acc_x/y/z` and `ang_vel_x/y/z` at the IMU’s native rate, from `raw/imu.csv`                                                                                                                                                                                            |

### Label types and why

**No FiftyOne sample-level label fields (`Detections`, `Keypoints`,
`Classification`, etc.) are attached.** Two reasons:

1. The source dataset has no object-level annotations at all — it is a raw
   multi-sensor SLAM/reconstruction benchmark. Its “ground truth” is TLS 3D
   models and 6-DoF trajectories, used to compute benchmark metrics, not
   per-sample labels.
2. Each sample is a continuous 3.5–13.7 minute recording rather than a
   single frame, so there is no single fixed-length list that a
   sample-level label field could hold.

The one label-like quantity that does exist per episode — the 6-DoF
trajectory — is therefore embedded as an MCAP transform stream (`/tf`)
inside the same timeline as the sensor data, decoded live by the multimodal
viewer, exactly like the sensor topics themselves. This also lets the
viewer place point clouds and camera frustums correctly in the world frame
during playback. It was logged as `foxglove.FrameTransform` rather than as
a static pose list because the viewer consumes transforms to resolve frames
at each timestamp.

All remaining per-sample fields are primitives, not labels: identifiers
(`sequence_id`, `location`, `recording_date`, `run_number`), MCAP inventory
(`duration_s`, `message_count`, `channel_count`, `topics`, `schemas`), and
boolean capability flags (`has_*`). The flags exist so episodes can be
filtered in the grid without opening every MCAP first, e.g.
`dataset.match(F("has_rgb_lidar"))`. They are derived from **schema names,
not topic substrings**, except `has_imu` (hardcoded `True`, since the IMU
rides on a generic JSON channel) and the `has_depth`/`has_depth_overlay`/
`has_surface_normals`/`has_rgb_lidar` flags (derived from topic prefixes,
since those products share the generic image/point-cloud schemas).

### Parsing decisions

- **One sample = one full episode.** The source Hugging Face repository is
  roughly 1.3 TB, so 6 of the 24 sequences were selected — one per site,
  choosing the smallest sequence at each site by download size — and each
  was authored as a complete recording rather than a short window.
- **Built from the per-file raw artifacts, not the ROS bags.** The source
  ships every sequence as both ROS 1 `.bag` and ROS 2 `.db3` (2–15 GB each)
  *and* as individual files. This build reads the individual files
  (`raw/images.zip`, `raw/imu.csv`,
  `processed/vilens-slam/undist-clouds.zip`,
  `processed/vilens-slam/slam-poses.csv`), which avoids the bag conversion
  entirely and gives the motion-undistorted, image-synchronised clouds the
  devkit’s own tooling expects.
- **Devkit frame naming.** Frames are `base`, `lidar`, `cam_front`,
  `cam_left`, `cam_right`, matching the labels in `configs/sensor.yaml`
  rather than inventing ROS-style names.
- **Cloud points are in the body frame, so they are logged as `base`.** The
  undistorted PCD files’ `VIEWPOINT` header equals the matching
  `slam-poses.csv` row, and the devkit’s own `undistort_sequence.py` writes
  an identity viewpoint — i.e. the point data itself is already expressed in
  the sensor/body frame at capture time. Verified directly: a mid-sequence
  cloud whose pose is 65 m from the origin has its point centroid at the
  origin.
- **Camera extrinsics follow the devkit’s own composition.**
  `T_base_cam = T_base_lidar @ inv(T_cam_lidar)`, the same expression used
  in the devkit’s `scripts/reconstruction_benchmark/main.py`, with the
  per-camera `T_cam_lidar` and rig `T_base_lidar` read from
  `configs/sensor.yaml`. Cameras are logged as direct time-varying
  `world` → `cam_*` transforms at each SLAM keyframe, so the viewer never
  has to chain through a single-sample static transform to resolve `world`.
- **Depth/overlay/normal maps are ports of the devkit’s own functions**
  (`oxspires_tools.depth.projection.encode_points_as_depthmap`,
  `depth.utils.get_overlay`, `depth.surface_normal.compute_normalmap`):
  euclidean depth scaled by 256 into uint16, nearest point winning per
  pixel, HSV colormap overlay with radius-2 circles, and normal maps that
  flip normals toward the camera and reserve `(128,128,128)` for empty
  pixels. They were reimplemented in numpy/OpenCV rather than called
  directly because the devkit’s projection module imports `open3d`, which
  has no wheel for this machine’s platform.
- **Hidden-point removal is skipped.** It is the one step of the devkit’s
  depth pipeline that genuinely requires `open3d`; the devkit itself exposes
  this as a supported `--skip_hpr` flag. For a single LiDAR keyframe sweep
  this mainly affects points on the far side of thin occluders.
- **The projection field-of-view cone is 160°**, taken from the devkit’s
  `camera_fov: 160.0` in `configs/sensor.yaml`, which is the value its own
  projection code uses to filter points before `cv2.fisheye.projectPoints`.
  Note this differs from the per-camera field of view quoted in the paper
  (126° × 92.4°); the devkit value was kept for fidelity to the devkit’s
  output.
- **Image-to-cloud pairing uses a 25 ms threshold**
  (`max_time_diff_camera_and_pose: 0.025` from `configs/sensor.yaml`),
  matching each cloud to the nearest image per camera. Cameras that fall
  outside the threshold for a given keyframe simply get no derived frame
  there, which is why per-camera depth/overlay/normal counts differ slightly
  within an episode.
- **`/lidar_rgb` is only built where the source clouds lack colour.** The
  `keble-college-02` and `observatory-quarter-02` undistorted clouds carry a
  packed PCL `rgb` field instead of `intensity` (colourised upstream by the
  dataset authors); for those, the colour is decoded into `red`/`green`/
  `blue` on `/lidar` itself and the derived `/lidar_rgb` is skipped as a
  strictly worse duplicate — camera projection can only colour the 56–84% of
  points that fall inside a camera’s view, whereas the upstream colour
  covers the whole sweep.
- **The IMU rides on a generic JSON channel**, not `foxglove.Imu`, because
  `imu.csv` provides only linear acceleration and angular velocity — no
  orientation or covariance. Consequence: it is inspectable in the Plot and
  Message tiles but is not reported under `schemas` as a recognised IMU
  schema, hence the hardcoded `has_imu`.
- **Real capture timestamps throughout.** Every message is logged at its
  source timestamp (Unix epoch nanoseconds) with no rebasing, so
  `duration_s` reflects the true recording span and all streams share one
  coherent clock.
- **Clouds without a matching SLAM pose are still logged.**
  `blenheim-palace-05` has 361 undistorted clouds but only 340 poses in
  `slam-poses.csv`; all 361 clouds appear on `/lidar`, while `/tf` carries
  340 keyframes, so the surplus clouds render at the last resolved
  transform.
- **Not included in this build:** the raw 10 Hz LiDAR sweeps
  (`raw/lidar-clouds.zip`), the COLMAP SfM outputs
  (`processed/colmap/`), the TLS ground-truth 3D models
  (`ground_truth_map/`), the TLS-registered ground-truth and refined
  trajectory variants (`gt-tum.txt`, `hba-tum.txt`, `colmap-tum.txt`), the
  TLS-rendered ground-truth depth images, and the reconstruction and
  novel-view-synthesis benchmark artifacts. Note also that the source
  repository does not currently ship `processed/trajectory/` or
  `processed/colmap/` for the New College sequences at all.

## Dataset Creation

### Curation Rationale

The full source repository is roughly 1.3 TB — every sequence ships raw
images, raw LiDAR sweeps, ROS 1 and ROS 2 bags, VILENS-SLAM outputs and
COLMAP outputs, on top of per-site TLS reference maps (about 76 GB) and
benchmark result artifacts. Exhaustive coverage is impractical for a
lightweight FiftyOne showcase, so this build optimises for **site
diversity at minimum download**: one sequence per landmark, choosing the
smallest available sequence at each site, giving all six architectural
settings the paper describes (Bodleian Library ≈37,000 m²; Christ Church
College ≈26,000 m²; Keble College ≈18,000 m²; New College ≈18,000 m²;
Blenheim Palace ≈14,000 m²; Radcliffe Observatory Quarter ≈12,000 m²).

Within each episode the priority was the opposite of trimming: episodes are
complete recordings, and the streams were chosen to reproduce what the
devkit itself visualises — hence the inclusion of all three
`generate_depth.py` products rather than depth maps alone.

### Source Data

#### Data Collection and Processing

Per the paper: each sequence was collected by walking with the **Frontier**
handheld perception unit mounted in a backpack. The unit carries three
colour fisheye cameras facing forward, left and right — a customised
Alphasense Core Development Kit from Sevensense Robotics AG — each
1440×1080 (1.6 MP) global shutter with a 126° × 92.4° field of view and
roughly 36° of overlap between adjacent cameras, running at 20 Hz with
auto-exposure enabled. A cellphone-grade IMU inside the Alphasense Core
runs at 400 Hz and is hardware-synchronised to the three cameras by a
Sevensense FPGA. A 64-channel Hesai QT64 LiDAR (10 Hz, 104° field of view,
60 m maximum range, ±3 cm typical accuracy) is mounted on top of the
cameras.

Synchronisation is both hardware and software: the Alphasense device clock
and the Hesai LiDAR are synchronised to the unit’s host computer using
Precision Time Protocol (sub-microsecond accuracy); the cameras’ exposure
intervals are aligned about their midpoints so the image triplets share one
timestamp; and because the QT64 scans continuously, each point cloud is
motion-corrected with IMU preintegration using VILENS and undistorted to the
time of the next camera frame. The result is that every node in the SLAM
pose graph has three camera images and one undistorted LiDAR cloud at an
identical timestamp.

Calibration used the equidistant (Kannala-Brandt) model for the fisheye
lenses: camera intrinsics and inter-camera extrinsics with Kalibr
(sub-pixel reprojection residuals, 0.22–0.23 px mean), IMU noise from an
eight-hour Allan variance sequence, per-camera camera-IMU extrinsics with
Kalibr, and a single SE(3) camera-bundle-to-LiDAR transform with DiffCal.

Reference data: a Leica RTC360 TLS (360° × 300° field of view, 130 m range,
1.9 mm point accuracy at 10 m and 5.3 mm at 40 m, colourised from 432
megapixel imagery) scanned each site; scans were registered with Leica
Cyclone REGISTER 360 Plus to 3–7 mm average cloud-to-cloud error and merged
into a 1 cm colourised map. Ground-truth trajectories were then produced by
ICP-registering each undistorted LiDAR cloud to that merged map with an
offline version of VILENS, reaching approximately 1–2 cm accuracy.
Processed outputs released alongside the raw data include the VILENS-SLAM
trajectory and undistorted clouds, and COLMAP SfM results computed over
images spaced 1 m apart (about 1 Hz at walking pace).

For this repackaging: the six sequences’ individual-file artifacts were
downloaded from the public Hugging Face repository, extracted, and packed
into one `.mcap` file per episode with the `foxglove-sdk`, computing the
depth/overlay/normal products with ports of the devkit’s own projection
code. Every episode was then verified against its own source files — image
counts per camera, IMU row count, cloud count, and expected transform count
all had to match exactly, with per-channel timestamp monotonicity checked —
before ingest. No sensor data was synthesized, relabeled, or altered beyond
the conversions documented in [Parsing decisions](#parsing-decisions).

#### Who are the source data producers?

The Oxford Robotics Institute, Department of Engineering Science,
University of Oxford, with the Group of Automation, Robotics and Computer
Vision (AUROVA) at the University of Alicante — specifically Yifu Tao,
Miguel Ángel Muñoz-Bañón, Lintong Zhang, Jiahao Wang, Lanke Frank Tarimo
Fu, and Maurice Fallon. The paper additionally acknowledges Tobit
Flatscher, Ayoung Kim, Matias Mattamala, Christina Kassab, Haedam Oh,
Jianeng Wang, and Dongjae Lee for help with sensing, calibration,
collection, post-processing and proofreading.

### Annotations

#### Annotation process

There is no annotation process: the Oxford Spires Dataset contains no
manually labelled, object-level annotations. What the source dataset calls
ground truth is produced automatically — reconstruction ground truth is the
registered and merged Leica RTC360 TLS point cloud per site, and
localisation ground truth is the trajectory obtained by ICP-registering each
motion-undistorted LiDAR cloud against that TLS map using an offline version
of VILENS. Neither of those artifacts is included in this build; the pose
track logged here is the VILENS-SLAM *estimate* from `slam-poses.csv` (see
[Out-of-Scope Use](#out-of-scope-use)).

#### Who are the annotators?

No human annotators were involved. All ground-truth quantities in the source
dataset are the output of automated survey-grade scanning, registration, and
SLAM/ICP post-processing performed by the dataset authors.

#### Personal and Sensitive Information

The recordings were made in and around public and semi-public university and
palace grounds in Oxford, so incidental pedestrians can appear in the camera
imagery. The ROS bag filenames in the source repository contain
`blurred_filtered` (e.g.
`..._blurred_filtered_compressed.db3`), which indicates that a privacy
filtering step — presumably face and licence-plate blurring — was applied
upstream to the bag streams before public release.

[More Information Needed] on the specifics: neither the paper nor the devkit
documents an anonymisation procedure or tool, and it is not documented
whether the `raw/images.zip` JPEG archives used to build these episodes
carry the same blurring as the bag streams. This repackaging performs no
additional processing, re-identification, or redaction beyond what the
Oxford Robotics Institute already released publicly.

## Citation

**BibTeX:**

```bibtex
@article{tao2025spires,
  title   = {The Oxford Spires Dataset: Benchmarking Large-Scale
             LiDAR-Visual Localisation, Reconstruction and Radiance Field
             Methods},
  author  = {Tao, Yifu and Mu{\~n}oz-Ba{\~n}{\'o}n, Miguel {\'A}ngel and
             Zhang, Lintong and Wang, Jiahao and Fu, Lanke Frank Tarimo and
             Fallon, Maurice},
  journal = {The International Journal of Robotics Research},
  year    = {2025},
  doi     = {10.1177/02783649251369905}
}
```

**APA:**

Tao, Y., Muñoz-Bañón, M. Á., Zhang, L., Wang, J., Fu, L. F. T., & Fallon,
M. (2025). The Oxford Spires Dataset: Benchmarking large-scale LiDAR-visual
localisation, reconstruction and radiance field methods. *The International
Journal of Robotics Research*.

## More Information

This repository is an independently-curated, derived subset of the official
Oxford Spires Dataset, repackaged as MCAP for FiftyOne’s multimodal
support. It is not an official Oxford Robotics Institute artifact, and it is
subject to the source dataset’s non-commercial CC BY-NC-SA 4.0 licence.

For the full dataset (all 24 sequences, raw LiDAR sweeps, ROS 1/ROS 2 bags,
COLMAP outputs, per-site TLS reference maps, TLS-registered ground-truth
trajectories, rendered ground-truth depth images, and the localisation /
reconstruction / novel-view-synthesis benchmark artifacts and evaluation
code), see:

- https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/
- https://huggingface.co/datasets/ori-drs/oxford_spires_dataset
- https://github.com/ori-drs/oxford_spires_dataset and its
  [wiki](https://github.com/ori-drs/oxford_spires_dataset/wiki)

Viewing these episodes requires FiftyOne 1.19 or newer for multimodal media
support. The 6 MCAP files total 37 GB.

## Dataset Card Authors

Harpreet Sahota ([@harpreetsahota](https://huggingface.co/harpreetsahota))
— MCAP repackaging and this card. Original dataset producers are listed
under [Dataset Description](#dataset-description).

## Dataset Card Contact

Harpreet Sahota — https://huggingface.co/harpreetsahota
