<!-- # hard line break macro for HTML -->

<a id="dataset-zoo-coco-2017"></a>

# COCO-2017


<div class="available-in">
    <div class="available-in-row">
        <span class="available-in-label">Available in:</span>
        <span class="available-in-pill available-in-pill--oss">Open Source</span><span class="available-in-pill available-in-pill--enterprise">Enterprise</span>
    </div>
    <div class="available-in-row">
        <span class="available-in-versions">Introduced in <a href="../../release-notes.html#fiftyone-0-3-0">FiftyOne 0.3.0</a> &middot; <a href="../../release-notes.html#fiftyone-enterprise-1-0">FiftyOne Enterprise 1.0</a></span>
    </div>
    
</div>

COCO is a large-scale object detection, segmentation, and captioning
dataset.

This version contains images, bounding boxes, and segmentations for the 2017
version of the dataset.

#### NOTE
With support from the [COCO team](https://cocodataset.org/#download),
FiftyOne is a recommended tool for downloading, visualizing, and evaluating
on the COCO dataset!

Check out [this guide](../../integrations/coco.md#coco) for more details on using FiftyOne to
work with COCO.

**Notes**

- COCO defines 91 classes but the data only uses 80 classes
- Some images from the train and validation sets don’t have annotations
- The test set does not have annotations
- COCO 2014 and 2017 use the same images, but the splits are different

**Details**

- Dataset name: `coco-2017`
- Dataset source: [http://cocodataset.org/#home](http://cocodataset.org/#home)
- Dataset license: CC-BY-4.0
- Dataset size: 25.20 GB
- Tags: `image, detection, segmentation`
- Supported splits: `train, validation, test`
- ZooDataset class:
  [`COCO2017Dataset`](../../api/fiftyone.zoo.datasets.base.md#fiftyone.zoo.datasets.base.COCO2017Dataset)

**Full split stats**

- Train split: 118,287 images
- Test split: 40,670 images
- Validation split: 5,000 images

**Partial downloads**

FiftyOne provides parameters that can be used to efficiently download specific
subsets of the COCO dataset to suit your needs. When new subsets are specified,
FiftyOne will use existing downloaded data first if possible before resorting
to downloading additional data from the web.

The following parameters are available to configure a partial download of
COCO-2017 by passing them to
[`load_zoo_dataset()`](../../api/fiftyone.zoo.datasets.md#fiftyone.zoo.datasets.load_zoo_dataset):

- **split** (*None*) and **splits** (*None*): a string or list of strings,
  respectively, specifying the splits to load. Supported values are
  `("train", "test", "validation")`. If neither is provided, all available
  splits are loaded
- **label_types** (*None*): a label type or list of label types to load.
  Supported values are `("detections", "segmentations")`. By default, only
  detections are loaded
- **classes** (*None*): a string or list of strings specifying required
  classes to load. If provided, only samples containing at least one instance
  of a specified class will be loaded
- **image_ids** (*None*): a list of specific image IDs to load. The IDs can
  be specified either as `<split>/<image-id>` strings or `<image-id>`
  ints or strings. Alternatively, you can provide the path to a TXT
  (newline-separated), JSON, or CSV file containing the list of image IDs to
  load in either of the first two formats
- **include_id** (*False*): whether to include the COCO ID of each sample in
  the loaded labels
- **include_license** (*False*): whether to include the COCO license of each
  sample in the loaded labels, if available. The supported values are:
  - `"False"` (default): don’t load the license
  - `True`/`"name"`: store the string license name
  - `"id"`: store the integer license ID
  - `"url"`: store the license URL
- **only_matching** (*False*): whether to only load labels that match the
  `classes` or `attrs` requirements that you provide (True), or to load
  all labels for samples that match the requirements (False)
- **num_workers** (*None*): the number of processes to use when downloading
  individual images. By default, `multiprocessing.cpu_count()` is used
- **shuffle** (*False*): whether to randomly shuffle the order in which
  samples are chosen for partial downloads
- **seed** (*None*): a random seed to use when shuffling
- **max_samples** (*None*): a maximum number of samples to load per split. If
  `label_types` and/or `classes` are also specified, first priority will
  be given to samples that contain all of the specified label types and/or
  classes, followed by samples that contain at least one of the specified
  labels types or classes. The actual number of samples loaded may be less
  than this maximum value if the dataset does not contain sufficient samples
  matching your requirements

#### NOTE
See
[`COCO2017Dataset`](../../api/fiftyone.zoo.datasets.base.md#fiftyone.zoo.datasets.base.COCO2017Dataset) and
[`COCODetectionDatasetImporter`](../../api/fiftyone.utils.coco.md#fiftyone.utils.coco.COCODetectionDatasetImporter)
for complete descriptions of the optional keyword arguments that you can
pass to [`load_zoo_dataset()`](../../api/fiftyone.zoo.datasets.md#fiftyone.zoo.datasets.load_zoo_dataset).

**Example usage**

Python

CLI

```python
import fiftyone as fo
import fiftyone.zoo as foz

#
# Load 50 random samples from the validation split
#
# Only the required images will be downloaded (if necessary).
# By default, only detections are loaded
#

dataset = foz.load_zoo_dataset(
    "coco-2017",
    split="validation",
    max_samples=50,
    shuffle=True,
)

session = fo.launch_app(dataset)

#
# Load segmentations for 25 samples from the validation split that
# contain cats and dogs
#
# Images that contain all `classes` will be prioritized first, followed
# by images that contain at least one of the required `classes`. If
# there are not enough images matching `classes` in the split to meet
# `max_samples`, only the available images will be loaded.
#
# Images will only be downloaded if necessary
#

dataset = foz.load_zoo_dataset(
    "coco-2017",
    split="validation",
    label_types=["segmentations"],
    classes=["cat", "dog"],
    max_samples=25,
)

session.dataset = dataset

#
# Download the entire validation split and load both detections and
# segmentations
#
# Subsequent partial loads of the validation split will never require
# downloading any images
#

dataset = foz.load_zoo_dataset(
    "coco-2017",
    split="validation",
    label_types=["detections", "segmentations"],
)

session.dataset = dataset
```

```shell
#
# Load 50 random samples from the validation split
#
# Only the required images will be downloaded (if necessary).
# By default, only detections are loaded
#

fiftyone zoo datasets load coco-2017 \
    --split validation \
    --kwargs \
        max_samples=50

fiftyone app launch coco-2017-validation-50

#
# Load segmentations for 25 samples from the validation split that
# contain cats and dogs
#
# Images that contain all `classes` will be prioritized first, followed
# by images that contain at least one of the required `classes`. If
# there are not enough images matching `classes` in the split to meet
# `max_samples`, only the available images will be loaded.
#
# Images will only be downloaded if necessary
#

fiftyone zoo datasets load coco-2017 \
    --split validation \
    --kwargs \
        label_types=segmentations \
        classes=cat,dog \
        max_samples=25

fiftyone app launch coco-2017-validation-25

#
# Download the entire validation split and load both detections and
# segmentations
#
# Subsequent partial loads of the validation split will never require
# downloading any images
#

fiftyone zoo datasets load coco-2017 \
    --split validation \
    --kwargs \
        label_types=detections,segmentations

fiftyone app launch coco-2017-validation
```

![coco-2017-validation](images/dataset_zoo/coco-2017-validation.png)
