<!-- # hard line break macro for HTML -->

<a id="evaluating-segmentations"></a>

# Evaluating Segmentations


<div class="available-in">
    <div class="available-in-row">
        <span class="available-in-label">Available in:</span>
        <span class="available-in-pill available-in-pill--oss">Open Source</span><span class="available-in-pill available-in-pill--enterprise">Enterprise</span>
    </div>
    <div class="available-in-row">
        <span class="available-in-versions">Introduced in <a href="../../release-notes.html#fiftyone-0-7-3">FiftyOne 0.7.3</a> &middot; <a href="../../release-notes.html#fiftyone-enterprise-1-0">FiftyOne Enterprise 1.0</a></span>
    </div>
    
</div>

You can use the
[`evaluate_segmentations()`](../../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.evaluate_segmentations)
method to evaluate the predictions of a semantic segmentation model stored in a
[`Segmentation`](../../api/fiftyone.core.labels.md#fiftyone.core.labels.Segmentation) field of your dataset.

By default, the full segmentation masks will be evaluated at a pixel level, but
you can specify other evaluation strategies such as evaluating only boundary
pixels (see below for details).

Invoking
[`evaluate_segmentations()`](../../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.evaluate_segmentations)
returns a [`SegmentationResults`](../../api/fiftyone.utils.eval.segmentation.md#fiftyone.utils.eval.segmentation.SegmentationResults) instance that provides a variety of methods for
generating various aggregate evaluation reports about your model.

In addition, when you specify an `eval_key` parameter, a number of helpful
fields will be populated on each sample that you can leverage via the
[FiftyOne App](../app.md#fiftyone-app) to interactively explore the strengths and
weaknesses of your model on individual samples.

#### NOTE
You can [store mask targets](../using_datasets.md#storing-mask-targets) for your
[`Segmentation`](../../api/fiftyone.core.labels.md#fiftyone.core.labels.Segmentation) fields on your dataset so that you can view semantic labels
in the App and avoid having to manually specify the set of mask targets
each time you run
[`evaluate_segmentations()`](../../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.evaluate_segmentations)
on a dataset.

## Simple evaluation (default)

By default,
[`evaluate_segmentations()`](../../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.evaluate_detections)
will perform pixelwise evaluation of the segmentation masks, treating each
pixel as a multiclass classification.

Here are some things to keep in mind:

- If the size of a predicted mask does not match the ground truth mask, it is
  resized to match the ground truth.
- You can specify the optional `bandwidth` parameter to evaluate only along
  the contours of the ground truth masks. By default, the entire masks are
  evaluated.

You can explicitly request that this strategy be used by setting the `method`
parameter to `"simple"`.

When you specify an `eval_key` parameter, the accuracy, precision, and recall
of each sample is recorded in top-level fields of each sample:

```text
 Accuracy: sample.<eval_key>_accuracy
Precision: sample.<eval_key>_precision
   Recall: sample.<eval_key>_recall
```

#### NOTE
The mask values `0` and `#000000` are treated as a background class
for the purposes of computing evaluation metrics like precision and
recall.

The example below demonstrates segmentation evaluation by comparing the
masks generated by two DeepLabv3 models (with
[ResNet50](../../model_zoo/models/deeplabv3_resnet50_coco_torch.md#model-zoo-deeplabv3-resnet50-coco-torch) and
[ResNet101](../../model_zoo/models/deeplabv3_resnet101_coco_torch.md#model-zoo-deeplabv3-resnet101-coco-torch) backbones):

```python
import fiftyone as fo
import fiftyone.zoo as foz

# Load a few samples from COCO-2017
dataset = foz.load_zoo_dataset(
    "quickstart",
    dataset_name="segmentation-eval-demo",
    max_samples=10,
    shuffle=True,
)

# The models are trained on the VOC classes
CLASSES = (
    "background,aeroplane,bicycle,bird,boat,bottle,bus,car,cat,chair,cow," +
    "diningtable,dog,horse,motorbike,person,pottedplant,sheep,sofa,train," +
    "tvmonitor"
)
dataset.default_mask_targets = {
    idx: label for idx, label in enumerate(CLASSES.split(","))
}

# Add DeepLabv3-ResNet101 predictions to dataset
model = foz.load_zoo_model("deeplabv3-resnet101-coco-torch")
dataset.apply_model(model, "resnet101")

# Add DeepLabv3-ResNet50 predictions to dataset
model = foz.load_zoo_model("deeplabv3-resnet50-coco-torch")
dataset.apply_model(model, "resnet50")

print(dataset)

# Evaluate the masks w/ ResNet50 backbone, treating the masks w/ ResNet101
# backbone as "ground truth"
results = dataset.evaluate_segmentations(
    "resnet50",
    gt_field="resnet101",
    eval_key="eval_simple",
)

# Get a sense for the per-sample variation in likeness
print("Accuracy range: (%f, %f)" % dataset.bounds("eval_simple_accuracy"))
print("Precision range: (%f, %f)" % dataset.bounds("eval_simple_precision"))
print("Recall range: (%f, %f)" % dataset.bounds("eval_simple_recall"))

# Print a classification report
results.print_report()

# Visualize results in the App
session = fo.launch_app(dataset)
```

![evaluate-segmentations](images/evaluation/evaluate_segmentations.gif)

#### NOTE
The easiest way to analyze models in FiftyOne is via the
[Model Evaluation panel](../app.md#app-model-evaluation-panel)!
