<!-- # hard line break macro for HTML -->

<a id="brain-similarity"></a>

# Similarity Search


<div class="available-in">
    <div class="available-in-row">
        <span class="available-in-label">Available in:</span>
        <span class="available-in-pill available-in-pill--oss">Open Source</span><span class="available-in-pill available-in-pill--enterprise">Enterprise</span>
    </div>
    <div class="available-in-row">
        <span class="available-in-versions">Introduced in <a href="../release-notes.html#fiftyone-0-9-0">FiftyOne 0.9.0</a> &middot; <a href="../release-notes.html#fiftyone-enterprise-1-0">FiftyOne Enterprise 1.0</a></span>
    </div>
    
</div>

The FiftyOne Brain provides a
[`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity) method that
you can use to index the images or object patches in a dataset by similarity.

Once you’ve indexed a dataset by similarity, you can use the
[`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
view stage to programmatically sort your dataset by similarity to any image(s)
or object patch(es) of your choice in your dataset. In addition, the App
provides a convenient [point-and-click interface](app.md#app-similarity) for
sorting by similarity with respect to an index on a dataset.

#### NOTE
Did you know? You can
[search by natural language](#brain-similarity-text) using similarity
indexes!

## Embedding methods

Like [embeddings visualization](../brain/index.md#brain-embeddings-visualization),
similarity leverages deep embeddings to generate an index for a dataset.

The `embeddings` and `model` parameters of
[`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity) support a
variety of ways to generate embeddings for your data:

- Provide nothing, in which case a default general purpose model is used to
  index your data
- Provide a [`Model`](../api/fiftyone.core.models.md#fiftyone.core.models.Model) instance or the name of any model from the
  [Model Zoo](../model_zoo/index.md#model-zoo) that supports embeddings
- Provide your own precomputed embeddings in array form
- Provide the name of a [`VectorField`](../api/fiftyone.core.fields.md#fiftyone.core.fields.VectorField) or [`ArrayField`](../api/fiftyone.core.fields.md#fiftyone.core.fields.ArrayField) of your dataset in
  which precomputed embeddings are stored

<a id="brain-similarity-backends"></a>

## Similarity backends

By default, all similarity indexes are served using a builtin
[scikit-learn](https://scikit-learn.org) backend, but you can pass the
optional `backend` parameter to
[`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity) to switch to
another supported backend:

- **sklearn** (*default*): a [scikit-learn](https://scikit-learn.org) backend
- **qdrant**: a [Qdrant backend](../integrations/qdrant.md#qdrant-integration)
- **redis**: a [Redis backend](../integrations/redis.md#redis-integration)
- **pinecone**: a [Pinecone backend](../integrations/pinecone.md#pinecone-integration)
- **mongodb**: a [MongoDB backend](../integrations/mongodb.md#mongodb-integration)
- **elasticsearch**: a [Elasticsearch backend](../integrations/elasticsearch.md#elasticsearch-integration)
- **pgvector**: a [PostgreSQL Pgvector backend](../integrations/pgvector.md#pgvector-integration)
- **mosaic**: a [Databricks Mosaic AI backend](../integrations/mosaic.md#mosaic-integration)
- **milvus**: a [Milvus backend](../integrations/milvus.md#milvus-integration)
- **lancedb**: a [LanceDB backend](../integrations/lancedb.md#lancedb-integration)

```python
import fiftyone.brain as fob

results = fob.compute_similarity(
    dataset,
    backend="sklearn",  # "sklearn", "qdrant", "redis", etc
    brain_key="...",
    ...
)
```

#### NOTE
Refer to [this section](#brain-similarity-api) for more information
about creating, managing and deleting similarity indexes.

<a id="brain-image-similarity"></a>

## Image similarity

This section demonstrates the basic workflow of:

- Indexing an image dataset by similarity
- Using the App’s [image similarity](app.md#app-image-similarity) UI to query
  by visual similarity
- Using the SDK’s
  [`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
  view stage to programmatically query the index

To index a dataset by image similarity, pass the [`Dataset`](../api/fiftyone.core.dataset.md#fiftyone.core.dataset.Dataset) or [`DatasetView`](../api/fiftyone.core.view.md#fiftyone.core.view.DatasetView) of
interest to [`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity)
along with a name for the index via the `brain_key` argument.

Next load the dataset in the App and select some image(s). Whenever there is
an active selection in the App, a [similarity icon](app.md#app-image-similarity)
will appear above the grid, enabling you to sort by similarity to your current
selection.

You can use the [Similarity Search panel](app.md#app-similarity-search-panel) for
advanced search options, run management, and search history.

```python
import fiftyone as fo
import fiftyone.brain as fob
import fiftyone.zoo as foz

dataset = foz.load_zoo_dataset("quickstart")

# Index images by similarity
fob.compute_similarity(
    dataset,
    model="clip-vit-base32-torch",
    brain_key="img_sim",
)

session = fo.launch_app(dataset)
```

#### NOTE
In the example above, we specify a [zoo model](../model_zoo/index.md#model-zoo) with which
to generate embeddings, but you can also provide
[precomputed embeddings](#brain-similarity-api).

![image-similarity](images/brain/brain-image-similarity.gif)

Alternatively, you can use the
[`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
view stage to programmatically [construct a view](using_views.md#using-views) that
contains the sorted results:

```python
# Choose a random image from the dataset
query_id = dataset.take(1).first().id

# Programmatically construct a view containing the 15 most similar images
view = dataset.sort_by_similarity(query_id, k=15, brain_key="img_sim")

session.view = view
```

#### NOTE
Performing a similarity search on a [`DatasetView`](../api/fiftyone.core.view.md#fiftyone.core.view.DatasetView) will **only** return
results from the view; if the view contains samples that were not included
in the index, they will never be included in the result.

This means that you can index an entire [`Dataset`](../api/fiftyone.core.dataset.md#fiftyone.core.dataset.Dataset) once and then perform
searches on subsets of the dataset by
[constructing views](using_views.md#using-views) that contain the images of
interest.

#### NOTE
For large datasets, you may notice longer load times the first time you use
a similarity index in a session. Subsequent similarity searches will use
cached results and will be faster!

<a id="brain-object-similarity"></a>

## Object similarity

This section demonstrates the basic workflow of:

- Indexing a dataset of objects by similarity
- Using the App’s [object similarity](app.md#app-object-similarity) UI to
  query by visual similarity
- Using the SDK’s
  [`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
  view stage to programmatically query the index

You can index any objects stored on datasets in [`Detection`](../api/fiftyone.core.labels.md#fiftyone.core.labels.Detection), [`Detections`](../api/fiftyone.core.labels.md#fiftyone.core.labels.Detections),
[`Polyline`](../api/fiftyone.core.labels.md#fiftyone.core.labels.Polyline), or [`Polylines`](../api/fiftyone.core.labels.md#fiftyone.core.labels.Polylines) format. See [this section](using_datasets.md#using-labels) for
more information about adding labels to your datasets.

To index by object patches, simply pass the [`Dataset`](../api/fiftyone.core.dataset.md#fiftyone.core.dataset.Dataset) or [`DatasetView`](../api/fiftyone.core.view.md#fiftyone.core.view.DatasetView) of
interest to [`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity)
along with the name of the patches field and a name for the index via the
`brain_key` argument.

Next load the dataset in the App and switch to
[object patches view](app.md#app-object-patches) by clicking the patches icon
above the grid and choosing the label field of interest from the dropdown.

Now whenever you have selected one or more patches in the App, a
[similarity icon](app.md#app-object-similarity) will appear above the grid,
enabling you to sort by similarity to your current selection.

You can also use the [Similarity Search panel](app.md#app-similarity-search-panel) for
advanced search options, run management, and search history.

```python
import fiftyone as fo
import fiftyone.brain as fob
import fiftyone.zoo as foz

dataset = foz.load_zoo_dataset("quickstart")

# Index ground truth objects by similarity
fob.compute_similarity(
    dataset,
    patches_field="ground_truth",
    model="clip-vit-base32-torch",
    brain_key="gt_sim",
)

session = fo.launch_app(dataset)
```

#### NOTE
In the example above, we specify a [zoo model](../model_zoo/index.md#model-zoo) with which
to generate embeddings, but you can also provide
[precomputed embeddings](#brain-similarity-api).

![object-similarity](images/brain/brain-object-similarity.gif)

Alternatively, you can directly use the
[`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
view stage to programmatically [construct a view](using_views.md#using-views) that
contains the sorted results:

```python
# Convert to patches view
patches = dataset.to_patches("ground_truth")

# Choose a random patch object from the dataset
query_id = patches.take(1).first().id

# Programmatically construct a view containing the 15 most similar objects
view = patches.sort_by_similarity(query_id, k=15, brain_key="gt_sim")

session.view = view
```

#### NOTE
Performing a similarity search on a [`DatasetView`](../api/fiftyone.core.view.md#fiftyone.core.view.DatasetView) will **only** return
results from the view; if the view contains objects that were not included
in the index, they will never be included in the result.

This means that you can index an entire [`Dataset`](../api/fiftyone.core.dataset.md#fiftyone.core.dataset.Dataset) once and then perform
searches on subsets of the dataset by
[constructing views](using_views.md#using-views) that contain the objects of
interest.

#### NOTE
For large datasets, you may notice longer load times the first time you use
a similarity index in a session. Subsequent similarity searches will use
cached results and will be faster!

<a id="brain-similarity-text"></a>

## Text similarity


<div class="available-in">
    <div class="available-in-row">
        <span class="available-in-label">Available in:</span>
        <span class="available-in-pill available-in-pill--oss">Open Source</span><span class="available-in-pill available-in-pill--enterprise">Enterprise</span>
    </div>
    <div class="available-in-row">
        <span class="available-in-versions">Introduced in <a href="../release-notes.html#fiftyone-0-20-0">FiftyOne 0.20.0</a> &middot; <a href="../release-notes.html#fiftyone-enterprise-1-2">FiftyOne Enterprise 1.2</a></span>
    </div>
    
</div>

When you create a similarity index powered by the
[CLIP model](../model_zoo/models/clip_vit_base32_torch.md#model-zoo-clip-vit-base32-torch), you can also search by
arbitrary natural language queries
[natively in the App](app.md#app-text-similarity), including via the
[Similarity Search panel](app.md#app-similarity-search-panel)!

Image similarity

Object similarity

```python
import fiftyone as fo
import fiftyone.brain as fob
import fiftyone.zoo as foz

dataset = foz.load_zoo_dataset("quickstart")

# Index images by similarity
image_index = fob.compute_similarity(
    dataset,
    model="clip-vit-base32-torch",
    brain_key="img_sim",
)

session = fo.launch_app(dataset)
```

You can verify that an index supports text queries by checking that it
`supports_prompts`:

```python
# If you have already loaded the index
print(image_index.config.supports_prompts)  # True

# Without loading the index
info = dataset.get_brain_info("img_sim")
print(info.config.supports_prompts)  # True
```

```python
import fiftyone as fo
import fiftyone.brain as fob
import fiftyone.zoo as foz

dataset = foz.load_zoo_dataset("quickstart")

# Index ground truth objects by similarity
object_index = fob.compute_similarity(
    dataset,
    patches_field="ground_truth",
    model="clip-vit-base32-torch",
    brain_key="gt_sim",
)

session = fo.launch_app(dataset)
```

You can verify that an index supports text queries by checking that it
`supports_prompts`:

```python
# If you have already loaded the index
print(object_index.config.supports_prompts)  # True

# Without loading the index
info = dataset.get_brain_info("gt_sim")
print(info.config.supports_prompts)  # True
```

![text-similarity](images/brain/brain-text-similarity.gif)

You can also perform text queries via the SDK by passing a prompt directly to
[`sort_by_similarity()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.sort_by_similarity)
along with the `brain_key` of a compatible similarity index:

Image similarity

Object similarity

```python
# Perform a text query
query = "kites high in the air"
view = dataset.sort_by_similarity(query, k=15, brain_key="img_sim")

session.view = view
```

```python
# Convert to patches view
patches = dataset.to_patches("ground_truth")

# Perform a text query
query = "cute puppies"
view = patches.sort_by_similarity(query, k=15, brain_key="gt_sim")

session.view = view
```

#### NOTE
In general, any custom model that is made available via the
[model zoo interface](../model_zoo/api.md#model-zoo-add) that implements the
[`PromptMixin`](../api/fiftyone.core.models.md#fiftyone.core.models.PromptMixin) interface can
support text similarity queries!

#### NOTE
Pass the model’s **name** rather than a model instance when creating an
index that you want to query by text. An index created from a
[`Model`](../api/fiftyone.core.models.md#fiftyone.core.models.Model) instance has no way to reload
that model later, so its text queries will not work in the App or in future
Python sessions:

```python
# Text queries work in the App and in future sessions
fob.compute_similarity(
    dataset, model="clip-vit-base32-torch", brain_key="sim"
)

# Text queries work only in the current session
model = foz.load_zoo_model("clip-vit-base32-torch")
fob.compute_similarity(dataset, model=model, brain_key="sim")
```

#### NOTE
You can provide precomputed `embeddings` **and** a `model` name, which
retains text query support for embeddings that you generated yourself:

```python
fob.compute_similarity(
    dataset,
    embeddings=embeddings,  # precomputed
    model="clip-vit-base32-torch",
    brain_key="sim",
)
```

This is the supported path for data whose embeddings cannot be generated by
applying a model to the samples’ media, since
[`Model`](../api/fiftyone.core.models.md#fiftyone.core.models.Model) declares a `media_type` of
either `"image"` or `"video"`.

<a id="brain-similarity-api"></a>

## Similarity API

This section describes how to setup, create, and manage similarity indexes in
detail.

### Changing your similarity backend

You can use a specific backend for a particular similarity index by passing the
`backend` parameter to
[`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity):

```python
index = fob.compute_similarity(..., backend="<backend>", ...)
```

Alternatively, you can change your default similarity backend for an entire
session by setting the `FIFTYONE_BRAIN_DEFAULT_SIMILARITY_BACKEND` environment
variable.

```shell
export FIFTYONE_BRAIN_DEFAULT_SIMILARITY_BACKEND=<backend>
```

Finally, you can permanently change your default similarity backend by
updating the `default_similarity_backend` key of your
[brain config](../brain/index.md#brain-config) at `~/.fiftyone/brain_config.json`:

```text
{
    "default_similarity_backend": "<backend>",
    "similarity_backends": {
        "<backend>": {...},
        ...
    }
}
```

### Configuring your backend

Similarity backends may be configured in a variety of backend-specific ways,
which you can see by inspecting the parameters of a backend’s associated
[`SimilarityConfig`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityConfig) class.

The relevant classes for the builtin similarity backends are:

- **sklearn**: [`fiftyone.brain.internal.core.sklearn.SklearnSimilarityConfig`](../api/fiftyone.brain.internal.core.sklearn.md#fiftyone.brain.internal.core.sklearn.SklearnSimilarityConfig)
- **qdrant**: `fiftyone.brain.internal.core.qdrant.QdrantSimilarityConfig`
- **redis**: [`fiftyone.brain.internal.core.redis.RedisSimilarityConfig`](../api/fiftyone.brain.internal.core.redis.md#fiftyone.brain.internal.core.redis.RedisSimilarityConfig)
- **pinecone**: [`fiftyone.brain.internal.core.pinecone.PineconeSimilarityConfig`](../api/fiftyone.brain.internal.core.pinecone.md#fiftyone.brain.internal.core.pinecone.PineconeSimilarityConfig)
- **mongodb**: [`fiftyone.brain.internal.core.mongodb.MongoDBSimilarityConfig`](../api/fiftyone.brain.internal.core.mongodb.md#fiftyone.brain.internal.core.mongodb.MongoDBSimilarityConfig)
- **elasticsearch**: [`fiftyone.brain.internal.core.elasticsearch.ElasticsearchSimilarityConfig`](../api/fiftyone.brain.internal.core.elasticsearch.md#fiftyone.brain.internal.core.elasticsearch.ElasticsearchSimilarityConfig)
- **pgvector**: [`fiftyone.brain.internal.core.pgvector.PgVectorSimilarityConfig`](../api/fiftyone.brain.internal.core.pgvector.md#fiftyone.brain.internal.core.pgvector.PgVectorSimilarityConfig)
- **mosaic**: [`fiftyone.brain.internal.core.mosaic.MosaicSimilarityConfig`](../api/fiftyone.brain.internal.core.mosaic.md#fiftyone.brain.internal.core.mosaic.MosaicSimilarityConfig)
- **milvus**: [`fiftyone.brain.internal.core.milvus.MilvusSimilarityConfig`](../api/fiftyone.brain.internal.core.milvus.md#fiftyone.brain.internal.core.milvus.MilvusSimilarityConfig)
- **lancedb**: [`fiftyone.brain.internal.core.lancedb.LanceDBSimilarityConfig`](../api/fiftyone.brain.internal.core.lancedb.md#fiftyone.brain.internal.core.lancedb.LanceDBSimilarityConfig)

You can configure a similarity backend’s parameters for a specific index by
simply passing supported config parameters as keyword arguments each time you
call [`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity):

```python
index = fob.compute_similarity(
    ...
    backend="qdrant",
    url="http://localhost:6333",
)
```

Alternatively, you can more permanently configure your backend(s) via your
[brain config](../brain/index.md#brain-config).

### Creating an index

The [`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity) method
provides a number of different syntaxes for initializing a similarity index.
Let’s see some common patterns on the quickstart dataset:

```python
import fiftyone as fo
import fiftyone.brain as fob
import fiftyone.zoo as foz

dataset = foz.load_zoo_dataset("quickstart")
```

#### Default behavior

With no arguments, embeddings will be automatically computed for all images or
patches in the dataset using a default model and added to a new index in your
default backend:

Image similarity

Object similarity

```python
tmp_index = fob.compute_similarity(dataset, brain_key="tmp")

print(tmp_index.config.method)  # 'sklearn'
print(tmp_index.config.model)  # 'mobilenet-v2-imagenet-torch'
print(tmp_index.total_index_size)  # 200

dataset.delete_brain_run("tmp")
```

```python
tmp_index = fob.compute_similarity(
    dataset,
    patches_field="ground_truth",   # field containing objects of interest
    brain_key="tmp",
)

print(tmp_index.config.method)  # 'sklearn'
print(tmp_index.config.model)  # 'mobilenet-v2-imagenet-torch'
print(tmp_index.total_index_size)  # 1232

dataset.delete_brain_run("tmp")
```

#### Custom model, custom backend, add embeddings later

With the syntax below, we’re specifying a similarity backend of our choice,
specifying a custom model from the [Model Zoo](../model_zoo/index.md#model-zoo) to use to
generate embeddings, and using the `embeddings=False` syntax to create
the index without initially adding any embeddings to it:

Image similarity

Object similarity

```python
image_index = fob.compute_similarity(
    dataset,
    model="clip-vit-base32-torch",  # custom model
    embeddings=False,               # add embeddings later
    backend="sklearn",              # custom backend
    brain_key="img_sim",
)

print(image_index.total_index_size)  # 0
```

```python
object_index = fob.compute_similarity(
    dataset,
    patches_field="ground_truth",   # field containing objects of interest
    model="clip-vit-base32-torch",  # custom model
    embeddings=False,               # add embeddings later
    backend="sklearn",              # custom backend
    brain_key="gt_sim",
)

print(object_index.total_index_size)  # 0
```

#### Precomputed embeddings

You can pass precomputed image or object embeddings to
[`compute_similarity()`](../api/fiftyone.brain.md#fiftyone.brain.compute_similarity) via the
`embeddings` argument:

Image similarity

Object similarity

```python
model = foz.load_zoo_model("clip-vit-base32-torch")
embeddings = dataset.compute_embeddings(model)

tmp_index = fob.compute_similarity(
    dataset,
    model="clip-vit-base32-torch",  # store model's name for future use
    embeddings=embeddings,          # precomputed image embeddings
    brain_key="tmp",
)

print(tmp_index.total_index_size)  # 200

dataset.delete_brain_run("tmp")
```

```python
model = foz.load_zoo_model("clip-vit-base32-torch")
embeddings = dataset.compute_patch_embeddings(model, "ground_truth")

tmp_index = fob.compute_similarity(
    dataset,
    patches_field="ground_truth",   # field containing objects of interest
    model="clip-vit-base32-torch",  # store model's name for future use
    embeddings=embeddings,          # precomputed patch embeddings
    brain_key="tmp",
)

print(tmp_index.total_index_size)  # 1232

dataset.delete_brain_run("tmp")
```

### Adding embeddings to an index

You can use
[`add_to_index()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.add_to_index)
to add new embeddings or overwrite existing embeddings in an index at any time:

Image similarity

Object similarity

```python
image_index = dataset.load_brain_results("img_sim")
print(image_index.total_index_size)  # 0

view1 = dataset[:100]
view2 = dataset[100:]

#
# Approach 1: use the index to compute embeddings for `view1`
#

embeddings, sample_ids, _ = image_index.compute_embeddings(view1)
image_index.add_to_index(embeddings, sample_ids)
print(image_index.total_index_size)  # 100

#
# Approach 2: manually compute embeddings for `view2`
#

model = image_index.get_model()  # the index's model
embeddings = view2.compute_embeddings(model)
sample_ids = view2.values("id")
image_index.add_to_index(embeddings, sample_ids)
print(image_index.total_index_size)  # 200

# Must save after edits when using the sklearn backend
image_index.save()
```

When working with object embeddings, you must provide the sample ID and
label ID for each embedding you add to the index:

```python
import numpy as np

object_index = dataset.load_brain_results("gt_sim")
print(object_index.total_index_size)  # 0

view1 = dataset[:100]
view2 = dataset[100:]

#
# Approach 1: use the index to compute embeddings for `view1`
#

embeddings, sample_ids, label_ids = object_index.compute_embeddings(view1)
object_index.add_to_index(embeddings, sample_ids, label_ids=label_ids)
print(object_index.total_index_size)  # 471

#
# Approach 2: manually compute embeddings for `view2`
#

# Manually load the index's model
model = object_index.get_model()

# Compute patch embeddings
_embeddings = view2.compute_patch_embeddings(model, "ground_truth")
_label_ids = dict(zip(*view2.values(["id", "ground_truth.detections.id"])))

# Organize into correct format
embeddings = []
sample_ids = []
label_ids = []
for sample_id, patch_embeddings in _embeddings.items():
    patch_ids = _label_ids[sample_id]
    if not patch_ids:
        continue

    for embedding, label_id in zip(patch_embeddings, patch_ids):
        embeddings.append(embedding)
        sample_ids.append(sample_id)
        label_ids.append(label_id)

object_index.add_to_index(
    np.stack(embeddings),
    np.array(sample_ids),
    label_ids=np.array(label_ids),
)
print(object_index.total_index_size)  # 1232

# Must save after edits when using the sklearn backend
object_index.save()
```

#### NOTE
When using the default `sklearn` backend, you must manually call
[`save()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.save) after
adding or removing embeddings from an index in order to save the index to
the database. This is not required when using external vector databases
like [Qdrant](../integrations/qdrant.md#qdrant-integration).

#### NOTE
Did you know? If you provided the name of a [zoo model](../model_zoo/index.md#model-zoo)
when creating the similarity index, you can use
[`get_model()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.get_model)
to load the model later. Or, you can use
[`compute_embeddings()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.compute_embeddings)
to conveniently generate embeddings for new samples/objects using the
index’s model.

### Retrieving embeddings in an index

You can use
[`get_embeddings()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.get_embeddings)
to retrieve the embeddings for any or all IDs of interest from an existing
index:

Image similarity

Object similarity

```python
ids = dataset.take(50).values("id")
embeddings, sample_ids, _ = image_index.get_embeddings(sample_ids=ids)

print(embeddings.shape)  # (50, 512)
print(sample_ids.shape)  # (50,)
```

When working with object embeddings, you can provide either sample IDs or
label IDs for which you want to retrieve embeddings:

```python
from fiftyone import ViewField as F

ids = (
    dataset
    .filter_labels("ground_truth", F("label") == "person")
    .values("ground_truth.detections.id", unwind=True)
)

embeddings, sample_ids, label_ids = object_index.get_embeddings(label_ids=ids)

print(embeddings.shape)  # (378, 512)
print(sample_ids.shape)  # (378,)
print(label_ids.shape)  # (378,)
```

### Removing embeddings from an index

You can use
[`remove_from_index()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.remove_from_index)
to delete embeddings from an index by their ID:

Image similarity

Object similarity

```python
ids = dataset.take(50).values("id")

image_index.remove_from_index(sample_ids=ids)
print(image_index.total_index_size)  # 150

# Must save after edits when using the sklearn backend
image_index.save()
```

When working with object embeddings, you can provide either sample IDs or
label IDs for which you want to delete embeddings:

```python
from fiftyone import ViewField as F

ids = (
    dataset
    .filter_labels("ground_truth", F("label") == "person")
    .values("ground_truth.detections.id", unwind=True)
)

object_index.remove_from_index(label_ids=ids)
print(object_index.total_index_size)  # 854

# Must save after edits when using the sklearn backend
object_index.save()
```

#### NOTE
When using the default `sklearn` backend, you must manually call
[`save()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.save) after
adding or removing embeddings from an index in order to save the index to
the database.

This is not required when using external vector databases like
[Qdrant](../integrations/qdrant.md#qdrant-integration).

### Deleting an index

When working with backends like [Qdrant](../integrations/qdrant.md#qdrant-integration) that
leverage external vector databases, you can call
[`cleanup()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.cleanup) to delete
the external index/collection:

Image similarity

Object similarity

```python
# First delete the index from the backend (if applicable)
image_index.cleanup()

# Now delete the index from your dataset
dataset.delete_brain_run("img_sim")
```

```python
# First delete the index from the backend (if applicable)
object_index.cleanup()

# Now delete the index from your dataset
dataset.delete_brain_run("gt_sim")
```

#### NOTE
Calling
[`cleanup()`](../api/fiftyone.brain.similarity.md#fiftyone.brain.similarity.SimilarityIndex.cleanup) has
no effect when working with the default sklearn backend. The index is
deleted only when you call
[`delete_brain_run()`](../api/fiftyone.core.collections.md#fiftyone.core.collections.SampleCollection.delete_brain_run).

<a id="brain-similarity-applications"></a>

## Applications

How can similarity be used in practice? A common pattern is to mine your
dataset for similar examples to certain images or object patches of interest,
e.g., those that represent failure modes of a model that need to be studied in
more detail or underrepresented classes that need more training examples.

Here are a few of the many possible applications:

- Pruning [near-duplicate images](../brain/index.md#brain-near-duplicates) from your
  training dataset
- Identifying failure patterns of a model
- Finding examples of target scenarios in your data lake
- Mining hard examples for your evaluation pipeline
- Recommending samples from your data lake for classes that need additional
  training data
