Embeddings¶
Beta API
The Embeddings API is in beta. Its interface may change in future releases without a deprecation period.
LightlyStudio embeds your data automatically on ingestion. To supply your own
embeddings — either computed on the fly or loaded from a precomputed store —
subclass the capability interfaces for the inputs you can embed and register the
embedder with register_default_embedder. The
registration must happen before you load a dataset or before the GUI is started.
See the Embeddings page for more details.
register_default_embedder¶
Public API for configuring embedders.
register_default_embedder ¶
register_default_embedder(
embedder: Embedder, for_capabilities: Set[Capability] | None = None
) -> None
Register an embedder as the default for the capabilities it implements.
Beta
Call this in either of two cases:
- Before creating a dataset, so that ingestion embeds with
embedder. This draws on the capabilities for the data being added, such as embedding images, image crops, or videos. - Before starting the GUI, so that search embeds queries with
embedder. This draws on the capabilities for the query kinds, such as embedding text, or images for reverse-image search. Search resolves the embedder by the collection's stored embedding space, soembeddermust share that space to take effect.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
embedder
|
Embedder
|
The embedder to register. Its embedding space is read from
|
required |
for_capabilities
|
Set[Capability] | None
|
Capabilities for which this embedder becomes the default choice. By default, all implemented capabilities are updated. An empty set registers the embedder without changing the defaults. Requested capabilities the embedder does not implement are ignored. |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the embedder implements no capability, or if it shares a space with an already registered embedder but does not match its spec. |
Capability interfaces¶
An embedder subclasses Embedder through one interface per input it can embed.
Subclass only the capabilities your model provides.
Embedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
Embedder ¶
Bases: ABC
Base class for every embedder.
Beta
Subclasses add one abstract method per input they support. A concrete model implements the subclasses for the capabilities it provides.
An embedder must accept calls from more than one thread, because the server runs a request in a worker thread. Hold a lock in the methods when the model cannot.
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
ImagePathEmbedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
ImagePathEmbedder ¶
Bases: Embedder
Embeds images read from an fsspec path (local, s3://, ...).
Beta
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embed_images
abstractmethod
¶
embed_images(paths: list[str]) -> EmbeddingResult
Embed a batch of images given by path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
paths
|
list[str]
|
fsspec paths of the images to embed. |
required |
Returns:
| Type | Description |
|---|---|
EmbeddingResult
|
The embeddings and the indices of the inputs they cover. |
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
ImageCropPathEmbedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
ImageCropPathEmbedder ¶
Bases: Embedder
Embeds crops of stored images, each an image path plus a pixel box.
Beta
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embed_image_crops
abstractmethod
¶
embed_image_crops(crops: list[ImageCrop]) -> EmbeddingResult
Embed a batch of image crops given by path and pixel box.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crops
|
list[ImageCrop]
|
The crops to embed. |
required |
Returns:
| Type | Description |
|---|---|
EmbeddingResult
|
The embeddings and the indices of the inputs they cover. |
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
VideoPathEmbedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
VideoPathEmbedder ¶
Bases: Embedder
Embeds videos read from an fsspec path.
Beta
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embed_videos
abstractmethod
¶
embed_videos(paths: list[str]) -> EmbeddingResult
Embed a batch of videos given by path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
paths
|
list[str]
|
fsspec paths of the videos to embed. |
required |
Returns:
| Type | Description |
|---|---|
EmbeddingResult
|
The embeddings and the indices of the inputs they cover. |
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
ImagePILEmbedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
ImagePILEmbedder ¶
Bases: Embedder
Embeds images represented as PIL images.
Beta
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embed_images_pil
abstractmethod
¶
embed_images_pil(images: list[Image]) -> EmbeddingResult
Embed a batch of PIL images.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
images
|
list[Image]
|
The PIL images to embed. |
required |
Returns:
| Type | Description |
|---|---|
EmbeddingResult
|
The embeddings and the indices of the inputs they cover. |
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
TextEmbedder¶
Capability-split embedder interfaces.
Defines the Embedder base class and one abstract subclass per input a model
can embed: images and crops by path, videos, PIL images, text, and images by
bytes. A concrete model implements only the capabilities it supports, and
callers pick an embedder by the capability they need.
An embedder subclasses only the interfaces that it implements. The class list is the advertisement. An embedder cannot advertise a capability and then not serve it.
TextEmbedder ¶
Bases: Embedder
Embeds text queries.
Beta
ready
property
¶
ready: bool
Whether the model is loaded and can answer requests.
The default suits a model that loads when the object exists. Override the
property when the weights load in the background. While the value is False,
/v1/describe reports ready: false and the embed endpoints answer 503.
The property covers the weights only. embedding_space_spec answers either way.
embed_text
abstractmethod
¶
embed_text(texts: list[str]) -> EmbeddingResult
Embed a batch of text strings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
texts
|
list[str]
|
The strings to embed. |
required |
Returns:
| Type | Description |
|---|---|
EmbeddingResult
|
The embeddings and the indices of the inputs they cover. |
embedding_space_spec
abstractmethod
¶
embedding_space_spec() -> EmbeddingSpaceSpec
Describe the embedding space this embedder produces.
The method must answer as soon as the object exists, also while ready is
False. /v1/describe reports the space to a client that waits for the
model. The identity of a space comes from the architecture, so a model knows it
before the weights arrive.
Returns:
| Type | Description |
|---|---|
EmbeddingSpaceSpec
|
Metadata identifying the embedding space, stored so the same space can be recognized across LightlyStudio runs. |
Supporting types¶
EmbeddingSpaceSpec¶
Shared value types for the embedding paths.
Holds the model-agnostic types passed to and returned by embedders:
EmbeddingSpaceSpec, EmbeddingResult and ImageCrop. They live in their
own module so any caller can depend on them without importing the Embedder
classes.
EmbeddingSpaceSpec
dataclass
¶
EmbeddingSpaceSpec(space_key: str, dimension: int)
Identity and shape of the embedding space an embedder produces.
Beta
Stored in the database so the same embedding space can be recognized across LightlyStudio runs.
dimension
instance-attribute
¶
dimension: int
Length of each embedding vector this embedder produces.
space_key
instance-attribute
¶
space_key: str
Stable identifier for the embedding space.
Two embedders that share a space_key are treated as producing the same
embedding space, so their vectors are comparable. Change it whenever the
produced vectors become incomparable, e.g. for a model version change. Can be
any string, e.g. your-company/model-family@version.
EmbeddingResult¶
Shared value types for the embedding paths.
Holds the model-agnostic types passed to and returned by embedders:
EmbeddingSpaceSpec, EmbeddingResult and ImageCrop. They live in their
own module so any caller can depend on them without importing the Embedder
classes.
EmbeddingResult
dataclass
¶
EmbeddingResult(embeddings: NDArray[float32], kept_indices: list[int])
Embeddings for the inputs that could be read, plus which inputs they cover.
Beta
An embedder skips broken inputs (files it cannot read or decode) instead of
failing the whole batch, so embeddings can have fewer rows than the input
list. kept_indices gives the position of each row in the input list, in
input order. Use it to line up the embeddings with any per-input data you keep
on the side, such as sample IDs.
ImageCrop¶
Shared value types for the embedding paths.
Holds the model-agnostic types passed to and returned by embedders:
EmbeddingSpaceSpec, EmbeddingResult and ImageCrop. They live in their
own module so any caller can depend on them without importing the Embedder
classes.
ImageCrop
dataclass
¶
ImageCrop(filepath: str, x: int, y: int, width: int, height: int)
A rectangular region of an image to embed, given in pixel coordinates.
Beta