Skip to main content
Version: 2026.3

Semantic Search in Studio (Experimental)

caution

This is an experimental feature, built on the experimental Embeddings feature. The request contract, behavior, and APIs may still change in breaking ways between releases.

The bundle adds semantic (kNN) search to Studio's existing grid and search endpoints. A bpt.semanticSearch column filter travels inside the regular request body; a tagged Studio filter reads it and pushes the bundle's kNN search modifier onto the query. Because the kNN clause joins the same bool query as every other filter, semantic search is hybrid by construction: classic column filters, full-text search, and workspace permissions all still apply.

Engine support is capability-gated. OpenSearch is always supported. Elasticsearch is supported on indices created on 8.12.1 or newer.

Covered endpoints (all POST):

  • /pimcore-studio/api/assets/grid
  • /pimcore-studio/api/data-objects/grid/{classId}
  • /pimcore-studio/api/search/assets
  • /pimcore-studio/api/search/data-objects

Request contract​

Add an entry to the ordinary filters.columnFilters array:

{
"filters": {
"columnFilters": [
{
"key": "semanticSearch",
"type": "bpt.semanticSearch",
"filterValue": {
"query": "red sports car",
"orderByRelevance": true,
"target": "image"
}
}
]
}
}
  • query — the natural-language search text. Truncated at 1000 characters before embedding. A bare string filterValue is accepted as shorthand for { "query": ... }.
  • orderByRelevance — default true: the sort list is cleared so the engine orders by _score. Send false when the user explicitly picked a column sort.
  • target — optional, "image" or "text", selects which usage the query embeds against. Defaults to "image" for assets. Data objects only ever search "text"; sending target: "image" on a data object query is rejected with a 422. Sending an unrecognized target value on either element type is rejected the same way.

Asset queries default to the image_search usage (visual thumbnail embeddings); sending target: "text" searches semantic_text_search instead, embedding the asset's filename, path and string metadata. Data object queries always search semantic_text_search, embedding a chunked text representation of the object. Exactly one target is scored per request — semantic search never scores both vectors at once. Both are combined freely with other filters — e.g. a system.fulltext filter plus the semantic filter returns the intersection, relevance-ordered.

Result-set semantics​

  • With orderByRelevance: true the sort list is cleared — the engine defaults to _score descending, and GDI's OrderByPageNumber handler early-returns (no page inversion, no count round-trip). With orderByRelevance: false the explicit column sort is preserved and the set is still k-bounded; note that a non-empty sort list makes OrderByPageNumber issue a _count request that includes the kNN clause.

Degrade behaviour​

If the inference service cannot embed the query (outage, timeout), the semantic clause is silently skipped, and the request returns plain (non-semantic) results; the server logs Semantic search degraded: query embedding unavailable, clause skipped. at warning level. Query embeddings are cached for 24h, so a repeated query may keep working through a short outage.

Global semantic search endpoint​

Besides the grid/search filter above, the bundle exposes a standalone GET endpoint that runs a kNN search directly, without going through the ordinary grid/search request body:

GET /pimcore-studio/api/bundle/backend-power-tools/semantic-search/search
curl -G 'https://<host>/pimcore-studio/api/bundle/backend-power-tools/semantic-search/search' \
--data-urlencode 'query=red sports car' \
--data-urlencode 'target=image' \
--data-urlencode 'page=1' \
--data-urlencode 'pageSize=50' \
-H 'Authorization: Bearer <token>'

Request contract​

  • query (required) — 1–1000 characters. Trimmed, then truncated to 1000 chars before embedding; an empty result after trimming is rejected.
  • target (optional) — "image" or "text". Any other value is rejected. Defaults to "text" if the parameter is omitted entirely.
  • page (optional) — defaults to 1.
  • pageSize (optional) — defaults to 50.

Unlike the grid/search filter, this endpoint has no per-element-type default for target (the filter defaults assets to "image"); the global endpoint defaults to "text" when the parameter is left out, and accepts either value explicitly.

Target semantics and coverage​

  • target=image embeds the query against the image_search usage and returns assets only (visual thumbnail embeddings).
  • target=text embeds the query against the semantic_text_search usage and returns assets (filename, path, string metadata) and embedding-enabled data object classes (chunked text representation) together. Because both element types share the same usage/vector space for "text", a single query embedding produces one merged, comparable ranking across them — there is no separate per-type scoring pass to reconcile.

Candidate pool (k)​

Each request runs a single kNN clause with k = QueryInputLimits::MAX_K (150), applied per index. Because the element-search alias spans every embedding-bearing index, the effective global candidate pool before pagination is 150 × <number of embedding-bearing indices>, not a single global top-150.

Status codes​

  • 400 (SearchException) — the inference service could not embed the query (outage, timeout).
  • 404 (NotFoundException) — the embeddings feature is disabled, or the current search engine doesn't support the query-clause kNN capability.
  • 422 (InvalidArgumentException) — query is missing/empty after trimming, or target isn't image/text.

Degrade behaviour — hard fail, not a silent skip​

This is the opposite of the grid/search filter's degradation behavior above: the filter silently drops the semantic clause and falls back to plain results on an embedding outage, but this endpoint hard-fails with 400 instead. The endpoint's only purpose is a ranked semantic result set — there is no non-semantic fallback result to return — so an outage is surfaced to the caller rather than swallowed.

Availability data​

GET /pimcore-studio/api/settings contributes a backend_power_tools_semantic_search block (via SemanticSearchSettingsProvider) so the Studio UI only offers a smart-search mode where embeddings can actually answer it.

The key is absent entirely when embeddings are disabled or the search engine doesn't support query-clause kNN — key presence is the feature flag.

"backend_power_tools_semantic_search": {
"targets": {
"image": {
"enabled": true,
"element_types": ["asset"],
"asset_types": ["image"]
},
"text": {
"enabled": true,
"element_types": ["asset", "object"],
"asset_types": ["image"],
"object_classes": ["Car", "News"]
}
}
}
  • targets.image — the "image" search mode. enabled requires image in assets.types and a mapped image_search usage.
  • targets.text — the "text" search mode. enabled requires a mapped semantic_text_search usage AND at least one text-bearing surface: a non-empty object class allowlist, or at least one asset type in assets.types. element_types is built from whichever surfaces are actually on: "asset" when assets.types is non-empty, "object" when the class allowlist is non-empty. asset_types lists the configured assets.types (e.g. ["image", "document", "text"]) and is present only when "asset" is in element_types. Document and text assets are searchable here but never enter duplicate detection. object_classes always reflects the configured allowlist, regardless of whether "object" is currently in element_types.