Semantic Search in Studio (Experimental)
This is an experimental feature, built on the experimental Embeddings feature. The request contract, behavior, and APIs may still change in breaking ways between releases.
The bundle adds semantic (kNN) search to Studio's existing grid and search endpoints.
A bpt.semanticSearch column filter travels inside the regular request
body; a tagged Studio filter reads it and pushes the bundle's kNN search modifier onto the query.
Because the kNN clause joins the same bool query as every other filter, semantic search is hybrid
by construction: classic column filters, full-text search, and workspace permissions all still
apply.
Engine support is capability-gated. OpenSearch is always supported. Elasticsearch is supported on indices created on 8.12.1 or newer.
Covered endpoints (all POST):
/pimcore-studio/api/assets/grid/pimcore-studio/api/data-objects/grid/{classId}/pimcore-studio/api/search/assets/pimcore-studio/api/search/data-objects
Request contract
Add an entry to the ordinary filters.columnFilters array:
{
"filters": {
"columnFilters": [
{
"key": "semanticSearch",
"type": "bpt.semanticSearch",
"filterValue": {
"query": "red sports car",
"orderByRelevance": true,
"target": "image"
}
}
]
}
}
query— the natural-language search text. Truncated at 1000 characters before embedding. A bare stringfilterValueis accepted as shorthand for{ "query": ... }.orderByRelevance— defaulttrue: the sort list is cleared so the engine orders by_score. Sendfalsewhen the user explicitly picked a column sort.target— optional,"image"or"text", selects which usage the query embeds against. Defaults to"image"for assets. Data objects only ever search"text"; sendingtarget: "image"on a data object query is rejected with a 422. Sending an unrecognizedtargetvalue on either element type is rejected the same way.
Asset queries default to the image_search usage (visual thumbnail embeddings); sending
target: "text" searches semantic_text_search instead, embedding the asset's filename, path
and string metadata. Data object queries always search semantic_text_search, embedding a
chunked text representation of the object. Exactly one target is scored per request — semantic
search never scores both vectors at once. Both are combined freely with other filters — e.g. a
system.fulltext filter plus the semantic filter returns the intersection, relevance-ordered.
Result-set semantics
- With
orderByRelevance: truethe sort list is cleared — the engine defaults to_scoredescending, and GDI'sOrderByPageNumberhandler early-returns (no page inversion, no count round-trip). WithorderByRelevance: falsethe explicit column sort is preserved and the set is still k-bounded; note that a non-empty sort list makesOrderByPageNumberissue a_countrequest that includes the kNN clause.
Degrade behaviour
If the inference service cannot embed the query (outage, timeout), the semantic clause is
silently skipped, and the request returns plain (non-semantic) results; the server logs
Semantic search degraded: query embedding unavailable, clause skipped. at warning level.
Query embeddings are cached for 24h, so a repeated query may keep working through a short outage.
Global semantic search endpoint
Besides the grid/search filter above, the bundle exposes a standalone GET endpoint that runs a kNN search directly, without going through the ordinary grid/search request body:
GET /pimcore-studio/api/bundle/backend-power-tools/semantic-search/search
curl -G 'https://<host>/pimcore-studio/api/bundle/backend-power-tools/semantic-search/search' \
--data-urlencode 'query=red sports car' \
--data-urlencode 'target=image' \
--data-urlencode 'page=1' \
--data-urlencode 'pageSize=50' \
-H 'Authorization: Bearer <token>'
Request contract
query(required) — 1–1000 characters. Trimmed, then truncated to 1000 chars before embedding; an empty result after trimming is rejected.target(optional) —"image"or"text". Any other value is rejected. Defaults to"text"if the parameter is omitted entirely.page(optional) — defaults to1.pageSize(optional) — defaults to50.
Unlike the grid/search filter, this endpoint has no per-element-type default for target (the
filter defaults assets to "image"); the global endpoint defaults to "text" when the parameter
is left out, and accepts either value explicitly.
Target semantics and coverage
target=imageembeds the query against theimage_searchusage and returns assets only (visual thumbnail embeddings).target=textembeds the query against thesemantic_text_searchusage and returns assets (filename, path, string metadata) and embedding-enabled data object classes (chunked text representation) together. Because both element types share the same usage/vector space for"text", a single query embedding produces one merged, comparable ranking across them — there is no separate per-type scoring pass to reconcile.
Candidate pool (k)
Each request runs a single kNN clause with k = QueryInputLimits::MAX_K (150), applied per
index. Because the element-search alias spans every embedding-bearing index, the effective
global candidate pool before pagination is 150 × <number of embedding-bearing indices>, not a
single global top-150.
Status codes
400(SearchException) — the inference service could not embed the query (outage, timeout).404(NotFoundException) — the embeddings feature is disabled, or the current search engine doesn't support the query-clause kNN capability.422(InvalidArgumentException) —queryis missing/empty after trimming, ortargetisn'timage/text.
Degrade behaviour — hard fail, not a silent skip
This is the opposite of the grid/search filter's degradation behavior above: the filter silently drops the semantic clause and falls back to plain results on an embedding outage, but this endpoint hard-fails with 400 instead. The endpoint's only purpose is a ranked semantic result set — there is no non-semantic fallback result to return — so an outage is surfaced to the caller rather than swallowed.
Availability data
GET /pimcore-studio/api/settings contributes a backend_power_tools_semantic_search block
(via SemanticSearchSettingsProvider) so the Studio UI only offers a smart-search mode where
embeddings can actually answer it.
The key is absent entirely when embeddings are disabled or the search engine doesn't support query-clause kNN — key presence is the feature flag.
"backend_power_tools_semantic_search": {
"targets": {
"image": {
"enabled": true,
"element_types": ["asset"],
"asset_types": ["image"]
},
"text": {
"enabled": true,
"element_types": ["asset", "object"],
"asset_types": ["image"],
"object_classes": ["Car", "News"]
}
}
}
targets.image— the"image"search mode.enabledrequiresimageinassets.typesand a mappedimage_searchusage.targets.text— the"text"search mode.enabledrequires a mappedsemantic_text_searchusage AND at least one text-bearing surface: a non-empty object class allowlist, or at least one asset type inassets.types.element_typesis built from whichever surfaces are actually on:"asset"whenassets.typesis non-empty,"object"when the class allowlist is non-empty.asset_typeslists the configuredassets.types(e.g.["image", "document", "text"]) and is present only when"asset"is inelement_types. Document and text assets are searchable here but never enter duplicate detection.object_classesalways reflects the configured allowlist, regardless of whether"object"is currently inelement_types.