Getting Started
This page assumes the bundle is installed and the Generic Data Index is running on OpenSearch or Elasticsearch. Embeddings are off by default: they add fields to your search index and send content to an external inference service, so nothing happens until you opt in.
1. Pick a model
The bundle ships no model and has no built-in default endpoint — you choose the model and where it runs. Two decisions matter:
- Modality. A text model covers semantic search and duplicate detection over data objects (and the text of assets). Image search needs an image model, paired with a text model that embeds the incoming query into the same latent space — see the cross-modal section of Inference Service.
- Language coverage. If your audience searches in more than the default language, pick a multilingual text model. Retrofitting this later means re-embedding everything. Localized fields are currently embedded using Pimcore's system default language. Because the locale determines the composed text, changing the system default language changes every object's identity key, so the affected vectors re-embed on the next indexing pass.
If you want a starting point rather than an evaluation: the models the bundle's own benchmarks use, and the measurements behind that choice, are in Defaults and Benchmarks. Treat them as a reasonable default, not a recommendation for your data — the Evaluation Tool exists to compare candidates on your content, and it works before you enable the feature.
2. Serve the model over HTTP
The bundle never loads or downloads a model. It only makes HTTP requests to an inference service that you run or subscribe to — self-hosted or a third-party API, whichever you prefer. Out of the box it speaks the OpenAI-compatible embeddings shape, which most embedding servers and hosted APIs already expose.
Inference Service covers your hosting options, how to get the model weights, and the exact contract an endpoint must fulfil. Two points decide whether your first run succeeds:
- Vectors must be L2-normalized (unit length). Unnormalized vectors do not error — they silently rank partly by vector length and quietly ruin result quality.
- The first backfill is your heaviest inference burst. Embedding a whole catalogue on CPU-only
hardware can take hours, which is why the model
timeoutdefaults to 120 seconds.
If no available service fits your model or preprocessing needs, you can implement your own provider — see Customization.
3. Configure the bundle
The primary way to configure embeddings is the Studio editor — open it from the main navigation at Automation & Integration → Pimcore Embeddings → Embedding Configuration (admin users only) — which persists to the settings store and needs no deployment to change.
Every option is documented in the Configuration reference.
4. Recreate the search indices
The vector fields are part of the index mapping, so the indices must be rebuilt once after enabling.
On OpenSearch the required index.knn setting is added automatically while embeddings are enabled.
bin/console generic-data-index:update:index -r
5. Run the workers
Embedding generation is asynchronous. Two transports must be consumed — the index queue and the embedding queue:
bin/console messenger:consume pimcore_generic_data_index_queue
bin/console messenger:consume pimcore_backend_power_tools_embedding_queue
Run the embedding queue in its own worker (or several): embedding batches are slow compared to structural indexing, and sharing a worker lets them starve normal index updates. In production these belong under a process supervisor.
Neither transport draining is proof a re-embed run has finished — the job that starts a run only dispatches work and returns immediately. For an actual progress and queue-depth view, use the status endpoint described in Progress and Cancellation.
6. Backfill existing elements
New and edited elements are embedded automatically from now on. Existing content needs one pass:
bin/console bpt:embeddings:reindex
If you're upgrading and the composed text format itself changed (see Caching), expect this pass to regenerate every vector once, even for content you didn't touch — that's expected, not a sign something is broken.
7. Verify
Ask the reconciler whether anything is still missing its vectors — the authoritative check:
bin/console bpt:embeddings:reindex --reconcile
missing: 0 for every field means every eligible element carries vectors. "Eligible" is the element
scope plus what can actually produce input: an allowlisted data object class (Concrete objects and
their variants, never folders), image assets whose thumbnail rasterizes to a usable raster,
and document/text assets of a configured type. A document whose text cannot be
extracted (a scanned PDF, say) is still embedded from its identifying context (name, path,
metadata), so it is never reported as missing. Then confirm quality end-to-end with a semantic
query through Studio's search, or run the Evaluation Tool against
label files.
Troubleshooting the first run
| Symptom | Likely cause |
|---|---|
| No vectors appear at all, no errors | Most often the element scope: objects.classes is empty (no object is embedded) or assets.types is empty (no asset is embedded). Then check enabled: true and the usages block. |
| A newly added class indexes but never gets vectors | Its index mapping has not been updated yet. Press Re-embed in Studio (it verifies/updates the mapping first) or run generic-data-index:update:index -c <classDefinitionId>; elements indexed in the meantime were auto-marked for backfill and heal on the next re-embed or backfill drain. |
| A document has vectors but content queries never find it | Only its name, path and metadata were embedded: a scanned/image-only PDF has no text layer and there is no OCR, or assets.document.process_text is disabled so no PDF text is extracted at all. |
| Search returns results but ranking looks random, scores nearly identical | The service returns unnormalized vectors. See serving invariant 1. |
| Vectors missing for some elements only | Generation failed for those; they are marked for backfill. Re-run bpt:embeddings:reindex and check the worker logs. |
| A worker starts and exits immediately without output | A stale worker-restart signal. Clear it with bin/console cache:pool:clear cache.messenger.restart_workers_signal, then consume again. |
| How do I know a re-embed finished | Poll the status endpoint or wait for the Pimcore notification sent to the user who started the run — see Progress and Cancellation. |
| Everything re-embeds on every reindex | An identity-key input is changing — see Caching. |
index.knn / unknown setting errors on Elasticsearch | index.knn is OpenSearch-only; the bundle omits it on Elasticsearch. Recreate the index after switching engines. |
Turning it off again
pimcore_backend_power_tools:
embeddings:
enabled: false
This stops all embedding behaviour in one place — mapping, generation, indexing, and kNN querying — regardless of the configured models and usages. Already-stored vectors stay in the index until it is recreated; the commands refuse to run, and the kNN search modifier becomes a no-op instead of calling the inference service.