Configuration
All options are editable through the Studio editor — Automation & Integration → Pimcore Embeddings → Embedding Configuration in the main navigation (admin users only) — which is the primary way to configure embeddings and persists to the settings store without a deployment.
Where the configuration is read from and written to is itself configurable — see Configuration storage.
Where a default was a measured decision rather than a convention, the numbers behind it are in Defaults and Benchmarks.
Configuration storage (config_location)
The embeddings configuration is a Pimcore location-aware configuration: read source and write target are chosen per installation, not hard-wired.
pimcore_backend_power_tools:
config_location:
embeddings:
write_target:
type: settings-store # default; also: symfony-config, disabled
read_target:
type: ~ # default; also: settings-store, symfony-config
| Setup | Behaviour |
|---|---|
| Defaults (as shipped) | Reads container config first, settings store as fallback — a hand-written pimcore_backend_power_tools.embeddings yaml node therefore takes over and the Studio editor turns read-only. Saves go to the settings store. |
write_target.type: symfony-config | Studio saves write yaml files into var/config/backend_power_tools_embeddings/, which are loaded back into the container config — a version-controllable workflow. Writing is only permitted in debug mode, so production editors are read-only under this target. |
write_target.type: disabled | The Studio editor is read-only everywhere; the configuration is fully deploy-managed. |
read_target.type: settings-store | Reads come ONLY from the settings store — a yaml embeddings node is ignored entirely (no shadowing). |
read_target.type: symfony-config | Reads come ONLY from yaml/container config; the settings store is ignored. |
Pointing read_target and write_target at different locations makes the editor write where
reads never look — saves appear to vanish. Pick a matching pair unless you know exactly why not.
Example
The configuration the benchmark was performed on, reduced to the models it actually used — one multilingual text model, and the two towers of a cross-modal image model. Every option left at its default is omitted, so nothing here is redundant:
pimcore_backend_power_tools:
embeddings:
enabled: true # master switch; off by default
assets:
types: [image, document] # asset types to embed; empty by default
objects:
classes: # ONLY these classes are embedded; empty = none
- Car
- News
models:
nomic-text-v2: # free id, referenced by `usages` below
modality: text # embeds text (objects + asset text)
endpoint: 'https://your-inference-host/text-embedding'
model: 'nomic-ai/nomic-embed-text-v2-moe' # multilingual, Apache-2.0
dimension: 768 # required; sizes the kNN field, validates responses
max_text_chars: 1500 # one input at most this long (512-token window, German/English)
chunk_chars: 1000 # longer text is split into chunks of this size
prefixes: # instruction prefixes this model is trained to expect
document: 'search_document: ' # prepended when embedding indexed content
query: 'search_query: ' # prepended when embedding a search query
timeout: 1800 # rose from 120s when the benchmark ran on CPU-only inference
siglip2-text: # TEXT tower: embeds the incoming image-search query
modality: text
endpoint: 'https://your-inference-host/text-embedding'
model: 'google/siglip2-base-patch16-256' # multilingual, Apache-2.0
dimension: 768
timeout: 600
shared_space: 'siglip2-base' # must match the vision tower exactly
siglip2-vision: # VISION tower: embeds the indexed images
modality: image # input is a base64 thumbnail
endpoint: 'https://your-inference-host/image-embedding'
model: 'google/siglip2-base-patch16-256' # multilingual, Apache-2.0
dimension: 768
timeout: 600
shared_space: 'siglip2-base' # same space as siglip2-text
usages: # which model serves which capability
semantic_text_search: nomic-text-v2 # shorthand: same model for documents and queries
object_dedup: nomic-text-v2
image_search:
model: siglip2-vision # embedded the indexed images
query_model: siglip2-text # embeds the text query (cross-modal)
image_dedup: siglip2-vision
excluded_fields: # field keys never embedded, at any nesting level
- internalNotes
Top level
| Option | Default | Description |
|---|---|---|
enabled | false | Master switch for all embedding behaviour (mapping, generation, indexing, kNN querying). Off by default: embeddings add index fields and call an external service, so they must be opted into — set it to true and configure at least one model and usage. |
excluded_fields | [] | Field keys never embedded, by both production indexing and the evaluation tool, so the two cannot disagree about an element's text. Changing it needs a reindex. |
store_embedded_text | false | Debug aid: also store the composed element text in an emb_<model>_source sibling field — without the model's instruction prefix, which generation prepends afterwards. Text inputs only; image inputs are never stored. Costs storage — keep off in production. |
evaluation_mode | false | When enabled, every text representation is mapped, generated and stored per element so the Evaluation Tool can compare them. Off, only the primary representation is stored. Changing it needs index recreation + reindex. |
Cache table location and storage engine are deploy-time-only settings, not part of this editable
block — see embedding_cache below.
assets and objects — the element scope
Which elements are embedded? Deny by default on both sides: enabling embeddings alone embeds nothing, because sending a catalogue to an inference service should be a decision, not a side effect.
| Option | Default | Description |
|---|---|---|
assets.types | [] | Asset types to embed: image, document, text. Empty embeds no asset. document and text produce text vectors only. |
objects.classes | [] | Data object class names to embed, e.g. [Car, News]. Empty means no data object is embedded. Changing this needs an index recreation + reindex. |
Document and text assets are searchable once embedded, but they never enter duplicate
detection — asset dedup requires image in assets.types.
An unknown assets.types value is rejected at container build for a yaml override, and dropped
with a warning when read from the settings store.
Widening the scope (adding a class, or adding an entry to assets.types) needs the affected index's mapping
updated before it can carry vectors — pressing Re-embed in Studio does this automatically; the CLI
equivalent (generic-data-index:update:index -c <classDefinitionId>) remains for scripted setups. Until
the mapping exists, the newly scoped elements index without their vector fields (a write-path guard
omits a field the live mapping doesn't carry yet, rather than fail) and are marked for backfill, so
their vectors are filled in on the next re-embed or backfill drain — no manual reindex of those
elements is needed.
Document and text extraction limits
Document and text assets are embedded from their identifying context (name, path, metadata) plus
whatever text can be extracted from the file. The amount of content is governed by one knob,
chunking.max_input_chars: the extractors stop reading once they have delivered that many
characters, because the chunker would discard anything beyond it anyway. A few fixed guards
protect the indexing worker's memory on top.
| Source | Extraction route | Content bound | Memory guard (fixed) |
|---|---|---|---|
PDF, Word, PowerPoint and other office documents (Asset\Document) | Pimcore core text extraction (pdftotext/Ghostscript for PDF, Gotenberg or LibreOffice for office formats); needs pimcore.assets.document.process_text enabled | the composed text is cut at chunking.max_input_chars | none in the bundle |
Spreadsheets and CSV (.xlsx, .xlsm, .ods, .xls, .csv, on Asset\Document or Asset\Text) | PhpSpreadsheet, read-data-only, emitted row by row with the header once per sheet | rows are emitted until the text reaches chunking.max_input_chars; a notice names the asset, the sheet and the row it stopped at | a workbook over 32 MB uncompressed is skipped with a warning; at most 256 columns per row and 200,000 cells per workbook are loaded |
Plain text (Asset\Text: .txt, .md, .json, .xml, …) | file read | the first 4 bytes per character of the cap (2 MB at the default), cut at a UTF-8 character boundary; the chunker logs the truncation to the cap | the same byte budget |
| Scanned or image-only PDFs | no OCR anywhere in the chain | — | only name, path and metadata are embedded |
Why one knob. Rows and characters only line up for one row width: a 20,000-row cap would cut a narrow ID list at 400,000 characters, below the default cap, while a 256-column sheet reaches the cap after about 40 rows. Stopping at the character cap is the same rule for every shape. The flip side is inherent to row-major text: a very wide sheet yields only its first few dozen rows, and the sheets after the one that filled the cap are not read at all.
Why the memory guards are what they are. PhpSpreadsheet parses a workbook's shared-string table and each sheet's XML in full before any read filter runs, at several times the XML size in memory, so a highly compressible upload has to be refused by its uncompressed size up front: 32 MB keeps that parse near 200 MB. The read filter cannot see cell values, so it cannot apply the character cap while the file is parsed; it bounds what is loaded instead. PhpSpreadsheet holds roughly 1 KB per loaded cell (a 283-row demo workbook peaked at 20 MB), so 200,000 cells keep the loaded workbook near 200 MB as well, inside a 512 MB PHP memory limit, and the column cap stops a single vast row from spending that budget on cells no row-major text would ever include. Extraction never fails to index: an unreadable or malformed file logs one warning, and the asset is embedded from its identifying context alone.
The image thumbnail
An image model embeds a Pimcore image thumbnail of the asset, never the original file, and names it
with thumbnail. Please define a thumbnail based on the dimensions your molde requires for the best results:
pimcore:
assets:
image:
thumbnails:
definitions:
embeddings_siglip2:
name: embeddings_siglip2
description: 'Image input for embedding generation (SigLIP 2 base, 256x256)'
format: JPEG
quality: 90
items:
- method: cover
arguments:
width: 256
height: 256
- Match the model's native input.
siglip2-base-patch16-256takes 256×256; feeding it 224 makes the service upscale and lose detail. A larger thumbnail is downscaled correctly, so err upward. - Per model. Two image models can use two definitions — useful when one wants 256 and another 384.
- Editing it re-embeds. The definition's pixel-affecting settings are part of the identity key, so a resize exactly invalidates that model's images. See Caching.
models
One entry per embedding model, keyed by a free id — dashes in the id you choose are folded to
underscores internally (a model configured as nomic-text-v2 is addressed as nomic_text_v2 by
CLI flags like --regenerate --model=<id>, see Caching):
| Option | Default | Description |
|---|---|---|
modality | (required) | text or image. |
endpoint | (required) | Base URL of the HTTP inference service. Must be an absolute http:// or https:// URL — a protocol-relative or relative value is rejected at validation, so a client profile's credentials (see below) can never attach to an ambiguous host. Not part of the identity key — see model. |
model | (required) | Model name the inference service should load. Not part of the identity key: swapping it under the same config id does not invalidate stored vectors — bump serving_version alongside it (see Caching). |
provider | default | Which registered embedding provider talks to the endpoint. Not part of the identity key — see model. |
dimension | (required) | Vector dimension the model returns. It sizes the kNN field mapping — immutable once the index exists — and is the expectation every provider response is validated against, so a misconfigured endpoint is rejected instead of poisoning the index. Must be ≥ 1. Not part of the identity key — changing it re-embeds nothing by itself. |
max_text_chars | 1500 (min 200) | Longest text, in characters, sent to the model as one input; longer text is split into chunk_chars chunks. Like dimension, a model property the bundle cannot discover through an inference endpoint: provider, service and model are all replaceable and none exposes a tokenizer, so you convert the model's token window to characters for your content language (see Sizing). Text beyond the real window is truncated by the inference service without notice. The default fits a 512-token model (nomic-embed-text-v2-moe) for English and German. Ignored for image models. Not part of the identity key. |
chunk_chars | 1000 (min 100, ≤ max_text_chars) | Chunk length for text over max_text_chars; overlap is 15 % of it. A retrieval-granularity convention (~250 English tokens), not a model property. Larger values cut the vector count of long documents roughly in proportion; smaller values match specific queries more sharply. Not part of the identity key, but changing it changes chunk texts, so that model's long texts re-embed. |
timeout | 120 | Request timeout in seconds — embedding batches on CPU-only services can be slow. |
thumbnail | (required for image models) | Pimcore image thumbnail configuration used as the model input. |
prefixes.document / prefixes.query | '' | Instruction prefixes some models are trained to expect (e.g. search_document: / search_query: for Nomic; E5 uses passage: / query: ). The bundle prepends them before the request, so the endpoint must not add its own. Only prefixes.document is part of the stored-vector identity key, so editing it re-embeds that model's text vectors; prefixes.query only re-keys the query cache and re-embeds nothing. |
serving_version | '' | Version of the inference service's preprocessing for this model, folded into the identity key. Bump when serving-side preparation changes (tokenizer fix, normalization) — such changes are invisible to the text-based key, so stored vectors would otherwise stay silently stale. Empty keeps keys byte-identical to before the option existed. |
shared_space | null | Name of the latent space the model embeds into. Two models may only be combined in one search (cross-modal, e.g. text query → image field) when both declare the same space — equal dimensions are not sufficient and would return confident nonsense. |
auth_token | null | Optional bearer token for the endpoint. Literal value only — this field is runtime-editable (Studio or yaml), so %env(...)% placeholders are rejected at validation. For an env-backed secret, configure embedding_client_profiles instead (see below). |
options | [] | Free-form provider-specific settings, passed through to the provider untouched. The default provider reads options.params (extra request body params) and options.headers (custom HTTP headers — see below). Not part of the identity key: if a change here alters how the service embeds (a pooling mode, a truncation flag, a version header selecting different preprocessing), stored vectors go silently stale — bump serving_version alongside it. |
quantization.mode | on_disk | OpenSearch kNN storage mode; on_disk trades a little latency for far less graph memory. Ignored on Elasticsearch. |
quantization.compression_level | 32x | 1x (float32) … 32x (binary, the default). Changing it needs a reindex (no re-embed — vectors carry over via index recreation). |
Three combinations are measured against the quality bars in Defaults and Benchmarks and exposed as named presets in the Studio editor:
| Preset (Studio label) | mode | compression_level |
|---|---|---|
| Binary, disk-backed (recommended) | on_disk | 32x |
| Int8, in memory | in_memory | 4x |
| Full precision (float32) | in_memory | 1x |
The Studio editor exposes exactly one Quantization select — not separate raw mode and
compression-level pickers. A stored mode/compression_level pair that matches none of the three
presets above.
Sizing max_text_chars and chunk_chars
Both are characters because the bundle has no tokenizer: provider, endpoint and model are replaceable and none exposes one. You convert the model's token window yourself:
usable_tokens = window_tokens − tokens(prefixes.document) − 10 % of window_tokens
max_text_chars = usable_tokens × chars_per_token(language)
chunk_chars = 250 × chars_per_token(language) # ~250 tokens is the retrieval convention
window_tokens is on the model card (nomic-embed-text-v2-moe 512, nomic-embed-text-v1.5 8192,
Qwen3-Embedding 32k). The 10 % margin covers the spread of real texts around the language average.
For the default model and the Nomic prefix: 512 − 6 − 51 ≈ 455 usable tokens.
| Content language | Chars per token (approx.) | max_text_chars (455 usable tokens) | chunk_chars (~250 tokens) |
|---|---|---|---|
| English | 4 | 1800 | 1000 |
| German, French, Spanish, Italian | 3.5 | 1500 | 875 |
| Russian, Greek | 3 | 1350 | 750 |
| Arabic, Hebrew, Hindi | 2 | 900 | 500 |
| Chinese, Japanese, Korean, Thai | 1 | 450 | 250 |
The defaults sit at two different points of that range on purpose. max_text_chars = 1500 is the formula
at 3.5 characters per token, so the whole-text budget fits German and French, not only English.
chunk_chars = 1000 keeps the established retrieval size of 250 tokens at 4 characters per token
(English); for German it is about 285 tokens, still inside the 200 to 500 token band where passage
retrieval is flat. Size everything at 3.5 and the table's 875 is the stricter, equally valid choice. Both
defaults are an assumption, not a measurement: English has spare room, Russian is at the edge, denser
scripts need lower values. Any other value from the same formula is equally valid — the formula is the
argument, not the number.
One model has one budget: for mixed-language content, size for the densest language you index; the other languages then produce more, smaller chunks, which costs vectors but loses no text. Nothing in the bundle can detect a budget that is too large — the service truncates silently — so err low.
Custom request headers
Endpoints that don't authenticate with a bearer token (or need extra routing/tenant headers) can be
served without a custom provider via options.headers. Configured headers are added to every request
and replace a same-named default, compared case-insensitively — so an Authorization entry wins
over the auth_token bearer default. Like auth_token, values are literal only: %env(...)%
placeholders in options.headers are rejected at validation, the same as for auth_token. Never put
a literal secret in config either — use an embedding_client_profiles entry (below) for anything
env-backed:
models:
azure-text:
modality: 'text'
endpoint: 'https://my-resource.openai.azure.com/openai/deployments/embed/embeddings?api-version=2024-02-01'
model: 'text-embedding-3-small'
dimension: 1536
options:
headers:
api-key: 'temporary-value' # rotate/replace via embedding_client_profiles instead
Headers are transport details, not part of the vector identity key. If a header changes what the
service computes (e.g. an API-version header selecting different preprocessing), bump
serving_version alongside it. Header values may carry credentials — the bundle never logs them, and
neither should your own tooling.
Credentials & client profiles
The auth_token and options.headers fields above live in the editable configuration (Studio or
the read-only YAML override) — both are runtime-changeable by anyone with access to the configuration
screen. To keep an environment variable from being readable through that surface, stored values are
literal-only: a %env(...)% placeholder anywhere in auth_token or options.headers is rejected
at validation, whether it comes from Studio or from YAML.
%env(...)% in pimcore_backend_power_tools.embeddings.models.*.auth_token or .options.headers is
always rejected. This is a deliberate security control (env-var exfiltration through the editable
config)..
Env-backed secrets belong in embedding_client_profiles instead — a separate, deploy-time-only
Symfony configuration node (config/packages/*.yaml, not the embeddings block, and not editable
through Studio). Each profile attaches its auth_token/headers to a request only when the
request's endpoint matches the profile's endpoint_pattern — an anchored regex tested against the
full request URL, the same semantics as Symfony's ScopingHttpClient:
pimcore_backend_power_tools:
embedding_client_profiles:
my_inference_host:
endpoint_pattern: 'https://inference\.example\.com/.*' # anchored regex, escape literal dots
auth_token: '%env(MY_INFERENCE_TOKEN)%' # container-resolved — safe here
headers:
X-Tenant: 'catalog'
Rules, in order of precedence:
- A stored literal token always wins. If the model's own
auth_tokenis set, the matching profile'sauth_tokenis never used — profiles are a fallback, not an override. - Profile headers only fill gaps. A profile header is added only when the model doesn't already
define a header of the same name (case-insensitive); it never replaces one the model a set via
options.headers. - No match, no credentials. If no profile's
endpoint_patternmatches the request URL, nothing fromembedding_client_profilesis attached. - Non-absolute endpoints never receive profile credentials. Only a request to an absolute
http:///https://URL can match a profile at all — see theendpointrequirement above.
Stored tokens are write-only via the API
The Studio editor's underlying API never returns a stored auth_token value. A GET exposes only
hasAuthToken: boolean (whether a token is stored, not what it is); a PUT/save distinguishes three
cases for the authToken field:
| Sent value | Effect |
|---|---|
omitted / null | Keep the currently stored token for that model key unchanged. |
"" (empty string) | Clear the stored token. |
| any other value | Overwrite the stored token with that value. |
Renaming a model key does not inherit its old token. The keep-on-null behavior is looked up by
the model's key in the update payload — a model saved under a new key is treated as a new model, so
its stored token starts empty even if the old key had one. Re-enter the token after a rename.
The Test-configuration endpoint (used to try a model from the editor before saving) follows the same
rule: when the request sends authToken: null, it falls back to the token already stored for that
model key rather than testing with no auth at all.
embedding_cache
Cache table location and storage engine are, like embedding_client_profiles above, a separate,
deploy-time-only Symfony configuration node (config/packages/*.yaml, not the embeddings
block) — never persisted to the settings store and never editable through Studio:
pimcore_backend_power_tools:
embedding_cache:
database: '' # optional schema for the cache table; empty = Pimcore's main schema
table_storage_engine: '' # explicit engine; empty = auto-detect at table creation
| Option | Default | Description |
|---|---|---|
database | '' | Optional schema name on the same MySQL/MariaDB server that hosts the embedding vector cache table. Empty (the default) keeps the table in Pimcore's main schema. Must match ^[A-Za-z0-9_]*$. Qualifies the cache table name on every read/write — injected straight into the repository. Cross-schema DDL is not something the Installer or a Doctrine migration on this connection can express, so both skip creating the table and print the manual CREATE TABLE statement instead — see Provisioning the table. |
table_storage_engine | '' | Storage engine for the embedding vector cache table. Empty (the default) auto-detects one at table-creation time — see Storage engine. An explicit value is verified to exist on the server and bypasses auto-detection, including its replication guard. Must match ^[A-Za-z0-9_]*$. Only takes effect at table-creation time, via the Installer (fresh install) or the bundle's migrations (existing installation); neither ever ALTERs an existing table onto a different engine. |
usages
What each search capability embeds with. The default four usages are semantic_text_search,
object_dedup, image_search and image_dedup. A usage name is validated against the set of
registered usage definitions — see
Registering a Custom Embedding Usage to add one
of your own.
usages:
semantic_text_search: my-text-model # shorthand: documents AND queries use this model
image_search:
model: my-vision-model # embedded the indexed images
query_model: my-vision-text-model # embeds the incoming text query (cross-modal)
query_model defaults to model; when it differs, both must declare the same shared_space.
chunking
| Option | Default | Description |
|---|---|---|
max_input_chars | 500000 (min 2000) | Cost/scope guard: composed text beyond this is truncated before chunking (truncation is logged), bounding how many embedding requests one element can cause. About 590 chunks at the default. |
How text is split. Per model, in characters (models.<id>.max_text_chars, models.<id>.chunk_chars,
see Sizing). A composed text at or below max_text_chars is embedded
whole and byte-identical (no whitespace normalisation), so short elements' identity keys never depend on
chunking. Longer text is split into chunks of chunk_chars with a 15 % overlap, cut at sentence or word
boundaries, so roughly one chunk per chunk_chars minus overlap (850 characters with the defaults). With the
defaults a 30,000-character document yields about 35 chunks.
Changing it. Nothing about search quality argues for or against the cap: nested kNN ranks by the best-matching chunk, and a long document occupies one result slot however many chunks it has. What the cap controls is cost, per element and per text field:
| Cap | Chunks | Nested vectors with evaluation mode (3 formats) | Inference requests (96 texts each) | Stored _source (768-dim floats) |
|---|---|---|---|---|
| 100,000 | ~118 | ~354 | 2 | ~2 MB |
| 500,000 | ~588 | ~1,764 | 7 | ~9 MB |
| 1,000,000 | ~1,176 | ~3,528 | 13 | ~18 MB |
The hard ceilings are the search engine's: OpenSearch allows 10,000 nested objects per document by
default (index.mapping.nested_objects.limit, not overridden by the generic data index), and one
bulk request must stay under http.max_content_length (100 MB by default). The default of 500,000
is the practical upper bound: it keeps a document's _source in the tens of megabytes even with
evaluation mode on, and stays well clear of the nested-object limit. Lower it to save inference
and index cost on installations with many long documents; above it, prefer splitting the source
document.
cache_eviction
Controls bpt:embeddings:reindex --evict (see Caching):
| Option | Default | Description |
|---|---|---|
max_age_days | 90 (min 1) | Cache rows untouched for longer than this are pruned first. |
max_rows | 1000000 (min 1) | If the table is still over this row count after age-based pruning, the least-recently-used rows are removed down to the cap. |
Vector storage, reuse and index recreation
Vectors live in the search index documents and in a durable MySQL/MariaDB cache table, and are
reused via version-aware identity keys, so re-indexing unchanged content costs no inference calls.
serving_version, the cache table, and the ways to force a re-embed are all covered in
Caching.
How the defaults were chosen
The measured reasoning behind the default model, text representation, chunking and compression choices lives in Defaults and Benchmarks.