[Store] Forward platform options from Retriever to the vectorizer - #2367
wachterjohannes wants to merge 1 commit into
Conversation
|
were you using the Retriever directly in user land or via SimilaritySearch? |
| This mirrors the ``platform_options`` key of :class:`Symfony\\AI\\Store\\Indexer\\DocumentProcessor`, | ||
| which does the same for the indexing side. |
There was a problem hiding this comment.
i think we can drop that part here, not sure it adds something
| // Configures the query embedding, not the store query, so it must not reach the store: | ||
| // bridges like MongoDB merge the options they get straight into their query payload. | ||
| $platformOptions = $options['platform_options'] ?? []; | ||
| unset($options['platform_options']); |
There was a problem hiding this comment.
i'm wondering if we should restore the platform_options after the store->query call so the PostQuery has knowledge of it - WDYT?
ac21b71 to
a7715ff
Compare
|
Directly. Through Which is the uncomfortable half of the answer, because the case this fixes is a RAG one: with Cohere the query is embedded as Dropped the On restoring The asymmetry with If a post-query listener ever needs the embedding mode, I would add a dedicated accessor on the event rather than smuggle it back through the store options. |
a7715ff to
89bc2d0
Compare
Retriever::createQuery()calls$this->vectorizer->vectorize($query)with no options, so nothing on the query side can reach the platform that does the embedding.That is a problem for embedding models with asymmetric input modes. Cohere embeds a stored document and a search query differently, so that a short question lands next to the longer passage that answers it, and expects to be told which of the two it is looking at. Its default is
search_document, which is right for indexing and wrong for retrieval. Today aRetrieverembeds the query as if it were a document, and there is no option to change that. The results still look plausible, which is what makes it easy to miss: the two sides are simply misaligned.The indexing side already has the answer.
DocumentProcessor::process()reads aplatform_optionskey out of its options and forwards it to the vectorizer. This does the same for the query side, so both halves of a RAG pipeline are configured the same way:Notes:
platform_optionsis consumed by theRetrieverand stripped from the options handed to the store, since it configures the embedding and not the store query. That matters: the MongoDB bridge merges the options it receives straight into its$vectorSearchstage, so leaving the key in would send an unknown field to the vector database. Every other option is passed to the store unchanged.PreQueryEventlistener can setplatform_optionsas well, since listeners can already rewrite the options.vectorize()is called with an empty array, exactly as before.Tests cover the vector and hybrid paths, the listener case, that plain store options do not leak into the vectorizer, and that
platform_optionsdoes not leak into the store.src/store: 422 tests green, PHPStan clean, CS fixer run.