Skip to content

fix: send the query embedding as a vector in external pgvector retrieval - #6

Closed
macodev00 wants to merge 96 commits into
devfrom
fix/pgvector-external-cast
Closed

macodev00 wants to merge 96 commits into
devfrom
fix/pgvector-external-cast

Conversation

@macodev00

@macodev00 macodev00 commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Fork CI branch only — not an upstream PR.

Rebases fix/pgvector-external-cast onto latest open-webui/open-webui dev for issue open-webui#26663 (prior closed PR open-webui#28363).

Change

_retrieve_pgvector in backend/open_webui/retrieval/external.py wraps the query embedding in Vector(...) so pgvector receives a vector instead of float8[].

Single commit, single file. Unrelated ruff drive-bys were dropped; the touched file already passes ruff format --check.

Changelog Entry

Fixed

  • External knowledge retrieval against pgvector no longer fails with operator does not exist: vector <=> double precision[].
Open in Web Open in Cursor 

@cursor
cursor Bot force-pushed the fix/pgvector-external-cast branch 2 times, most recently from f63cd42 to 9adb611 Compare September 21, 2026 06:21
andreadegiovine and others added 28 commits September 21, 2026 08:04
* i18n: update and improve it-IT

* i18n: update and improve it-IT

---------

Co-authored-by: andreadegiovine <andreadegiovine@gmail.com>
Staan rejects any count other than its fixed page size of 10, so every search 400ed with the default result count of 3. The API always returns 10 results; the local results[:count] slice already applies the configured count, so the request parameter is simply dropped.
…pen-webui#29757)

Retiring a model leaves every chat that used it stuck. The chat still stores the old model id, so it loads with nothing usable selected and refuses to send until the user picks a replacement by hand, in every old conversation. There is no admin-side way to move those chats across.

Loading a chat now drops model ids that no longer exist, and an empty selection falls into the fallback this function already applies to new chats: the user's default model, then the admin-configured default, then the first available model. A chat that still has one live model keeps it. The filter runs before the single-model permission clamp so a restricted user keeps a live model the chat already has instead of losing it to the clamp.

The model list cannot distinguish a retired model from one whose connection is momentarily unreachable, so during such an outage a chat pinned to that connection will open on the fallback model and record it when saved. Models defined in the workspace are unaffected, since they stay in the list while their connection is down.
…home page (open-webui#29762)

Typing a message, uploading a file or toggling web search in a chat started from the home page was lost on reload: the input came back empty and the uploaded files were gone for good.

Drafts are kept per chat in sessionStorage, but the key was built from the route prop, which stays empty for these chats because creating the chat only swaps the URL with history.replaceState and never re-runs the route load. Drafts were written under the new-chat key while a reload of /c/<id> read the per-chat key. Falling back to the active chat id lines the two up, the same fallback loadChat already uses for this reason.

The Placeholder submit handler also cleared the bare key rather than the current one, which left a stale draft behind once a chat's messages had all been deleted.

Verified against the base commit on a local instance: on base the draft lands under chat-input and is lost on reload, with the change it lands under chat-input-<id> and both the text and the uploaded file come back. Image attachments stay excluded from drafts by design.

The model selection half of open-webui#29760 has a different cause and is not fixed here.

Refs open-webui#29760
…f the URL host (open-webui#29691)

The native edit_image tool fails with "400: [ERROR: Error loading image]" whenever the model hands it an absolute URL for an image Open WebUI already stores. Such a URL is treated as local only when its host string matches the incoming request's host exactly, so a default-port form, a container name or any host the model composed itself falls through to an outbound HTTP fetch instead. That fetch asks /api/v1/files/{id}/content without a session, gets a 401, and the user sees the generic 400.

Match the file URL on its path and let the existing local branch resolve it. Fetching that endpoint over the network can never succeed for a local or a remote instance, because it requires an authenticated user, so the host comparison only decided which way the request failed. Access control is unchanged: the local branch still goes through get_file_content_by_id, which enforces owner, admin or shared access.

Fixes open-webui#29220
open-webui#29262)

* fix: correct recurrence rule parsing for schedules and calendar events

An automation set to repeat a limited number of times, say ten or a hundred, was treated as a one-shot and reported no repeat interval, because any count whose digits began with a one matched a text check for the one-shot case. The scheduler already answers that question correctly by asking the rule for its next two occurrences, so the text check is gone and the count is read as the number it is.

A recurrence rule that carries its start date on the same line as the repeat text kept that date when the automation was parsed, so the schedule ran from whatever date the rule happened to carry and ignored the start the user picked. The filter that drops the start date now splits the rule on any whitespace, the same way the rule parser itself does, so both agree on where one part of the rule ends and the next begins.

The same mismatch on the calendar path anchored a recurring event to the date inside its rule, so occurrences showed up before the event had begun and at the wrong time of day. That filter splits the rule the same way now, and the series starts at the event's own start.

All three come from one place, recurrence rules being matched and cut as text. Rules written across several lines, which is what the schedule and calendar editors produce, behave exactly as before.

* refac: name rrule token vars parts to match calendar.py
…ages (open-webui#29623)

The remote chat-image fetch now asks for an identity-encoded response and skips one that comes back content-encoded.

An image whose host stores and echoes a `Content-Encoding` regardless of what the client asks for (an S3 or MinIO object uploaded with that metadata) is no longer inlined; the message is forwarded with the original URL instead, the same way an unreachable image already behaves.
…rver (open-webui#28185)

Every client in a note re-broadcasts each update it receives, with a full content snapshot attached, and the server appends each echo to the document log and writes the note again. Traffic and note writes therefore scale with the number of people who have the note open: every extra participant adds one more full echo of every keystroke.

Remote updates are now applied with the `'server'` origin the state path already uses, which the local listener ignores, so the echo stops. The echo did carry one thing worth keeping: the receiving client holds the merged document, which the sender had not seen yet, so each receiver now sends a content-only message once the edits settle, debounced 500ms, and flushes a pending one when the editor is torn down. The backend accepts an update message with no `update` field for that case, where it previously raised and dropped the save.

Replaying keystrokes at 120ms with a real Yjs document, socket messages fall 48.8% with two clients, 65.6% with three and 79.2% with five, with byte counts tracking the same on notes up to 50KB. Server-side update appends drop by a factor of the client count. The cost is one snapshot upload per receiving client per typing pause, which the server's own debounce then collapses into a single extra note write however many people are watching. A snapshot has to come from a client because the server cannot rebuild the markdown, HTML and JSON shape the note record stores. It also moves the merge 500ms later than the echo delivered it, so if two edits cross on the wire and every editor then loses its connection and closes inside that window, the note keeps what it was last sent and one of the two edits is lost. A client still connected at teardown flushes its pending snapshot, and any other editor left in the note closes the gap.
…open-webui#28184)

Long conversations get progressively more expensive to stream into. Every message-level write reloads the entire chat JSON, walks every string in it for a null-byte sanitisation pass and rewrites the whole column, and reading a single message loads and validates the whole chat too. The websocket event emitter reads and then writes, so one message event round-trips the full conversation twice to change a few hundred bytes.

Reading one message now selects only history.messages out of the JSON column and indexes it in Python, so the rest of the chat is neither loaded nor model-validated. Writing one message now sanitises only what is actually entering the chat, plus the title column that is mirrored off the blob, instead of re-walking a conversation that was already sanitised when it was written. Legacy rows are still healed on read by get_chat_by_id, which is unchanged.

One behaviour change: chats written before the null-byte sanitisation existed can still hold null bytes in the stored JSON. Those rows used to be rewritten clean as a side effect of any message write, and are now cleaned when the chat is read instead. The visible consequence is that the message write and delete endpoints echo back the updated chat, and for such a legacy row that echo now carries the raw null bytes rather than stripped ones, until the next read of that chat heals it. A GET of the chat is unaffected, and title, the one text column that PostgreSQL cannot store a null byte in, is still sanitised on every write.

Measured on SQLite with a 10.8 MB chat (3000 messages): a single-message upsert is 227 ms before the branch and 173 ms after, and the legacy single-message read is 135 ms before and 79 ms after. The whole-blob scan the delete path used to run costs 5.40 ms on a 1.8 MB chat and 38.79 ms on the 10.8 MB one, and is gone.

Ref open-webui#28169

### Contributor License Agreement

<!--
🚨 DO NOT DELETE THE TEXT BELOW 🚨
Keep the "Contributor License Agreement" confirmation text intact.
Deleting it will trigger the CLA-Bot to INVALIDATE your PR.

Your PR will NOT be reviewed or merged until you check the box below confirming that you have read and agree to the terms of the CLA.
-->

- [x] By submitting this pull request, I confirm that I have read and fully agree to the [Contributor License Agreement (CLA)](https://github.com/open-webui/open-webui/blob/main/CONTRIBUTOR_LICENSE_AGREEMENT), and I am providing my contributions under its terms.

> [!NOTE]
> Deleting the CLA section will lead to immediate closure of your PR and it will not be merged in.
…#28176)

With the Redis websocket manager the shared model pool is a `RedisDict`, so resolving a model id fetches from Redis, and more than one of those fetches pulls the whole pool. Every chat request pays that latency and that traffic, and the cost grows with the number of models configured: at 120 models a request moves several hundred kilobytes to ask about models that have not changed since the last request.

The pool now keeps a per-worker cache of the hash and refetches it only when the signature key that `set()` already maintains changes. A cache may only trust a signature that describes the bytes it fetched, so every write invalidates the signature and no signature value is ever issued twice. Without a fresh token each time, a pool that flaps back to an earlier state returns to an earlier digest, so a reader that races the write keeps serving the old pool. `delete_many()` was a second hole in that: it updated a dead attribute and never touched the signature, so readers kept serving models it had already removed. A TTL cache would have been simpler, but it serves a pool it already knows may be stale and costs the same single round trip this check costs.

Replaying the real access pattern, a request that finds the pool unchanged moves about 1 KB no matter how many models are configured, over the same number of round trips as before, except after a write that leaves no signature, which holds reads at two commands until the next non-empty `set()` writes one. A request that first sees a changed pool costs one extra command, and so does rewriting the pool. Values stay cached serialized, so `items()` and `values()` still decode on every call, and each worker holds the whole pool in memory.
Co-Authored-By: Classic298 <27028174+Classic298@users.noreply.github.com>
…webui#30308)

Web search via searchapi.io could come back empty or near-empty with no
hint of why: an invalid or expired API key turned into an empty result
set instead of an error, the google_news engine splits its results
between organic_results and top_stories and only the first block was
read, and google links came back as google.com/goto redirects the web
loader cannot fetch, so citations pointed at a redirect blob.

The search now reads both result blocks, asks google engines for
resolved destination links, raises on HTTP errors, carries a 30s request
timeout, skips result rows without a link, and logs the response body at
debug instead of dumping every search at info.

Fixes open-webui#30305
silentoplayz and others added 27 commits September 23, 2026 23:32
…open-webui#30450)

open-webui#30426 stopped concurrent requests from refreshing the same OAuth session twice, but its lock only lives inside one process. With several uvicorn workers or replicas, two requests on different workers still send the same refresh token, a rotating provider rejects the second with invalid_grant, and the session gets deleted, so the user's OAuth session is logged out again.

When Redis is configured, which multi-worker and multi-replica deployments require, the refresh now takes a Redis lock per session instead of the in-process one. Single-process deployments without Redis keep the in-process lock. The waiter re-reads the session inside the lock as before and uses the token that was just stored.

It uses redis-py's own async lock because the existing RedisLock is synchronous and never waits. The Sentinel proxy now passes `lock` through unwrapped like `pipeline` and `pubsub`; otherwise it returned a coroutine and every refresh behind Sentinel would fail.

Tested with separate OS processes on one sqlite DB, a real Redis and a rotating mock provider: 2 and 5 processes (and 5 processes x 3 requests) now cause 1 refresh, every caller gets the new token and the session is kept (before: one refresh per process, session deleted every run). Single refresh, failed refresh, valid token and the single-process path without Redis are unchanged.

Follow-up to open-webui#30426, refs open-webui#30416
Folder file entries are also checked against the user making the change.
Generates the default WEBUI_SECRET_KEY file with secrets.token_bytes, matching how the start scripts read from the OS random source.
Removing members from a group or DM channel now also removes their sessions from the channel room, matching how access grant changes are handled.
open-webui#30390)

On Admin Settings > Models the drag handle was disabled as soon as a search, view or tag filter was active, so on a long list the only way to move a model was to clear everything and hunt for it by eye.

Dragging now works in any filtered list. The move is applied to the full order: the dragged model is placed directly after the visible model it was dropped below (or directly before the one it was dropped above), and every model hidden by the filter keeps its place. That anchoring is what makes reordering a subset safe, which is why the filters previously blocked it.

The tag filter now filters the loaded list client-side like search and view already do. Before, it reloaded the page data and rebuilt the order from only the tagged models, so saving under a tag would have dropped every other model from the order, and switching tags discarded unsaved moves. Export keeps its existing tag-filtered behaviour.

Verified in the browser with search, view (enabled/disabled) and tag filters, dragging up and down, multiple moves before one save and switching tags with unsaved moves; the saved order always contains every model.

Closes open-webui#29634
…pen-webui#30385)

Attach Webpage fails with 403 Forbidden on sites that reject the bare aiohttp user agent, Wikipedia among them, even when USER_AGENT is set. The web loader sends USER_AGENT, but the request that runs first to decide whether the URL is a page or a file does not, so the attachment fails before the loader is ever reached.

The pre-check now sends USER_AGENT as the request User-Agent when it is set. With it unset the request is unchanged and keeps the aiohttp default.

Verified against the real _fetch_url with https://en.wikipedia.org/wiki/OpenAI: 403 before, page detected after; USER_AGENT unset still returns the same 403 as before, and a direct PDF URL is still detected as a file.

Fixes open-webui#29617
…eter (open-webui#30382)

Asked for a PowerPoint or Word file, the model had no library for it, so it followed the prompt's "use an alternative approach" and wrote the OOXML zip by hand. The sandbox reported success and Office refused to open the result.

python-pptx and python-docx are now vendored the same way as openpyxl: their wheels (plus xlsxwriter) go through the PyPI wheel path, lxml joins the Pyodide distribution list so its wasm wheel is cached, and importing pptx or docx installs them from the bundled wheels. No prompt change is needed, since the app installs on import.

static/pyodide grows by about 3 MB (58 to 61 MB) and verifyBundledWheels() passes.

Verified in headless Chromium with pypi.org, files.pythonhosted.org and the jsDelivr CDN blocked: both packages install from the local wheels only, and a deck and a document built in the sandbox reopen with the native libraries. openpyxl, seaborn, black, pandas, matplotlib and requests still install offline. Without the change both installs fail offline.

Fixes open-webui#30361
…ui#30379)

Automations lost their model's tool bindings (including MCP servers), default features such as web search, default filters and terminal on the first run after a restart. The model then answered that it had no tools. Later runs and a manual Regenerate worked. Automations on the base models cache were not affected.

The run read the model from app.state.MODELS before anything had loaded it. After a restart or a connection settings save, that cache stays empty until a browser loads the model list or a chat completion runs. The completion runs only after the automation has already built its request.

execute_automation now loads the models when the cache is empty, using the same guard chat_completion uses, before either the chat or the channel target reads it. This also fixes channel automations showing the raw model ID in place of the model name on a cold cache.

Verified end to end on a restarted instance with a mock upstream: before, both chat and channel runs reached the pipeline without tool_ids or features. After, both carry the model's tools and web search, and the upstream receives the tool.

Fixes open-webui#27694
…pen-webui#30381)

Any Python code that only mentioned "matplotlib" (a comment, a string, or an importlib.util.find_spec("matplotlib") check) failed with ModuleNotFoundError before a single line of it ran. The plt.show() patch was applied on a plain substring match, but matplotlib is only installed when the code actually imports it, so the patch's own import crashed the run. The user's try/except could not catch it, because their code never started. The same thing broke the code editor's Python formatter on any code that imports matplotlib, since that run only installs black.

The patch now also requires matplotlib to be present in the runtime's loaded packages, in both the worker and the sandboxed iframe host. Real matplotlib code still gets inline PNG output from plt.show(), and code that merely mentions matplotlib runs unchanged.

Gating on the callers' import regex instead was also tested and still fails the formatter case, because there the code containing the import sits inside a string while only black is installed.

Fixes open-webui#29894
…-webui#30383)

Adding a second model to a chat turned off Web Search, Image Generation and Code Interpreter even when every selected model has them as Default Features, so the comparison ran without them. Changing the selection clears the feature toggles and then reapplies the model defaults, but defaults were only reapplied for a single model.

In compare mode a default feature is now turned on when every selected model supports it and has it as a default, since the toggles are shared by all models in the comparison. It only ever turns features on, so flags passed in the URL (?web-search=true) are not overwritten. Single-model defaults are unchanged.

The input reset now waits for the model-selection update to finish before applying defaults. Before, the Web Search value actually sent could disagree with the toggle on screen: a single model with Web Search on by default showed the toggle on but sent it off, and switching to a model without it showed it off but still sent it on (see open-webui#29326). The toggle and the request now match.

Fixes open-webui#30310
…ayback silent on iOS (open-webui#30373)

On iOS and iPadOS, auto-playback of a finished reply is never heard and the speaker button stays stuck in "speaking", so the first tap only stops a playback that never started. WebKit rejects play() started from a network event, and the audio queue ignored that rejection, so its state never reset. Call mode on the same devices was silent too.

A rejected play() now resets the queue, returns the button to idle and shows a toast (an aborted play from stop or a message switch is ignored). The first user tap or keypress plays a 10 ms silent clip on the shared audio element, which WebKit then allows to play later without a gesture; if that attempt fails it retries on the next gesture. Call mode plays unmuted: WebKit pauses an element that is unmuted after play() outside a gesture, even once unlocked.

The unlock and call mode change follow the reporter's on-device tests (iPhone iOS 27, iPad iPadOS 26.6.2). Verified in Chromium with autoplay restricted: the base queue wedges and drops later chunks; with the fix it reports the error, recovers, the unlock plays once and never interrupts audio already playing, and chunks queued during the unlock still play.

Fixes open-webui#30262
…webui#30384)

MCP tool servers using OAuth 2.1 with dynamic client registration authorized without any scope when the authorization server left scope out of its registration response, which RFC 7591 allows (Atlassian and Notion do). Consent completed and the tool showed as connected, but the issued token lacked the scopes the resource requires, so every tool call was refused. Discovered scopes and the custom OAuth Scopes field were both affected.

The stored client now falls back to the scope sent in the registration request when the response has none. A scope the server does return is kept as is.

Connections registered before this fix already have a null scope stored. The protected resource metadata recovery that static-credential clients already use now also runs for dynamically registered clients, so those connections pick up the discovered scopes on the next load without registering again.

Fixes open-webui#29967
…#30419)

When an MCP tool returns an image or audio item, the file is uploaded to storage, but the whole MCP item, including its base64 payload, was also passed as upload metadata. That metadata is persisted in the file table's meta column, so every such result was stored twice: once in storage and once as base64 in the database, growing the DB and every file query that loads meta.

The MCP path now passes only chat_id, message_id and session_id, the same metadata the non-MCP tool image path already stores. Nothing reads the removed key.

Fixes open-webui#30411
…n-webui#30420)

When a ComfyUI workflow ends in the core "Save Image (Advanced)" node, ComfyUI finishes the job and saves the image, but Open WebUI returns an empty result, so the chat shows nothing. Image editing workflows such as the Qwen Image Edit template use this node by default.

Open WebUI only collects images from output nodes of type SaveImage and PreviewImage. This adds SaveImageAdvanced to that list. The node reports its files in the same format as SaveImage, so the rest of the download and storage path works unchanged. Generation and editing share this code, so both are fixed.

Fixes open-webui#30404
…cleanup (open-webui#30394)

The patterns that remove details blocks and inline images from task messages now stop at the next details tag, bracket or parenthesis.
…30421)

In voice mode, pressing M after a turn typed an "m" into the chat box instead of muting. After each voice message the chat input pulled keyboard focus back to itself, and the call overlay ignores M while focus is in a text field so typing still works.

The chat input now leaves focus alone while a call is open: clearing it after a send, the remount on the first message of a new chat and the editor's autofocus all skip focusing during a call. Clicking into the box and typing during a call still works, and outside a call the input focuses exactly as before.

Checked in a browser with a fake microphone: M mutes after the first, second and third turn, in new and existing chats; "m" typed into a focused input during a call still types.

Fixes open-webui#30406
…-webui#30393)

Each '%'-separated segment of the LIKE pattern is now matched with an atomic group.
… UI Scale on mobile (open-webui#30374)

On a phone with UI Scale at 1.2x or higher, the model selector in the chat composer covers the Integrations button, so it cannot be seen or tapped. From 1.3x it covers the + button too, which leaves no way to attach files or open integrations.

The model selector's width cap is rem based, so it grows with UI Scale, and the left button group could shrink to nothing, so the selector took the space first. The + and Integrations buttons now sit in a group that keeps its width, and only the toggled chips stay in the scrolling strip. When space runs out the model name truncates, and chips scroll as before.

Measured in the real app at a 360px mobile viewport: before, Integrations is covered at 1.2x and both buttons at 1.3x and 1.5x; after, both are tappable at every scale, with the model name truncated (114px at 1.3x). Button positions at 1x and on desktop are unchanged to the pixel, with and without chips.

Fixes open-webui#29989
@cursor
cursor Bot force-pushed the fix/pgvector-external-cast branch from 22f7b4b to e02af3d Compare September 24, 2026 06:23
@macodev00

Copy link
Copy Markdown
Owner Author

Skipping while Classic298 holds competing upstream PR open-webui#31112 for open-webui#26663. Dropping stale draft from queue.

@macodev00 macodev00 closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants