Skip to content

fix: decode base64 embeddings as little-endian floats - #3711

Open
HostX0 wants to merge 1 commit into
openai:mainfrom
HostX0:fix/embedding-base64-byte-order
Open

HostX0 wants to merge 1 commit into
openai:mainfrom
HostX0:fix/embedding-base64-byte-order

Conversation

@HostX0

@HostX0 HostX0 commented Aug 21, 2026 •

Copy link
Copy Markdown
  • I understand that this repository is auto-generated and my pull request may not be merged

Root cause

The base64 embedding representation contains little-endian float32 bytes, but the handwritten response parser used native byte order in both its NumPy and stdlib paths. On big-endian Python platforms, that silently produces incorrect vector values. The existing fixture was also native-endian, so it followed the test host instead of the wire format and could not expose the defect.

Fix

  • Decode the NumPy path with the explicit <f4 dtype.
  • Byte-swap the stdlib array("f") path on big-endian hosts.
  • Construct fixture bytes with struct.pack("<3f", ...) so the test represents the wire format on every host.
  • Cover the stdlib big-endian branch without requiring big-endian CI hardware.

The branch is rebased onto current main and preserves #3757: NumPy availability is still checked once per response, while each encoded vector is decoded only once.

Validation

On the current head with Python 3.12:

  • Pydantic v2: .venv/bin/pytest -q -n 0 tests/lib/test_embeddings.py — 60 passed
  • Pydantic v1.10.26: the same focused command — 60 passed
  • .venv/bin/ruff check src/openai/lib/_parsing/_embeddings.py tests/lib/test_embeddings.py — passed
  • .venv/bin/ruff format --check src/openai/lib/_parsing/_embeddings.py tests/lib/test_embeddings.py — passed
  • git diff --check — passed

@HostX0
HostX0 marked this pull request as ready for review August 21, 2026 13:38
@HostX0
HostX0 requested a review from a team as a code owner August 21, 2026 13:38
@HostX0
HostX0 force-pushed the fix/embedding-base64-byte-order branch from 6ad2fec to ac63c61 Compare August 29, 2026 18:24
@HostX0
HostX0 force-pushed the fix/embedding-base64-byte-order branch from ac63c61 to 1fc4841 Compare August 30, 2026 16:36

@sylvesterkaczmarek sylvesterkaczmarek left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks right. The wire bytes are little-endian, so <f4 makes the NumPy path explicit and byteswap() gives the stdlib path the same behaviour on big-endian hosts. Using struct.pack("<3f", ...) also makes the fixture independent of the test machine.

@github-actions

Copy link
Copy Markdown
Contributor

Castiron custom code

✅ No new custom-code files detected.

34 mixed files remain; 0 existing customizations changed.

Compared b19c2161b1ea → 1fc4841bdcfb. Generated baselines verified.

34 existing customizations unchanged
  • api.md
  • scripts/castiron/README.md
  • scripts/castiron/custom_code_report.py
  • scripts/castiron/test_custom_code_report.py
  • src/openai/init.py
  • src/openai/_client.py
  • src/openai/resources/audio/transcriptions.py
  • src/openai/resources/audio/translations.py
  • src/openai/resources/beta/beta.py
  • src/openai/resources/beta/responses/responses.py
  • src/openai/resources/beta/threads/runs/runs.py
  • src/openai/resources/beta/threads/threads.py
  • src/openai/resources/chat/completions/completions.py
  • src/openai/resources/embeddings.py
  • src/openai/resources/files.py
  • src/openai/resources/realtime/realtime.py
  • src/openai/resources/responses/responses.py
  • src/openai/resources/uploads/uploads.py
  • src/openai/resources/vector_stores/file_batches.py
  • src/openai/resources/vector_stores/files.py
  • src/openai/resources/videos.py
  • src/openai/resources/webhooks/init.py
  • src/openai/resources/webhooks/webhooks.py
  • src/openai/types/chat/init.py
  • src/openai/types/chat/chat_completion_message_tool_call.py
  • src/openai/types/fine_tuning/fine_tuning_job_integration.py
  • src/openai/types/responses/init.py
  • src/openai/types/responses/response.py
  • src/openai/types/responses/response_function_web_search.py
  • src/openai/types/responses/response_function_web_search_param.py
  • src/openai/types/responses/tool.py
  • src/openai/types/responses/tool_param.py
  • src/openai/types/webhooks/init.py
  • tests/api_resources/test_videos.py

A changed generated baseline means this report cannot reliably identify which handwritten lines changed.

Inspect the custom-code diff

Download the exact patch produced by this run (requires repository access):

gh run download 36599625795 --repo openai/openai-python \
  --name castiron-custom-code-36599625795-1 --dir /tmp/castiron-custom-code-36599625795-1
git apply --stat /tmp/castiron-custom-code-36599625795-1/custom-code.patch
cat /tmp/castiron-custom-code-36599625795-1/custom-code.patch

Or reproduce it from an SDK checkout containing the vendored reporter:

git fetch --no-tags origin b19c2161b1eac80fbf1f6f67a64a50af99c53356 1fc4841bdcfb1a5d21810f47d355945170486b3a
python3 scripts/castiron/custom_code_report.py report \
  --base b19c2161b1eac80fbf1f6f67a64a50af99c53356 \
  --head 1fc4841bdcfb1a5d21810f47d355945170486b3a --fetch --require-head-hash --public \
  --out /tmp/castiron-custom-code-1fc4841bdcfb
cat /tmp/castiron-custom-code-1fc4841bdcfb/custom-code.patch

This is the current full custom patch for mixed files, not an attribution of only the handwritten lines changed by this PR.

Full report and patch

@sylvesterkaczmarek sylvesterkaczmarek left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed current head 1fc4841 after the refreshed generated-baseline validation. The wire format is explicitly little-endian in both paths: NumPy uses <f4, while the stdlib array path byteswaps only on big-endian hosts. The fixture is machine-independent via struct.pack("<3f", ...), and the big-endian regression exercises the byteswap branch. Castiron reports no new custom-code files or changed existing customizations, and there are no unresolved review threads. No blocker from my review.

@jnohclee-rgb

Copy link
Copy Markdown

AI-assisted independent numeric check at 1fc4841bdcfb1a5d21810f47d355945170486b3a against parent b19c2161b1eac80fbf1f6f67a64a50af99c53356 on Python 3.14.7 / NumPy 2.5.3. I simulated big-endian host interpretation for both parser branches, with network access denied.

Unlike a test double that returns expected VALUES after observing byteswap(), this harness derives every result from the supplied bytes:

  • stdlib adapter stores initializer bytes, reverses each four-byte word when byteswap() is called, then uses struct.unpack(">Nf", ...);
  • NumPy adapter interprets native "float32" as ">f4", while passing the PR's explicit "<f4" through to real NumPy.

Little-endian fixture construction: base64.b64encode(struct.pack("<3f", 0.125, -2.5, 3.75)).

Both simulated branches on the parent produced [8.688050478813866e-44, 1.1748486324899266e-41, 4.0267712670837943e-41]; both on this head produced [0.125, -2.5, 3.75]. A second fixture [-0.0, 0.0, 1.0] recovered the values and negative-zero sign on both branches. Explicit encoding_format="base64" preserved the original encoded string on both revisions. Response identity, model and usage metadata were retained in all six cases per revision.

This adds byte-derived numerical evidence for the big-endian branch alongside the existing call-order regression test. It is a simulation on a little-endian machine, not execution on big-endian hardware, a live API test, or a full-suite result. All vectors and metadata were fictional.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants