Skip to content

fix(streaming): preserve first SSE event after BOM - #3819

Open
Hughhhhcoder wants to merge 1 commit into
openai:mainfrom
Hughhhhcoder:codex/openai-python-sse-bom
Open

Hughhhhcoder wants to merge 1 commit into
openai:mainfrom
Hughhhhcoder:codex/openai-python-sse-bom

Conversation

@Hughhhhcoder

Copy link
Copy Markdown
Contributor
  • I understand that this repository is auto-generated and my pull request may not be merged

Changes being requested

Strip the single leading UTF-8 byte-order mark allowed by the SSE format before interpreting the first line.

Today SSEDecoder decodes each completed line independently. A valid stream beginning with a BOM therefore presents its first field as \ufeffdata instead of data, so the decoder silently drops the first event. The WHATWG event-stream interpretation algorithm explicitly strips one leading UTF-8 BOM.

The decoder now tracks the first line of each stream and removes only a leading BOM there. The regression covers synchronous and asynchronous decoding with the three BOM bytes both contiguous and split across input chunks.

Additional context & links

Validation:

  • uv run --locked --all-extras pytest tests/test_sse_framing.py — 48 passed
  • ./scripts/test-pydantic-v1 tests/test_sse_framing.py — 48 passed
  • ./scripts/lint — Ruff, Pyright, Mypy, and import checks passed
  • Castiron custom-code budget check — passed (6,626 / 10,000 lines)
  • git diff --check origin/main...HEAD — passed

@Hughhhhcoder
Hughhhhcoder requested a review from a team as a code owner September 8, 2026 06:33
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-08T06:36:21.400904Z 3601840 PR opened
🔒 Security Review ✅ Completed 2026-09-08T06:39:47.959125Z 3601840 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@jnohclee-rgb

Copy link
Copy Markdown

Independent offline verification (AI-assisted): I independently tested LF/CRLF/CR streams with no BOM, one leading BOM, and two leading BOMs at every two-part byte split position (291 cases each for sync and async decoding). Main passed 194/291 per decoder; this head passed 291/291 per decoder. The 97 baseline failures were the single-leading-BOM cases. A BOM inside event content remained intact. The double-BOM control checks that only one leading BOM is ignored, consistent with the WHATWG event-stream interpretation. These are fresh decoder instances with local byte iterators, not network tests.

Compared main 919b6236382f3f9f76d8453efb9f9ad03cff1126 with this PR head 3601840f2e6e59b352facccd0109d17e9871dc40 using Python 3.14.7 and a shared existing dependency environment (not a historical lockfile matrix). Source trees were isolated; runs denied network access. This is bounded additional evidence, not full SDK or merge approval.

Reproducer, runnable separately against each source tree:

import asyncio,json
from openai._streaming import SSEDecoder
import openai._streaming as source
async def main():
 rows=[]
 for eol in [b'\n',b'\r\n',b'\r']:
  for label,prefix,first in [('plain',b'','first'),('bom',b'\xef\xbb\xbf','first'),('double_bom',b'\xef\xbb\xbf'*2,None)]:
   payload=prefix+b'data: first'+eol+eol+b'data: second'+eol+eol
   expected=([first] if first else [])+['second']
   for split in range(len(payload)+1):
    chunks=[payload[:split],payload[split:]]
    actual=[x.data for x in SSEDecoder().iter_bytes(iter(chunks))]
    async def chunks_async():
     for chunk in chunks:yield chunk
    other=[x.data async for x in SSEDecoder().aiter_bytes(chunks_async())]
    rows.append({'line_ending':eol.decode(),'prefix':label,'split':split,'sync_pass':actual==expected,'async_pass':other==expected,'actual':actual,'expected':expected})
 # A BOM inside payload content must remain content, not be stripped.
 payload=b'data: first\n\ndata: \xef\xbb\xbfsecond\n\n'
 interior=[x.data for x in SSEDecoder().iter_bytes(iter([payload]))]
 assert interior==['first','\ufeffsecond']
 print(json.dumps({'source_module':source.__file__,'split_cases':len(rows),'sync_passed':sum(r['sync_pass'] for r in rows),'async_passed':sum(r['async_pass'] for r in rows),'interior_bom_preserved':True,'cases':rows,'scope':'Fresh decoder per stream, all two-part split positions; double-leading BOM strips at most one per WHATWG. Not network or full SDK proof.'}))
asyncio.run(main())

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Castiron custom code

Evaluated main: 4e152cdefe1844c2d5d78653310e9b9c0195c44e.

✅ No new custom-code files detected.

36 mixed files remain; 0 existing customizations changed.

Compared be928151372e → 3601840f2e6e. Generated baselines verified.

36 existing customizations unchanged
  • api.md
  • scripts/castiron/README.md
  • scripts/castiron/custom_code_report.py
  • scripts/castiron/test_custom_code_report.py
  • src/openai/init.py
  • src/openai/_client.py
  • src/openai/resources/audio/transcriptions.py
  • src/openai/resources/audio/translations.py
  • src/openai/resources/beta/beta.py
  • src/openai/resources/beta/responses/responses.py
  • src/openai/resources/beta/threads/runs/runs.py
  • src/openai/resources/beta/threads/threads.py
  • src/openai/resources/chat/completions/completions.py
  • src/openai/resources/embeddings.py
  • src/openai/resources/files.py
  • src/openai/resources/realtime/realtime.py
  • src/openai/resources/responses/responses.py
  • src/openai/resources/uploads/uploads.py
  • src/openai/resources/vector_stores/file_batches.py
  • src/openai/resources/vector_stores/files.py
  • src/openai/resources/videos.py
  • src/openai/resources/webhooks/init.py
  • src/openai/resources/webhooks/webhooks.py
  • src/openai/types/chat/init.py
  • src/openai/types/chat/chat_completion_message_tool_call.py
  • src/openai/types/fine_tuning/fine_tuning_job_integration.py
  • src/openai/types/responses/init.py
  • src/openai/types/responses/response.py
  • src/openai/types/responses/response_function_web_search.py
  • src/openai/types/responses/response_function_web_search_param.py
  • src/openai/types/responses/responses_client_event.py
  • src/openai/types/responses/responses_client_event_param.py
  • src/openai/types/responses/tool.py
  • src/openai/types/responses/tool_param.py
  • src/openai/types/webhooks/init.py
  • tests/api_resources/test_videos.py

A changed generated baseline means this report cannot reliably identify which handwritten lines changed.

Inspect the custom-code diff

Download the exact patch produced by this run (requires repository access):

gh run download 37738615493 --repo openai/openai-python \
  --name castiron-custom-code-37738615493-1 --dir /tmp/castiron-custom-code-37738615493-1
git apply --stat /tmp/castiron-custom-code-37738615493-1/custom-code.patch
cat /tmp/castiron-custom-code-37738615493-1/custom-code.patch

Or reproduce it from an SDK checkout containing the vendored reporter:

git fetch --no-tags origin 4e152cdefe1844c2d5d78653310e9b9c0195c44e 3601840f2e6e59b352facccd0109d17e9871dc40
python3 scripts/castiron/custom_code_report.py report \
  --base 4e152cdefe1844c2d5d78653310e9b9c0195c44e \
  --head 3601840f2e6e59b352facccd0109d17e9871dc40 --fetch --require-head-hash --public \
  --out /tmp/castiron-custom-code-3601840f2e6e
cat /tmp/castiron-custom-code-3601840f2e6e/custom-code.patch

This is the current full custom patch for mixed files, not an attribution of only the handwritten lines changed by this PR.

Full report and patch

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants