Repository navigation
fix(streaming): preserve first SSE event after BOM - #3819
Hughhhhcoder wants to merge 1 commit into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Independent offline verification (AI-assisted): I independently tested LF/CRLF/CR streams with no BOM, one leading BOM, and two leading BOMs at every two-part byte split position (291 cases each for sync and async decoding). Main passed 194/291 per decoder; this head passed 291/291 per decoder. The 97 baseline failures were the single-leading-BOM cases. A BOM inside event content remained intact. The double-BOM control checks that only one leading BOM is ignored, consistent with the WHATWG event-stream interpretation. These are fresh decoder instances with local byte iterators, not network tests. Compared main Reproducer, runnable separately against each source tree: import asyncio,json
from openai._streaming import SSEDecoder
import openai._streaming as source
async def main():
rows=[]
for eol in [b'\n',b'\r\n',b'\r']:
for label,prefix,first in [('plain',b'','first'),('bom',b'\xef\xbb\xbf','first'),('double_bom',b'\xef\xbb\xbf'*2,None)]:
payload=prefix+b'data: first'+eol+eol+b'data: second'+eol+eol
expected=([first] if first else [])+['second']
for split in range(len(payload)+1):
chunks=[payload[:split],payload[split:]]
actual=[x.data for x in SSEDecoder().iter_bytes(iter(chunks))]
async def chunks_async():
for chunk in chunks:yield chunk
other=[x.data async for x in SSEDecoder().aiter_bytes(chunks_async())]
rows.append({'line_ending':eol.decode(),'prefix':label,'split':split,'sync_pass':actual==expected,'async_pass':other==expected,'actual':actual,'expected':expected})
# A BOM inside payload content must remain content, not be stripped.
payload=b'data: first\n\ndata: \xef\xbb\xbfsecond\n\n'
interior=[x.data for x in SSEDecoder().iter_bytes(iter([payload]))]
assert interior==['first','\ufeffsecond']
print(json.dumps({'source_module':source.__file__,'split_cases':len(rows),'sync_passed':sum(r['sync_pass'] for r in rows),'async_passed':sum(r['async_pass'] for r in rows),'interior_bom_preserved':True,'cases':rows,'scope':'Fresh decoder per stream, all two-part split positions; double-leading BOM strips at most one per WHATWG. Not network or full SDK proof.'}))
asyncio.run(main()) |
Castiron custom codeEvaluated main: ✅ No new custom-code files detected. 36 mixed files remain; 0 existing customizations changed. Compared 36 existing customizations unchanged
A changed generated baseline means this report cannot reliably identify which handwritten lines changed. Inspect the custom-code diffDownload the exact patch produced by this run (requires repository access): gh run download 37738615493 --repo openai/openai-python \
--name castiron-custom-code-37738615493-1 --dir /tmp/castiron-custom-code-37738615493-1
git apply --stat /tmp/castiron-custom-code-37738615493-1/custom-code.patch
cat /tmp/castiron-custom-code-37738615493-1/custom-code.patchOr reproduce it from an SDK checkout containing the vendored reporter: git fetch --no-tags origin 4e152cdefe1844c2d5d78653310e9b9c0195c44e 3601840f2e6e59b352facccd0109d17e9871dc40
python3 scripts/castiron/custom_code_report.py report \
--base 4e152cdefe1844c2d5d78653310e9b9c0195c44e \
--head 3601840f2e6e59b352facccd0109d17e9871dc40 --fetch --require-head-hash --public \
--out /tmp/castiron-custom-code-3601840f2e6e
cat /tmp/castiron-custom-code-3601840f2e6e/custom-code.patchThis is the current full custom patch for mixed files, not an attribution of only the handwritten lines changed by this PR. |
Changes being requested
Strip the single leading UTF-8 byte-order mark allowed by the SSE format before interpreting the first line.
Today
SSEDecoderdecodes each completed line independently. A valid stream beginning with a BOM therefore presents its first field as\ufeffdatainstead ofdata, so the decoder silently drops the first event. The WHATWG event-stream interpretation algorithm explicitly strips one leading UTF-8 BOM.The decoder now tracks the first line of each stream and removes only a leading BOM there. The regression covers synchronous and asynchronous decoding with the three BOM bytes both contiguous and split across input chunks.
Additional context & links
BOM + data:first + blank line + data:secondyielded onlysecondfirst, thensecondValidation:
uv run --locked --all-extras pytest tests/test_sse_framing.py— 48 passed./scripts/test-pydantic-v1 tests/test_sse_framing.py— 48 passed./scripts/lint— Ruff, Pyright, Mypy, and import checks passedgit diff --check origin/main...HEAD— passed