Skip to content

fix(lib): handle null and empty text in parse_text #3966

Description

@Kuldeeep18

Confirm this is an issue with the Python library and not an underlying OpenAI API

  • This is an issue with the Python library

Describe the bug

In src/openai/lib/_parsing/_responses.py, parse_text() attempts to deserialize text into structured output (model_parse_json or json.loads) without verifying that text is non-empty / non-null:

def parse_text(text: str, text_format: type[TextFormatT] | Omit, *, phase: str | None) -> TextFormatT | None:
    if phase not in (None, "final_answer"):
        return None

    if not is_given(text_format):
        return None

    if is_basemodel_type(text_format):
        return cast(TextFormatT, model_parse_json(text_format, text))

    return cast(TextFormatT, json.loads(text))

When an output message contains empty or null text (for instance, in streaming response.output_text.done events where no tokens were emitted yet, message parts with content: [{"type": "output_text", "text": ""}], or custom endpoints/mocked transports), parse_text() raises:

  • TypeError: argument 'data': 'NoneType' object cannot be converted to 'PyString' when text is None
  • ValidationError: JSON decode error: EOF while parsing a value (or json.decoder.JSONDecodeError) when text == ""

Consistency with _completions.py and #3851:

  1. In src/openai/lib/_parsing/_completions.py:195, maybe_parse_content() already guards against falsy content before attempting parsing:
    if has_rich_response_format(response_format) and message.content and not message.refusal:
        return _parse_content(response_format, message.content)
    return None
  2. In fix(lib): treat null message content as empty in parse_response #3851, null output.content was handled by treating it as empty. Adding an early guard if not text: return None in parse_text() provides the same defensive handling for output text.

Proposed fix

In src/openai/lib/_parsing/_responses.py:

def parse_text(text: str | None, text_format: type[TextFormatT] | Omit, *, phase: str | None) -> TextFormatT | None:
    if phase not in (None, "final_answer"):
        return None

    if not is_given(text_format):
        return None

    if not text:
        return None

    if is_basemodel_type(text_format):
        return cast(TextFormatT, model_parse_json(text_format, text))

    return cast(TextFormatT, json.loads(text))

Ready branch & tests

The change (+5/-2) and comprehensive sync/async unit tests covering both null and empty text in static parsing and streaming (+44 lines) are tested and ready in fork branch:
👉 https://github.com/Kuldeeep18/openai-python/tree/fix/responses-parse-null-text
(Commit: 6fd8a7b)

To Reproduce

from pydantic import BaseModel
from openai.lib._parsing._responses import parse_text

class Result(BaseModel):
    answer: str

# 1. Null text:
parse_text(None, Result, phase="final_answer")
# -> TypeError: argument 'data': 'NoneType' object cannot be converted to 'PyString'

# 2. Empty text:
parse_text("", Result, phase="final_answer")
# -> pydantic_core._pydantic_core.ValidationError: JSON decode error: EOF while parsing a value

OS

All platforms (cross-platform library parsing issue)

Python version

Python 3.10+

Library version

Latest main (v3.x)

Activity

  1. Kuldeeep18 commented on Sep 29, 2026

    @Kuldeeep18
    Author

    Thanks for verifying and confirming this @dajiaohuang!

    Regarding the contribution route: code under src/openai/lib/ is handwritten (not auto-generated by Stainless), and community PRs in src/openai/lib/ have been accepted before (e.g. #3851).

    I previously submitted the complete fix and unit tests in PR #3897, which already has independent regression proof from community reviewers.

    cc @marcuswood-oai — following your suggestion on #3897, this issue captures the minimal reproductions for None and "" text along with the consistency rationale with _completions.py:195. Whenever you have a chance to take a look, please let us know if reopening #3897 or having a collaborator carry it forward works best.

  2. added a commit that references this issue on Oct 8, 2026
    9109efc
  3. XZQ173 commented on Oct 8, 2026

    @XZQ173

    I put together a fix for this on a fork before noticing the contribution policy (PRs are limited to collaborators), so leaving the analysis and patch here instead of opening a PR — issue comments are welcome per CONTRIBUTING.

    Root cause

    parse_text() in src/openai/lib/_parsing/_responses.py deserializes text without checking that it is non-empty:

    if is_basemodel_type(text_format):
        return cast(TextFormatT, model_parse_json(text_format, text))
    
    if is_dataclass_like_type(text_format):
        ...
        return pydantic.TypeAdapter(text_format).validate_json(text)

    So:

    • text is None → TypeError: argument 'data': 'NoneType' object cannot be converted to 'PyString'
    • text == "" → ValidationError: Invalid JSON: EOF while parsing a value (or json.JSONDecodeError)

    Proposed fix

    Add an early guard after the phase / text_format checks:

    if not text:
        return None

    and widen the annotation to str | None, since callers may legitimately pass None.

    Why this is consistent with the rest of the library

    1. _completions.maybe_parse_content() already guards against falsy content before parsing:
      if has_rich_response_format(response_format) and message.content and not message.refusal:
          return _parse_content(response_format, message.content)
      return None
    2. fix(lib): treat null message content as empty in parse_response #3851 treated null output.content as empty. The guard extends the same defensive handling to output text, which is what streaming hits first: a response.output_text.done event can carry no tokens yet.

    Returning None keeps the TextFormatT | None contract — "no structured result" is the correct answer for empty text, rather than an exception.

    Test

    Regression test (passes with the fix, fails without — the 4 failures reproduce exactly the TypeError / JSON decode errors above):

    @pytest.mark.parametrize("text", [None, ""], ids=["none", "empty"])
    @pytest.mark.parametrize("phase", [None, "final_answer"], ids=["no-phase", "final-answer"])
    def test_parse_text_returns_none_for_empty_text(text, phase):
        assert parse_text(text, Answer, phase=phase) is None
    
    
    def test_parse_text_still_parses_non_empty_text():
        assert parse_text('{"answer": 4}', Answer, phase="final_answer") == Answer(answer=4)

    Verified against tests/lib/responses/ (760 passed) and tests/lib/chat/ (76 passed); the 9 failures in tests/lib/streaming/agents/test_files.py are pre-existing in my environment (symlink/trio) and unrelated.

    Patch and tests are on Xx-173/openai-python@fix/parse-text-empty-guard if a maintainer wants to pick them up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions