Description
The encoder and decoder disagree about whether "1_000" is a number.
- Encoder: decides whether a string needs quotes by checking it against
NUMERIC_REGEX. "1_000" doesn't match that pattern, so it is written unquoted as 1_000.
- Decoder:
is_numeric_literal() checks a token by calling Python's float(token). Since PEP 515, float() accepts underscores between digits, so 1_000 passes the check. decoder.py then converts it with int(token), which accepts underscores too.
As a result, the string "1_000" is encoded as 1_000 and decoded as the integer 1000. Nothing raises an error, and the value's type and content both change silently. It was first seen with identifiers that contain underscores between digits. They came back as integers and broke equality checks further down the pipeline.
Reproduction Steps
import toon_format
data = "1_000"
encoded = toon_format.encode(data)
decoded = toon_format.decode(encoded)
print(f"{data!r} -> {encoded!r} -> {decoded!r}")
Expected Behavior
decode(encode(x)) == x, so "1_000" should come back as the string "1_000". The decoder should read 1_000 as a string, because underscores aren't part of the TOON number grammar.
Actual Behavior
'1_000' -> '1_000' -> 1000
No error or warning is raised. The value's type changes from str to int, and in cases like "1_0" → 10 the value itself changes.
Related decoder results:
decode('0_0') returns '0_0', but only because the leading-zero check runs before float().
decode('1__0') returns '1__0', because float() rejects double underscores.
Environment
toon_format 0.9.0b1
- CPython 3.14.7
- macOS 26.6.2 (arm64)
Additional Context
- Our workaround: we kept
toon_format for encoding, because we need options like custom delimiters, and decode with toons 0.9.0, which keeps 1_000 as a string.
Description
The encoder and decoder disagree about whether
"1_000"is a number.NUMERIC_REGEX."1_000"doesn't match that pattern, so it is written unquoted as1_000.is_numeric_literal()checks a token by calling Python'sfloat(token). Since PEP 515,float()accepts underscores between digits, so1_000passes the check.decoder.pythen converts it withint(token), which accepts underscores too.As a result, the string
"1_000"is encoded as1_000and decoded as the integer1000. Nothing raises an error, and the value's type and content both change silently. It was first seen with identifiers that contain underscores between digits. They came back as integers and broke equality checks further down the pipeline.Reproduction Steps
Expected Behavior
decode(encode(x)) == x, so"1_000"should come back as the string"1_000". The decoder should read1_000as a string, because underscores aren't part of the TOON number grammar.Actual Behavior
No error or warning is raised. The value's type changes from
strtoint, and in cases like"1_0"→10the value itself changes.Related decoder results:
decode('0_0')returns'0_0', but only because the leading-zero check runs beforefloat().decode('1__0')returns'1__0', becausefloat()rejects double underscores.Environment
toon_format0.9.0b1Additional Context
toon_formatfor encoding, because we need options like custom delimiters, and decode withtoons0.9.0, which keeps1_000as a string.