Skip to content

docs: state that a root string starting with U+FEFF is quoted - #60

Merged
johannschopplich merged 2 commits into
mainfrom
docs/root-bom-quoting
Oct 2, 2026
Merged

johannschopplich merged 2 commits into
mainfrom
docs/root-bom-quoting

Conversation

@johannschopplich

@johannschopplich johannschopplich commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

An unquoted root string that starts with U+FEFF begins the document with a byte-order mark, which §12 already forbids encoders to emit – but §7.2 still allowed it unquoted. Encoders that follow §7.2 lose the first character on decode: toon-format/toon#339, and the Rust port has the same bug.

§7.2 now lists the case, with one encode fixture.

@johannschopplich johannschopplich changed the title docs: require quoting a root string that starts with a byte-order mark docs: state that a root string starting with U+FEFF is quoted Oct 1, 2026
sjawhar-agent Bot pushed a commit to sjawhar/toon-rust that referenced this pull request Oct 1, 2026
On line 1 a leading U+FEFF is a byte-order mark that decoders remove
(§12), so an unquoted root string starting with it lost that character
on decode: "\u{FEFF}abc" came back as "abc", "\u{FEFF}8" as the number
8, and "\u{FEFF}" as an empty object. Quote it instead, as the §7.2 rule
added by toon-format/spec#60 requires.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab
@johannschopplich
johannschopplich merged commit 6831ede into main Oct 2, 2026
1 check passed
@johannschopplich
johannschopplich deleted the docs/root-bom-quoting branch October 3, 2026 17:24
sjawhar-agent Bot pushed a commit to sjawhar/toon-rust that referenced this pull request Oct 4, 2026
On line 1 a leading U+FEFF is a byte-order mark that decoders remove
(§12), so an unquoted root string starting with it lost that character
on decode: "\u{FEFF}abc" came back as "abc", "\u{FEFF}8" as the number
8, and "\u{FEFF}" as an empty object. Quote it instead, as the §7.2 rule
added by toon-format/spec#60 requires.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab
shreyasbhat0 pushed a commit to toon-format/toon-rust that referenced this pull request Oct 5, 2026
* chore: bump spec submodule to v4.1.1

* feat!: rewrite decoder as line-based parser for spec v4.1

Implements comment pre-pass, BOM removal, keyed tabular form, nested
field groups, the normative number grammar, quoted-token boundaries,
and v4.1 strict/non-strict error semantics. Removes key folding, path
expansion, and the non-spec decode options (delimiter, coerce_types).

* feat!: mandate spec v4.1 encoder forms

Emits keyed tabular form, nested field groups, and the key: [] / []
empty-array forms; quotes hash-leading, leading-plus, and non-ASCII
keys per SPEC 7.2/7.3; escapes controls as \uXXXX. Replaces the
streaming encoder with in-memory wrappers: v4.1 headers depend on
whole-subtree analysis, so bounded-memory streaming is no longer
implementable.

* docs: update README, crate docs, and TUI guide for spec v4.1

* fix!: harden the v4.1 implementation from adversarial review

Bounds recursion on all three unbounded axes (nested field groups in
headers, nested-uniform column detection, encode normalization) at the
documented MAX_DEPTH instead of aborting on stack overflow; fixes i64
saturation corruption at 2^63 and normalizes integral exponent forms
across the full i64/u64 domain; scans blank-line spans in O(log n);
preserves a content CR before a CRLF terminator.

Deletes dead public API orphaned by the rewrite (LengthMismatch,
UnexpectedEof, InvalidDelimiter, InvalidCharacter error variants and
unused validators), records keyed tabular and nested field groups in
layout metadata, rejects indentSize 0 in the encoder like the decoder,
and wires --indent into CLI decode mode.

* style: apply cargo +nightly fmt

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab

* fix: quote a root string starting with U+FEFF

On line 1 a leading U+FEFF is a byte-order mark that decoders remove
(§12), so an unquoted root string starting with it lost that character
on decode: "\u{FEFF}abc" came back as "abc", "\u{FEFF}8" as the number
8, and "\u{FEFF}" as an empty object. Quote it instead, as the §7.2 rule
added by toon-format/spec#60 requires.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab

* fix: decode out-of-range integral tokens as strings

An integral token outside the i64/u64 domain fell through to f64, so
-9223372036854775809 decoded as the integer -9223372036854775808 and
18446744073709551616 as 1.8446744073709552e19. Restore the 0.5.0 string
fallback, the lossless-first policy §4 recommends.

The one exception is a token that is the canonical spelling of an f64:
the encoder writes an integral f64 past that domain, such as 1e20, as a
plain integer (§2 allows no exponent below 1e21), and such a token must
decode as the same float for the value to round-trip.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab

* test: compare decode fixture output in key order

serde_json Value equality ignores object key order even with
preserve_order, so the decode fixtures did not check the key order §2
and §9.3 require. Compare the serializations instead.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab

* test: drop numeric cases the spec fixtures cover

-1E+03, -0 and 1.5 repeat cases in the decode/numbers.json fixture.

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab

* docs: point parse_number_token at the crate-level numeric policy

Omp-Session: 01a0f927-0d99-702b-9358-76909e4395ab
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant