Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions .out-of-scope/binary-encoding.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Binary Encoding

TOON is a text format. A binary TOON wire format is out of scope, in the spec and in the reference library and CLI.

## Why this is out of scope

TOON saves tokens, not bytes, and a model reads text, so a binary encoding would always be decoded back to text first. Compression fails for the same reason: the model tokenizes the decompressed text ([toon#125](https://github.com/toon-format/toon/issues/125#issuecomment-3528565342)).

> Adding a second "Binary TOON" wire format would require a separate spec, cross‑language implementations, and long‑term maintenance, while not really helping the primary use case (LLMs still need text, so this would always be decoded before use). For general binary serialization there are already well‑established options like CBOR/MessagePack.
> – [toon#201](https://github.com/toon-format/toon/pull/201#issuecomment-3566643061)

A binary variant can live as its own package.

## Prior requests

- toon#201 – Binary TOON
17 changes: 17 additions & 0 deletions .out-of-scope/custom-delimiters.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Custom Delimiters

The delimiter set is comma, tab, and pipe (§11). Other delimiters – non-ASCII symbols like `✦`, semicolons, or user-defined separators – are out of scope.

## Why this is out of scope

The header declares the delimiter from a closed set (§6: `delimsym = HTAB / "|"`, absent means comma), and every decoder splits and every encoder quotes against that set (§7.2, §11). Each new delimiter means new grammar and new quoting rules in every implementation:

> I want to keep the spec to comma/tab/pipe to stay ASCII-only for maximum interop, editor/terminal safety, and predictable tokenization. Non-ASCII delimiters like ✦ would add complexity and may be token-inefficient on some models.
> – [toon#133](https://github.com/toon-format/toon/issues/133#issuecomment-3532885149)

Tab and pipe already cover readability.

## Prior requests

- toon#133 – `✦` as delimiter
- toon discussion #139 – one-line rows with `|` and `;`
19 changes: 19 additions & 0 deletions .out-of-scope/encoder-data-rewrites.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Encoder-Side Data Rewrites

Encoders don't rewrite the data to fit a more compact form. That rules out filling missing fields with `null` so semi-uniform arrays qualify for tabular form, dropping `null`, empty-string, or empty-array fields, splitting semi-uniform arrays into base and extras tables, and abbreviating, renaming, or reordering keys.

## Why this is out of scope

Each of these changes the value, so `decode(encode(x))` stops returning `x` under §2's JSON-model equality: objects compare by their ordered key sequence, a missing key and a key set to `null` are different values, and no decoder can restore a dropped field or undo a renamed key. Whether `null` and "absent" mean the same thing, or which short names a prompt can live with, is a property of your data, not of the format:

> TOON is a **transport format**: it encodes the JSON data model as text, and that's where its responsibility ends. Splitting semi-uniform arrays into base + extras tables is a **data transformation step** that belongs in the application layer, not inside the encoder.
> – [toon#292](https://github.com/toon-format/toon/pull/292#issuecomment-4162837784)

[toon-python#35](https://github.com/toon-format/toon-python/pull/35#issuecomment-4073056071) sent key abbreviation to higher-level wrappers as well. Normalize before encoding where your data allows it.

## Prior requests

- toon#344 – `fillna` option without an explicit replacer
- toon#292 – pre-encoding normalization for semi-uniform arrays
- spec#48 – `ignoreNullOrEmpty` and `excludeEmptyArrays` (the mixed columnar half is in `mixed-columnar-arrays.md`)
- toon-python#25, toon-python#35 – "semantic optimization": field abbreviation and semantic key ordering
2 changes: 1 addition & 1 deletion .out-of-scope/inline-annotations.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Inline comments intersect every quoting rule in the spec at once: §7.2 (string

The underlying need is real but is served without grammar changes:

- **Hand-authored documents** (config files, prompt schemas): v4 adopts *full-line* `#` comments – a line whose first non-whitespace character is `#` is stripped before parsing. In a hand-authored prompt the model reads those lines, so field-adjacent guidance works by placing a comment line above the field.
- **Hand-authored documents** (config files, prompt schemas): v4 adopts *full-line* `#` comments – a line whose first character after leading spaces is `#` is stripped before parsing. In a hand-authored prompt the model reads those lines, so field-adjacent guidance works by placing a comment line above the field.
- **Programmatic per-field annotations that must survive encode**: use a `_note:` key convention – annotations as data – or the `rawString` replacer primitive (toon#308/toon#321) for controlled raw output.

The distinction matters: decode-side comments are never *emitted* by encoders (JSON has no comments to encode), so anything that must flow through `encode()` has to be data, not comment syntax.
Expand Down
4 changes: 2 additions & 2 deletions .out-of-scope/key-folding.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,11 +8,11 @@ Measured and verified against the reference implementation before removal:

- **Zero benefit on real data.** 0.00% token savings on all six of the project's own benchmark datasets – folding only fires on single-key chains, which realistic tabular/nested data lacks. The published benchmark numbers never enabled it.
- **Wire ambiguity by construction.** `encode({'a.b.c': 1})` and `encode({a:{b:{c:1}}}, {keyFolding:'safe'})` produce the byte-identical document, so `expandPaths` cannot distinguish a folded chain from a genuine dotted literal key – it silently corrupts the latter. Two distinct JSON values collapsing to one wire document breaks TOON's determinism and round-trip guarantees.
- **Streaming-incompatible by design.** Expansion requires the fully materialized tree (§13.4 applied it after all parsing); the reference streaming decoder throws on it. It was the only feature in the spec that prevented streaming decode.
- **Streaming-incompatible by design.** Expansion requires the fully materialized tree (§13.4 applied it after all parsing); the reference streaming decoder threw on it. It was the only feature in the spec that prevented streaming decode.
- **Security surface.** The `IdentifierSegment` pattern admitted `__proto__`, enabling prototype pollution through path expansion (fixed in the v2.x line; the v4 spec makes prototype-key handling normative in §15 independent of folding).
- **Round-trip asymmetry.** Both options defaulted to `off`, so correctness depended on out-of-band option agreement between producer and consumer.

Folded documents remain valid TOON forever: dotted keys decode as literal keys (a MUST since v3.x). Users who need dotted-path notation can expand in userland after decode – with the caveat that userland expansion must guard `__proto__`/`constructor`/`prototype`.
Folded documents remain valid TOON forever: dotted keys decode as literal keys (a MUST since v4.0). Users who need dotted-path notation can expand in userland after decode – with the caveat that userland expansion must guard `__proto__`/`constructor`/`prototype`.

## Prior requests

Expand Down
13 changes: 13 additions & 0 deletions .out-of-scope/lenient-quoted-strings.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Lenient Quoted Strings

A token that starts with `"` must end at its closing quote (§7.4, "Quoted-token boundary"). Relaxing that rule so that LLM output like `text: "hello" said Alice` decodes as a string is out of scope.

## Why this is out of scope

v4.1 made characters after the closing quote an error in strict and non-strict mode alike (§7.4, §14.2). Without the rule, a decoder has to guess whether a leading `"` opens a quoted string or is literal data, and two decoders can guess differently. A decoder that repairs input silently also can't tell a model's mistake from data.

Malformed model output is the application's to handle: show the escaped form in the prompt, re-prompt on the strict decode error, or repair before decoding.

## Prior requests

- spec discussion #41 – quoted string specification
18 changes: 18 additions & 0 deletions .out-of-scope/mixed-columnar-arrays.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Mixed Columnar Arrays

Tabular rows carry primitive cells only. Proposals that let a row continue with nested content – spill lines under a row, nested tabular blocks per row, or an encoder option like `objectArrayLayout: "columnar"` that opts into such a layout – are out of scope.

## Why this is out of scope

The tabular form is useful because it is constrained: one line per element and one primitive cell per leaf field, so row boundaries are trivial and a strict decoder checks row count and width (§9.3, §14.1). Nested lines under a row turn rows into trees:

> This proposal removes that constraint and turns rows into tree-structured entities – effectively creating a second tree-encoding path alongside the existing list/object syntax.
> – [spec#21](https://github.com/toon-format/spec/issues/21#issuecomment-4163086929)

A layout option also conflicts with §1.4: the form follows from the value's shape and position, not from encoder preference. Uniform nested objects already fit the header as nested field groups (§9.3); arrays whose elements differ stay in list form (§9.4).

## Prior requests

- spec#21 – hybrid tabular arrays with nested row content
- spec#48 – mixed columnar arrays and `objectArrayLayout` (the `ignoreNullOrEmpty` and `excludeEmptyArrays` half is in `encoder-data-rewrites.md`)
- spec#47 – draft spec text for spec#48
20 changes: 20 additions & 0 deletions .out-of-scope/multi-document-streams.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Multi-Document Streams

A TOON document holds one value. There is no multi-document syntax – no `---` separators, no "TOON Lines" counterpart to JSON Lines – and the reference library stays single-document.

## Why this is out of scope

TOON's root form spans the whole document: once a root array or keyed tabular root is complete, strict decoders MUST reject further content (§5, §14.2). A separator would need a new line class and new root-form rules in every implementation, for structure TOON already expresses:

> JSON Lines is a *framing* format (a sequence of independent JSON values), while TOON is deliberately defined as "one JSON value per document", with no multi‑document syntax or directives.
> – [toon#120](https://github.com/toon-format/toon/issues/120#issuecomment-3558763808)

The only structure a JSONL stream carries is an ordered sequence, so encode it as one document with an array root. Separators and batch encode/decode are framing utilities for a companion package built on `encode` and `decode` ([toon#163](https://github.com/toon-format/toon/issues/163#issuecomment-3559409605)).

## Prior requests

- toon#120 – JSON Lines support
- toon#121 – JSON Lines support (implementation PR)
- toon discussion #119 – what about JSON Lines?
- toon#163 – streaming API with document separators and batch processing
- toon#176 – streaming API for large datasets (implementation PR)
18 changes: 18 additions & 0 deletions .out-of-scope/non-json-data-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Non-JSON Data Models

TOON encodes the JSON data model and nothing else (§2). Modes or extensions for XML (attributes, namespaces), HTML/CSS/JS markup, or GraphQL are out of scope.

## Why this is out of scope

A second data model inside the same syntax makes one document decode differently depending on a flag:

> **Ambiguity.** The same TOON document would decode differently depending on a `mode` flag. A key like `id: 123` is a plain key-value pair in JSON mode but an XML attribute in XML mode. That breaks TOON's promise of deterministic, unambiguous encoding.
> – [spec#29](https://github.com/toon-format/spec/pull/29#issuecomment-4163086242)

It would also change the grammar under existing parsers – namespace prefixes put a colon into the key, where §5.2 ends it – and add normative sections every implementation carries even if it never sees XML. Convert to JSON first and encode that; markup already has dedicated tools ([toon discussion #220](https://github.com/toon-format/toon/discussions/220#discussioncomment-15070508)). A format for another data model is a separate spec, not a TOON mode.

## Prior requests

- spec#29 – XML support
- toon discussion #220 – HTML, CSS, and JS as TOON
- toon#3 – GraphQL
20 changes: 20 additions & 0 deletions .out-of-scope/optional-array-length.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Optional Array Length

Every array header declares its length: `key[N]:`. Dropping `[N]`, making it optional for tabular arrays, allowing an unknown length (`[-]`), or writing `items[]{…}:` for one element is out of scope.

## Why this is out of scope

The length is load-bearing: §6 makes it mandatory, encoders MUST emit the actual count (§13.1), and strict decoders MUST error when rows, items, or entries don't add up (§14.1) – that check is how truncated or injected data gets noticed (§15).

> The `[N]` header is a fundamental invariant in TOON: every array declares its length explicitly. That property underpins strict‑mode validation (count checks, truncation detection), the "structure awareness / structural validation" behavior in the benchmarks, and a big part of TOON's differentiation from "CSV + indent." Making `[N]` optional for tabular arrays would weaken those guarantees and introduce a second, non‑canonical syntax for the same structure.
> – [spec#11](https://github.com/toon-format/spec/pull/11#issuecomment-3569473088)

It isn't backward compatible either: strict decoders reject `key[]:` (§6). A producer that can't know `N` up front counts first or chunks the data into several arrays ([spec#15](https://github.com/toon-format/spec/issues/15#issuecomment-3569512191)); hand-edited data is better edited as JSON and converted, so `[N]` is generated ([toon#145](https://github.com/toon-format/toon/issues/145#issuecomment-3536756237)).

## Prior requests

- spec#15 – arrays of unknown size
- spec#11 – optional implicit length for tabular arrays
- spec#19 – `items[]{…}` for length-1 arrays (the empty-array half was accepted as `key: []`)
- toon#135 – UTOON, an unbounded TOON variant
- toon#145 – why arrays need an item count
16 changes: 16 additions & 0 deletions .out-of-scope/query-language.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Query and Patch Language

TOON defines no query or patch language – no "TOON-path" and no TOON-native patch operations.

## Why this is out of scope

A decoded TOON document is a plain JSON value (§2), so every standard that operates on JSON applies unchanged, and a TOON-specific language would duplicate them:

> TOON does not define its own query or patch language. It is intentionally just a compact, line oriented concrete syntax for the JSON data model.
> – [spec#16](https://github.com/toon-format/spec/issues/16#issuecomment-3589490562)

Decode, then query with JSONPath, JMESPath, or jq and update with JSON Pointer or JSON Patch.

## Prior requests

- spec#16 – TOON-path
21 changes: 21 additions & 0 deletions .out-of-scope/references-and-aliases.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# References, Aliases, and Dictionary Encoding

TOON has no references between values: no `$ID` aliases defined once and reused, no `@tables` foreign-key resolution, no per-column enumerations (`role(admin,user)` with cells `0`/`1`), and no key maps (`_map`) that shorten field names.

## Why this is out of scope

Each of these turns a cell into a pointer, and the JSON data model TOON encodes (§2) has none:

> JSON has no concept of references, foreign keys, or inter-document linking. Introducing `@tables` would make TOON a different kind of format: a relational data language with its own resolution semantics, hydration modes, and cycle-detection requirements.
> – [spec#27](https://github.com/toon-format/spec/issues/27#issuecomment-3941522699)

Resolving references means buffering the whole document, which breaks streaming decode; literal values that look like references (`$100`) would suddenly need quoting ([spec#36](https://github.com/toon-format/spec/issues/36#issuecomment-4163086597)); and one JSON value could encode with different aliases, where §1.4 lets only its shape and position pick the rendering. The gain is small, too: values like `admin` are already single tokens, and a model has to map `0` back to `admin` ([toon discussion #216](https://github.com/toon-format/toon/discussions/216#discussioncomment-15053977)).

Repetition is a data-modeling problem: normalize before encoding – encode the foreign key instead of the embedded object, or put the dictionary into the data as an ordinary array.

## Prior requests

- spec#27 – relational references (`@tables`)
- spec#36 – value aliasing with `$ID` references
- spec discussion #14 – `_map` key compression
- toon discussion #216 – per-column enumerations and base-62 indices
21 changes: 21 additions & 0 deletions .out-of-scope/schema-language.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Schema Language

TOON does not define a schema language, embed JSON Schema, or require decoders to validate against a schema.

## Why this is out of scope

TOON is already JSON Schema compatible at the data-model level: a schema that validates a JSON value validates the decoded TOON document too. What these requests add is schema syntax inside TOON, with decoders required to resolve it:

> What the RFC is really asking for is a *schema authoring syntax* embedded in TOON, plus normative requirements on decoders to resolve `$ref`, evaluate `allOf`, etc. That's a much larger surface area, and it would couple the core spec to every future JSON Schema draft revision. I'd rather keep SPEC.md scoped to one job: *how JSON values are serialized as text*.
> – [spec#7](https://github.com/toon-format/spec/issues/7#issuecomment-3941530685)

Header syntax like `@SchemaName` would also break every conforming parser (§6). Decode, then validate with the tooling you already use; a JSON Schema document itself encodes as TOON like any other value. A companion document or package one layer above the format – never normative in `SPEC.md` – is welcome ([toon#126](https://github.com/toon-format/toon/pull/126#issuecomment-3571906444)).

## Prior requests

- spec#7 – make TOON JSON Schema compatible
- spec#17 – JSON Schema in TOON format
- toon#103, toon#109 – describe object schemas in TOON
- toon#126 – optional schema validation system with `@SchemaName` headers
- toon discussion #127 – schema validation for type-safe data handling
- toon discussion #140 – TOON Schema proof of concept
16 changes: 16 additions & 0 deletions .out-of-scope/trailing-newline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# Trailing Newline

Encoders don't end a document with a newline (§12, §13.1). Requests to make a trailing newline part of the encoder output are out of scope.

## Why this is out of scope

An encoder returns the serialized value, not a file:

> TOON follows the same convention as `JSON.stringify()`: the output is the serialized data structure itself, not a file format with specific storage requirements.
> – [toon#23](https://github.com/toon-format/toon/issues/23#issuecomment-3464114715)

A newline wouldn't make concatenated documents valid either. Decoders SHOULD accept a trailing newline (§12), so add one when you write to disk.

## Prior requests

- toon#23 – trailing newline at end of output
Loading
Loading