diff --git a/SPEC.md b/SPEC.md index 676f3ec..007377a 100644 --- a/SPEC.md +++ b/SPEC.md @@ -280,8 +280,8 @@ TOON is a deterministic, line-oriented, indentation-based notation. - Else if the document has exactly one non-blank line and it is neither a valid array header nor a key-value line (quoted or unquoted key), decode a single primitive (examples: `hello`, `42`, `true`). - Otherwise, decode an object. - An empty document (no non-blank lines after comment removal, §5.1) decodes to an empty object `{}`. A document consisting only of comment and blank lines is therefore `{}`. - - The root form spans the whole document: once a root array, an empty root array (`[]`), or a keyed tabular root object is complete, no further non-comment, non-blank line may follow. In strict mode, decoders MUST error on such trailing content (§14.2) – it MUST NOT be silently discarded. In non-strict mode, decoders MAY ignore it. (A root object extends to the last line of the document, so this case does not arise for object roots.) - - In strict mode, if there are two or more non-blank depth-0 lines that are neither headers nor key-value lines, the document is invalid. Example of invalid input (strict mode): + - The root form spans the whole document: once a root array, an empty root array (`[]`), or a keyed tabular root object is complete, no further non-comment, non-blank line may follow. In strict mode, decoders MUST error on such trailing content (§14.2) – it MUST NOT be silently discarded. In non-strict mode, decoders MAY ignore it, except scalar lines, which are an error in any mode (§5.2). (A root object extends to the last line of the document, so this case does not arise for object roots.) + - If there are two or more non-blank depth-0 lines that are neither headers nor key-value lines, the document is invalid in strict and non-strict mode alike (§14.2). Example of invalid input: ``` hello world @@ -393,7 +393,7 @@ In quoted strings and keys, codepoints are encoded according to the following ta | CR (U+000D) | MUST emit `\r` | MUST decode `\r` → CR | | HTAB (U+0009) | MUST emit `\t` | MUST decode `\t` → HTAB | | Other U+0000–U+001F controls | MUST emit `\uXXXX` (lowercase hex SHOULD) | MUST decode `\uXXXX` (case-insensitive hex) | -| U+D800–U+DFFF lone surrogates | (not produced by valid encoders) | MUST reject when decoded from `\uXXXX` | +| U+D800–U+DFFF surrogates | (not produced by valid encoders) | MUST reject when decoded from `\uXXXX`, lone or paired | | Other BMP codepoints (U+0020–U+D7FF, U+E000–U+FFFF) | SHOULD emit literal UTF-8; MAY emit `\uXXXX` | MUST accept either form | | Supplementary scalar values (U+10000–U+10FFFF) | MUST emit as literal UTF-8 | MUST accept literal UTF-8; surrogate `\uXXXX` escapes MUST be rejected (see row above) | @@ -441,7 +441,7 @@ Keys requiring quoting per the above rules MUST be quoted in all contexts, inclu Decoding of value tokens follows §4 (unquoted type inference, quoted strings, numeric rules). This section adds key-specific requirements: - Quoted keys MUST be unescaped per §7.1; any other escape MUST error. -- Keys (quoted or unquoted) MUST be followed by ":"; missing colon MUST error (see also §14.2). +- Keys (quoted or unquoted) MUST be followed by ":", optionally after spaces (§12); missing colon MUST error (see also §14.2). - Unquoted key token (normative): an unquoted key token is the text before the first unquoted colon of a key-value line (§5.2) or entry row (§9.5), with surrounding spaces trimmed (§12); the text before a header's bracket segment; or a field name in a field list (§6). Decoders MUST accept any such token as a literal key, in strict and non-strict mode alike, even when it does not match §7.3's unquoted-key pattern: `foo-bar: 1`, `foo-bar[2]: 1,2`, and `items[1]{2key}:` are valid input. §7.3 constrains what encoders may emit unquoted, not what decoders accept. - Quoted-token boundary (normative): a token whose first character, after the trimming of §12, is `"` MUST be a complete quoted token – its closing `"` MUST be the token's last character. This applies wherever a token is extracted; any character after the closing quote MUST error. It overrides §4's "Otherwise → string" fallback. - Symmetrically for values: an unquoted value token that an encoder would have been required to quote (§7.2) is not an error. Decoders, strict mode included, MUST decode it per §4 – unless another rule of this specification assigns the token structural meaning (§5.2, §6, §9.1). Example: `key: -x` decodes to the string `-x`. §7.2 governs encoder output; it adds no decoder-side rejection. @@ -458,7 +458,7 @@ Decoding of value tokens follows §4 (unquoted type inference, quoted strings, n - Lines in an object body are classified per §5.2; the rules below cover its key-value class. - A line "key:" with nothing after the colon at depth d opens an object; subsequent lines at depth > d belong to that object until the depth decreases to ≤ d. - In strict mode, the first line of a non-empty nested scope MUST be at exactly depth d+1; a depth increase of more than one level relative to the enclosing scope MUST error (§14.2). Conforming encoders never produce depth jumps; §10's depth model governs fields carried on a list-item hyphen line. - - A line deeper than the content depth of its enclosing scope whose preceding line did not open a scope belongs to no scope (e.g., a depth d+1 line directly under a depth-d primitive field). In strict mode, decoders MUST error (§14.2) – such lines MUST NOT be silently discarded. In non-strict mode, decoders MAY skip them. + - A line deeper than the content depth of its enclosing scope whose preceding line did not open a scope belongs to no scope (e.g., a depth d+1 line directly under a depth-d primitive field). In strict mode, decoders MUST error (§14.2) – such lines MUST NOT be silently discarded. In non-strict mode, decoders MAY skip them, except scalar lines, which are an error in any mode (§5.2). - A bare `key:` (no value after the colon) MUST decode as an empty or nested object, not an empty array. Empty arrays use the explicit `key: []` form (§9.1). - Lines "key: value" at the same depth are sibling fields. - Duplicate sibling keys at the same depth: see §14.3 for strict/non-strict behavior. @@ -536,7 +536,7 @@ When tabular requirements are not met (encoding; including any column that is ne - Each element is rendered as a list item at depth +1 under the header: - Primitive: `- ` - Primitive array: `- [M]: v1…` - - Array of objects or non-uniform array: `- [M]:` on the hyphen line, followed by the nested array's list items at depth +1 relative to the hyphen line (i.e. +2 from the outer array header). Items are encoded recursively per §9.1–§9.4 as each item's shape requires; tabular form (§9.3) is not available in this position (a keyless fields-bearing header is valid only at the document root, §6) – encoders MUST use list form. + - Any other array (of objects, of arrays, or mixed): `- [M]:` on the hyphen line, followed by the nested array's list items at depth +1 relative to the hyphen line (i.e. +2 from the outer array header). Items are encoded recursively per §9.1–§9.4 as each item's shape requires; tabular form (§9.3) is not available in this position (a keyless fields-bearing header is valid only at the document root, §6) – encoders MUST use list form. - Object: formatted per §10 (objects as list items). Decoding: @@ -629,7 +629,7 @@ For an object appearing as a list item: - Encoders MUST NOT emit a trailing newline at the end of the document. - Decoding: - Byte-order mark: a single U+FEFF at the very start of the document is a byte-order mark, not content – decoders MUST remove it before any processing in §5.1 and this section. A U+FEFF anywhere else is content. Encoders MUST NOT emit one. - - Line terminators: a CR (U+000D) at the end of a line is part of the line terminator, not of the line's content – decoders MUST exclude it before any processing in §5.1 and this section, thereby accepting CRLF input. A CR anywhere else in a line is content. + - Line terminators: a single CR (U+000D) at the end of a line is part of the line terminator, not of the line's content – decoders MUST exclude it before any processing in §5.1 and this section, thereby accepting CRLF input. A CR anywhere else in a line is content. - Strict mode: - The number of leading spaces on a line MUST be an exact multiple of indentSize; otherwise MUST error. - Tabs used as indentation MUST error (see §7.1 for tabs in quoted strings and as the HTAB delimiter). @@ -637,7 +637,7 @@ For an object appearing as a list item: - Depth MAY be computed as floor(indentSpaces / indentSize). - Implementations MAY accept tab characters in indentation. When they do, leading tabs are indentation and MUST be removed from the line's content before classification (§5.2). Depth computation for tabs is implementation-defined and MUST be documented. - Trailing spaces: trailing spaces (U+0020) at the end of a line are not part of the line's content. Decoders MUST strip them after the CR exclusion above and before line classification (§5.2); a line whose content is `-` followed only by spaces is therefore the bare marker for an empty-object list item (§9.4, §10), not a list item carrying an empty token. - - Token trimming: when a token is extracted – a key token before a key-value colon or an entry key's colon (§7.4, §9.5), or a value token after a key-value colon, after an array-header colon, or around each delimiter-separated token – decoders MUST trim surrounding spaces, exactly U+0020, no other characters. Any other whitespace (e.g., NBSP, or HTAB outside its delimiter role) is part of the token; internal semantics follow quoting rules. This trimming does not apply between a key and its bracket segment, where whitespace is a header syntax error (§6). + - Token trimming: when a token is extracted – a key token before a key-value colon or an entry key's colon (§7.4, §9.5), a field entry in a field list (§6), or a value token after a key-value colon, after an array-header colon, or around each delimiter-separated token – decoders MUST trim surrounding spaces, exactly U+0020, no other characters. Any other whitespace (e.g., NBSP, or HTAB outside its delimiter role) is part of the token; internal semantics follow quoting rules. This trimming does not apply between a key and its bracket segment, where whitespace is a header syntax error (§6). - Comment lines are removed before any check in this section applies (§5.1). - Blank lines: - A line whose content trims to empty is blank, regardless of leading-space count; the indentation checks above do not apply to blank lines. diff --git a/tests/fixtures/decode/arrays-nested.json b/tests/fixtures/decode/arrays-nested.json index 69e3a40..7a549e9 100644 --- a/tests/fixtures/decode/arrays-nested.json +++ b/tests/fixtures/decode/arrays-nested.json @@ -301,6 +301,21 @@ ] }, "specSection": "9.4" + }, + { + "name": "keeps list items under a legacy [0] header in non-strict mode", + "input": "a[0]:\n - x", + "expected": { + "a": [ + "x" + ] + }, + "options": { + "strict": false + }, + "specSection": "14.1", + "note": "A legacy [0] is a declared length, so its content is decoded (§14.1)", + "minSpecVersion": "4.1" } ] } diff --git a/tests/fixtures/decode/arrays-primitive.json b/tests/fixtures/decode/arrays-primitive.json index fd5f610..f84fb8c 100644 --- a/tests/fixtures/decode/arrays-primitive.json +++ b/tests/fixtures/decode/arrays-primitive.json @@ -167,6 +167,47 @@ }, "specSection": "9.1", "note": "An escaped quote does not close a quoted key, so the [2] inside it is not a bracket segment" + }, + { + "name": "keeps inline values under a legacy [0] header in non-strict mode", + "input": "a[0]: 1,2", + "expected": { + "a": [ + 1, + 2 + ] + }, + "options": { + "strict": false + }, + "specSection": "14.1", + "note": "A legacy [0] is a declared length, so its content is decoded (§14.1)", + "minSpecVersion": "4.1" + }, + { + "name": "keeps every inline value when the count mismatches in non-strict mode", + "input": "tags[1]: a,b", + "expected": { + "tags": [ + "a", + "b" + ] + }, + "options": { + "strict": false + }, + "specSection": "14.1", + "minSpecVersion": "4.1" + }, + { + "name": "decodes the inline element [] as a string, not an empty array", + "input": "a[1]: []", + "expected": { + "a": [ + "[]" + ] + }, + "specSection": "9.3" } ] } diff --git a/tests/fixtures/decode/arrays-tabular.json b/tests/fixtures/decode/arrays-tabular.json index 64830aa..269bffa 100644 --- a/tests/fixtures/decode/arrays-tabular.json +++ b/tests/fixtures/decode/arrays-tabular.json @@ -230,6 +230,66 @@ ] }, "specSection": "9.3" + }, + { + "name": "trims spaces around field names", + "input": "items[1]{ a , b }:\n 1,2", + "expected": { + "items": [ + { + "a": 1, + "b": 2 + } + ] + }, + "specSection": "12", + "minSpecVersion": "4.1" + }, + { + "name": "materializes a nested field group without cells as an empty object in non-strict mode", + "input": "items[1]{a,n{x}}:\n 1", + "expected": { + "items": [ + { + "a": 1, + "n": {} + } + ] + }, + "options": { + "strict": false + }, + "specSection": "14.1", + "minSpecVersion": "4.1" + }, + { + "name": "drops surplus cells when the width mismatches in non-strict mode", + "input": "items[1]{a}:\n 1,2", + "expected": { + "items": [ + { + "a": 1 + } + ] + }, + "options": { + "strict": false + }, + "specSection": "14.1", + "minSpecVersion": "4.1" + }, + { + "name": "parses unquoted colon after the delimiter in tabular row as data", + "input": "items[1]{id,note}:\n 1,a:b", + "expected": { + "items": [ + { + "id": 1, + "note": "a:b" + } + ] + }, + "specSection": "9.3" } ] } diff --git a/tests/fixtures/decode/blank-lines.json b/tests/fixtures/decode/blank-lines.json index b8eb4f9..74e04cf 100644 --- a/tests/fixtures/decode/blank-lines.json +++ b/tests/fixtures/decode/blank-lines.json @@ -211,6 +211,22 @@ "input": "m[2:]{v}:\n\n a: 1\n b: 2", "expected": { "m": { "a": { "v": 1 }, "b": { "v": 2 } } }, "specSection": "12" + }, + { + "name": "throws on blank line between a nested header and its first row inside a list item", + "input": "outer[2]:\n - inner[1]{a}:\n\n 1\n - x", + "expected": null, + "shouldError": true, + "specSection": "14.2", + "note": "The blank is between the inner header and its first row but inside the outer array's span (§12)" + }, + { + "name": "throws on blank line between list items after a nested object", + "input": "a[2]:\n - x:\n y: 1\n\n - 2", + "expected": null, + "shouldError": true, + "specSection": "14.2", + "note": "The blank is after the nested object's content but inside the outer array's span (§12)" } ] } diff --git a/tests/fixtures/decode/indentation-errors.json b/tests/fixtures/decode/indentation-errors.json index 96097d0..2c32119 100644 --- a/tests/fixtures/decode/indentation-errors.json +++ b/tests/fixtures/decode/indentation-errors.json @@ -196,6 +196,18 @@ "shouldError": true, "specSection": "14.2", "note": "A scalar line is valid only at root primitive position" + }, + { + "name": "throws on orphan scalar line under a primitive field in non-strict mode", + "input": "a: 1\n hello", + "expected": null, + "shouldError": true, + "options": { + "strict": false + }, + "specSection": "8", + "note": "The non-strict skip for over-indented lines excludes scalar lines (§5.2)", + "minSpecVersion": "4.1" } ] } diff --git a/tests/fixtures/decode/objects.json b/tests/fixtures/decode/objects.json index af52a61..a53710c 100644 --- a/tests/fixtures/decode/objects.json +++ b/tests/fixtures/decode/objects.json @@ -538,6 +538,53 @@ ] }, "specSection": "15" + }, + { + "name": "materializes __proto__ entry key as an ordinary own key", + "input": "u[1:]{x}:\n __proto__: 1", + "expected": { + "u": { + "__proto__": { + "x": 1 + } + } + }, + "specSection": "15" + }, + { + "name": "parses an unquoted key followed by spaces before the colon", + "input": "a : 1", + "expected": { + "a": 1 + }, + "specSection": "7.4" + }, + { + "name": "parses a quoted key followed by spaces before the colon", + "input": "\"a\" : 1", + "expected": { + "a": 1 + }, + "specSection": "7.4", + "minSpecVersion": "4.1" + }, + { + "name": "parses a hyphen-leading line outside a list as a key-value line", + "input": "- a: 1", + "expected": { + "- a": 1 + }, + "specSection": "5.2" + }, + { + "name": "keeps keys that differ only in normalization form distinct", + "input": "\u00e9: 1\ne\u0301: 2", + "expected": { + "\u00e9": 1, + "e\u0301": 2 + }, + "specSection": "16", + "minSpecVersion": "4.1" } ] } diff --git a/tests/fixtures/decode/root-form.json b/tests/fixtures/decode/root-form.json index 6378932..c5fc94c 100644 --- a/tests/fixtures/decode/root-form.json +++ b/tests/fixtures/decode/root-form.json @@ -87,6 +87,18 @@ "strict": true }, "specSection": "5" + }, + { + "name": "throws on a trailing bare token after a root array in non-strict mode", + "input": "[2]: 1,2\njunk", + "expected": null, + "shouldError": true, + "options": { + "strict": false + }, + "specSection": "5", + "note": "The non-strict leniency for trailing content excludes scalar lines (§5.2)", + "minSpecVersion": "4.1" } ] } diff --git a/tests/fixtures/decode/validation-errors.json b/tests/fixtures/decode/validation-errors.json index 5e20b3e..a55d2b8 100644 --- a/tests/fixtures/decode/validation-errors.json +++ b/tests/fixtures/decode/validation-errors.json @@ -566,6 +566,71 @@ "shouldError": true, "specSection": "7.4", "minSpecVersion": "4.1" + }, + { + "name": "throws on two primitives at root depth in non-strict mode", + "input": "hello\nworld", + "expected": null, + "shouldError": true, + "options": { + "strict": false + }, + "specSection": "5", + "minSpecVersion": "4.1" + }, + { + "name": "throws on a surrogate pair escape", + "input": "val: \"\\ud83d\\ude00\"", + "expected": null, + "shouldError": true, + "specSection": "7.1" + }, + { + "name": "throws on a key-value line at row depth inside a tabular array", + "input": "items[2]{a,b}:\n 1,2\n c: 3,4", + "expected": null, + "shouldError": true, + "specSection": "9.3", + "note": "The colon precedes the delimiter, so the line ends the rows instead of becoming a second row" + }, + { + "name": "throws on a header delimiter mismatch that row widths do not expose", + "input": "items[1|]{a,b}:\n 1,2", + "expected": null, + "shouldError": true, + "options": { + "strict": true + }, + "specSection": "6" + }, + { + "name": "throws on a hyphen without a following space at item depth", + "input": "items[2]:\n - a\n -b", + "expected": null, + "shouldError": true, + "specSection": "5.2" + }, + { + "name": "throws on tabular rows at depth +1 under a list-item first-field header", + "input": "x[1]:\n - a[1]{p}:\n 1", + "expected": null, + "shouldError": true, + "specSection": "10" + }, + { + "name": "throws on invalid escape sequence in a quoted key", + "input": "\"a\\x\": 1", + "expected": null, + "shouldError": true, + "specSection": "7.4" + }, + { + "name": "throws on a [1] header with nothing after the colon", + "input": "a[1]:", + "expected": null, + "shouldError": true, + "specSection": "9.1", + "note": "Nothing after the colon is list form with zero items, not one empty inline value" } ] } diff --git a/tests/fixtures/decode/whitespace.json b/tests/fixtures/decode/whitespace.json index 3667076..f527885 100644 --- a/tests/fixtures/decode/whitespace.json +++ b/tests/fixtures/decode/whitespace.json @@ -145,12 +145,22 @@ "specSection": "12" }, { - "name": "keeps a carriage return inside a line as content", - "input": "a: x\ry", + "name": "keeps a second carriage return before CRLF as content", + "input": "a: x\r\r\nb: 1", "expected": { - "a": "x\ry" + "a": "x\r", + "b": 1 }, "specSection": "12" + }, + { + "name": "keeps a second leading byte-order mark as content", + "input": "\ufeff\ufeffa: 1", + "expected": { + "\ufeffa": 1 + }, + "specSection": "12", + "minSpecVersion": "4.1" } ] } diff --git a/tests/fixtures/encode/arrays-nested.json b/tests/fixtures/encode/arrays-nested.json index bc2abc4..55e0715 100644 --- a/tests/fixtures/encode/arrays-nested.json +++ b/tests/fixtures/encode/arrays-nested.json @@ -139,6 +139,22 @@ ], "expected": "[1]:\n - [1]:\n - [1]: 1", "specSection": "9.4" + }, + { + "name": "uses list form for an array mixing primitives and objects in list-item position", + "input": { + "a": [ + [ + 1, + { + "x": 1 + } + ], + 2 + ] + }, + "expected": "a[2]:\n - [2]:\n - 1\n - x: 1\n - 2", + "specSection": "9.4" } ] } diff --git a/tests/fixtures/encode/arrays-objects.json b/tests/fixtures/encode/arrays-objects.json index c462ded..bc46a04 100644 --- a/tests/fixtures/encode/arrays-objects.json +++ b/tests/fixtures/encode/arrays-objects.json @@ -199,6 +199,28 @@ }, "expected": "items[1]:\n - a:\n b: 1", "specSection": "10" + }, + { + "name": "uses tabular form for a later field of a list-item object", + "input": { + "items": [ + { + "id": 1, + "users": [ + { + "id": 1, + "name": "Ada" + }, + { + "id": 2, + "name": "Bob" + } + ] + } + ] + }, + "expected": "items[1]:\n - id: 1\n users[2]{id,name}:\n 1,Ada\n 2,Bob", + "specSection": "10" } ] }