fix(ci): the invisible-character gate never matched anything - #48
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🔇 Additional comments (2)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe workflow updates its invisible-character scan to use Unicode code-point patterns. It also forces ChangesInvisible-character gate
Estimated code review effort: 2 (Simple) | ~5 minutes Merge Risk: ⚪ Minimal · up to This localized workflow change corrects invisible-character detection and adds handling for control characters and NUL-bearing files. No actionable merge-blocking risk remains beyond normal checks and review. Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Description checkExplanation The description explains the root cause, the code-point and grep changes, and the verification performed. It does not use every template heading or complete the checklist, but it provides the key required information. Full details: Linked Issues checkExplanation The change addresses Unicode code-point escapes and grep -a in the CI gate. The provided file summary does not confirm the required C0 control-character range, excluding TAB, LF, and CR, so the core detection objective appears incomplete [ Full details: Docstring CoverageExplanation No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
The PR successfully corrects the invisible-character detection gate by migrating to PCRE codepoint syntax and enabling binary-safe processing. While the regex logic is now functionally correct, the implementation within the GitHub Workflow is inefficient and lacks the enforcement mechanisms typical of a 'gate'—it currently logs findings without failing the build. Additionally, although the fix addresses a previously silent failure, there are no automated tests included to prevent future regressions of this linter logic. Codacy analysis indicates the project remains up to standards, but improvements in error handling and execution efficiency are recommended.
About this PR
- There are no automated regression tests for the linter logic itself. Consider adding a set of fixture files containing the targeted invisible characters to the repository's test suite to ensure the gate remains functional in the future.
Test suggestions
- Detection of a Non-Breaking Space (U+00A0) in a YAML file
- Detection of a Null byte (U+0000) inside a source file
- Detection of a C0 control character (e.g., Backspace \x08) in a workflow file
- Verification that valid whitespace (TAB, LF, CR) does not trigger the gate
- Detection of Byte Order Mark (BOM, U+FEFF) at the start of a file
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Detection of a Non-Breaking Space (U+00A0) in a YAML file
2. Detection of a Null byte (U+0000) inside a source file
3. Detection of a C0 control character (e.g., Backspace \x08) in a workflow file
4. Verification that valid whitespace (TAB, LF, CR) does not trigger the gate
5. Detection of Byte Order Mark (BOM, U+FEFF) at the start of a file
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
🟡 MEDIUM RISK
Suggestion: The implementation of the grep command should be optimized for efficiency and reliability: 1) Replace -exec ... {} \; with -exec ... {} + to batch file processing and remove the redundant -r flag since find already handles recursion. 2) Avoid suppressing stderr (2>/dev/null) because grep -P in a UTF-8 locale will exit with an error on invalid sequences; these should be visible to prevent silent skips of files. 3) To function as a true 'gate', the step should exit with a non-zero code if FINDINGS are detected.
Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.
Root cause
The pattern used UTF-8 byte sequences (
\xc2\xa0) whilegrep -Pmatches characters. Bytesc2 a0are one character U+00A0;\xc2\xa0asks for two, U+00C2 then U+00A0 — never present.Only
\x00worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.Fixed
\x01-\x08,\x0B,\x0C,\x0E-\x1Fadded (TAB/LF/CR excluded)grep -a— without it grep skips any NUL-bearing file as binaryThe C0 range matters: a stray backspace byte made a workflow unparseable in
developer-ecosystem, so it never ran — and this linter called it clean.Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.