Skip to content

fix(ci): the invisible-character gate never matched anything - #157

Open
hyperpolymath wants to merge 1 commit into
mainfrom
fix/empty-linter-pattern-never-matched
Open

fix(ci): the invisible-character gate never matched anything#157
hyperpolymath wants to merge 1 commit into
mainfrom
fix/empty-linter-pattern-never-matched

Conversation

@hyperpolymath

Copy link
Copy Markdown
Owner

Measured 2026-08-27: this gate caught 0 of 6 invisible-character test cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi override or word joiner.

Root cause

The pattern used UTF-8 byte sequences (\xc2\xa0) while grep -P matches characters. Bytes c2 a0 are one character U+00A0; \xc2\xa0 asks for two, U+00C2 then U+00A0 — never present.

grep -P '\xc2\xa0'  ->  miss
grep -P '\x{a0}'    ->  MATCH

Only \x00 worked, being single-byte in both readings. The gate ran, passed, and could not see what it exists to see.

Fixed

  • codepoint escapes in place of byte sequences
  • C0 controls \x01-\x08,\x0B,\x0C,\x0E-\x1F added (TAB/LF/CR excluded)
  • grep -a — without it grep skips any NUL-bearing file as binary

The C0 range matters: a stray backspace byte made a workflow unparseable in developer-ecosystem, so it never ran — and this linter called it clean.

Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.

Verified: YAML re-parsed, and the corrected pattern was confirmed to catch a real NBSP before the change was kept.

MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.

ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.

  grep -P '\xc2\xa0'  ->  miss
  grep -P '\x{a0}'    ->  MATCH

Only \x00 worked, being single-byte in both readings.

FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.

The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.

Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
@coderabbitai

coderabbitai Bot commented Aug 27, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: a4e51289-b9bf-4a9d-b951-abb2a86e5e8c

📥 Commits

Reviewing files that changed from the base of the PR and between 54cd6aa and bac9152.

📒 Files selected for processing (1)
  • .github/workflows/dogfood-gate.yml

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.

📜 Recent review details
🔇 Additional comments (2)
.github/workflows/dogfood-gate.yml (2)

141-141: LGTM!


130-141: 🎯 Functional Correctness

No leading-BOM check is required for this reason.

GNU grep does not strip a leading BOM before matching. The byte pattern \xEF\xBB\xBF matches both leading and embedded BOMs.


📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Improved detection of invisible and control characters during automated checks.
    • Ensured text files are scanned consistently, including files containing non-standard characters.

Walkthrough

The workflow updates invisible-character detection. The pattern now matches Unicode code points and selected control characters. The grep command now treats files as text, including files that contain NUL bytes.

Changes

Invisible-character gate

Layer / File(s) Summary
Correct invisible-character matching and scanning
.github/workflows/dogfood-gate.yml
The PATTERNS regex now uses Unicode code-point escapes and control-character ranges. The grep command adds -a to process files as text.

Estimated code review effort: 1 (Trivial) | ~5 minutes

Merge Risk: ⚪ Minimal · up to bac91

This is a localized CI pattern correction, and no actionable merge-blocking risk remains after normal checks and review.

Suggested reviewers: metadatastician

Poem

A rabbit checks the hidden marks,
With careful hops through text,
Code points now reveal their shapes,
And NULs no longer vex.
The gate scans files as text today,
Then thumps: “The check is fixed!”

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The pull request updates the codepoint pattern, adds C0 controls, and uses grep -a. However, issue #70 also requires a separate leading-BOM check and matching updates to the compiled linter and config… Add the separate byte-wise leading-BOM check. Update stdlib/ByteDetector.affine and config.ncl with the same C0-control logic. Verify that the CI gate and compiled linter remain consistent.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the CI gate defect and the intended fix.
Description check ✅ Passed The description explains the invisible-character gate defect and the changes intended to correct it.
Out of Scope Changes check ✅ Passed The reported changes are limited to the CI invisible-character gate and are related to issue #70. No unrelated changes are identified.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Full details: Linked Issues check

Explanation

The pull request updates the codepoint pattern, adds C0 controls, and uses grep -a. However, issue #70 also requires a separate leading-BOM check and matching updates to the compiled linter and configuration. The provided changes summary shows only dogfood-gate.yml was changed.

Full details: Docstring Coverage

Explanation

No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sonarqubecloud

Copy link
Copy Markdown

@gitar-bot

gitar-bot Bot commented Aug 27, 2026

Copy link
Copy Markdown

Important

You are using the Gitar free plan. Upgrade to unlock code review, CI analysis, auto-apply, custom automations, and more.

Gitar

@codacy-production

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

AI Reviewer: first review requested successfully. AI can make mistakes. Always validate suggestions.

Run reviewer

TIP This summary will be updated as you push new changes.

@codacy-production codacy-production Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

The PR fixes the invisible-character gate which was previously ineffective due to incorrect escape sequences. However, a major technical gap remains: when using PCRE (grep -P) for Unicode codepoints, the engine must be explicitly set to UTF-8 mode using the (*UTF) prefix. Without this, the linter will continue to miss multi-byte characters like the Non-Breaking Space (U+00A0) in UTF-8 encoded files.

While Codacy reports that the changes are up to standards, the implementation lacks automated regression tests. Without a sample file containing forbidden characters to verify the linter's output, it is difficult to ensure the gate remains functional after future CI environment updates. Finally, the search logic can be optimized to reduce overhead in the CI pipeline.

About this PR

  • To prevent future regressions, consider adding a small automated test or a 'known-fail' sample file to the repository. This would allow the CI to verify that the invisible-character linter is actually capable of detecting the targeted characters in the current environment.

Test suggestions

  • Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0)
  • Missing recommended test scenario: Verify detection of Byte Order Mark (U+FEFF)
  • Missing recommended test scenario: Verify detection of C0 control characters like Backspace (\x08)
  • Missing recommended test scenario: Verify scanning of files containing NUL bytes using the -a flag
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Missing recommended test scenario: Verify detection of Non-Breaking Space (U+00A0)
2. Missing recommended test scenario: Verify detection of Byte Order Mark (U+FEFF)
3. Missing recommended test scenario: Verify detection of C0 control characters like Backspace (\x08)
4. Missing recommended test scenario: Verify scanning of files containing NUL bytes using the -a flag

TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback

# non-breaking spaces, null bytes, and other invisible Unicode in source files.
set +e
PATTERNS='\xc2\xa0|\xe2\x80\x8b|\xe2\x80\x8c|\xe2\x80\x8d|\xef\xbb\xbf|\xc2\xad|\xe2\x80\x8e|\xe2\x80\x8f|\xe2\x80\xaa|\xe2\x80\xab|\xe2\x80\xac|\xe2\x80\xad|\xe2\x80\xae|\x00'
PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 HIGH RISK

Enable UTF-8 mode in the PCRE engine to ensure Unicode code points are correctly matched against UTF-8 encoded files.

Suggested change
PATTERNS='\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}'
PATTERNS='(*UTF)\x00|[\x01-\x08\x0B\x0C\x0E-\x1F]|\x{a0}|\x{ad}|\x{200b}|\x{200c}|\x{200d}|\x{200e}|\x{200f}|\x{202a}|\x{202b}|\x{202c}|\x{202d}|\x{202e}|\x{2060}|\x{feff}'

-o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \
-o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \
-exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null
-exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚪ LOW RISK

Suggestion: Optimize the search by passing multiple files to grep at once and removing the redundant recursive flag.

Suggested change
-exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null
-exec grep -aPl "$PATTERNS" {} + > /tmp/empty-lint-results.txt 2>/dev/null

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant