Skip to content

Repository files navigation

demarkify

a fast, non-dependent tool/library to detect, decode, and remove hidden unicode-based ai tracking markers, steganography, invisible characters, homoglyphs, and even byte anomalies from any text.

npm version npm downloads GitHub release license zero deps


what is demarkify?

LLMS, ai text generation models, and tracking systems often integrate subtle, invisible watermarks into generated/output text. these watermarks range from zero width unicode spaces and tag chars to directional overrides, invisible math operators, homoglyph char subs (like replacing latin letters with near-identical cyrillic or greek glyphs), and trailing whitespace binary modulations.

demarkify comes into play by completely eliminating these hidden markers by performing deep byte-level inspection, decoding any hidden payloads, and reconstructs completely clean text fully from scratch.


what demarkify has to offer

  • pure native nodejs / ts results in being extremely fast, lightweight, and secure.
  • deep watermark and steganography detection
  • zero width and invis codepoints: ZWSP (U+200B), ZWNJ (U+200C), ZWJ (U+200D), BOM (U+FEFF), Word joiner (U+2060), soft hyphen (U+00AD), combining grapheme joiner (U+034F), hangul fillers (U+3164, U+FFA0, U+115F, U+1160), mongolian vowel separator (U+180E), invisible operators (U+2061-U+2064), etc.
  • procedural unicode injection detection: catches unlisted \p{Cf} (Format) and \p{Co} (Private Use Area) injection markers.
  • detects and decodes hidden payload chars in the U+E0000 / U+E007F range used for prompt injection or llm output tracking
  • zerowidth and whitespace binary stego: detects and decodes binary payloads embedded into zerowidth char sequences, hangul filler patterns, or trailing spaces/tabs.
  • directional overrides & isolates: LTR/RTL marks, embeddings, overrides (U+202AU+202E, U+2066/U+2069)
  • homoglyphs and confusable scripts: identifies lookalike chars swapped in from cyrillic, greek, fullwidth, or mathematical alphanumeric blocks (fraktur, bold, sans, double-struck, monospace)
  • normalize non breaking spaces (U+00A0), braille blanks (U+2800), thin spaces, em/en spaces, line/paragraph separators (U+2028/U+2029), and next-line markers (U+0085).
  • non-printable control chars and byte anomalies
  • complete scratch reconstruction: instead of naive regex replacement, text is parsed and rebuilt cleanly with canonical unicode normalizing (NFC/NFKC) and uniform line endings
  • versatile inputs: supports direct string arguments, stdin pipes, individual files, and even entire directories, recursively
  • CI/CD ready: --check flag returns exit code 1 if watermarks and such are detected, making it a viable option for git precommits.

install

option 1: global cli via pm

# using npm
npm install -g demarkify

# using pnpm
pnpm add -g demarkify

# using bun
bun add -g demarkify

option 2: local project dep

npm install demarkify
# or
pnpm add demarkify

option 3: standalone binaries (nodejs not required)

precompiled standalone binaries are available for download on the GitHub releases page:

platform arch download
linux x86_64 demarkify-v1.1.0-linux-x64.tar.gz
linux ARM64 / AArch64 demarkify-v1.1.0-linux-arm64.tar.gz
macOS Apple Silicon (M1/M2/M3/M4) demarkify-v1.1.0-darwin-arm64.tar.gz
macOS Intel x86_64 demarkify-v1.1.0-darwin-x64.tar.gz
windows x86_64 demarkify-v1.1.0-windows-x64.zip

cli usage

basic usage

# clean text from stdin and output clean text to stdout
cat ai_output.txt | demarkify > clean.txt

# clean a string snippet directly
demarkify "Hello\u200BWorld"

# check if file has watermarks (dryrun)
demarkify --check document_example.md

# sanitize a file in-place
demarkify -w document.md

# sanitize a file to a new destination
demarkify document.md -o document.clean.md

# recursively clean an entire directory in-place
demarkify -w ./src/

programmatic api (ts / js)

you can import and use demarkify directly in your njs or ts projects:

import {
  demarkify,
  detectWatermarks,
  sanitizeText,
  processFile,
  processDirectory
} from 'demarkify';

// inspecting text for watermarks
const detection = detectWatermarks('prompt text\u200Bwith hidden data');
console.log(detection.clean); // false
console.log(detection.totalWatermarks); // 1
console.log(detection.matches);

// clean text from scratch
const result = sanitizeText('some\u200B dirty\u00A0text', {
  normalizeSpaces: true,
  stripZeroWidth: true,
  unicodeNormalization: 'NFC'
});
console.log(result.cleanText); // "some dirty text"
console.log(result.bytesSaved); // 4

// process files or directories
const fileResult = processFile('./report.md', { normalizeHomoglyphs: true }, true);
console.log(fileResult.changed); // true if modified

testing

# run unit and integ tests
pnpm test

# build ts to dist/
pnpm run build

license

MIT / 2026 (READ ./LICENSE FOR MORE INFORMATIOn)

About

a fast, non-dependent tool/library to detect, decode, and remove hidden unicode-based ai tracking markers, steganography, invisible characters, homoglyphs, and even byte anomalies from any text.

Topics

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages