Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 9 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,14 +45,21 @@ jobs:
node-version: "20"
- name: Install anydoc for the Windows path test
run: npm install -g @firecrawl/anydoc@0.2.4
- name: Spaces, percent signs, and output collisions
- name: Windows CSV and DOCX paths, content, and output collisions
shell: pwsh
run: |
New-Item -ItemType Directory -Force 'test-input/a', 'test-input/b', 'test-output' | Out-Null
New-Item -ItemType Directory -Force 'test-input/a', 'test-input/b', 'test-input/docx-source/_rels', 'test-input/docx-source/word', 'test-output' | Out-Null
Set-Content -LiteralPath 'test-input/a/a %TEMP% b.csv' -Value "name,value`na,1" -NoNewline
Set-Content -LiteralPath 'test-input/a/dup.csv' -Value "name,value`na,2" -NoNewline
Set-Content -LiteralPath 'test-input/b/dup.csv' -Value "name,value`nb,3" -NoNewline
Set-Content -LiteralPath 'test-input/docx-source/[Content_Types].xml' -Value '<?xml version="1.0" encoding="UTF-8"?><Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"><Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/><Default Extension="xml" ContentType="application/xml"/><Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/></Types>' -NoNewline
Set-Content -LiteralPath 'test-input/docx-source/_rels/.rels' -Value '<?xml version="1.0" encoding="UTF-8"?><Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships"><Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/></Relationships>' -NoNewline
Set-Content -LiteralPath 'test-input/docx-source/word/document.xml' -Value '<?xml version="1.0" encoding="UTF-8"?><w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body><w:p><w:r><w:t>DOC2MD_WINDOWS_DOCX_OK</w:t></w:r></w:p><w:sectPr/></w:body></w:document>' -NoNewline
[System.IO.Compression.ZipFile]::CreateFromDirectory((Resolve-Path 'test-input/docx-source').Path, [System.IO.Path]::GetFullPath('test-input/a/a %TEMP% b.docx'))
node skills/doc2md/scripts/convert.js test-input/a test-input/b -o test-output
if ($LASTEXITCODE -ne 0) { exit $LASTEXITCODE }
if (-not (Test-Path -LiteralPath 'test-output/a %TEMP% b.csv.md')) { throw 'Special-character path was lost' }
$docxOutput = 'test-output/a %TEMP% b.docx.md'
if (-not (Test-Path -LiteralPath $docxOutput)) { throw 'DOCX path was lost' }
if (-not (Select-String -LiteralPath $docxOutput -Pattern 'DOC2MD_WINDOWS_DOCX_OK' -Quiet)) { throw 'DOCX text was not converted' }
if ((Get-ChildItem -LiteralPath test-output -Filter 'dup.csv.*.md').Count -ne 2) { throw 'Output collision lost a file' }
10 changes: 5 additions & 5 deletions README.en.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,25 +63,25 @@ Without `-o`, each `<name>.<ext>.md` is written next to the source — the exten
| OpenDocument | `.odt` `.ods` `.odp` |
| Other | `.rtf` `.epub` `.csv` `.pdf` |

Scans and image-only PDFs with no text layer aren't read (the engine has no OCR) — they go to "SKIP" and the batch doesn't fail. Run OCR first for those.
In local mode, scans and image-only PDFs without a text layer go to "SKIP" and the batch continues. Run OCR first to convert them. Upstream anydoc 0.2.4 also offers `--ocr hosted`, which sends the whole document to Firecrawl Parse; doc2md does not enable that mode.

## Outcome categories

- **OK** — anydoc returned exit code `0`.
- **ERROR** — couldn't even launch `npx`/anydoc (not on PATH, code 126/127): an environment failure, not "the file is unreadable" — OCR won't help.
- **SKIP** — anydoc ran but the file didn't convert: almost always a scan / encrypted / corrupt file.
- **SKIP** — anydoc returned exit code `3`: PDF pages need OCR. Encrypted or corrupt files are errors, not OCR candidates.

> anydoc's exit-code contract isn't formally documented; the wrapper treats it conservatively (see `SKILL.md`).
> These exit codes are listed by `anydoc --help` in version 0.2.4 (see `SKILL.md`).

## Tested / not tested

Tested (2026-08-06, Cowork sandbox, Node 22, Linux/macOS): `.csv` and `.docx` with Cyrillic, name-collision folder, paths with spaces, missing paths, recursive folders — for both `convert.js` and `convert.sh`.

**Not tested live:** `.pptx`, `.xlsx`, `.pdf`, `.epub`, a real client document batch, and a live Windows run. The Windows code path is reasoned about, not verified. Verified a format or OS? [Open an issue](https://github.com/ilyautov/doc2md/issues).
**Automated Windows CI verified (2026-10-07):** CSV and a real DOCX whose names contain spaces and `%TEMP%` convert successfully; the test checks extracted DOCX text and output collisions. **Not verified on a user's Windows machine:** skill installation or a real client batch. `.pptx`, `.xlsx`, `.pdf`, and `.epub` also remain untested live. Verified a format or environment? [Open an issue](https://github.com/ilyautov/doc2md/issues).

## Source

A wrapper over [firecrawl/anydoc](https://github.com/firecrawl/anydoc) (MIT, Rust). The upstream also ships a minimal single-file skill; doc2md adds batch processing, an honest OK/SKIP/ERROR summary, OCR routing, a cross-platform Node version and docs. Benchmark numbers are the vendor's own — see [SOURCES.md](SOURCES.md).
A wrapper over [firecrawl/anydoc](https://github.com/firecrawl/anydoc) (MIT, Rust). The upstream also ships a minimal single-file skill; doc2md adds batch processing, an honest OK/SKIP/ERROR summary, explicit identification of PDFs needing OCR, a cross-platform Node version and docs. Benchmark numbers are the vendor's own — see [SOURCES.md](SOURCES.md).

## Author

Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,7 +114,7 @@ bash skills/doc2md/scripts/convert.sh договор.docx
| OpenDocument | `.odt` `.ods` `.odp` |
| Прочее | `.rtf` `.epub` `.csv` `.pdf` |

Сканы и PDF-картинки без текстового слоя не читаются (в движке нет OCR) — попадают в «ПРОПУСК», batch на этом не падает. Для них нужен отдельный OCR-проход перед конвертацией (например, скилл `pdf`, шаг «сделать PDF searchable»).
В локальном режиме сканы и PDF-картинки без текстового слоя попадают в «ПРОПУСК», batch на этом не падает. Для них нужен отдельный OCR-проход перед конвертацией (например, скилл `pdf`, шаг «сделать PDF searchable»). У anydoc 0.2.4 есть отдельный `--ocr hosted`, который отправляет весь документ в Firecrawl Parse; doc2md этот режим не включает.

## Что скрипт делает, а что нет

Expand All @@ -132,11 +132,11 @@ bash skills/doc2md/scripts/convert.sh договор.docx
- Несуществующий путь в аргументах — пропускается с понятным сообщением, batch не падает.
- `convert.js` — та же батарея тестов повторена и прошла (рекурсивная папка, коллизия имён, кириллица, вложенная подпапка, missing path).

**Не тестировано вживую:** `.pptx`, `.xlsx`, `.pdf`, `.epub`, боевая пачка документов клиента. **Живой запуск новой версии на Windows пока не сделан** — на macOS проверены имена с пробелом и `%TEMP%`, коллизии в `-o`, повреждённый `.docx` и отсутствие значения `-o`. Windows-подтверждение из issue относится к прежнему коду. Проверил новую версию на Windows — [сообщи результат](https://github.com/ilyautov/doc2md/issues).
**Автоматически проверено на Windows (CI, 07.10.2026):** CSV и настоящий DOCX с пробелами и `%TEMP%` в имени конвертируются; тест проверяет текст DOCX и отсутствие перезаписи при совпадении имён. **Не проверено на пользовательском Windows-компьютере:** установка скилла и боевая пачка документов. Также пока не проверены вживую `.pptx`, `.xlsx`, `.pdf`, `.epub`. Проверил на своих файлах — [сообщи результат](https://github.com/ilyautov/doc2md/issues).

## Чем отличается от upstream-скилла anydoc

`firecrawl/anydoc` публикует свой минимальный skill (`npx skills add firecrawl/anydoc`) — конвертирует один файл за раз, без пакетной обработки, без сводки по batch и без русской документации. doc2md достраивает именно это: пачка/папка на вход, честная сводка `OK`/`ПРОПУСК`/`ОШИБКА`, маршрутизация сканов на OCR, кросс-платформенная Node-версия, документация на русском.
`firecrawl/anydoc` публикует свой минимальный skill (`npx skills add firecrawl/anydoc`) — конвертирует один файл за раз, без пакетной обработки, без сводки по batch и без русской документации. doc2md достраивает именно это: пачка/папка на вход, честная сводка `OK`/`ПРОПУСК`/`ОШИБКА`, явная пометка PDF, которым нужен OCR, кросс-платформенная Node-версия, документация на русском.

## Источник

Expand Down
2 changes: 1 addition & 1 deletion SOURCES.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ doc2md — тонкая обёртка над **[firecrawl/anydoc](https://githu
(MIT, написан на Rust). Вся собственно конвертация форматов в Markdown
выполняется anydoc; doc2md добавляет поверх пакетную обработку (файл / список /
папка), построчный статус и итоговую сводку `OK` / `ПРОПУСК` / `ОШИБКА`,
маршрутизацию сканов на OCR и русскую документацию.
явную пометку PDF, которым нужен OCR, и русскую документацию.

- Репозиторий движка: https://github.com/firecrawl/anydoc
- npm-пакет: [`@firecrawl/anydoc`](https://www.npmjs.com/package/@firecrawl/anydoc)
Expand Down
4 changes: 2 additions & 2 deletions skills/doc2md/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ npx -y @firecrawl/anydoc report.docx -o report.docx.md

- Конвертирует файл(ы), печатает построчный статус (`OK` / `ПРОПУСК` / `ОШИБКА`) и итоговую сводку: сколько успешно, сколько пропущено, сколько ошибок вызова.
- **Различие ПРОПУСК vs ОШИБКА.** `OK` — код anydoc `0` и результат записан. `ПРОПУСК` — код `3`, PDF нужны OCR-страницы. Любой другой сбой чтения, конвертации, записи или вызова — `ОШИБКА` и ненулевой код всего пакета. В строку выводится stderr от anydoc.
- **Сканы и PDF-картинки без текстового слоя anydoc не читает** — это не баг, а ограничение движка (нет OCR внутри). Такие файлы попадают в «ПРОПУСК», batch на этом не падает. Для них: сначала OCR через скилл `pdf` (шаг «сделать PDF searchable»), потом повторный прогон.
- **Сканы и PDF-картинки без текстового слоя локальный режим не читает.** Такие файлы попадают в «ПРОПУСК», batch на этом не падает. Для них: сначала OCR через скилл `pdf` (шаг «сделать PDF searchable»), потом повторный прогон. У anydoc 0.2.4 есть `--ocr hosted`, но он отправляет весь документ в Firecrawl Parse; эта обёртка его не включает.
- Зашифрованные/защищённые паролем файлы — `ОШИБКА` с объяснением от anydoc; не отправляй их на OCR.
- Не удаляет и не трогает исходники. Повторный запуск может обновить ранее созданный `.md`.

Expand All @@ -86,7 +86,7 @@ npx -y @firecrawl/anydoc report.docx -o report.docx.md
- `scripts/convert.js` (кросс-платформенная Node-версия): та же батарея тестов повторена и прошла — рекурсивный обход папки, коллизия имён, кириллица в `.docx`/`.csv`, вложенная подпапка, несуществующий путь.
- Синтаксис-гейт обеих версий (`node --check`, `bash -n`) и smoke-тесты (`--help`, несуществующий путь) гоняются в CI на каждый PR.

**Не тестировал новую версию на Windows:** нет Windows-окружения. На macOS проверены имена с пробелом и `%TEMP%`, коллизии в `-o`, повреждённый `.docx` и отсутствие значения `-o`. Windows-подтверждение из issue относится к прежнему коду. При первом боевом использовании на Windows — перепроверить и дополнить эту секцию. Также не тестировались `.pptx`, `.xlsx`, `.pdf`, `.epub` и боевая пачка документов клиента.
**Автоматический Windows CI прошёл 07.10.2026:** CSV и настоящий DOCX с пробелами и `%TEMP%` в имени, извлечение текста из DOCX и коллизии в `-o`. Установку скилла и боевую пачку на пользовательском Windows-компьютере ещё не проверяли; `.pptx`, `.xlsx`, `.pdf` и `.epub` тоже не тестировались вживую.

## Источник

Expand Down
Loading