Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 61 additions & 0 deletions Docs/NativeConverterBackends.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# Native and FOSS Converter Backend Plan

This document records the likely native/FOSS backends for the reserved document formats in `DocumentFormat`. These are **not package dependencies yet**; the current release still returns `unsupportedFormat` for PDF, DOCX, PPTX, and XLSX. The goal is to keep the app's import UI honest while making the implementation path explicit.

## Short answer: are these simple to add?

They are straightforward to integrate as Swift Package / Apple-framework building blocks, but they are not all equally "drop-in" as Markdown converters:

| Format | Proposed backend | Integration effort | Why |
| --- | --- | --- | --- |
| PDF | Apple's PDFKit | Low for embedded text; medium when OCR fallback is included. | PDFKit is built into Apple platforms and can extract text from many PDFs, but scanned/image-only PDFs still need page rendering plus Vision OCR. |
| DOCX | ZIPFoundation + OOXML parsing | Medium. | DOCX is a ZIP of XML parts, but useful Markdown needs document body parsing, relationships, styles, numbering, tables, hyperlinks, and images. |
| PPTX | ZIPFoundation + OOXML parsing | Medium-high. | PPTX uses the same OpenXML ZIP structure, but slide ordering, shapes, notes, and layout-driven reading order make Markdown extraction more involved than DOCX. |
| XLSX | CoreXLSX | Medium-low for worksheet tables; medium for richer workbooks. | CoreXLSX already parses XLSX structure in Swift, but Markdown output still needs shared strings, sheet selection, empty-cell handling, formulas, merged cells, and table shaping decisions. |

## Backend source acknowledgements

When one of these backends is implemented, the PR that adds it should also add the dependency/framework acknowledgement to this table and to any required license notice files.

| Backend | Source | License / status | Intended use |
| --- | --- | --- | --- |
| Apple PDFKit | <https://developer.apple.com/documentation/pdfkit> | Apple system framework; no SwiftPM dependency. | PDF text extraction and optional page rendering for OCR fallback on Apple platforms. |
| ZIPFoundation | <https://github.com/weichsel/ZIPFoundation> | MIT-licensed Swift package. | ZIP container access for DOCX and PPTX OpenXML parts. |
| CoreXLSX | <https://github.com/CoreOffice/CoreXLSX> | Apache-2.0-licensed Swift package. | Read-only parsing of XLSX workbooks and worksheets. |
| Office Open XML structure | <https://ecma-international.org/publications-and-standards/standards/ecma-376/> | Published standard. | Format reference for DOCX/PPTX/XLSX XML parts and relationships. |

## Recommended implementation order

1. **PDF text extraction with PDFKit**
- Add a `PDFConverter` behind `#if canImport(PDFKit)`.
- Extract embedded page text first.
- If a page has no embedded text, optionally render the page and reuse the Vision OCR path already used by image ingestion.
- Keep non-Apple platforms returning `unsupportedFormat` unless a separate cross-platform PDF backend is added.

2. **Shared OpenXML ZIP infrastructure**
- Add ZIPFoundation as a SwiftPM dependency only when DOCX or PPTX work starts.
- Build a small internal helper for reading XML parts, relationships, content types, and document metadata from OpenXML packages.
- Use this helper for DOCX first, then PPTX.

3. **DOCX converter**
- Parse `word/document.xml` in document order.
- Resolve relationships for hyperlinks and images.
- Map paragraphs, headings, lists, tables, emphasis, and links to Markdown.
- Add fixtures that cover common Word exports rather than only hand-written XML.

4. **PPTX converter**
- Parse slide order from `ppt/presentation.xml` and slide relationship parts.
- Extract text from shapes, grouped shapes, speaker notes, and tables.
- Use simple slide-section Markdown first; improve layout ordering later.

5. **XLSX converter with CoreXLSX**
- Add CoreXLSX as a SwiftPM dependency when XLSX implementation begins.
- Convert each selected worksheet to Markdown tables.
- Decide how to handle formulas, empty rows/columns, merged cells, dates, and multiple sheets.

## Dependency policy

- Do not add ZIPFoundation or CoreXLSX to `Package.swift` until a converter actually uses them.
- Prefer conditional compilation for Apple-only frameworks such as PDFKit and Vision.
- Keep unsupported formats visible in `DocumentFormat` so apps can show useful import affordances and friendly `unsupportedFormat` errors.
- Add license acknowledgements in the same PR that introduces any FOSS dependency.
21 changes: 21 additions & 0 deletions LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2026 SwiftMarkItDown contributors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
17 changes: 16 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,18 @@ SwiftMarkItDown is the start of a native Swift/iOS document-to-Markdown pipeline

## What works today

The current MVP is intentionally small and deterministic so it can run fully on-device:
The current MVP is intentionally small and deterministic. These formats have converters in the default `MarkItDown` pipeline:

| Input family | Extensions / aliases | Content-type hints | Conversion behavior | Platform availability |
| --- | --- | --- | --- | --- |
| Plain text | `txt`, `text` | `text/plain` | Decodes text and normalizes blank lines. | All package platforms. |
| Markdown | `md`, `markdown` | `text/markdown`, `text/x-markdown` | Treats Markdown as text-like input and normalizes blank lines. | All package platforms. |
| HTML | `html`, `htm` | `text/html`, `application/xhtml+xml` | Converts common headings, inline emphasis, links, code, paragraphs, and list items; ignores document `<head>`, script, and style content. | All package platforms. |
| CSV | `csv` | `text/csv`, `application/csv` | Converts rows to GitHub-Flavored Markdown tables, including quoted fields and escaped pipes. | All package platforms. |
| JSON | `json` | `application/json`, `text/json` | Converts objects and arrays to nested Markdown bullets with stable key ordering. | All package platforms. |
| Images | `png`, `jpg`, `jpeg`, `heic`, `heif`, `tif`, `tiff`, `gif` | `image/png`, `image/jpeg`, `image/heic`, `image/heif`, `image/tiff`, `image/gif` | Uses Apple Vision OCR and returns recognized text lines as Markdown text. GIF OCR uses the decoded first image. | Apple platforms that provide Vision, CoreGraphics, and ImageIO. Other platforms recognize the formats but return `unsupportedFormat`. |

The package also includes:

- `txt` and `md` passthrough with text decoding and blank-line cleanup.
- `html` to Markdown for common headings, inline emphasis, links, code, paragraphs, and list items, with document `<head>`, script, and style content ignored.
Expand Down Expand Up @@ -68,6 +79,10 @@ Scripts/smoke-test.sh

GitHub Actions runs the same checks on pushes to `main`, pull requests, and manual workflow dispatches.

## License

SwiftMarkItDown is available under the [MIT License](LICENSE).

## Roadmap

1. Expand the text/HTML/CSV/JSON converters with richer Markdown normalization and metadata extraction.
Expand Down
Loading