diff --git a/Docs/NativeConverterBackends.md b/Docs/NativeConverterBackends.md new file mode 100644 index 0000000..61c8e79 --- /dev/null +++ b/Docs/NativeConverterBackends.md @@ -0,0 +1,61 @@ +# Native and FOSS Converter Backend Plan + +This document records the likely native/FOSS backends for the reserved document formats in `DocumentFormat`. These are **not package dependencies yet**; the current release still returns `unsupportedFormat` for PDF, DOCX, PPTX, and XLSX. The goal is to keep the app's import UI honest while making the implementation path explicit. + +## Short answer: are these simple to add? + +They are straightforward to integrate as Swift Package / Apple-framework building blocks, but they are not all equally "drop-in" as Markdown converters: + +| Format | Proposed backend | Integration effort | Why | +| --- | --- | --- | --- | +| PDF | Apple's PDFKit | Low for embedded text; medium when OCR fallback is included. | PDFKit is built into Apple platforms and can extract text from many PDFs, but scanned/image-only PDFs still need page rendering plus Vision OCR. | +| DOCX | ZIPFoundation + OOXML parsing | Medium. | DOCX is a ZIP of XML parts, but useful Markdown needs document body parsing, relationships, styles, numbering, tables, hyperlinks, and images. | +| PPTX | ZIPFoundation + OOXML parsing | Medium-high. | PPTX uses the same OpenXML ZIP structure, but slide ordering, shapes, notes, and layout-driven reading order make Markdown extraction more involved than DOCX. | +| XLSX | CoreXLSX | Medium-low for worksheet tables; medium for richer workbooks. | CoreXLSX already parses XLSX structure in Swift, but Markdown output still needs shared strings, sheet selection, empty-cell handling, formulas, merged cells, and table shaping decisions. | + +## Backend source acknowledgements + +When one of these backends is implemented, the PR that adds it should also add the dependency/framework acknowledgement to this table and to any required license notice files. + +| Backend | Source | License / status | Intended use | +| --- | --- | --- | --- | +| Apple PDFKit | | Apple system framework; no SwiftPM dependency. | PDF text extraction and optional page rendering for OCR fallback on Apple platforms. | +| ZIPFoundation | | MIT-licensed Swift package. | ZIP container access for DOCX and PPTX OpenXML parts. | +| CoreXLSX | | Apache-2.0-licensed Swift package. | Read-only parsing of XLSX workbooks and worksheets. | +| Office Open XML structure | | Published standard. | Format reference for DOCX/PPTX/XLSX XML parts and relationships. | + +## Recommended implementation order + +1. **PDF text extraction with PDFKit** + - Add a `PDFConverter` behind `#if canImport(PDFKit)`. + - Extract embedded page text first. + - If a page has no embedded text, optionally render the page and reuse the Vision OCR path already used by image ingestion. + - Keep non-Apple platforms returning `unsupportedFormat` unless a separate cross-platform PDF backend is added. + +2. **Shared OpenXML ZIP infrastructure** + - Add ZIPFoundation as a SwiftPM dependency only when DOCX or PPTX work starts. + - Build a small internal helper for reading XML parts, relationships, content types, and document metadata from OpenXML packages. + - Use this helper for DOCX first, then PPTX. + +3. **DOCX converter** + - Parse `word/document.xml` in document order. + - Resolve relationships for hyperlinks and images. + - Map paragraphs, headings, lists, tables, emphasis, and links to Markdown. + - Add fixtures that cover common Word exports rather than only hand-written XML. + +4. **PPTX converter** + - Parse slide order from `ppt/presentation.xml` and slide relationship parts. + - Extract text from shapes, grouped shapes, speaker notes, and tables. + - Use simple slide-section Markdown first; improve layout ordering later. + +5. **XLSX converter with CoreXLSX** + - Add CoreXLSX as a SwiftPM dependency when XLSX implementation begins. + - Convert each selected worksheet to Markdown tables. + - Decide how to handle formulas, empty rows/columns, merged cells, dates, and multiple sheets. + +## Dependency policy + +- Do not add ZIPFoundation or CoreXLSX to `Package.swift` until a converter actually uses them. +- Prefer conditional compilation for Apple-only frameworks such as PDFKit and Vision. +- Keep unsupported formats visible in `DocumentFormat` so apps can show useful import affordances and friendly `unsupportedFormat` errors. +- Add license acknowledgements in the same PR that introduces any FOSS dependency. diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000..556c6c2 --- /dev/null +++ b/LICENSE @@ -0,0 +1,21 @@ +MIT License + +Copyright (c) 2026 SwiftMarkItDown contributors + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. diff --git a/README.md b/README.md index 4394adb..81f28b1 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,18 @@ SwiftMarkItDown is the start of a native Swift/iOS document-to-Markdown pipeline ## What works today -The current MVP is intentionally small and deterministic so it can run fully on-device: +The current MVP is intentionally small and deterministic. These formats have converters in the default `MarkItDown` pipeline: + +| Input family | Extensions / aliases | Content-type hints | Conversion behavior | Platform availability | +| --- | --- | --- | --- | --- | +| Plain text | `txt`, `text` | `text/plain` | Decodes text and normalizes blank lines. | All package platforms. | +| Markdown | `md`, `markdown` | `text/markdown`, `text/x-markdown` | Treats Markdown as text-like input and normalizes blank lines. | All package platforms. | +| HTML | `html`, `htm` | `text/html`, `application/xhtml+xml` | Converts common headings, inline emphasis, links, code, paragraphs, and list items; ignores document ``, script, and style content. | All package platforms. | +| CSV | `csv` | `text/csv`, `application/csv` | Converts rows to GitHub-Flavored Markdown tables, including quoted fields and escaped pipes. | All package platforms. | +| JSON | `json` | `application/json`, `text/json` | Converts objects and arrays to nested Markdown bullets with stable key ordering. | All package platforms. | +| Images | `png`, `jpg`, `jpeg`, `heic`, `heif`, `tif`, `tiff`, `gif` | `image/png`, `image/jpeg`, `image/heic`, `image/heif`, `image/tiff`, `image/gif` | Uses Apple Vision OCR and returns recognized text lines as Markdown text. GIF OCR uses the decoded first image. | Apple platforms that provide Vision, CoreGraphics, and ImageIO. Other platforms recognize the formats but return `unsupportedFormat`. | + +The package also includes: - `txt` and `md` passthrough with text decoding and blank-line cleanup. - `html` to Markdown for common headings, inline emphasis, links, code, paragraphs, and list items, with document ``, script, and style content ignored. @@ -68,6 +79,10 @@ Scripts/smoke-test.sh GitHub Actions runs the same checks on pushes to `main`, pull requests, and manual workflow dispatches. +## License + +SwiftMarkItDown is available under the [MIT License](LICENSE). + ## Roadmap 1. Expand the text/HTML/CSV/JSON converters with richer Markdown normalization and metadata extraction.