Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,14 @@ on:

jobs:
test:
name: Swift tests and smoke tests
runs-on: macos-15
name: Swift tests and smoke tests (${{ matrix.os }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os:
- macos-15
- ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v5
Expand Down
14 changes: 7 additions & 7 deletions Docs/NativeConverterBackends.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
# Native and FOSS Converter Backend Plan

This document records the likely native/FOSS backends for the reserved document formats in `DocumentFormat`. These are **not package dependencies yet**; the current release still returns `unsupportedFormat` for PDF, DOCX, PPTX, and XLSX. The goal is to keep the app's import UI honest while making the implementation path explicit.
This document records the native/FOSS backend status for the reserved document formats in `DocumentFormat`. PDFKit-backed embedded-text PDF extraction is now implemented on Apple platforms; DOCX, PPTX, and XLSX still return `unsupportedFormat` until their native converter modules are implemented. The goal is to keep the app's import UI honest while making the implementation path explicit.

## Short answer: are these simple to add?

They are straightforward to integrate as Swift Package / Apple-framework building blocks, but they are not all equally "drop-in" as Markdown converters:

| Format | Proposed backend | Integration effort | Why |
| --- | --- | --- | --- |
| PDF | Apple's PDFKit | Low for embedded text; medium when OCR fallback is included. | PDFKit is built into Apple platforms and can extract text from many PDFs, but scanned/image-only PDFs still need page rendering plus Vision OCR. |
| PDF | Apple's PDFKit | Shipped for embedded text; medium remaining work for OCR fallback. | PDFKit is built into Apple platforms and now extracts embedded text from PDFs, but scanned/image-only PDFs still need page rendering plus Vision OCR. |
| DOCX | ZIPFoundation + OOXML parsing | Medium. | DOCX is a ZIP of XML parts, but useful Markdown needs document body parsing, relationships, styles, numbering, tables, hyperlinks, and images. |
| PPTX | ZIPFoundation + OOXML parsing | Medium-high. | PPTX uses the same OpenXML ZIP structure, but slide ordering, shapes, notes, and layout-driven reading order make Markdown extraction more involved than DOCX. |
| XLSX | CoreXLSX | Medium-low for worksheet tables; medium for richer workbooks. | CoreXLSX already parses XLSX structure in Swift, but Markdown output still needs shared strings, sheet selection, empty-cell handling, formulas, merged cells, and table shaping decisions. |
Expand All @@ -26,11 +26,11 @@ When one of these backends is implemented, the PR that adds it should also add t

## Recommended implementation order

1. **PDF text extraction with PDFKit**
- Add a `PDFConverter` behind `#if canImport(PDFKit)`.
- Extract embedded page text first.
- If a page has no embedded text, optionally render the page and reuse the Vision OCR path already used by image ingestion.
- Keep non-Apple platforms returning `unsupportedFormat` unless a separate cross-platform PDF backend is added.
1. **PDF text extraction with PDFKit** — shipped for embedded text
- `PDFConverter` is compiled behind `#if canImport(PDFKit)`.
- Embedded page text is extracted first.
- Remaining work: if a page has no embedded text, optionally render the page and reuse the Vision OCR path already used by image ingestion.
- Non-Apple platforms continue returning `unsupportedFormat` unless a separate cross-platform PDF backend is added.

2. **Shared OpenXML ZIP infrastructure**
- Add ZIPFoundation as a SwiftPM dependency only when DOCX or PPTX work starts.
Expand Down
10 changes: 6 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ The current MVP is intentionally small and deterministic. These formats have con
| CSV | `csv` | `text/csv`, `application/csv` | Converts rows to GitHub-Flavored Markdown tables, including quoted fields and escaped pipes. | All package platforms. |
| JSON | `json` | `application/json`, `text/json` | Converts objects and arrays to nested Markdown bullets with stable key ordering. | All package platforms. |
| Images | `png`, `jpg`, `jpeg`, `heic`, `heif`, `tif`, `tiff`, `gif` | `image/png`, `image/jpeg`, `image/heic`, `image/heif`, `image/tiff`, `image/gif` | Uses Apple Vision OCR and returns recognized text lines as Markdown text. GIF OCR uses the decoded first image. | Apple platforms that provide Vision, CoreGraphics, and ImageIO. Other platforms recognize the formats but return `unsupportedFormat`. |
| PDF | `pdf` | `application/pdf` | Uses Apple PDFKit to extract embedded page text into Markdown paragraphs and records page-count metadata. | Apple platforms that provide PDFKit. Other platforms recognize PDF but return `unsupportedFormat`. |

The package also includes:

Expand All @@ -22,12 +23,13 @@ The package also includes:
- `csv` to GitHub-Flavored Markdown tables, including quoted fields and escaped pipes.
- `json` to nested Markdown bullets with stable key ordering.
- Apple-platform image OCR for `png`, `jpg`/`jpeg`, `heic`, `tiff`, and `gif` inputs using Vision text recognition, returning recognized lines as Markdown text.
- Apple-platform PDF text extraction for embedded-text PDFs using PDFKit.
- A CLI wrapper for local/manual conversion checks.
- A SwiftUI iOS demo app for editing sample input and converting it to Markdown in the simulator.

Image OCR is available when the package is built on platforms that provide Vision, CoreGraphics, and ImageIO. On other platforms, image formats are recognized but return `unsupportedFormat`.

PDF, DOCX, PPTX, and XLSX are represented in the format model but still return `unsupportedFormat` until their native converter modules are implemented.
DOCX, PPTX, and XLSX are represented in the format model but still return `unsupportedFormat` until their native converter modules are implemented. PDF conversion is implemented on Apple platforms that provide PDFKit; non-Apple platforms still return `unsupportedFormat` for PDF.

## Repository layout

Expand Down Expand Up @@ -87,7 +89,7 @@ SwiftMarkItDown is available under the [MIT License](LICENSE).

1. Expand the text/HTML/CSV/JSON converters with richer Markdown normalization and metadata extraction.
2. Improve OCR layout reconstruction for headings, lists, tables, and multi-column scans.
3. Add a ZIP/OpenXML package reader as shared infrastructure for DOCX, PPTX, and XLSX.
4. Implement DOCX paragraph, heading, table, hyperlink, and image-reference extraction.
5. Add PDFKit/Vision-backed PDF text and OCR extraction for Apple platforms behind conditional compilation.
3. Add OCR fallback for scanned/image-only PDF pages on Apple platforms.
4. Add a ZIP/OpenXML package reader as shared infrastructure for DOCX, PPTX, and XLSX.
5. Implement DOCX paragraph, heading, table, hyperlink, and image-reference extraction.
6. Evolve the demo into a more complete iOS MVP with document picker import, share/export flows, progress reporting, and a pluggable backend escape hatch for heavyweight conversions.
10 changes: 8 additions & 2 deletions Scripts/smoke-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -70,11 +70,17 @@ run_case empty.html
run_case empty.csv

run_error_case empty.json "swift-markitdown: JSON input is empty."
run_error_case sample.pdf "swift-markitdown: No converter is registered for pdf."
if swift -e 'import PDFKit' >/dev/null 2>&1; then
run_case sample.pdf
run_case empty.pdf
else
run_error_case sample.pdf "swift-markitdown: No converter is registered for pdf."
run_error_case empty.pdf "swift-markitdown: No converter is registered for pdf."
fi

run_error_case sample.docx "swift-markitdown: No converter is registered for docx."
run_error_case sample.pptx "swift-markitdown: No converter is registered for pptx."
run_error_case sample.xlsx "swift-markitdown: No converter is registered for xlsx."
run_error_case empty.pdf "swift-markitdown: No converter is registered for pdf."
run_error_case empty.docx "swift-markitdown: No converter is registered for docx."
run_error_case empty.pptx "swift-markitdown: No converter is registered for pptx."
run_error_case empty.xlsx "swift-markitdown: No converter is registered for xlsx."
Expand Down
44 changes: 44 additions & 0 deletions Sources/SwiftMarkItDown/Converters/PDFConverter.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
import Foundation
#if canImport(PDFKit)
import PDFKit
#endif

/// Extracts Markdown-ready text from PDF documents using Apple PDFKit when available.
public struct PDFConverter: DocumentConverter {
public let supportedFormats: Set<DocumentFormat> = [.pdf]

public init() {}

public func convert(_ request: ConversionRequest, format: DocumentFormat) throws -> MarkdownDocument {
#if canImport(PDFKit)
guard let document = PDFDocument(data: request.data) else {
throw ConversionError.malformedInput("The input could not be parsed as a PDF document.")
}

var extractedPageCount = 0
let pages = (0..<document.pageCount).map { index in
let text = document.page(at: index)?.string?.smid_trimmedBlankLines ?? ""
if !text.isEmpty {
extractedPageCount += 1
}
return text
}

let markdown = pages
.filter { !$0.isEmpty }
.joined(separator: "\n\n")
.smid_trimmedBlankLines

return MarkdownDocument(
markdown: markdown,
sourceFormat: format,
metadata: [
"pageCount": String(document.pageCount),
"extractedTextPageCount": String(extractedPageCount)
]
)
#else
throw ConversionError.unsupportedFormat(format)
#endif
}
}
3 changes: 2 additions & 1 deletion Sources/SwiftMarkItDown/MarkItDown.swift
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,8 @@ public struct MarkItDown: Sendable {
HTMLConverter(),
CSVConverter(),
JSONConverter(),
ImageOCRConverter()
ImageOCRConverter(),
PDFConverter()
]
}

Expand Down
2 changes: 2 additions & 0 deletions Tests/Expected/sample.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
SwiftMarkItDown PDF
Mock document for CI
Binary file modified Tests/Fixtures/empty.pdf
Binary file not shown.
Binary file modified Tests/Fixtures/sample.pdf
Binary file not shown.
63 changes: 58 additions & 5 deletions Tests/SwiftMarkItDownTests/SwiftMarkItDownTests.swift
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@ import Foundation
import Testing
@testable import SwiftMarkItDown

#if canImport(PDFKit)
import PDFKit
#endif

#if canImport(Vision) && canImport(CoreGraphics) && canImport(CoreText) && canImport(ImageIO)
import CoreGraphics
import CoreText
Expand Down Expand Up @@ -61,13 +65,52 @@ struct SwiftMarkItDownTests {
}


@Test("throws for reserved but unimplemented formats")
@Test("throws for reserved but unimplemented non-PDF formats")
func throwsForUnimplementedFormats() throws {
let request = ConversionRequest(data: Data(), fileName: "paper.pdf")
let cases: [(String, DocumentFormat)] = [
("document.docx", .docx),
("deck.pptx", .pptx),
("workbook.xlsx", .xlsx)
]

for (fileName, format) in cases {
let request = ConversionRequest(data: Data(), fileName: fileName)
#expect(throws: ConversionError.unsupportedFormat(format)) {
try MarkItDown().convert(request)
}
}
}

#if canImport(PDFKit)
@Test("converts embedded-text PDF fixtures to Markdown")
func convertsPDF() throws {
let document = try MarkItDown().convert(contentsOf: fixtureURL("sample.pdf"))
let expected = try String(contentsOf: expectedURL("sample.md"), encoding: .utf8)
.smid_testTrimmedTrailingNewline

#expect(document.markdown == expected)
#expect(document.sourceFormat == .pdf)
#expect(document.metadata["pageCount"] == "1")
#expect(document.metadata["extractedTextPageCount"] == "1")
}

@Test("converts blank PDF fixtures to empty Markdown")
func convertsBlankPDF() throws {
let document = try MarkItDown().convert(contentsOf: fixtureURL("empty.pdf"))

#expect(document.markdown == "")
#expect(document.sourceFormat == .pdf)
#expect(document.metadata["pageCount"] == "1")
#expect(document.metadata["extractedTextPageCount"] == "0")
}
#else
@Test("throws unsupported for PDFs when PDFKit is unavailable")
func throwsForPDFsWhenPDFKitIsUnavailable() throws {
#expect(throws: ConversionError.unsupportedFormat(.pdf)) {
try MarkItDown().convert(request)
try MarkItDown().convert(contentsOf: fixtureURL("sample.pdf"))
}
}
#endif

#if canImport(Vision) && canImport(CoreGraphics) && canImport(CoreText) && canImport(ImageIO)
@Test("uses Vision OCR to convert rendered images to Markdown")
Expand Down Expand Up @@ -105,11 +148,9 @@ struct SwiftMarkItDownTests {
@Test("throws unsupported for every reserved document path, including empty documents")
func throwsUnsupportedForReservedDocumentPaths() throws {
let cases: [(String, DocumentFormat)] = [
("sample.pdf", .pdf),
("sample.docx", .docx),
("sample.pptx", .pptx),
("sample.xlsx", .xlsx),
("empty.pdf", .pdf),
("empty.docx", .docx),
("empty.pptx", .pptx),
("empty.xlsx", .xlsx)
Expand All @@ -136,6 +177,18 @@ private func fixtureURL(_ fileName: String) -> URL {
.appendingPathComponent(fileName)
}

private func expectedURL(_ fileName: String) -> URL {
URL(fileURLWithPath: FileManager.default.currentDirectoryPath)
.appendingPathComponent("Tests/Expected")
.appendingPathComponent(fileName)
}

private extension String {
var smid_testTrimmedTrailingNewline: String {
trimmingCharacters(in: .newlines)
}
}

private func blankPNGFixtureData() throws -> Data {
#if canImport(Vision) && canImport(CoreGraphics) && canImport(ImageIO)
let width = 100
Expand Down