Skip to content

Latest commit

 

History

History
388 lines (301 loc) · 14.8 KB

File metadata and controls

388 lines (301 loc) · 14.8 KB

Development Guide

This guide covers the project architecture, how to run tests, add new converters, and contribute code.


Project structure

filemorph/
├── app/
│   ├── main.py                  # FastAPI app: middleware, routers, UI route
│   ├── compat.py                # Frozen (.exe) vs. source detection + path helpers
│   ├── api/
│   │   ├── deps.py              # API-Key authentication dependency
│   │   └── routes/
│   │       ├── convert.py       # POST /api/v1/convert
│   │       ├── compress.py      # POST /api/v1/compress
│   │       ├── formats.py       # GET  /api/v1/formats
│   │       ├── health.py        # GET  /api/v1/health, /ready
│   │       ├── auth.py          # Cloud Edition: register / login / account
│   │       ├── keys.py          # Cloud Edition: dashboard API-key CRUD
│   │       ├── billing.py       # Cloud Edition: Stripe checkout + webhook
│   │       └── cockpit.py       # Cloud Edition: admin routes
│   ├── core/
│   │   ├── config.py            # Settings loaded from .env via pydantic-settings
│   │   ├── security.py          # API key generation, hashing, validation
│   │   ├── rate_limit.py        # slowapi limiter + failed-key budget
│   │   ├── quotas.py            # Per-tier size / concurrency / output caps
│   │   └── audit.py             # Tamper-evident audit-log hash chain
│   ├── converters/
│   │   ├── base.py              # BaseConverter base class
│   │   ├── registry.py          # Converter registry (@register decorator)
│   │   ├── image.py             # Image conversions (Pillow + pillow-heif + pillow-avif-plugin)
│   │   ├── document.py          # Document conversions (docx, pdf, txt, md)
│   │   ├── video.py             # Video conversions (ffmpeg-python)
│   │   ├── audio.py             # Audio conversions (ffmpeg-python)
│   │   └── spreadsheet.py       # Spreadsheet conversions (openpyxl, csv, json)
│   ├── compressors/
│   │   ├── image.py             # Image quality / target-size compression (Pillow)
│   │   ├── video.py             # Video CRF compression (ffmpeg)
│   │   └── pdf.py               # PDF compression
│   ├── db/                      # SQLAlchemy models + async engine (Cloud Edition)
│   ├── ee/                      # Commercial-licensed add-ons (PII redaction) — inert by default
│   ├── models/
│   │   └── schemas.py           # Pydantic response schemas
│   ├── static/                  # CSS, JavaScript and vendored assets (Chart.js)
│   └── templates/               # Jinja2 HTML templates
├── alembic/                     # Cloud Edition schema migrations
├── tests/
├── scripts/
│   ├── generate_api_key.py      # CLI key generator
│   ├── promote_admin.py         # Promote a registered user to the admin role
│   ├── first_run.py             # Called by Docker entrypoint on first start
│   └── build-tailwind.sh        # Rebuilds the self-hosted Tailwind bundle
├── data/
│   └── api_keys.json            # Hashed API keys (gitignored)
├── run.py                       # Entry point for direct Python runs
├── dev.ps1                      # Windows developer startup script (auto-setup + server)
├── create-shortcut.ps1          # Creates a Desktop shortcut for dev.ps1
├── start.bat                    # Windows launcher: Docker mode
├── start.sh                     # Linux/macOS launcher: Docker mode
├── entrypoint.sh                # Docker container entrypoint (first-run key setup)
├── docker-compose.yml           # Community Edition (default)
├── docker-compose.cloud.yml     # Cloud Edition overlay (Postgres, JWT, billing)
├── docker-compose.office.yml    # High-fidelity docx→pdf overlay (LibreOffice)
├── requirements.txt             # Direct dependencies (source of truth for versions)
└── requirements.lock            # Hash-pinned lockfile — what the image actually installs

Development setup

Windows (recommended — one command)

git clone https://github.com/MrChengLen/FileMorph.git
cd FileMorph
.\dev.ps1

dev.ps1 creates the virtual environment, installs dependencies, generates an API key on first run, and starts uvicorn with --reload. Code changes are picked up automatically without restarting the server.

dev.ps1 installs only requirements.txt (the runtime dependencies). To run tests or lint locally, also install the dev tools it doesn't cover (pytest, ruff, pip-audit, …):

.venv\Scripts\pip.exe install -r requirements-dev.txt

Linux / macOS

git clone https://github.com/MrChengLen/FileMorph.git
cd FileMorph

python3 -m venv .venv   # Python 3.11 or newer
source .venv/bin/activate

pip install -r requirements-dev.txt
cp .env.example .env
python scripts/generate_api_key.py

uvicorn app.main:app --reload

Live reload

With --reload, uvicorn watches the project directory for file changes and restarts the server process automatically. No manual restart needed when editing Python files.


Running tests

pytest tests/ -v

Run a single test file:

pytest tests/test_convert_image.py -v

Run with coverage (optional — pytest-cov isn't in requirements-dev.txt, install it separately: pip install pytest-cov):

pytest tests/ --cov=app --cov-report=term-missing

Linting and formatting

FileMorph uses ruff for both linting and formatting.

# Check for issues
ruff check .

# Auto-fix fixable issues
ruff check --fix .

# Format code
ruff format .

CI will fail if either check fails — run both before pushing.


How the converter registry works

The registry maps (source_format, target_format) pairs to converter classes. Converters register themselves using the @register decorator:

# app/converters/registry.py
from app.converters.registry import register
from app.converters.base import BaseConverter


@register(("txt", "pdf"))
class TxtToPdfConverter(BaseConverter):
    def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path:
        # ... conversion logic ...
        return output_path

The register decorator accepts one or more (src, tgt) tuples:

@register(("heic", "jpg"), ("heif", "jpg"))
class HeicToJpgConverter(BaseConverter): ...

The _ensure_loaded() function in registry.py imports all converter modules at startup so their @register decorators run. Add new modules to _ensure_loaded() when creating a new converter file.

Note: the AI / PII-redaction feature under app/ee/ is not a registry converter. It is commercial-licensed (LicenseRef-FileMorph-Commercial), lives outside the @register plugin model, is lazy-imported only by the gated app/api/routes/ai.py, and is inert unless AI_OPERATIONS_ENABLED is set — don't try to @register a redaction "format". See docs/pii-redaction.md.


Adding a new converter

Step 1 — Create or open the relevant converter file

Choose the appropriate file based on the category:

  • app/converters/image.py — images
  • app/converters/document.py — documents
  • app/converters/video.py — video
  • app/converters/audio.py — audio
  • app/converters/spreadsheet.py — data files

For a completely new category, create a new file.

Step 2 — Implement the converter

# Example: adding EPUB → TXT support in a new app/converters/ebook.py

from pathlib import Path
from app.converters.base import BaseConverter
from app.converters.registry import register


@register(("epub", "txt"))
class EpubToTxtConverter(BaseConverter):
    def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path:
        import ebooklib
        from ebooklib import epub
        from bs4 import BeautifulSoup

        book = epub.read_epub(str(input_path))
        texts = []
        for item in book.get_items_of_type(ebooklib.ITEM_DOCUMENT):
            soup = BeautifulSoup(item.get_content(), "html.parser")
            texts.append(soup.get_text())

        output_path.write_text("\n\n".join(texts), encoding="utf-8")
        return output_path

Step 3 — Register the module in _ensure_loaded()

Open app/converters/registry.py and add the import:

def _ensure_loaded() -> None:
    global _loaded
    if _loaded:
        return
    _loaded = True
    import app.converters.audio  # noqa: F401
    import app.converters.document  # noqa: F401
    import app.converters.ebook  # noqa: F401  ← add this
    import app.converters.image  # noqa: F401
    import app.converters.spreadsheet  # noqa: F401
    import app.converters.video  # noqa: F401

Step 4 — Add dependencies

Add any new Python packages to requirements.txt, then recompile requirements.lock — the image installs only from the lockfile, and CI's lockfile-drift job fails until the two match (command under "Reproducible builds" in self-hosting.md). Add system-level dependencies to Dockerfile.

Step 5 — Write tests

# tests/test_convert_ebook.py


def test_epub_to_txt(client, auth_headers, tmp_path):
    # Create a minimal EPUB for testing
    epub_path = tmp_path / "sample.epub"
    # ... create test epub ...

    with epub_path.open("rb") as f:
        res = client.post(
            "/api/v1/convert",
            headers=auth_headers,
            files={"file": ("sample.epub", f, "application/epub+zip")},
            data={"target_format": "txt"},
        )
    assert res.status_code == 200
    assert len(res.content) > 0

Step 6 — Update the format lists

Formats are also listed by hand in several places — docs/formats.md is one — and tests compare them with the registry, so the build fails until they agree. The full list, with the test that pins each place, is under "Parity places" below.


Adding a new compressor

Compressors live in app/compressors/. Each compressor is a plain function (not a class):

# app/compressors/image.py (existing)


def compress_image(input_path: Path, output_path: Path, quality: int = 85) -> Path: ...

Add the new format to the _SUPPORTED_FORMATS list in the relevant compressor file, and import + call it from app/api/routes/compress.py.

Parity places to update in the same PR, for either a new converter or a new compressor format — each is pinned by a test, so a missed one fails the build:

  • docs/formats.md (the From → To tables and the Audio/Video lists), the README (the drop-zone mockup and the "Supported Formats" table), the homepage FAQ answer "Which file formats can I convert?" (EN + DE), the "FileMorph converts …" sentence in /llms.txt, and the "Convert …" entries of the JSON-LD featureList (app/core/jsonld.py) — tests/test_format_lists_match_registry.py compares each with the registry.
  • The homepage drop-zone captions #supported-convert and #supported-compress (app/templates/partials/convert_tool.html, EN + DE) — tests/test_homepage_drop_zone_modes.py compares them with /api/v1/formats.
  • _HOMEPAGE_ADVERTISED in tests/test_format_registry.py — the test's own list of the formats the homepage names. Add the format there; the test only checks that every listed format is registered, so a missing entry goes unnoticed.
  • _FORMAT_CATEGORY in app/api/routes/pages.py — tests/test_formats_categories.py fails for a source format without a category (it would land in the "Other" bucket on /formats).
  • If the format supports exact-size compression, TARGET_SIZE_FORMATS in app/compressors/image.py and the matching TARGET_SIZE_FORMATS array in app/static/js/app.js — tests/test_target_size_formats_parity.py fails the build if the two disagree, and also pins the hand-written claims about which formats hit an exact target (homepage FAQ, /formats, /llms.txt, the OpenAPI form-field docs, the /tools card); the /compress page copy is pinned by tests/test_compress_page.py.
  • A test per new format/pair (Step 5 above / the equivalent for compressors).

API key internals

Key lifecycle:

  1. secrets.token_urlsafe(32) generates a 256-bit random key
  2. hashlib.sha256(key.encode()).hexdigest() produces the stored hash
  3. On each request, the submitted key is hashed and compared against stored hashes
  4. Plaintext keys never touch disk

All logic is in app/core/security.py.


Environment variables reference

Defined in app/core/config.py using pydantic-settings — a representative slice (the real class has ~40 fields: JWT, Stripe, SMTP, audit-log, concurrency, AI-redaction and office-engine settings besides these):

class Settings(BaseSettings):
    app_host: str = "0.0.0.0"
    app_port: int = 8000
    app_debug: bool = False
    app_version: str = "1.1.0"
    api_keys_file: str = ""  # resolved to data/api_keys.json if left empty
    max_upload_size_mb: int = 100
    cors_origins: str = "http://localhost:8000"

All settings can be overridden via environment variables or .env (uppercase, same names). .env.example is the source of truth for the full list — every variable there carries a one-line description; this section only shows the shape.


Making a release

main is branch-protected, so the version bump goes through a PR; the signed tag is cut afterwards on the merged commit. Release tags must be GPG-signed — .github/workflows/release.yml rejects unsigned tags (fail-closed), and only keys listed in release-signing.md verify.

  1. On a branch, bump the version in pyproject.toml and app/core/config.py (e.g. 1.1.0.dev0 → 1.1.0), and roll ## [Unreleased] in CHANGELOG.md into ## [X.Y.Z] — <date> with a fresh empty [Unreleased] above it.
  2. Open a PR, let CI go green, and merge it to main.
  3. Cut the signed tag on the merged commit. On Windows do this in Git Bash (the GPG agent isn't reachable from PowerShell):
    git checkout main && git pull origin main
    git tag -s vX.Y.Z -m "release vX.Y.Z"   # prompts for the GPG passphrase
    git verify-tag vX.Y.Z                     # must print "Good signature"
    git push origin vX.Y.Z

The tag push triggers release.yml (verifies the signature, publishes the GitHub Release with a source tarball, the CycloneDX SBOM and IMAGE_DIGEST.txt) and docker.yml (builds + cosign-signs the slim and office images to GHCR). See release-signing.md for key setup/rotation and docs/patch-policy.md for the versioning rules.