This guide covers the project architecture, how to run tests, add new converters, and contribute code.
filemorph/
├── app/
│ ├── main.py # FastAPI app: middleware, routers, UI route
│ ├── compat.py # Frozen (.exe) vs. source detection + path helpers
│ ├── api/
│ │ ├── deps.py # API-Key authentication dependency
│ │ └── routes/
│ │ ├── convert.py # POST /api/v1/convert
│ │ ├── compress.py # POST /api/v1/compress
│ │ ├── formats.py # GET /api/v1/formats
│ │ ├── health.py # GET /api/v1/health, /ready
│ │ ├── auth.py # Cloud Edition: register / login / account
│ │ ├── keys.py # Cloud Edition: dashboard API-key CRUD
│ │ ├── billing.py # Cloud Edition: Stripe checkout + webhook
│ │ └── cockpit.py # Cloud Edition: admin routes
│ ├── core/
│ │ ├── config.py # Settings loaded from .env via pydantic-settings
│ │ ├── security.py # API key generation, hashing, validation
│ │ ├── rate_limit.py # slowapi limiter + failed-key budget
│ │ ├── quotas.py # Per-tier size / concurrency / output caps
│ │ └── audit.py # Tamper-evident audit-log hash chain
│ ├── converters/
│ │ ├── base.py # BaseConverter base class
│ │ ├── registry.py # Converter registry (@register decorator)
│ │ ├── image.py # Image conversions (Pillow + pillow-heif + pillow-avif-plugin)
│ │ ├── document.py # Document conversions (docx, pdf, txt, md)
│ │ ├── video.py # Video conversions (ffmpeg-python)
│ │ ├── audio.py # Audio conversions (ffmpeg-python)
│ │ └── spreadsheet.py # Spreadsheet conversions (openpyxl, csv, json)
│ ├── compressors/
│ │ ├── image.py # Image quality / target-size compression (Pillow)
│ │ ├── video.py # Video CRF compression (ffmpeg)
│ │ └── pdf.py # PDF compression
│ ├── db/ # SQLAlchemy models + async engine (Cloud Edition)
│ ├── ee/ # Commercial-licensed add-ons (PII redaction) — inert by default
│ ├── models/
│ │ └── schemas.py # Pydantic response schemas
│ ├── static/ # CSS, JavaScript and vendored assets (Chart.js)
│ └── templates/ # Jinja2 HTML templates
├── alembic/ # Cloud Edition schema migrations
├── tests/
├── scripts/
│ ├── generate_api_key.py # CLI key generator
│ ├── promote_admin.py # Promote a registered user to the admin role
│ ├── first_run.py # Called by Docker entrypoint on first start
│ └── build-tailwind.sh # Rebuilds the self-hosted Tailwind bundle
├── data/
│ └── api_keys.json # Hashed API keys (gitignored)
├── run.py # Entry point for direct Python runs
├── dev.ps1 # Windows developer startup script (auto-setup + server)
├── create-shortcut.ps1 # Creates a Desktop shortcut for dev.ps1
├── start.bat # Windows launcher: Docker mode
├── start.sh # Linux/macOS launcher: Docker mode
├── entrypoint.sh # Docker container entrypoint (first-run key setup)
├── docker-compose.yml # Community Edition (default)
├── docker-compose.cloud.yml # Cloud Edition overlay (Postgres, JWT, billing)
├── docker-compose.office.yml # High-fidelity docx→pdf overlay (LibreOffice)
├── requirements.txt # Direct dependencies (source of truth for versions)
└── requirements.lock # Hash-pinned lockfile — what the image actually installs
git clone https://github.com/MrChengLen/FileMorph.git
cd FileMorph
.\dev.ps1dev.ps1 creates the virtual environment, installs dependencies, generates an API key on
first run, and starts uvicorn with --reload. Code changes are picked up automatically
without restarting the server.
dev.ps1 installs only requirements.txt (the runtime dependencies). To run
tests or lint locally, also install the dev tools it doesn't cover
(pytest, ruff, pip-audit, …):
.venv\Scripts\pip.exe install -r requirements-dev.txtgit clone https://github.com/MrChengLen/FileMorph.git
cd FileMorph
python3 -m venv .venv # Python 3.11 or newer
source .venv/bin/activate
pip install -r requirements-dev.txt
cp .env.example .env
python scripts/generate_api_key.py
uvicorn app.main:app --reloadWith --reload, uvicorn watches the project directory for file changes and
restarts the server process automatically. No manual restart needed when editing Python files.
pytest tests/ -vRun a single test file:
pytest tests/test_convert_image.py -vRun with coverage (optional — pytest-cov isn't in requirements-dev.txt,
install it separately: pip install pytest-cov):
pytest tests/ --cov=app --cov-report=term-missingFileMorph uses ruff for both linting and formatting.
# Check for issues
ruff check .
# Auto-fix fixable issues
ruff check --fix .
# Format code
ruff format .CI will fail if either check fails — run both before pushing.
The registry maps (source_format, target_format) pairs to converter classes.
Converters register themselves using the @register decorator:
# app/converters/registry.py
from app.converters.registry import register
from app.converters.base import BaseConverter
@register(("txt", "pdf"))
class TxtToPdfConverter(BaseConverter):
def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path:
# ... conversion logic ...
return output_pathThe register decorator accepts one or more (src, tgt) tuples:
@register(("heic", "jpg"), ("heif", "jpg"))
class HeicToJpgConverter(BaseConverter): ...The _ensure_loaded() function in registry.py imports all converter modules at startup
so their @register decorators run. Add new modules to _ensure_loaded() when creating
a new converter file.
Note: the AI / PII-redaction feature under
app/ee/is not a registry converter. It is commercial-licensed (LicenseRef-FileMorph-Commercial), lives outside the@registerplugin model, is lazy-imported only by the gatedapp/api/routes/ai.py, and is inert unlessAI_OPERATIONS_ENABLEDis set — don't try to@registera redaction "format". Seedocs/pii-redaction.md.
Choose the appropriate file based on the category:
app/converters/image.py— imagesapp/converters/document.py— documentsapp/converters/video.py— videoapp/converters/audio.py— audioapp/converters/spreadsheet.py— data files
For a completely new category, create a new file.
# Example: adding EPUB → TXT support in a new app/converters/ebook.py
from pathlib import Path
from app.converters.base import BaseConverter
from app.converters.registry import register
@register(("epub", "txt"))
class EpubToTxtConverter(BaseConverter):
def convert(self, input_path: Path, output_path: Path, **kwargs) -> Path:
import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup
book = epub.read_epub(str(input_path))
texts = []
for item in book.get_items_of_type(ebooklib.ITEM_DOCUMENT):
soup = BeautifulSoup(item.get_content(), "html.parser")
texts.append(soup.get_text())
output_path.write_text("\n\n".join(texts), encoding="utf-8")
return output_pathOpen app/converters/registry.py and add the import:
def _ensure_loaded() -> None:
global _loaded
if _loaded:
return
_loaded = True
import app.converters.audio # noqa: F401
import app.converters.document # noqa: F401
import app.converters.ebook # noqa: F401 ← add this
import app.converters.image # noqa: F401
import app.converters.spreadsheet # noqa: F401
import app.converters.video # noqa: F401Add any new Python packages to requirements.txt, then recompile requirements.lock — the image installs only from the lockfile, and CI's lockfile-drift job fails until the two match (command under "Reproducible builds" in self-hosting.md). Add system-level dependencies to Dockerfile.
# tests/test_convert_ebook.py
def test_epub_to_txt(client, auth_headers, tmp_path):
# Create a minimal EPUB for testing
epub_path = tmp_path / "sample.epub"
# ... create test epub ...
with epub_path.open("rb") as f:
res = client.post(
"/api/v1/convert",
headers=auth_headers,
files={"file": ("sample.epub", f, "application/epub+zip")},
data={"target_format": "txt"},
)
assert res.status_code == 200
assert len(res.content) > 0Formats are also listed by hand in several places — docs/formats.md is one — and tests compare them with the registry, so the build fails until they agree. The full list, with the test that pins each place, is under "Parity places" below.
Compressors live in app/compressors/. Each compressor is a plain function (not a class):
# app/compressors/image.py (existing)
def compress_image(input_path: Path, output_path: Path, quality: int = 85) -> Path: ...Add the new format to the _SUPPORTED_FORMATS list in the relevant compressor file,
and import + call it from app/api/routes/compress.py.
Parity places to update in the same PR, for either a new converter or a new compressor format — each is pinned by a test, so a missed one fails the build:
docs/formats.md(the From → To tables and the Audio/Video lists), the README (the drop-zone mockup and the "Supported Formats" table), the homepage FAQ answer "Which file formats can I convert?" (EN + DE), the "FileMorph converts …" sentence in/llms.txt, and the "Convert …" entries of the JSON-LDfeatureList(app/core/jsonld.py) —tests/test_format_lists_match_registry.pycompares each with the registry.- The homepage drop-zone captions
#supported-convertand#supported-compress(app/templates/partials/convert_tool.html, EN + DE) —tests/test_homepage_drop_zone_modes.pycompares them with/api/v1/formats. _HOMEPAGE_ADVERTISEDintests/test_format_registry.py— the test's own list of the formats the homepage names. Add the format there; the test only checks that every listed format is registered, so a missing entry goes unnoticed._FORMAT_CATEGORYinapp/api/routes/pages.py—tests/test_formats_categories.pyfails for a source format without a category (it would land in the "Other" bucket on/formats).- If the format supports exact-size compression,
TARGET_SIZE_FORMATSinapp/compressors/image.pyand the matchingTARGET_SIZE_FORMATSarray inapp/static/js/app.js—tests/test_target_size_formats_parity.pyfails the build if the two disagree, and also pins the hand-written claims about which formats hit an exact target (homepage FAQ,/formats,/llms.txt, the OpenAPI form-field docs, the/toolscard); the/compresspage copy is pinned bytests/test_compress_page.py. - A test per new format/pair (Step 5 above / the equivalent for compressors).
Key lifecycle:
secrets.token_urlsafe(32)generates a 256-bit random keyhashlib.sha256(key.encode()).hexdigest()produces the stored hash- On each request, the submitted key is hashed and compared against stored hashes
- Plaintext keys never touch disk
All logic is in app/core/security.py.
Defined in app/core/config.py using pydantic-settings — a representative
slice (the real class has ~40 fields: JWT, Stripe, SMTP, audit-log,
concurrency, AI-redaction and office-engine settings besides these):
class Settings(BaseSettings):
app_host: str = "0.0.0.0"
app_port: int = 8000
app_debug: bool = False
app_version: str = "1.1.0"
api_keys_file: str = "" # resolved to data/api_keys.json if left empty
max_upload_size_mb: int = 100
cors_origins: str = "http://localhost:8000"All settings can be overridden via environment variables or .env
(uppercase, same names). .env.example is the
source of truth for the full list — every variable there carries a
one-line description; this section only shows the shape.
main is branch-protected, so the version bump goes through a PR; the
signed tag is cut afterwards on the merged commit. Release tags must be
GPG-signed — .github/workflows/release.yml rejects unsigned tags (fail-closed),
and only keys listed in release-signing.md verify.
- On a branch, bump the version in
pyproject.tomlandapp/core/config.py(e.g.1.1.0.dev0→1.1.0), and roll## [Unreleased]inCHANGELOG.mdinto## [X.Y.Z] — <date>with a fresh empty[Unreleased]above it. - Open a PR, let CI go green, and merge it to
main. - Cut the signed tag on the merged commit. On Windows do this in Git Bash
(the GPG agent isn't reachable from PowerShell):
git checkout main && git pull origin main git tag -s vX.Y.Z -m "release vX.Y.Z" # prompts for the GPG passphrase git verify-tag vX.Y.Z # must print "Good signature" git push origin vX.Y.Z
The tag push triggers release.yml (verifies the signature, publishes the
GitHub Release with a source tarball, the CycloneDX SBOM and
IMAGE_DIGEST.txt) and docker.yml (builds + cosign-signs the slim and
office images to GHCR). See release-signing.md
for key setup/rotation and docs/patch-policy.md for the versioning rules.