Skip to content

feat(infra): add self-hosted container stack - #117

Open
flamboh wants to merge 3 commits into
stack/alchemy-cloudflarefrom
stack/alchemy-campus
Open

flamboh wants to merge 3 commits into
stack/alchemy-cloudflarefrom
stack/alchemy-campus

Conversation

@flamboh

@flamboh flamboh commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Note

🤖 Claude Opus 5.5 on behalf of Oliver

ELI5

This packages the dashboard as a container that runs on a machine the team controls. People open it through an SSH tunnel from their own computer. Nothing is exposed to the network.

Why

The team needs a self-hosted dashboard that is separate from the dev and Cloudflare deployments. The pipeline already writes SQLite, so running the dashboard next to that data avoids loading it into D1.

Flows to exercise

Setup: a Docker host you can reach with ssh <host> docker info without a password prompt, a data directory on it with <dataset-id>/netflow.sqlite, and CLOUDFLARE_API_TOKEN/CLOUDFLARE_ACCOUNT_ID exported (Alchemy stores state in Cloudflare and creates no Cloudflare resources).

export ATLANTIS_SELF_HOSTED_DOCKER_HOST=ssh://user@host.example.com   # add :2222 for a non-default SSH port
export ATLANTIS_SELF_HOSTED_DATA_DIR=/data/netflow/example
bun run deploy:self-hosted --stage <name>
  1. Deploy: the image builds on the host and the output prints a tunnel command. ssh <host> docker ps --filter name=atlantis-self-hosted-web shows the container as healthy.
  2. Connect: run the printed command, for example ssh -N -L 8080:127.0.0.1:8080 user@host.example.com. With a port in the URL, it prints -p 2222 user@host.example.com. Open http://localhost:8080 and the datasets from the data directory load.
  3. Redeploy: running the deploy again without changes does nothing. After a change under apps/web, the deploy replaces the container.
  4. Teardown: bun run destroy:self-hosted --stage <name> removes the container, the current image, and the Docker context. The data directory is not touched.

Decisions and edge cases

  • Access: the container publishes only on 127.0.0.1:<port> on the host. There is no app auth or TLS, and access goes through SSH port forwarding. Users' SSH keys must allow port forwarding.
  • Config validation: ATLANTIS_SELF_HOSTED_DOCKER_HOST must be an ssh:// URL. ATLANTIS_SELF_HOSTED_DATA_DIR must be an absolute path other than /. ATLANTIS_SELF_HOSTED_DATA_USER must be <uid>:<gid> (default 1000:1000). ATLANTIS_SELF_HOSTED_PORT defaults to 8080. Invalid values stop the deploy before anything changes.
  • Stage names: the self-hosted stack applies the same stage-name rule as the Cloudflare stack (feat(infra): deploy the dashboard to Cloudflare with Alchemy #116): lowercase letters, digits, and single hyphens. --stage alice_dev fails before anything is created, so two spellings can never share a container or Docker context.
  • Writable data mount: SQLite needs to create -shm/-wal files next to WAL-mode databases, even to read them. The app still opens every database read-only. The container runs as the data owner.
  • Image tag: the tag is a content hash. Dockerfile.dockerignore is an allowlist, and the hash walks only the allowlisted paths. Captures, target, and node_modules are never read, so an unreadable or huge directory elsewhere in the checkout cannot slow down or break a deploy. An ignore file that does not start with * is rejected.
  • Toolchain: Bun 1.3.11 is used everywhere (packageManager, shell.nix, the Dockerfile, CI). Node 24.18.1 is pinned in .node-version files, the Dockerfile, and the CI web job, because 24.19+ aborts when better-sqlite3 statements are garbage-collected (nodejs/node#65446). shell.nix and the CI e2e job stay on an earlier Node 24: Playwright 1.52 hangs while loading its config on Node 24.18.1, and moving to Playwright 1.53+ also means moving the nixpkgs browser pin.
  • Resource limits: memory is capped at 1g with a 768 MB heap. patches/alchemy@2.0.0-beta.79.patch adds only memory to Docker.Container, because upstream has no memory option.
  • Leftover images: images from earlier deploys keep their hash tags after destroy. The docs show how to prune them.
  • Not included: the pipeline image. The docs cover running scripts/netflow-db-docker.sh on the host instead.
  • CSRF: the app has no POST, form, or remote endpoints, so SvelteKit's CSRF check never runs. The dev docs say to set paths.origin or PROTOCOL_HEADER before adding any.

Verification

Automated (CI and local): format:check, lint and typecheck (web and infra), test:web, the new test:infra suite, build:web, test:e2e, landing lint, and build:landing. The infra tests cover tunnel-command generation (SSH aliases, user@host:port, IPv6, shell quoting, rejecting non-SSH URLs) and build-context hashing (allowlisted changes and build args change the hash; ignored subtrees and unreadable directories outside the allowlist don't). CI does not build or run the container.

Manual (container, done before the rename and hashing changes): the image built on two Linux Docker hosts. A test stage served a CSV-derived dataset through ssh -L. Pages and five API routes returned 200 and matched local SQLite byte for byte. The port was reachable only on the host's loopback (a request from another machine failed). An unchanged redeploy did nothing, and a source change replaced the container. Destroy left the host as it was and the stack state empty.

Remaining manual verification: run flows 1–4 once more under the new atlantis-self-hosted stack name. Include one ssh://user@host:port URL to check the printed tunnel command.


Claude Opus 5.5 · Claude Code (T3 Code)

@flamboh
flamboh added this pull request to stack #118 September 25, 2026 10:46
Run the SQLite dashboard as a Docker container on a self-hosted host,
managed by the atlantis-campus Alchemy stack. The container publishes
its port on the host loopback only; users connect through SSH port
forwarding. The stack creates no Cloudflare resources and keeps
Cloudflare only as the state store.

- ATLANTIS_DB_DRIVER=sqlite builds use @sveltejs/adapter-node.
- apps/web/Dockerfile builds with Bun and runs on Node 24.18.1, which
  avoids the Node 24.19+ better-sqlite3 GC abort (nodejs/node#65446).
- The image tag is a content hash of the build inputs, so a source
  change replaces the container and an unchanged deploy is a no-op.
- The data mount is writable because WAL-mode databases need their
  -shm/-wal sidecars; the dashboard still opens them read-only.
- LOCAL_DATA_DIR selects the dataset directory the dashboard scans.
…chain

- Rename the campus stack, scripts, env vars, and docs to self-hosted.
- Hash only the Dockerfile.dockerignore allowlist instead of walking the
  whole checkout before applying ignores.
- Emit `-p <port>` and quote arguments in the printed SSH tunnel command,
  and require an ssh:// Docker host URL.
- Pin Bun 1.3.11 and Node 24.18.1 in package.json, .node-version files,
  shell.nix, and CI.
- Add infra unit tests and run them in CI.
- Lead the operations docs with setup steps; move internals to
  architecture docs and drop real host names.
Playwright 1.52 hangs while loading its config on Node 24.18.1. Keep
shell.nix and the CI e2e job on the earlier Node 24 and document why;
Bun stays pinned to 1.3.11 everywhere.
@flamboh
flamboh force-pushed the stack/alchemy-campus branch from e57ea21 to ccb9892 Compare September 27, 2026 04:48
@flamboh flamboh changed the title feat(infra): add self-hosted campus container stack feat(infra): add self-hosted container stack Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant