Skip to content

feat: self-host launch readiness (llama.cpp / vLLM, e2e memory check, cleaner public signals) - #16

Merged
Kortexio merged 3 commits into
mainfrom
chore/self-host-launch-readiness
Sep 28, 2026
Merged

Kortexio merged 3 commits into
mainfrom
chore/self-host-launch-readiness

Conversation

@vitorcastro78

Copy link
Copy Markdown
Collaborator

What and why

The repo gets almost no traffic, and the traffic it does get sees a stale v0.1.0-beta, an Ollama-first README, and an Issues tab full of bot posts. This PR repositions the launch for people who self-host on llama.cpp / vLLM and makes the first "aha" provable in CI.

  • Engines: docker-compose.llamacpp.yml (CPU, pulls a small GGUF) and docker-compose.vllm.yml (GPU) with the flags that matter for tool calling (--jinja, --tool-call-parser).
  • Chat aha: scripts/aha-chat.sh / .ps1 sends two requests in one session. The second one carries only the new question, so the answer has to come from gateway memory.
  • e2e CI: e2e-llamacpp builds the gateway from the commit, starts llama-server with Qwen2.5-1.5B Q4_K_M, and runs aha-chat.sh. Runs on push to main, nightly, and PRs touching the chat path.
  • Public signals:
    • content-cadence writes to the run summary instead of opening issues.
    • sync-mirror only syncs one way (Kortexio to the backup) and gives a clear error when the token is bad.
    • release-please accepts RELEASE_PLEASE_TOKEN.
    • New Conventional Commit PR-title check, issue templates (bug report, engine compatibility report), and a PR template.
  • Docs: README leads with "Self-hosted memory gateway for your llama.cpp / vLLM server" and a three-command quickstart. docs/self-host.md gets an Engines section. docs/show-hn.md becomes a launch kit.

The first commit carries Release-As: 0.2.0-beta.

How it was tested

  • YAML of the compose overrides, workflows, and issue templates parses.
  • bash -n scripts/aha-chat.sh
  • e2e-llamacpp on this PR is the real end-to-end check; Docker was not available locally.

Needs an org admin

  1. Settings > Actions > General: enable "Allow GitHub Actions to create and approve pull requests", or add a RELEASE_PLEASE_TOKEN secret.
  2. Rotate the PEER_SYNC_TOKEN secret (fine-grained PAT, Contents: write on vitorcastro78/ContextMemory). The current one is not recognized by the GitHub API.
  3. Consider allowing squash merge only, so the checked PR title is what lands on main.

…memory check

Adds docker-compose.llamacpp.yml (CPU) and docker-compose.vllm.yml (GPU), scripts/aha-chat.sh/.ps1 (two-turn session recall where the second request carries only the new question), and an e2e-llamacpp workflow that runs it against a real llama-server.

Release-As: 0.2.0-beta
… open PRs

content-cadence writes to the run summary instead of opening issues; sync-mirror is one-way to the backup repo and fails with an actionable message when the token is invalid; release-please accepts RELEASE_PLEASE_TOKEN; PR titles must be Conventional Commits; issue and PR templates.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants