Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 27 additions & 7 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -5,18 +5,38 @@ ADMIN_PORT=5200
PUBLIC_API_URL=http://localhost:5100
ADMIN_PUBLIC_URL=http://localhost:5200

# Preferred host-level LLM base URL (any OpenAI-compatible or Ollama-compatible engine).
# When empty, Compose falls back to OLLAMA_ENDPOINT below.
# LLM_ENDPOINT=http://host.docker.internal:8000
# --- LLM engine -------------------------------------------------------------
# Bundled engines (no host LLM needed):
# llama.cpp: docker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build
# vLLM: docker compose -f docker-compose.yml -f docker-compose.vllm.yml up --build
# The overrides set LLM_ENDPOINT for you; the values below only matter for the plain compose file.

# DX default: Ollama on the host (used when LLM_ENDPOINT is unset)
# Model name sent to the engine. llama-server ignores it; vLLM uses it as --served-model-name;
# Ollama needs a pulled tag (e.g. qwen3.5:9b). Overrides default to local-model, plain compose to qwen3.5:9b.
# DEFAULT_LLM_MODEL=local-model

# llama.cpp override
# LLAMACPP_HF_MODEL=bartowski/Qwen2.5-3B-Instruct-GGUF:Q4_K_M
# LLAMACPP_CTX_SIZE=8192
# LLAMACPP_IMAGE=ghcr.io/ggml-org/llama.cpp:server

# vLLM override
# VLLM_MODEL=Qwen/Qwen2.5-7B-Instruct
# VLLM_MAX_MODEL_LEN=16384
# VLLM_TOOL_CALL_PARSER=hermes
# HF_TOKEN=

# Engine already running on the host (any OpenAI-compatible /v1 server):
# LLM_ENDPOINT=http://host.docker.internal:8080

# Ollama on the host (used only when LLM_ENDPOINT is empty):
OLLAMA_ENDPOINT=http://host.docker.internal:11434
DEFAULT_LLM_MODEL=qwen3.5:9b

# Optional: OpenAI / vLLM / ExLlamaSharp / any OpenAI-compatible host defaults
# OPENAI_ENDPOINT=http://host.docker.internal:8000
# Host defaults for apps with llmBackend openai / openai-compatible / vllm and no per-app endpoint
# OPENAI_ENDPOINT=https://api.openai.com
# OPENAI_API_KEY=

# --- Keys (change before any real use) -------------------------------------
MASTER_KEY=cm_master_dev_key_change_me
DEMO_APP_API_KEY=cm_live_dev_key_change_me

Expand Down
59 changes: 59 additions & 0 deletions .github/ISSUE_TEMPLATE/bug_report.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
name: Bug report
description: Something in the gateway, Admin, or MCP server does not work as documented.
labels: [bug]
body:
- type: textarea
id: what
attributes:
label: What happened
description: What you did, what you expected, and what you got instead. Do not paste API keys or private conversation content.
validations:
required: true
- type: textarea
id: repro
attributes:
label: Steps to reproduce
placeholder: |
1. docker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build
2. ./scripts/aha-chat.sh
3. ...
validations:
required: true
- type: dropdown
id: engine
attributes:
label: LLM engine
options:
- llama.cpp
- vLLM
- Ollama
- LM Studio
- OpenAI / Azure-compatible
- Other OpenAI-compatible
validations:
required: true
- type: input
id: model
attributes:
label: Model
placeholder: e.g. Qwen2.5-7B-Instruct Q4_K_M
- type: dropdown
id: deploy
attributes:
label: Deployment
options:
- Docker Compose (this repo)
- GHCR image
- dotnet run
- Kortexio Cloud
- type: input
id: version
attributes:
label: Version / commit
placeholder: v0.2.0-beta or commit SHA
- type: textarea
id: logs
attributes:
label: Relevant logs
description: Gateway logs around the failure (redact keys and message content).
render: text
5 changes: 5 additions & 0 deletions .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
blank_issues_enabled: true
contact_links:
- name: Kortexio Cloud / commercial license
url: https://kortexio.io
about: Hosted gateway, commercial terms, or enterprise deployment questions.
47 changes: 47 additions & 0 deletions .github/ISSUE_TEMPLATE/engine_report.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
name: Engine / model compatibility report
description: Tell us how a specific engine + model behaves behind ContextMemory (works, partially works, or fails).
labels: [engine-compat]
body:
- type: dropdown
id: engine
attributes:
label: Engine
options:
- llama.cpp
- vLLM
- Ollama
- LM Studio
- SGLang
- TGI
- Other OpenAI-compatible
validations:
required: true
- type: input
id: engine_flags
attributes:
label: Engine version and flags
placeholder: "llama-server b6xxx --jinja --ctx-size 8192"
validations:
required: true
- type: input
id: model
attributes:
label: Model and quantization
placeholder: Qwen2.5-7B-Instruct Q4_K_M
validations:
required: true
- type: checkboxes
id: results
attributes:
label: What works
options:
- label: scripts/aha-chat.sh passes (session memory)
- label: Streaming responses
- label: Agentic mode with wiki_search
- label: Agentic mode with MCP tools
- label: Sandbox tools (shell / python / node)
- type: textarea
id: notes
attributes:
label: Notes
description: Failures, loops, malformed tool calls, context overflows — anything we should add to the engine docs.
12 changes: 12 additions & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
## What and why

<!-- One or two sentences. PR title must be a Conventional Commit (feat:, fix:, docs:, ...). -->

## How it was tested

- [ ] `dotnet test tests/ContextMemory.Api.Tests/ContextMemory.Api.Tests.csproj`
- [ ] `./scripts/aha-chat.sh` against a local engine (if the chat path changed)

## Notes for reviewers

<!-- Contract changes, migrations, config keys, docs updated. -->
24 changes: 7 additions & 17 deletions .github/workflows/content-cadence.yml
Original file line number Diff line number Diff line change
@@ -1,24 +1,22 @@
name: content-cadence

# Suggests a short social hook (manual approve via issue).
# Does not auto-publish to social — safer for brand voice.
# Suggests a short social hook in the run summary (Actions tab).
# Does not auto-publish to social and does not open issues — the public tracker is for users.

on:
schedule:
- cron: "0 14 * * 2" # Tuesday 14:00 UTC
workflow_dispatch:

permissions:
issues: write
contents: read

jobs:
suggest:
if: github.repository == 'Kortexio/ContextMemory'
runs-on: ubuntu-latest
steps:
- name: Open hook suggestion issue
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- name: Write hook suggestion to the run summary
run: |
HOOKS=(
"Your coding agent forgets staging DB names. Fix that in five minutes."
Expand All @@ -31,16 +29,8 @@ jobs:
"Admin Playground: watch tool steps and HITL without a client."
)
HOOK="${HOOKS[$RANDOM % ${#HOOKS[@]}]}"
BODY=$(printf '%s\n\n%s\n\n%s\n' \
printf '%s\n\n%s\n\n%s\n' \
"## Suggested post" \
"$HOOK" \
"Copy to LinkedIn/dev.to after editing. Do **not** call this RAG. CTA: https://github.com/Kortexio/ContextMemory")
gh issue create \
--repo "$GITHUB_REPOSITORY" \
--title "Content cadence: social hook suggestion" \
--label "content-hook" \
--body "$BODY" || \
gh issue create \
--repo "$GITHUB_REPOSITORY" \
--title "Content cadence: social hook suggestion" \
--body "$BODY"
"Copy to LinkedIn/dev.to after editing. Do **not** call this RAG. CTA: https://github.com/Kortexio/ContextMemory" \
>> "$GITHUB_STEP_SUMMARY"
105 changes: 105 additions & 0 deletions .github/workflows/e2e-llamacpp.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
name: e2e-llamacpp

# End-to-end memory check against a real llama.cpp server (CPU, small GGUF):
# gateway image built from this commit + llama-server + scripts/aha-chat.sh.

on:
push:
branches: [main]
pull_request:
paths:
- "src/ContextMemory.Api/**"
- "src/ContextMemory.Adapters/**"
- "src/ContextMemory.Core/**"
- "src/ContextMemory.Infrastructure/**"
- "Dockerfile"
- "scripts/aha-chat.sh"
- ".github/workflows/e2e-llamacpp.yml"
schedule:
- cron: "0 4 * * *"
workflow_dispatch:
inputs:
hf_model:
description: "GGUF to test (<hf-repo>:<quant>)"
required: false
default: "bartowski/Qwen2.5-1.5B-Instruct-GGUF:Q4_K_M"

permissions:
contents: read

concurrency:
group: e2e-llamacpp-${{ github.ref }}
cancel-in-progress: true

env:
HF_MODEL: ${{ inputs.hf_model || 'bartowski/Qwen2.5-1.5B-Instruct-GGUF:Q4_K_M' }}
MODELS_DIR: ${{ github.workspace }}/.llama-models

jobs:
aha:
runs-on: ubuntu-latest
timeout-minutes: 40
steps:
- uses: actions/checkout@v4

- name: Cache GGUF
uses: actions/cache@v4
with:
path: .llama-models
key: llama-gguf-${{ env.HF_MODEL }}

- name: Start llama-server
run: |
mkdir -p "$MODELS_DIR" && chmod 777 "$MODELS_DIR"
docker network create cm-e2e
docker run -d --name llama-server --network cm-e2e -p 8081:8080 \
-v "$MODELS_DIR:/models" -e LLAMA_CACHE=/models \
ghcr.io/ggml-org/llama.cpp:server \
--host 0.0.0.0 --port 8080 --jinja --ctx-size 8192 -hf "$HF_MODEL"

- name: Build gateway image
run: docker build -t contextmemory-e2e -f Dockerfile .

- name: Wait for llama-server
run: |
for i in $(seq 1 120); do
if curl -fsS http://localhost:8081/health >/dev/null 2>&1; then echo "llama-server ready"; exit 0; fi
sleep 5
done
docker logs llama-server | tail -n 50
exit 1

- name: Start gateway
run: |
docker run -d --name cm-api --network cm-e2e -p 5100:8080 \
-e ContextMemory__PersistenceProvider=File \
-e ContextMemory__DataPath=/app/data \
-e ContextMemory__MasterKey=cm_master_e2e \
-e ContextMemory__LlmEndpoint=http://llama-server:8080 \
-e ContextMemory__DefaultLlmModel=local-model \
-e ContextMemory__Apps__demo-dev__ApiKey=cm_live_e2e \
-e ContextMemory__Apps__demo-dev__LlmBackend=openai-compatible \
-e ContextMemory__Apps__demo-dev__LlmEndpoint=http://llama-server:8080 \
-e ContextMemory__Apps__demo-dev__LlmModel=local-model \
contextmemory-e2e
for i in $(seq 1 60); do
if curl -fsS http://localhost:5100/health >/dev/null 2>&1; then echo "gateway ready"; exit 0; fi
sleep 3
done
docker logs cm-api | tail -n 80
exit 1

- name: Run chat aha
env:
CONTEXTMEMORY_BASE_URL: http://localhost:5100
CONTEXTMEMORY_API_KEY: cm_live_e2e
CONTEXTMEMORY_APP_ID: demo-dev
CONTEXTMEMORY_MODEL: local-model
CONTEXTMEMORY_TIMEOUT: "600"
run: bash scripts/aha-chat.sh

- name: Container logs
if: failure()
run: |
echo "::group::cm-api"; docker logs cm-api 2>&1 | tail -n 200; echo "::endgroup::"
echo "::group::llama-server"; docker logs llama-server 2>&1 | tail -n 100; echo "::endgroup::"
32 changes: 32 additions & 0 deletions .github/workflows/pr-title.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
name: pr-title

# release-please only counts Conventional Commits (feat:, fix:, docs:, chore:, ...).
# PRs are squash-merged with the PR title, so the title is what lands on main.

on:
pull_request_target:
types: [opened, edited, synchronize, reopened]

permissions:
pull-requests: read

jobs:
conventional:
runs-on: ubuntu-latest
steps:
- uses: amannn/action-semantic-pull-request@v5
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
types: |
feat
fix
perf
refactor
docs
test
build
ci
chore
revert
requireScope: false
3 changes: 3 additions & 0 deletions .github/workflows/release-please.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ permissions:

jobs:
release-please:
if: github.repository == 'Kortexio/ContextMemory'
runs-on: ubuntu-latest
outputs:
release_created: ${{ steps.release.outputs.release_created }}
Expand All @@ -21,3 +22,5 @@ jobs:
id: release
with:
release-type: simple
# GITHUB_TOKEN cannot open PRs unless the org allows it (Settings > Actions > General).
token: ${{ secrets.RELEASE_PLEASE_TOKEN || github.token }}
Loading
Loading