Skip to content

feat(llamacpp): unload idle models, report real GPU use, prune old builds, hold context to the trained length - #323

Open
dovvnloading wants to merge 7 commits into
mainfrom
feat/runtime-unload-and-hygiene
Open

dovvnloading wants to merge 7 commits into
mainfrom
feat/runtime-unload-and-hygiene

Conversation

@dovvnloading

@dovvnloading dovvnloading commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Problem

  • A loaded llama.cpp model kept its memory until Cortex exited. There was no way to release it, and no idle release.
  • Every runtime pin bump left the previous 100-200 MB llama.cpp build in the runtime folder for good.
  • With the default "auto" backend the roughly 100 MB Vulkan archive was downloaded and launched even on a machine with no Vulkan loader, and only then abandoned for the CPU build.
  • active_backend named the build that launched, not whether the GPU is used, so the "GPU" label could be wrong (with automatic layer offload, none or only some of the model can be on the GPU).
  • A requested context window above what the model was trained for was passed straight to -c, and there was no way to set cache-type, flash-attention or thread flags without editing source.

Root cause

Each was a missing piece rather than a defect in existing logic: stop() was wired only to app shutdown and nothing timed or counted use; BinaryFetcher only ever created <tag>-<backend> directories; _backend_order returned vulkan, cpu without asking whether the GPU build could load; the status recorded the attempted backend and never read the loader's own offload line; the argv was fixed and num_ctx was forwarded as requested.

Change

Five commits: four concerns, each with its tests, and one README sentence (plan items in the commit bodies):

  1. fix(llamacpp): prune superseded runtime builds after a verified fetch (RT-20). After the pinned build verifies, builds with a strictly lower build number are removed, once per process. Never the pinned build (either backend), a newer one (a rolled-back Cortex sharing the folder), .download- / .extract- in-progress names, a link or junction, or anything unrecognised. A build a program is running from is kept: every file is opened for writing without writing (Windows refuses that for a running image or loaded library), and the directory is renamed before it is deleted (refused while any file is held open). At most eight per pass, names only in the log, and never a reason for a launch to fail.
  2. feat(llamacpp): unload the local model on request or after an idle period (RT-07). POST /api/v1/llamacpp/unload, an "Unload model" button and an idle-minutes field in Settings > System, and LlamaCppSettings.idle_unload_minutes (default 30, 0 turns it off, re-read on every check). The chat client opens a request_scope around each chat and tokenize, so a generation longer than the idle period is never unloaded and the idle clock restarts when the last request ends. The idle check never waits behind a load or restart; the manual unload waits two seconds for the same lock and then answers 409; the route also answers 409 while a generation job is active. The reason ("unloaded after N minutes without use" / "at your request") stays on last_restart_reason until the next launch replaces it.
  3. fix(llamacpp): report real GPU use and skip a Vulkan build that cannot load (RT-08, RT-09). The loader's offload line is read (anchored, digits only, ignored if it echoes model text, last one before "listening" wins) and the counts are on the status as gpu_layers_offloaded / gpu_layers_total; nothing recognised means null (unknown), never 0. The System panel shows "GPU (Vulkan) · 24/33 layers on the GPU", says the GPU build is running with no layers on the GPU, or says the split was not reported. In "auto" the loader is probed first (vulkan-1.dll in System32, then the search path): with none, the CPU build is used directly, the Vulkan archive is not fetched, and backend_note says why. An explicit "vulkan" setting is honoured as before.
  4. feat(llamacpp): hold the context window to what the model was trained for and allow tuning flags (RT-22). num_ctx above the trained context (read from the GGUF header, remembered per file version; unreadable or implausible means no limit) is lowered, in both the reuse decision and the launch so the same request does not reload the model; context_note and the launch progress say so once. LlamaCppSettings.extra_args appends -ctk/-ctv, -fa, -t/-tb after the fixed launch contract; everything else is refused, and the flags that define the launch (model, context, host, port, key, GPU layers, slots, UI) are refused by name. Errors give a position, never the text. The list is part of the reuse decision, so a change restarts the model and gives a failing launch a fresh start.

Where the built change differs from the plan's premise: RT-20 as written would prune every other build; only strictly older ones are pruned, so a rolled-back Cortex keeps its newer build. RT-07 suggests a timestamp bumped from the chat client; a timestamp alone would unload a model in the middle of a long generation, so an in-flight counter (request_scope) is used and the timestamp only measures idleness. RT-22 places the clamp in the chat client; it is in the manager, because the reuse decision compares the requested window with the loaded one and ready_handle must apply the same limit. Whether the pinned llama-server already clamps on its own was not tested (see Limits).

Compatibility and rollback

  • New settings keys (llamacpp.idle_unload_minutes, llamacpp.extra_args) default when absent, so an existing stored document loads unchanged. Reading core/settings.py (not run): an older build validates with extra="forbid" and would refuse a document saved by this one; roll back the settings document (or delete the two keys) together with the code. To turn the behaviour off without reverting, set the idle period to 0.
  • Idle unload is on by default (30 minutes); the model then reloads on the next message. Previously it never unloaded.
  • The API gains one route and four optional status fields; contracts/ is regenerated. No existing field changes meaning.
  • Reverting the PR removes everything; runtime folders already pruned are re-downloaded on demand.

Checks

Measured on the tree rebased onto main at #316. Main then moved again (#318, frontend only) and was merged in (one conflict in frontend/src/app/App.tsx, resolved by keeping both sides' props); after that merge npm run typecheck, npm run lint and npm test -- --run were re-run (75 files, 948 tests passed) and the pre-push hook re-ran everything (below). The frontend counts in the next lines are from the #316 tree.

  • python -m ruff check backend tests tools main.py app_factory.py scripts: all checks passed.
  • python -m mypy: no issues in 98 source files.
  • python -m pytest -q (Python 3.14.0): 2948 passed, 3 warnings, 25 subtests passed. The third warning is an existing SyntaxWarning at tests/test_open_source_execution.py:189, not from this change.
  • python tools/generate_contracts.py --check: current for the final tree and, checked separately, for each of the five commits.
  • python tools/artifact_boundary_review.py --json --strict: 12 of 12 cases passed.
  • npm run typecheck and npm run lint: clean. npm test -- --run: 67 files, 869 tests passed. npm run test:coverage: thresholds met (statements 89.53%, branches 85.01%, functions 88.04%, lines 91.73%).
  • Python 3.12 (py -3.12), the six touched test files: 387 passed; 9 could not run because they write a synthetic GGUF with the gguf package, which is not installed for that interpreter (ModuleNotFoundError). The whole suite was not run on 3.12, and 3.10 and 3.13 were not run; CI covers them after merge.
  • Tests added: 99 Python test functions (several parametrized) and 34 frontend tests, in tests/test_llamacpp_{binary_fetcher,server_manager,status_api,extra_args,unload_api}.py, tests/test_chat_client_routing.py, frontend/src/features/models/ModelsPanel.runtime.test.tsx, frontend/src/features/settings/SettingsPanel.runtime.test.tsx, frontend/src/app/App.runtimeUnload.test.tsx and frontend/src/api/client.test.ts.
  • Do the new tests fail without the change? Not shown by stashing the whole change. Instead 31 hand-made mutations, each breaking one behaviour in a temporary copy of the source (prune skipped, in-use probe ignored, newer builds pruned, probe ignored, offload never recorded or first-line-wins or echoed text trusted, no clamp, ready_handle not clamped, extra args not appended or not part of reuse or not re-validated, idle ignoring requests in flight, manual unload ignoring requests, idle waiting behind a load, reuse or request end not restarting the idle clock, route ignoring an active generation, and so on), were each killed by at least one new test. One survived at first (link handling in pruning: the target was safe either way, so the test now also asserts the link itself is untouched) and was tightened. This is a hand-run check, not a mutation-testing tool.
  • Each commit was also checked on its own in an exported copy of its tree: ruff, mypy, contract check, its test files, and for the commits that touch the frontend, typecheck, lint and the touched frontend tests. That was done on main at feat(frontend): pick GGUF files from a repository, cancel and resume model downloads #314; the final rebase onto fix(api): small API defects - drop the raw message write, kind-checked /generations, can_cancel, client route check, one-query task list, 1 MiB body ceiling #316 (unrelated API files) applied without conflicts, and only the contract check was repeated per commit after it.
  • The pre-push hook (scripts/check.ps1 quick tier) passed all 11 checks on both pushes: 266.8 s on the five commits, and 285.3 s after the merge of feat(frontend): interaction polish A - toasts, chat delete with Undo, approval expiry, memory panel, new-chat guide, drafts #318 (backend tests 211.9 s, artifact-boundary review, contract check, tsc, eslint, vitest 40.9 s).
  • No flaky test was seen. The machine was heavily loaded and the runs were slow, but nothing failed on a rerun.

Security / data-loss / concurrency

  • Pruning deletes only directories matching b<digits>-(cpu|vulkan) with a lower build number, directly inside the runtime folder, not links, and never while a file in them is open for writing or held open; deletion is of a renamed copy so nothing half-deleted can be launched. A program that starts from an older build in the instant between the check and the rename is not detected.
  • extra_args is an allow-list applied twice (settings validation and again at launch); user words go after the fixed contract and none of them can be a contract flag. The API key still travels only in the environment.
  • Nothing from the child is stored except two integers (the offload counts) taken from an anchored line that is skipped when it echoes model text; a hostile model file that can print a whole log line could still spoof those numbers (same limit as the failure classifier). The new log lines carry build names and fixed reasons only; no prompts, responses or child output.
  • Unload and idle unload take the same slow-path lock as loading; an unload in progress makes a new request wait and then load again. Lock order (_ensure_lock before _state_lock) is unchanged. The watcher is a daemon thread named cortex-llama-idle-unload, stopped and joined in close().
  • The unload route requires the session token like every other route.

Limits

  • Not run against a real llama-server (no binary available here, and none was downloaded). The offload line format (load_tensors: offloaded N/M layers to GPU) and the flag spellings (-ctk -ctv -fa -t -tb, including -fa with or without a value) come from my knowledge of llama.cpp, not from the pinned b10311 build or its documentation. If the line is absent the status says unknown; a wrong flag fails the launch like any bad argument.
  • Whether the Vulkan build with a loader but no device exits or silently runs on the CPU is still unverified; the layer counts answer it at runtime when the line is present.
  • The prompt budget in services/generation.py / services/llm.py (not touched here) still sizes prompts against the requested num_ctx. When the window is lowered, a prompt between the two sizes reaches llama-server and is rejected there instead of being trimmed (Ollama already has the same mismatch). Follow-up: budget against loaded_context.
  • The composer's "Loaded · GPU" chip still derives from active_backend; the frontend change was kept to Settings > System, per this bundle's scope.
  • The idle period is enforced by a 30 second poll, so an unload can land up to 30 seconds late.

Plan item: RT-07, RT-08, RT-09, RT-20, RT-22

🤖 Generated with Claude Code

dovvnloading and others added 7 commits September 29, 2026 05:28
Every pin bump left the previous 100-200 MB llama.cpp build in the runtime
folder for good. After the pinned build has been verified, older builds are
now removed, once per process and never the cause of a failed launch.

Only a sibling named like one of this fetcher's own builds, with a strictly
lower build number, is touched: never the pinned release (either backend), a
newer build (a rolled-back Cortex sharing the data folder), an in-progress
download or extraction, a link, or anything unrecognised. A build a program
is running from is kept: every file is opened for writing without writing,
which Windows refuses for a running image or a loaded library, and the
directory is renamed before it is deleted, which is refused while any file in
it is held open. Names are logged, never contents; one pass removes at most
eight directories.

Plan item: RT-20

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…riod

A loaded model kept its memory until Cortex exited. There is now a manual
unload (POST /api/v1/llamacpp/unload and an "Unload model" button in System
settings) and an idle unload after a configurable number of minutes (default
30, 0 turns it off), read on every check so a change applies at once.

The manager counts requests in flight: the chat client opens a scope around
each call, so a generation longer than the idle period is never unloaded, and
the idle clock restarts when the last request ends. An unload never waits
behind a model load or restart (the idle check skips that round, the manual
one answers 409), and the route also answers 409 while a generation job is
active. The reason ("unloaded after N minutes without use" or "at your
request") stays on the status until the next launch replaces it.

Plan item: RT-07

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t load

active_backend named the build that launched, not whether the GPU is in use,
and the roughly 100 MB Vulkan archive was downloaded and launched on machines
with no Vulkan loader before falling back to the CPU build.

The manager now reads the loader's own offload line ("offloaded N/M layers to
GPU", anchored, digits only, ignored when it echoes model text, last one wins)
and reports the counts on the status; nothing recognised means unknown, not
zero, and the System panel says so, or says the GPU build is running with no
layers on the GPU. In "auto" the loader is probed first (vulkan-1.dll in
System32, then the search path): with none, the CPU build is used directly,
the Vulkan archive is not fetched, and the status carries a fixed note. An
explicit "vulkan" setting is still honoured as before.

Plan items: RT-08, RT-09

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… for and allow tuning flags

A requested context window above the model's trained context cost KV-cache
memory for positions the model never saw. It is now lowered to the trained
length read from the GGUF header (remembered per file version; an unreadable
or implausible value means no limit), the reuse decision and the launch both
see the lowered value so the same request does not reload the model, and the
status and the launch progress say so once.

A validated extra_args list in the llama.cpp settings appends cache-type
(-ctk, -ctv), flash-attention (-fa) and thread (-t, -tb) flags after the
fixed launch contract. Anything else is refused, and the flags that define
the launch (model, context, host, port, key, GPU layers, slots, UI) are
refused by name; errors give the position, never the text. The list is part
of the reuse decision, so a change restarts the model and gives a failing
launch a fresh start. Flag spellings are passed through as written and were
not run against a real llama-server.

Plan item: RT-22

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…xt limit

The runtime paragraph named the three GPU backend settings and nothing else.
It now says that "auto" skips Vulkan when the machine has no Vulkan loader,
that a loaded model is released after 30 idle minutes or with "Unload model"
(Settings > System, 0 keeps it loaded), and that the context window is held to
what the model was trained for.

Plan items: RT-07, RT-09, RT-22

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Resolves the one textual conflict in tests/test_chat_client_routing.py:
main and this branch both added a context-manager import; keep main's
'from contextlib import contextmanager' (plus Iterator) and use the bare
decorator in this branch's _ScopedProvider.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@dovvnloading

Copy link
Copy Markdown
Owner Author

Merged main into this branch (4edeae9)

origin/main (through #315: #313, #314, #316, #318, #322 and others) was merged in by merge commit, with no rebase or force-push. The branch is now mergeable.

How each conflict was resolved

Only one file conflicted textually.

  • tests/test_chat_client_routing.py: both sides added an import for a context manager. Main added from collections.abc import Iterator and from contextlib import contextmanager (for its _fake_ollama_listening helper). This branch added import contextlib (for the _ScopedProvider.request_scope test double). I kept main's two imports and changed the branch's decorator from @contextlib.contextmanager to @contextmanager, so no unused or duplicate import remains. Nothing else in the file changed.

Files git merged without a conflict, checked by hand

  • backend/cortex_backend/core/settings.py: both sides' settings are present with their own defaults and bounds. From main, generation.keep_alive_minutes (default 0, -1 to 1440). From this branch, llamacpp.idle_unload_minutes (default 30, 0 to 1440) and llamacpp.extra_args (allow-listed). The merge did not touch either.
  • contracts/openapi.json and contracts/cortex-api.ts: auto-merged. python tools/generate_contracts.py --check exits 0, so a regeneration would change nothing.
  • frontend/src/app/App.tsx and frontend/src/features/settings/SettingsPanel.tsx: main's model-download props (onDownloadGGUF, onListHuggingFaceFiles, model progress and cancel) and this branch's onUnloadModel and the idle/extra-args runtime props are all passed through.
  • README.md, backend/cortex_backend/api/routes.py and backend/cortex_backend/llamacpp/chat_client.py merged cleanly. The chat client keeps main's request-id and generation changes together with this branch's request_scope wrapper.

git diff origin/main --stat lists 29 files, all of which belong to this pull request. requirements*.lock.txt and frontend/package-lock.json are not in it. The pre-push hook rewrote the two requirements*.lock.txt files' line endings in the working tree, and I restored them before finishing.

Which tests cover both sides

  • Settings from both sides: tests/test_settings_compatibility.py (main's keep-alive default, bounds and "saved before it existed still loads" cases) and tests/test_llamacpp_unload_api.py (idle default, bounds, round trip through the settings API, rejection of 5000). Both files pass. No single test sets keep_alive_minutes and idle_unload_minutes in one document; they live in different sections (generation and llamacpp) and are covered separately.
  • Contract and API shape: tests/test_api_contract.py, tests/test_llamacpp_status_api.py and tests/test_llamacpp_unload_api.py.
  • Chat client with both changes: tests/test_chat_client_routing.py, which holds main's tests and this branch's request_scope tests.
  • Frontend: SettingsPanel.runtime.test.tsx, ModelsPanel.runtime.test.tsx and App.runtimeUnload.test.tsx (unload and idle props) sit next to main's existing model-download and settings tests, in one vitest run that passes.

Checks run on the merged tree

  • python -m ruff check backend tests tools main.py app_factory.py scripts: all checks passed.
  • python -m mypy: no issues in 98 source files.
  • python -m pytest -q (Python 3.14.0): 3109 passed, 3 warnings, 25 subtests passed (230.7 s). The warnings are the two existing deprecation notices and the existing SyntaxWarning at tests/test_open_source_execution.py:189.
  • python tools/generate_contracts.py --check: exit 0.
  • python tools/artifact_boundary_review.py --json --strict: 12 of 12 cases passed.
  • npm ci, then npm run typecheck, npm run lint: clean. npm test -- --run: 75 files, 948 tests passed.
  • Python 3.12 (py -3.12): the six test files this branch touches plus tests/test_settings_compatibility.py and tests/test_api_contract.py: 496 passed. The gguf package is not installed for 3.12, so a copy of the 3.14 package was put on PYTHONPATH. That means the synthetic-GGUF tests ran this time. The whole suite was not run on 3.12, and 3.10 and 3.13 were not run; CI covers them after merge.
  • Pre-push hook (scripts/check.ps1 quick tier): all 11 checks passed on the first attempt in 313.9 s (backend tests 248.7 s, vitest 40.9 s, 75 files and 948 tests). No reruns were needed and nothing flaked.

Rollback is unchanged from the description: revert this PR, or revert only the merge commit to return to the previous branch head.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant