feat(llamacpp): unload idle models, report real GPU use, prune old builds, hold context to the trained length - #323
Open
dovvnloading wants to merge 7 commits into
Open
dovvnloading wants to merge 7 commits into
dovvnloading wants to merge 7 commits into
Conversation
Every pin bump left the previous 100-200 MB llama.cpp build in the runtime folder for good. After the pinned build has been verified, older builds are now removed, once per process and never the cause of a failed launch. Only a sibling named like one of this fetcher's own builds, with a strictly lower build number, is touched: never the pinned release (either backend), a newer build (a rolled-back Cortex sharing the data folder), an in-progress download or extraction, a link, or anything unrecognised. A build a program is running from is kept: every file is opened for writing without writing, which Windows refuses for a running image or a loaded library, and the directory is renamed before it is deleted, which is refused while any file in it is held open. Names are logged, never contents; one pass removes at most eight directories. Plan item: RT-20 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…riod
A loaded model kept its memory until Cortex exited. There is now a manual
unload (POST /api/v1/llamacpp/unload and an "Unload model" button in System
settings) and an idle unload after a configurable number of minutes (default
30, 0 turns it off), read on every check so a change applies at once.
The manager counts requests in flight: the chat client opens a scope around
each call, so a generation longer than the idle period is never unloaded, and
the idle clock restarts when the last request ends. An unload never waits
behind a model load or restart (the idle check skips that round, the manual
one answers 409), and the route also answers 409 while a generation job is
active. The reason ("unloaded after N minutes without use" or "at your
request") stays on the status until the next launch replaces it.
Plan item: RT-07
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t load
active_backend named the build that launched, not whether the GPU is in use,
and the roughly 100 MB Vulkan archive was downloaded and launched on machines
with no Vulkan loader before falling back to the CPU build.
The manager now reads the loader's own offload line ("offloaded N/M layers to
GPU", anchored, digits only, ignored when it echoes model text, last one wins)
and reports the counts on the status; nothing recognised means unknown, not
zero, and the System panel says so, or says the GPU build is running with no
layers on the GPU. In "auto" the loader is probed first (vulkan-1.dll in
System32, then the search path): with none, the CPU build is used directly,
the Vulkan archive is not fetched, and the status carries a fixed note. An
explicit "vulkan" setting is still honoured as before.
Plan items: RT-08, RT-09
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… for and allow tuning flags A requested context window above the model's trained context cost KV-cache memory for positions the model never saw. It is now lowered to the trained length read from the GGUF header (remembered per file version; an unreadable or implausible value means no limit), the reuse decision and the launch both see the lowered value so the same request does not reload the model, and the status and the launch progress say so once. A validated extra_args list in the llama.cpp settings appends cache-type (-ctk, -ctv), flash-attention (-fa) and thread (-t, -tb) flags after the fixed launch contract. Anything else is refused, and the flags that define the launch (model, context, host, port, key, GPU layers, slots, UI) are refused by name; errors give the position, never the text. The list is part of the reuse decision, so a change restarts the model and gives a failing launch a fresh start. Flag spellings are passed through as written and were not run against a real llama-server. Plan item: RT-22 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…xt limit The runtime paragraph named the three GPU backend settings and nothing else. It now says that "auto" skips Vulkan when the machine has no Vulkan loader, that a loaded model is released after 30 idle minutes or with "Unload model" (Settings > System, 0 keeps it loaded), and that the context window is held to what the model was trained for. Plan items: RT-07, RT-09, RT-22 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Resolves the one textual conflict in tests/test_chat_client_routing.py: main and this branch both added a context-manager import; keep main's 'from contextlib import contextmanager' (plus Iterator) and use the bare decorator in this branch's _ScopedProvider. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Owner
Author
Merged main into this branch (4edeae9)
How each conflict was resolvedOnly one file conflicted textually.
Files git merged without a conflict, checked by hand
Which tests cover both sides
Checks run on the merged tree
Rollback is unchanged from the description: revert this PR, or revert only the merge commit to return to the previous branch head. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
active_backendnamed the build that launched, not whether the GPU is used, so the "GPU" label could be wrong (with automatic layer offload, none or only some of the model can be on the GPU).-c, and there was no way to set cache-type, flash-attention or thread flags without editing source.Root cause
Each was a missing piece rather than a defect in existing logic:
stop()was wired only to app shutdown and nothing timed or counted use;BinaryFetcheronly ever created<tag>-<backend>directories;_backend_orderreturnedvulkan, cpuwithout asking whether the GPU build could load; the status recorded the attempted backend and never read the loader's own offload line; the argv was fixed andnum_ctxwas forwarded as requested.Change
Five commits: four concerns, each with its tests, and one README sentence (plan items in the commit bodies):
fix(llamacpp): prune superseded runtime builds after a verified fetch(RT-20). After the pinned build verifies, builds with a strictly lower build number are removed, once per process. Never the pinned build (either backend), a newer one (a rolled-back Cortex sharing the folder),.download-/.extract-in-progress names, a link or junction, or anything unrecognised. A build a program is running from is kept: every file is opened for writing without writing (Windows refuses that for a running image or loaded library), and the directory is renamed before it is deleted (refused while any file is held open). At most eight per pass, names only in the log, and never a reason for a launch to fail.feat(llamacpp): unload the local model on request or after an idle period(RT-07).POST /api/v1/llamacpp/unload, an "Unload model" button and an idle-minutes field in Settings > System, andLlamaCppSettings.idle_unload_minutes(default 30, 0 turns it off, re-read on every check). The chat client opens arequest_scopearound eachchatandtokenize, so a generation longer than the idle period is never unloaded and the idle clock restarts when the last request ends. The idle check never waits behind a load or restart; the manual unload waits two seconds for the same lock and then answers 409; the route also answers 409 while a generation job is active. The reason ("unloaded after N minutes without use" / "at your request") stays onlast_restart_reasonuntil the next launch replaces it.fix(llamacpp): report real GPU use and skip a Vulkan build that cannot load(RT-08, RT-09). The loader's offload line is read (anchored, digits only, ignored if it echoes model text, last one before "listening" wins) and the counts are on the status asgpu_layers_offloaded/gpu_layers_total; nothing recognised means null (unknown), never 0. The System panel shows "GPU (Vulkan) · 24/33 layers on the GPU", says the GPU build is running with no layers on the GPU, or says the split was not reported. In "auto" the loader is probed first (vulkan-1.dllin System32, then the search path): with none, the CPU build is used directly, the Vulkan archive is not fetched, andbackend_notesays why. An explicit "vulkan" setting is honoured as before.feat(llamacpp): hold the context window to what the model was trained for and allow tuning flags(RT-22).num_ctxabove the trained context (read from the GGUF header, remembered per file version; unreadable or implausible means no limit) is lowered, in both the reuse decision and the launch so the same request does not reload the model;context_noteand the launch progress say so once.LlamaCppSettings.extra_argsappends-ctk/-ctv,-fa,-t/-tbafter the fixed launch contract; everything else is refused, and the flags that define the launch (model, context, host, port, key, GPU layers, slots, UI) are refused by name. Errors give a position, never the text. The list is part of the reuse decision, so a change restarts the model and gives a failing launch a fresh start.Where the built change differs from the plan's premise: RT-20 as written would prune every other build; only strictly older ones are pruned, so a rolled-back Cortex keeps its newer build. RT-07 suggests a timestamp bumped from the chat client; a timestamp alone would unload a model in the middle of a long generation, so an in-flight counter (
request_scope) is used and the timestamp only measures idleness. RT-22 places the clamp in the chat client; it is in the manager, because the reuse decision compares the requested window with the loaded one andready_handlemust apply the same limit. Whether the pinned llama-server already clamps on its own was not tested (see Limits).Compatibility and rollback
llamacpp.idle_unload_minutes,llamacpp.extra_args) default when absent, so an existing stored document loads unchanged. Readingcore/settings.py(not run): an older build validates withextra="forbid"and would refuse a document saved by this one; roll back the settings document (or delete the two keys) together with the code. To turn the behaviour off without reverting, set the idle period to 0.contracts/is regenerated. No existing field changes meaning.Checks
Measured on the tree rebased onto main at #316. Main then moved again (#318, frontend only) and was merged in (one conflict in
frontend/src/app/App.tsx, resolved by keeping both sides' props); after that mergenpm run typecheck,npm run lintandnpm test -- --runwere re-run (75 files, 948 tests passed) and the pre-push hook re-ran everything (below). The frontend counts in the next lines are from the #316 tree.python -m ruff check backend tests tools main.py app_factory.py scripts: all checks passed.python -m mypy: no issues in 98 source files.python -m pytest -q(Python 3.14.0): 2948 passed, 3 warnings, 25 subtests passed. The third warning is an existing SyntaxWarning attests/test_open_source_execution.py:189, not from this change.python tools/generate_contracts.py --check: current for the final tree and, checked separately, for each of the five commits.python tools/artifact_boundary_review.py --json --strict: 12 of 12 cases passed.npm run typecheckandnpm run lint: clean.npm test -- --run: 67 files, 869 tests passed.npm run test:coverage: thresholds met (statements 89.53%, branches 85.01%, functions 88.04%, lines 91.73%).py -3.12), the six touched test files: 387 passed; 9 could not run because they write a synthetic GGUF with theggufpackage, which is not installed for that interpreter (ModuleNotFoundError). The whole suite was not run on 3.12, and 3.10 and 3.13 were not run; CI covers them after merge.tests/test_llamacpp_{binary_fetcher,server_manager,status_api,extra_args,unload_api}.py,tests/test_chat_client_routing.py,frontend/src/features/models/ModelsPanel.runtime.test.tsx,frontend/src/features/settings/SettingsPanel.runtime.test.tsx,frontend/src/app/App.runtimeUnload.test.tsxandfrontend/src/api/client.test.ts.ready_handlenot clamped, extra args not appended or not part of reuse or not re-validated, idle ignoring requests in flight, manual unload ignoring requests, idle waiting behind a load, reuse or request end not restarting the idle clock, route ignoring an active generation, and so on), were each killed by at least one new test. One survived at first (link handling in pruning: the target was safe either way, so the test now also asserts the link itself is untouched) and was tightened. This is a hand-run check, not a mutation-testing tool.scripts/check.ps1quick tier) passed all 11 checks on both pushes: 266.8 s on the five commits, and 285.3 s after the merge of feat(frontend): interaction polish A - toasts, chat delete with Undo, approval expiry, memory panel, new-chat guide, drafts #318 (backend tests 211.9 s, artifact-boundary review, contract check, tsc, eslint, vitest 40.9 s).Security / data-loss / concurrency
b<digits>-(cpu|vulkan)with a lower build number, directly inside the runtime folder, not links, and never while a file in them is open for writing or held open; deletion is of a renamed copy so nothing half-deleted can be launched. A program that starts from an older build in the instant between the check and the rename is not detected.extra_argsis an allow-list applied twice (settings validation and again at launch); user words go after the fixed contract and none of them can be a contract flag. The API key still travels only in the environment._ensure_lockbefore_state_lock) is unchanged. The watcher is a daemon thread namedcortex-llama-idle-unload, stopped and joined inclose().Limits
load_tensors: offloaded N/M layers to GPU) and the flag spellings (-ctk -ctv -fa -t -tb, including-fawith or without a value) come from my knowledge of llama.cpp, not from the pinned b10311 build or its documentation. If the line is absent the status says unknown; a wrong flag fails the launch like any bad argument.services/generation.py/services/llm.py(not touched here) still sizes prompts against the requestednum_ctx. When the window is lowered, a prompt between the two sizes reaches llama-server and is rejected there instead of being trimmed (Ollama already has the same mismatch). Follow-up: budget againstloaded_context.active_backend; the frontend change was kept to Settings > System, per this bundle's scope.Plan item: RT-07, RT-08, RT-09, RT-20, RT-22
🤖 Generated with Claude Code