Repository navigation
Feat/model run profile - #489
Conversation
Plan and start now use one axis.model-run/v1 profile. With no new flags the argv stays llama-server, -m, the weights path, --port, the port, and --host 127.0.0.1. A plan-default port refuses until --port is passed. --n-gpu-layers is a string and requires measured VRAM on one discrete device.
The three nvidia-smi memory queries now start with index. A four-column row records Index and IndexSource nvidia-smi only when every numeric cell parses. Legacy three-column and two-column rows, Metal, and lspci leave Index nil. --main-gpu N is emitted only when that N matches an observed nvidia-smi index. The receipt quotes the split-mode none and row help. An omitted pin does not become 0.
…ening --ollama-model replaces --weights and is mutually exclusive with it. Load and unload curl 127.0.0.1:11434 /api/generate, then GET /api/ps. Axis does not exec ollama serve, does not send num_gpu or main_gpu, and does not kill comm=ollama. A generation whose engine is ollama unloads on that path and never reaches the llama-server process kill.
axis model start builds a profile and calls PlanStartProfile. PlanStart had no production caller, so the deadcode gate rejected the branch. Tests build that same default profile themselves.
toasterbook88
left a comment
There was a problem hiding this comment.
AXIS Grok gate
Verdict: PASS-WITH-RISK
PR: #489 Feat/model run profile
Head: d169143
Base: main @ cb202b8
URL: #489
Files (20): cmd/axis/model.go, model_ollama_test.go, model_run_profile.go, model_run_profile_test.go; internal/facts/local_gpu.go, remote.go, remote_bundle.go, gpu_index_test.go, gpu_query_test.go; internal/modellife/plan.go, plan_test.go, plan_profile_test.go, ollama.go, ollama_test.go; internal/modelplan/single_node.go, profile_test.go; internal/models/run_profile.go, run_profile_test.go, model_operation.go, types.go.
Invariant check:
- Advisory does not override the fact plane. axis.model-plan/v1 Selected is labeled advisory. Start rebuilds device fields from the snapshot, refuses a plan-default port, refuses n-gpu-layers without measured free VRAM on one discrete device, and emits --main-gpu only when that index is an observed nvidia-smi index. Omitted pin stays nil, not 0. Metal/lspci/legacy three-column rows leave Index nil.
- Cache stays explicit. Existing loadModelCommandSnapshot path; --live unchanged; receipts carry snapshot source and publication id. Load fact for Ollama is GET /api/ps after the curl, not OllamaInfo.Listening.
- HITL / dispatch lock not touched. No secret_read, ssh_bypass, fleet_exec, spawn_subagent, or self-authorize. Remote curl uses the existing node SSH seam, loopback 127.0.0.1:11434 only, shell-single-quoted JSON. Does not exec ollama serve, does not send num_gpu/main_gpu, does not kill comm=ollama.
- Agent self-model untouched.
Residuals:
- CI not green yet. Test & Build, govulncheck, Analyze (actions), and copilot-pull-request-reviewer are in progress. auto-merge skipped. Combined status success is Devin Review skip (trial expired), not the test suite.
- Ollama place and --ollama-model stop do not require StatusComplete or OllamaInfo.Listening before the SSH curl. Fail-closed on curl or missing /api/ps listing, but an incomplete node can still be targeted if resolve returns it. POST /api/generate can run tokens; it is not a pure load API.
- Receipt VRAMFreeMeasured comes from ObserveLaunchDevice (best discrete / unified), not from the pinned nvidia-smi index. A --main-gpu pin can name a different device than the VRAM figure on the receipt.
- --write-profile is 0644 on an operator path. Plan text does not print Selected.Refusals; JSON does. Start still refuses those profiles.
Operator action: do not merge on this comment. Wait for Test & Build / govulncheck / Analyze on d169143. No APPROVE from this gate.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
GPU fact framing, pinned-device validation, Ollama verification, receipt handling, and profile contracts contain correctness issues.
Review effort: Balanced
Findings: 7
Open (9)
Refresh daemon state after Ollama unload · New Honor profile refusals in Ollama execution · New Group fallback command before base64 encoding · New Accept Ollama's canonical :latest model name · New Resolve and measure the pinned GPU device · New Add snake_case YAML tags to model placement profile fields · New Deep-copy GPUInfo.Index in cloned snapshots · New Update documentation for Ollama lifecycle support · New Use Ollama-specific placement output formatting · New
What changed in this PR
Adds truth-backed model run profiles connecting placement plans to model lifecycle execution, including GPU pinning and Ollama placement.
Changes:
- Adds validated run profiles and llama-server flag projection.
- Collects NVIDIA GPU indices for explicit pinning.
- Adds Ollama load/unload lifecycle support and tests.
| File | Description |
|---|---|
internal/models/types.go |
Adds GPU index facts. |
internal/models/run_profile.go |
Defines run profiles and validation. |
internal/models/run_profile_test.go |
Tests profile behavior. |
internal/models/model_operation.go |
Extends operation receipts. |
internal/modelplan/single_node.go |
Adds selected launch profiles. |
internal/modelplan/profile_test.go |
Tests planned profiles. |
internal/modellife/plan.go |
Projects profiles into llama-server argv. |
internal/modellife/plan_test.go |
Adapts start-plan tests. |
internal/modellife/plan_profile_test.go |
Tests profile execution safeguards. |
internal/modellife/ollama.go |
Implements Ollama load/unload scripts. |
internal/modellife/ollama_test.go |
Tests Ollama scripts. |
internal/facts/remote.go |
Reuses GPU collection commands. |
internal/facts/remote_bundle.go |
Collects remote GPU indices. |
internal/facts/local_gpu.go |
Parses indexed NVIDIA facts. |
internal/facts/gpu_query_test.go |
Tests GPU query consistency. |
internal/facts/gpu_index_test.go |
Tests GPU index parsing. |
cmd/axis/model.go |
Integrates profiles and Ollama lifecycle. |
cmd/axis/model_run_profile.go |
Handles profile CLI input/output. |
cmd/axis/model_run_profile_test.go |
Tests profile CLI flows. |
cmd/axis/model_ollama_test.go |
Tests Ollama CLI lifecycle. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
mlx_lm.server is the observed launch tool. The stop guard matches that basename or the argv sequence -m then mlx_lm.server. It does not kill comm=python or comm=mlx_lm. Port-only stops stay on the llama-server check.
A device hold stores one GPU index and that device's MiB. Index 0 round-trips. Load drops an expired hold without adding it to the VRAM sums, and entry writes keep the hold slice.
… time After the llama-server probe succeeds, sample that node's listener RSS and store one ExecutionObservation. A failed sample warns and leaves the process up. Model plan drops a candidate whose fresh peak exceeds allocatable RAM.
Encode every remote nvidia-smi GPU row before base64. Match an untagged Ollama name only to the :latest tag. Refuse an Ollama profile with refusals before the load POST. Keep ModelRunProfile YAML keys in snake_case. Print the placed Ollama model on a successful text receipt. Refresh the daemon cache after a successful Ollama unload. Deep-copy GPUInfo.Index when cloning a snapshot. Gate n-gpu-layers on the pinned GPU's measured free VRAM. Name Ollama and MLX in the model help and two stale sentences. Wait briefly for SIGKILL before the MLX stop script asserts death.
|
Hold. |
axis model start <model> reads the snapshot and names one complete node whose runtime already has that model. A resident model is reported and left running. An Ollama server that is listening and lists the model is loaded on loopback, and that load is accepted only when the server returns done_reason load. Ollama SSH is refused unless the node is complete and the server is listening.
Hardware Validation red on
|
## Keep fixture stubs ahead of the host PATH (Hardware Validation is red on main) Three test fixtures added in #489 replace their child-process `PATH` with `stubdir:/usr/bin:/bin`. The scripts under test call real tools (`awk` for the nvidia-smi bundle parse and the classifier/observation checks), and on profile-based Linux runners those tools live only in the system profile — not under `/usr/bin`/`/bin`. Result: the last `Hardware Validation` pushes are red with `sh: line 1: awk: command not found` while the same tests pass on conventional hosts and locally (run/job refs: 37250216353, 37250828528, jobs 111576154161, 111577904135). Full root-cause comment is on the source PR (#489). ### Change - `cmd/axis/model_mlx_test.go` and `cmd/axis/model_observation_test.go` keep the stub directory **first** and append the inherited PATH after it, built with a single `sep := string(os.PathListSeparator)`; `HOME=` handling is unchanged. Fixture stubs keep shadowing real binaries. - `internal/facts/remote_bundle_gpu_test.go` replaces the manual `append(os.Environ(), "PATH=…")` with the package's existing `withSandboxedPATH(dir)` helper, which already orders stubs first and appends profile fallbacks. No new PATH implementation is introduced. ### Validation run - `go vet ./cmd/axis ./internal/facts` clean. - Focused suites green: `go test ./cmd/axis -run 'TestShellStopMLX|TestLlamaServerSample'` and `go test ./internal/facts -run 'TestRemoteBundleKeeps|TestResidentModels'`. - A batched `TestModelEvict*` run (`TestModelEvictDefaultsToLiveSnapshot`, `TestModelEvictLiveFalseUsesDaemonCache`, `TestModelEvictByGPUIndex`, `TestModelResumeByReceiptID`, `TestModelResumeUnitNameUsesSnapshotSupervisorAndPort`) first hit event-flush 5-second timeouts during a heavy-load window (workstation load average 116–160). A later batch re-run of the same five tests passed on this tree, and the same batched set passed on a pristine main-tip tree. No controlled load A/B was performed, so the timeouts are load-correlated with causality unproven; no test code was changed. - No workflow file is touched: the workflow's `Install runner dependencies` step already places the profile directory on PATH, and the hermetic-test harness leaves PATH to the caller. - This branch needs one `Hardware Validation` run on itself to confirm the self-hosted lane goes green with the fixture fix in place; local focused runs alone do not satisfy that. ### Notes - Test-only change; no product code touched (3 files, +14/−3). - Merging this before the next release tag keeps the hardware lane quiet. --------- Co-authored-by: cranium-agent <agent@local>


No description provided.