Skip to content

feat: per-task LLM model/provider override without restart (fixes #129) - #155

Open
Uzzoper wants to merge 14 commits into
google:mainfrom
Uzzoper:feat/per-task-llm-override
Open

Uzzoper wants to merge 14 commits into
google:mainfrom
Uzzoper:feat/per-task-llm-override

Conversation

@Uzzoper

@Uzzoper Uzzoper commented Sep 26, 2026 •

Copy link
Copy Markdown

Closes #129 with a small initial scope: a single optional override parameter, no settings UI, no global config mutation, no restart.

How it works

llm_model / llm_provider flow POST /api/run → enqueue_tasks (one shared normalize_llm_override() helper: trim/lowercase, blank → None, provider-requires-model fail-fast) → persisted queue item → worker --model/--provider flags → TaskRequestBuilder.with_llm_override() → per-task ArtemisContext → _resolve_endpoint precedence: task override > artemis.jsonc node > built-in default. get_llm() is the single funnel, so every model the task resolves is covered; the shared LLMConfig is never mutated.

Entry points

  • Console Prompt Dock: preset/provider dropdown fed by new GET /api/llm-options (providers, jsonc presets, configured default; degrades to free-text on 404) + custom free-text fields. Empty = server default; strictly per-task (never restored, cleared on successful submit).
  • Display: queue cards (override or scheduled default), history cards (model_info fallback), active dashboard, session detail — badges distinguish scheduled vs ran-with. Additive only, no restyle.
  • SDK: ArtemisClient.submit/run per-call params; daemon + batch shims (artemis run/batch --model/--provider).
  • MCP: mobile_run_task(llm_model, llm_provider) — named distinctly because existing model="Flash" means profile arch; background runner uses --llm-model/--llm-provider for the same reason.
  • Result: session row stores the keys in schemaless device_info (no migration); GET /api/sessions/{id} echoes them so TaskResult.llm_model/llm_provider resolves in production.

Also included (separate commit)

Fixes the 11 pre-existing ruff UP017/UP037/UP045 errors in playground/backend_manager keeping python-quality red on main — same mechanical fix as stalled #30/#67, needed for CI to go green here.

Validation (local, Python 3.12)

Deliberate scope decisions (for reviewers)

  • CLI/shim/validation/fallback-pinning kept: without them the SDK/MCP params would be dead or fail mid-task; fallback pinning is documented, not incidental.
  • Provider-only is rejected (fail-fast 422/ValueError): model IDs are provider-namespaced, so provider-alone can never resolve.
  • No LlmOverride cross-signature type yet; kept the pair convention (verification_level/explorer_mode travel the same way).

Follow-ups (out of scope)

  • Client-side provider hint on provider-without-model; shared TS card component; --llm-model alias on artemis run.

@google-cla

google-cla Bot commented Sep 26, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

…d_manager

Mechanical ruff check --fix plus format for the 11 errors keeping python-quality red on main (same as stalled google#30/google#67). Split out from the google#129 feature for scope discipline.
Implements google#129: optional llm_model/llm_provider override flowing POST /api/run -> task queue -> worker CLI -> per-task ArtemisContext -> _resolve_endpoint (precedence: task override > artemis.jsonc node > default). No global config mutation, no restart.
- Wire batch lane (--model/--provider through submit_batch_to_daemon
  and standalone run_batch_tasks, per-goal requests)
- Fail fast on provider-without-model (builder ValueError, API 422,
  MCP pre-trace) instead of dying mid-task
- Client: per-call override only (drop constructor defaults),
  provider lower-cased
- Document fallback pinning; comment idempotent-retry setdefault
- Tests: real-LLMConfig isolation, provider-only/invalid/Flash
  cases, run_task forwarding, batch override
Write llm_model/llm_provider into the schemaless device_info dict at
session creation (worker side, run_tuning precedent; no migration),
prefer them in model_service over the inferred/global model, and echo
them from GET /api/sessions/{id} so TaskResult.llm_model resolves in
production. Legacy rows fall back to current behavior.
Do not restore a previous override from localStorage: every task
starts from the server default unless set for that task. Stale keys
from earlier versions are dropped on load.
get_status() consults the session row override even when the worker
set an active profile (single row fetch, no new queries); the session
list prefers a stored override when no profile resolves. Rename
profile_sess_row to sess_row_for_model.
Queue cards display the requested override (llm_model/llm_provider
now flow through the status mapping); the session merge prefers the
stored model_info over the global active model; a successful submit
clears the override fields so the next task starts from defaults.
The second get_status() return path rebuilt model_info from the
connection profile alone; it now reuses the session row (or one
indexed lookup) and prefers a stored llm_model/llm_provider, with
trace lookups still skipped on this path.
History loops get the same conditional model badge as queue cards;
getTaskModelLabel falls back to the persisted session model_info when
no explicit override keys are present.
Pending tasks without an override have no session row yet, so the
badge had no data. getTaskModelLabel now chains explicit override
-> session model_info -> global activeModel, showing the model the
task will resolve at dispatch.
Single normalize_llm_override() helper (artemis/config/llm_override.py)
replaces six copies; TaskResult echoes llm_provider; new GET
/api/llm-options serves providers, presets and the configured default.
One shared getTaskModelLabel util; badges distinguish scheduled vs
ran-with via title/aria; Prompt Dock gains a preset/provider dropdown
backed by /api/llm-options with free-text fallback.
Rebuild the presets inventory test as a shape-only smoke test; fix the helper docstring coverage claim; stringify non-string override input in the client mirror.
@Uzzoper
Uzzoper force-pushed the feat/per-task-llm-override branch from 824560f to 58d8624 Compare September 26, 2026 17:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Allow selecting the agent LLM (Claude Haiku/Sonnet, GPT-4o, etc.) per task from the Web Console, without a server restart

1 participant