Skip to content

fix: provider-aware model routing for step-memory lenses - #168

Open
kypeli wants to merge 1 commit into
google:mainfrom
kypeli:johan/hopper-non-google-model
Open

kypeli wants to merge 1 commit into
google:mainfrom
kypeli:johan/hopper-non-google-model

Conversation

@kypeli

@kypeli kypeli commented Oct 4, 2026 •

Copy link
Copy Markdown

The Flash step summarizer (and, latently, the chunk-capsule lens) forced every explicitly configured model through get_google_llm(), which hardcodes provider=GOOGLE. A non-Gemini model ID such as "tensorx/deepseek/deepseek-v4-flash" was therefore sent verbatim to Google's Generative Language API, producing a terminal, non-retryable 404 per step:

VisualStepSummarizer: Error generating summary for step 1:
Error calling model 'tensorx/deepseek/deepseek-v4-flash'
(Not Found): 404 Not Found. {'message': '', 'status': 'Not Found'}

The error surfaced only at call time because construction errors were silently swallowed (the old try/except re-fell back to get_google_llm()), and 404 classifies as BAD_REQUEST (non-retryable, no fallback) in artemis.llm.reliability. Other nodes resolved the same model ID correctly: they route through _resolve_endpoint(), which maps the configured provider — "custom" nodes ride the OpenAI-compatible gateway (OPENAI_BASE_URL) where namespaced IDs are valid.

Add get_lens_llm(ctx, model_name, provider) to artemis.services.llm: a provider-aware, raw-model factory for step-memory lens calls. It returns the raw chat model (deliberately NOT wrapped in RobustChatModelWrapper) so the lenses keep their own bounded retry loops, 25s asyncio timeouts, and explicit "lens:*" usage metering. Provider resolution order:

  1. Explicit "provider" knob — validated via ModelProvider.from_string, so typos fail fast with the standard "Unknown LLM provider" error.
  2. Gemini model names (bare, or google/gemini-prefixed) route to GOOGLE, keeping every pre-existing config byte-for-byte compatible.
  3. Anything else inherits the LLM config's summarizer node provider (which itself inherits the global "default" block). Namespaced model IDs pass through verbatim — never parsed as provider prefixes.
  4. Resolution failures default to CUSTOM, matching how utils nodes behave when no config is live.

Also emit a loud warning when a non-Gemini model resolves to the Google Gemini API — the historical misroute produced no signal anywhere in the stack until the first 404.

Changes:

  • artemis/services/llm.py: add get_lens_llm, _inherit_lens_provider, and _warn_if_non_google_routes_to_google.
  • artemis/config/agent.py: add optional "provider" field (validated at load) to StepSummarizerConfig and MemoryChunkingConfig; expand model field descriptions with the routing rules. Old configs parse unchanged (fields default to None = auto-resolve).
  • artemis/agents/flash/summarizer.py: replace the get_google_llm / dead get_llm(is_utils=True) branches with a single get_lens_llm call; surface construction errors instead of silently swapping providers.
  • artemis/agents/flash/runner.py and artemis/memory/init.py: forward step_summarizer_cfg.provider (covers the Pro profile's shared step-memory service too).
  • artemis/memory/chunking.py: same fix for StepCapsuleLens — accept model_provider/fallback_model_provider and route through get_lens_llm. The capsule fallback now resolves the summarizer fallback block's OWN provider (like every other node's fallback via _resolve_endpoint(use_fallback=True)) instead of being restricted to google-only or inheriting the primary's provider; a Gemini model under a non-Google knob still auto-routes to GOOGLE.

Tests:

  • New tests/unit/services/test_llm_lens_routing.py: the full routing matrix (Gemini shortcut, google/gemini prefixes, explicit knob precedence, unknown-provider failure, summarizer-node inheritance, default-config fallback, CUSTOM-on-failure, non-Gemini→GOOGLE warning) plus an end-to-end assertion that the previously-404ing TensorX model builds a ChatOpenAI bound to the configured gateway base URL.
  • test_history_chunking.py: fallback resolution now returns (model, provider) and is provider-aware; new precedence tests for the fallback block's own provider vs. the chunking "provider" knob.
  • test_memory_config.py: the new provider knobs parse from config and reject unknown providers.
  • test_flash_runner.py: FlashRunner forwards step_summarizer provider.

Closes #167

The Flash step summarizer (and, latently, the chunk-capsule lens) forced
every explicitly configured model through get_google_llm(), which
hardcodes provider=GOOGLE. A non-Gemini model ID such as
"tensorx/deepseek/deepseek-v4-flash" was therefore sent verbatim to
Google's Generative Language API, producing a terminal, non-retryable
404 per step:

    VisualStepSummarizer: Error generating summary for step 1:
    Error calling model 'tensorx/deepseek/deepseek-v4-flash'
    (Not Found): 404 Not Found. {'message': '', 'status': 'Not Found'}

The error surfaced only at call time because construction errors were
silently swallowed (the old try/except re-fell back to
get_google_llm()), and 404 classifies as BAD_REQUEST (non-retryable, no
fallback) in artemis.llm.reliability. Other nodes resolved the same
model ID correctly: they route through _resolve_endpoint(), which maps
the configured provider — "custom" nodes ride the OpenAI-compatible
gateway (OPENAI_BASE_URL) where namespaced IDs are valid.

Add get_lens_llm(ctx, model_name, provider) to artemis.services.llm: a
provider-aware, raw-model factory for step-memory lens calls. It returns
the raw chat model (deliberately NOT wrapped in RobustChatModelWrapper)
so the lenses keep their own bounded retry loops, 25s asyncio timeouts,
and explicit "lens:*" usage metering. Provider resolution order:

1. Explicit "provider" knob — validated via ModelProvider.from_string,
   so typos fail fast with the standard "Unknown LLM provider" error.
2. Gemini model names (bare, or google/gemini-prefixed) route to
   GOOGLE, keeping every pre-existing config byte-for-byte compatible.
3. Anything else inherits the LLM config's summarizer node provider
   (which itself inherits the global "default" block). Namespaced model
   IDs pass through verbatim — never parsed as provider prefixes.
4. Resolution failures default to CUSTOM, matching how utils nodes
   behave when no config is live.

Also emit a loud warning when a non-Gemini model resolves to the Google
Gemini API — the historical misroute produced no signal anywhere in the
stack until the first 404.

Changes:
- artemis/services/llm.py: add get_lens_llm, _inherit_lens_provider,
  and _warn_if_non_google_routes_to_google.
- artemis/config/agent.py: add optional "provider" field (validated at
  load) to StepSummarizerConfig and MemoryChunkingConfig; expand model
  field descriptions with the routing rules. Old configs parse
  unchanged (fields default to None = auto-resolve).
- artemis/agents/flash/summarizer.py: replace the get_google_llm /
  dead get_llm(is_utils=True) branches with a single get_lens_llm
  call; surface construction errors instead of silently swapping
  providers.
- artemis/agents/flash/runner.py and artemis/memory/__init__.py:
  forward step_summarizer_cfg.provider (covers the Pro profile's
  shared step-memory service too).
- artemis/memory/chunking.py: same fix for StepCapsuleLens — accept
  model_provider/fallback_model_provider and route through
  get_lens_llm. The capsule fallback now resolves the summarizer
  fallback block's OWN provider (like every other node's fallback via
  _resolve_endpoint(use_fallback=True)) instead of being restricted to
  google-only or inheriting the primary's provider; a Gemini model
  under a non-Google knob still auto-routes to GOOGLE.

Tests:
- New tests/unit/services/test_llm_lens_routing.py: the full routing
  matrix (Gemini shortcut, google/gemini prefixes, explicit knob
  precedence, unknown-provider failure, summarizer-node inheritance,
  default-config fallback, CUSTOM-on-failure, non-Gemini→GOOGLE
  warning) plus an end-to-end assertion that the previously-404ing
  TensorX model builds a ChatOpenAI bound to the configured gateway
  base URL.
- test_history_chunking.py: fallback resolution now returns
  (model, provider) and is provider-aware; new precedence tests for
  the fallback block's own provider vs. the chunking "provider" knob.
- test_memory_config.py: the new provider knobs parse from config and
  reject unknown providers.
- test_flash_runner.py: FlashRunner forwards step_summarizer provider.

Note: config/artemis.jsonc is intentionally left out of this commit
(local model overrides are still in flux there); enable step_summarizer
with the TensorX model and no explicit provider to pick the fix up.
Two pre-existing config-coupled tests (test_config_package.py,
test_model_service.py) assert provider=="google" against the real
config file and fail while the local default block says "custom".
@kypeli

kypeli commented Oct 4, 2026

Copy link
Copy Markdown
Author

Disclaimer: LLMs were used to create this merge request, but I did review it manually, tested and verified the changes locally.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

VisualStepSummarizer hardcodes Google provider, breaking non-Gemini models in flash.step_summarizer.model (404 Not Found)

1 participant