Repository navigation
Conversation
The Flash step summarizer (and, latently, the chunk-capsule lens) forced
every explicitly configured model through get_google_llm(), which
hardcodes provider=GOOGLE. A non-Gemini model ID such as
"tensorx/deepseek/deepseek-v4-flash" was therefore sent verbatim to
Google's Generative Language API, producing a terminal, non-retryable
404 per step:
VisualStepSummarizer: Error generating summary for step 1:
Error calling model 'tensorx/deepseek/deepseek-v4-flash'
(Not Found): 404 Not Found. {'message': '', 'status': 'Not Found'}
The error surfaced only at call time because construction errors were
silently swallowed (the old try/except re-fell back to
get_google_llm()), and 404 classifies as BAD_REQUEST (non-retryable, no
fallback) in artemis.llm.reliability. Other nodes resolved the same
model ID correctly: they route through _resolve_endpoint(), which maps
the configured provider — "custom" nodes ride the OpenAI-compatible
gateway (OPENAI_BASE_URL) where namespaced IDs are valid.
Add get_lens_llm(ctx, model_name, provider) to artemis.services.llm: a
provider-aware, raw-model factory for step-memory lens calls. It returns
the raw chat model (deliberately NOT wrapped in RobustChatModelWrapper)
so the lenses keep their own bounded retry loops, 25s asyncio timeouts,
and explicit "lens:*" usage metering. Provider resolution order:
1. Explicit "provider" knob — validated via ModelProvider.from_string,
so typos fail fast with the standard "Unknown LLM provider" error.
2. Gemini model names (bare, or google/gemini-prefixed) route to
GOOGLE, keeping every pre-existing config byte-for-byte compatible.
3. Anything else inherits the LLM config's summarizer node provider
(which itself inherits the global "default" block). Namespaced model
IDs pass through verbatim — never parsed as provider prefixes.
4. Resolution failures default to CUSTOM, matching how utils nodes
behave when no config is live.
Also emit a loud warning when a non-Gemini model resolves to the Google
Gemini API — the historical misroute produced no signal anywhere in the
stack until the first 404.
Changes:
- artemis/services/llm.py: add get_lens_llm, _inherit_lens_provider,
and _warn_if_non_google_routes_to_google.
- artemis/config/agent.py: add optional "provider" field (validated at
load) to StepSummarizerConfig and MemoryChunkingConfig; expand model
field descriptions with the routing rules. Old configs parse
unchanged (fields default to None = auto-resolve).
- artemis/agents/flash/summarizer.py: replace the get_google_llm /
dead get_llm(is_utils=True) branches with a single get_lens_llm
call; surface construction errors instead of silently swapping
providers.
- artemis/agents/flash/runner.py and artemis/memory/__init__.py:
forward step_summarizer_cfg.provider (covers the Pro profile's
shared step-memory service too).
- artemis/memory/chunking.py: same fix for StepCapsuleLens — accept
model_provider/fallback_model_provider and route through
get_lens_llm. The capsule fallback now resolves the summarizer
fallback block's OWN provider (like every other node's fallback via
_resolve_endpoint(use_fallback=True)) instead of being restricted to
google-only or inheriting the primary's provider; a Gemini model
under a non-Google knob still auto-routes to GOOGLE.
Tests:
- New tests/unit/services/test_llm_lens_routing.py: the full routing
matrix (Gemini shortcut, google/gemini prefixes, explicit knob
precedence, unknown-provider failure, summarizer-node inheritance,
default-config fallback, CUSTOM-on-failure, non-Gemini→GOOGLE
warning) plus an end-to-end assertion that the previously-404ing
TensorX model builds a ChatOpenAI bound to the configured gateway
base URL.
- test_history_chunking.py: fallback resolution now returns
(model, provider) and is provider-aware; new precedence tests for
the fallback block's own provider vs. the chunking "provider" knob.
- test_memory_config.py: the new provider knobs parse from config and
reject unknown providers.
- test_flash_runner.py: FlashRunner forwards step_summarizer provider.
Note: config/artemis.jsonc is intentionally left out of this commit
(local model overrides are still in flux there); enable step_summarizer
with the TensorX model and no explicit provider to pick the fix up.
Two pre-existing config-coupled tests (test_config_package.py,
test_model_service.py) assert provider=="google" against the real
config file and fail while the local default block says "custom".
Author
|
Disclaimer: LLMs were used to create this merge request, but I did review it manually, tested and verified the changes locally. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Flash step summarizer (and, latently, the chunk-capsule lens) forced every explicitly configured model through get_google_llm(), which hardcodes provider=GOOGLE. A non-Gemini model ID such as "tensorx/deepseek/deepseek-v4-flash" was therefore sent verbatim to Google's Generative Language API, producing a terminal, non-retryable 404 per step:
The error surfaced only at call time because construction errors were silently swallowed (the old try/except re-fell back to get_google_llm()), and 404 classifies as BAD_REQUEST (non-retryable, no fallback) in artemis.llm.reliability. Other nodes resolved the same model ID correctly: they route through _resolve_endpoint(), which maps the configured provider — "custom" nodes ride the OpenAI-compatible gateway (OPENAI_BASE_URL) where namespaced IDs are valid.
Add get_lens_llm(ctx, model_name, provider) to artemis.services.llm: a provider-aware, raw-model factory for step-memory lens calls. It returns the raw chat model (deliberately NOT wrapped in RobustChatModelWrapper) so the lenses keep their own bounded retry loops, 25s asyncio timeouts, and explicit "lens:*" usage metering. Provider resolution order:
Also emit a loud warning when a non-Gemini model resolves to the Google Gemini API — the historical misroute produced no signal anywhere in the stack until the first 404.
Changes:
Tests:
Closes #167