Skip to content

Flash profile fails with OpenAI-compatible (non-Gemini) providers: step summarizer & memory chunker hard-code the Gemini API #138

Description

@yzf1997-china

Summary

When ARTEMIS is configured to use an OpenAI-compatible provider (e.g. a custom endpoint configured via OPENAI_API_KEY + OPENAI_BASE_URL, such as DeepSeek / Volcano Ark / vLLM / Ollama) instead of Google Gemini, running a Flash-profile task fails immediately with a ChatGoogleGenerativeAI validation error — even though the OpenAI configuration itself is confirmed correct.

Environment

  • OS: Windows
  • Device: Android phone connected via ADB (UIAutomator2 hierarchy backend)
  • Provider: custom OpenAI-compatible endpoint set in .env via OPENAI_API_KEY and OPENAI_BASE_URL; provider: "openai" + model: "<redacted-model>" in config/artemis.jsonc
  • Web console GET /api/system/model-config-env confirms the OpenAI key/base URL are set and the config file is loaded correctly

Steps to reproduce

  1. Configure .env:
    • OPENAI_API_KEY=<redacted>
    • OPENAI_BASE_URL=https://<endpoint>/<version>
  2. In config/artemis.jsonc, set the default LLM to "provider": "openai" with a model served by that endpoint.
  3. Start a Flash-profile task (web console or MCP mobile_run_task).
  4. The run fails immediately, before the first step executes.

Actual behavior

The trace stdout.log reports:

Error running automation: 1 validation error for ChatGoogleGenerativeAI
Value error, API key required for Gemini Developer API. Provide api_key parameter or set GOOGLE_API_KEY/GEMINI_API_KEY environment variable.
[type=value_error, input_value={'model': '<configured-model>', 'google_api_key': ''}, input_type=dict]

Expected behavior

The Flash pipeline should honor the configured provider — OpenAI-compatible endpoints are already supported by the router/ModelFactory (see artemis/llm/router.py, ModelProvider.OPENAI). If Flash genuinely requires Gemini, it should fail with a clear, actionable message instead of constructing a Gemini client with an empty API key.

Root cause

Two Flash-pipeline components hard-code the Gemini provider through get_google_llm(...) while passing the configured model name, bypassing the configured provider entirely:

  • artemis/agents/flash/summarizer.py — VisualStepSummarizer is constructed eagerly at runner init whenever agent.flash.step_summarizer.enabled is true (default), so the run dies before the first step.
  • artemis/memory/chunking.py — HistoryChunkManager._get_llm / _get_fallback_llm use get_google_llm.

Additionally, artemis/agents/flash/runner.py::_init_llm falls back to get_google_llm("gemini-2.5-flash") whenever config-based model resolution raises.

Proposed fix (verified locally)

Build these models through the provider-aware ModelFactory + ModelEndpoint, taking the provider from the LLM config (the summarizer node) instead of calling get_google_llm. After patching the two files, the same Flash task completed successfully (status: completed), and LLM usage was recorded as openai:<model>: N calls against the custom endpoint.

Additional notes

  • The Pro-profile judge nodes planner_validation / validator_pixel_safety_net still default to a Google lightweight model (lightweight_judge_default()); the default+nodes config format does not allow overriding them. Same class of issue in Pro.
  • With the non-Gemini model, the Flash step summarizer showed one intermittent per-step failure (auto-retried, then cancelled after the 30s flush timeout). Non-blocking for task completion, but worth addressing if background summarization must be robust across providers.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions