Summary
Running ARTEMIS against local / edge OpenAI-compatible endpoints (Ollama, llama.cpp-server, vLLM, LM Studio) per agent node is currently not configurable, and the shipped local-ollama preset is broken at startup. This blocks the roadmap item On-Device Lightweight VLMs (README §Roadmap).
Details
ModelFactory already routes ollama / vllm / custom providers to ChatOpenAI (artemis/llm/router.py, ~L369–386), but the config layer cannot reach it:
- Schema gap —
LLM (artemis/config/llm.py L71–105) has no api_base, api_key, max_tokens, timeout_seconds, or is_multimodal fields, even though ModelEndpoint (router.py L87–133) already defines them.
- Passthrough gap —
_resolve_endpoint() (artemis/services/llm.py ~L1043–1089) builds ModelEndpoint from the per-node config but silently drops api_base / api_key / max_tokens. The only way to reach a local server today is the global OPENAI_BASE_URL env var, which forces every agent node onto the same endpoint. A hybrid setup (local VLM for the flash explorer tier, cloud model for the planner) is impossible.
- Broken preset — the
local-ollama preset in config/artemis.jsonc (L44–48) uses provider: "openai" with no api_base, and validate_provider() raises without OPENAI_API_KEY, so the one preset advertising local usage fails at startup.
- Latent parser bug —
strip_json_comments() (artemis/utils/file.py) uses a regex that also strips // inside quoted strings, so any artemis.jsonc containing a URL value like "api_base": "http://localhost:11434/v1" gets corrupted at load time. This blocks any config-level fix for the above.
Relation to existing work
PR #128 covers admin-UI/env credential plumbing (CUSTOM_LLM_API_KEY etc.) and validate_provider branches for custom providers — complementary to this. It does not touch _resolve_endpoint() per-node passthrough, the preset, or the JSONC parser bug.
Proposed fix
- Add the five optional endpoint fields to the
LLM schema and forward them in _resolve_endpoint() (primary and fallback nodes).
- Let
validate_provider() skip cloud-key requirements for ollama / vllm / custom.
- Repair the
local-ollama preset and document local-endpoint variables in .env.example.
- Rewrite
strip_json_comments() as a string-aware scanner.
- Add provider-surface unit tests (
tests/unit/test_llm_router.py).
I have this implemented and will open a PR shortly.
Summary
Running ARTEMIS against local / edge OpenAI-compatible endpoints (Ollama, llama.cpp-server, vLLM, LM Studio) per agent node is currently not configurable, and the shipped
local-ollamapreset is broken at startup. This blocks the roadmap item On-Device Lightweight VLMs (README §Roadmap).Details
ModelFactoryalready routesollama/vllm/customproviders toChatOpenAI(artemis/llm/router.py, ~L369–386), but the config layer cannot reach it:LLM(artemis/config/llm.pyL71–105) has noapi_base,api_key,max_tokens,timeout_seconds, oris_multimodalfields, even thoughModelEndpoint(router.pyL87–133) already defines them._resolve_endpoint()(artemis/services/llm.py~L1043–1089) buildsModelEndpointfrom the per-node config but silently dropsapi_base/api_key/max_tokens. The only way to reach a local server today is the globalOPENAI_BASE_URLenv var, which forces every agent node onto the same endpoint. A hybrid setup (local VLM for the flash explorer tier, cloud model for the planner) is impossible.local-ollamapreset inconfig/artemis.jsonc(L44–48) usesprovider: "openai"with noapi_base, andvalidate_provider()raises withoutOPENAI_API_KEY, so the one preset advertising local usage fails at startup.strip_json_comments()(artemis/utils/file.py) uses a regex that also strips//inside quoted strings, so anyartemis.jsonccontaining a URL value like"api_base": "http://localhost:11434/v1"gets corrupted at load time. This blocks any config-level fix for the above.Relation to existing work
PR #128 covers admin-UI/env credential plumbing (
CUSTOM_LLM_API_KEYetc.) andvalidate_providerbranches for custom providers — complementary to this. It does not touch_resolve_endpoint()per-node passthrough, the preset, or the JSONC parser bug.Proposed fix
LLMschema and forward them in_resolve_endpoint()(primary and fallback nodes).validate_provider()skip cloud-key requirements forollama/vllm/custom.local-ollamapreset and document local-endpoint variables in.env.example.strip_json_comments()as a string-aware scanner.tests/unit/test_llm_router.py).I have this implemented and will open a PR shortly.