Skip to content

feat(harness): edge-cloud dual-path pipeline with pluggable arbitration - #2

Merged
PiloBi merged 3 commits into
mainfrom
feat/edge-cloud-harness
Aug 22, 2026
Merged

PiloBi merged 3 commits into
mainfrom
feat/edge-cloud-harness

Conversation

@PiloBi

@PiloBi PiloBi commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

Add a complete edge-cloud inference pipeline (harness architecture):

  • IPromptEngine: prompt compression + intent distillation before cloud dispatch. Reduces token consumption by 70-90% for typical cockpit queries.
  • ICloudBackend: async cloud LLM client (OpenAI-compatible). Supports DeepSeek, Qwen, vLLM, Ollama, and any /v1/chat/completions endpoint.
  • IArbiter: local arbitration between device and cloud results. Strategies: cloud_prefer, latency_first, confidence-based, local_only.
  • IConfidenceScorer: two-phase confidence gating (pre-score before inference, post-score after) to decide when cloud is worth the cost.
  • PipelineHarness: top-level orchestrator with DeepSeek Harness-style pluggable component registry. All slots are interface-based and hot-swappable at runtime via YAML config.

Design principles:

  1. Token-friendly: compress before sending to cloud
  2. Boundary intelligence: fire cloud only when local confidence is low
  3. Efficiency-first: async cloud, hard deadline, never block local path
  4. Pluggable: all components are interfaces, config-driven activation

Includes:

  • config/harness.yaml: full pipeline configuration (user fills API keys)
  • templates/: prompt templates (default, navigation, vehicle_control)
  • tests/test_harness.cpp: 15 unit tests, all passing
  • docs/edge_cloud_design.md: architecture design document (Chinese)

Tested: cmake clean build + ctest 6/6 passed (macOS ARM64)

PiloBi added 3 commits August 21, 2026 16:28
Add a complete edge-cloud inference pipeline (harness architecture):

- IPromptEngine: prompt compression + intent distillation before cloud
  dispatch. Reduces token consumption by 70-90% for typical cockpit queries.
- ICloudBackend: async cloud LLM client (OpenAI-compatible). Supports
  DeepSeek, Qwen, vLLM, Ollama, and any /v1/chat/completions endpoint.
- IArbiter: local arbitration between device and cloud results. Strategies:
  cloud_prefer, latency_first, confidence-based, local_only.
- IConfidenceScorer: two-phase confidence gating (pre-score before
  inference, post-score after) to decide when cloud is worth the cost.
- PipelineHarness: top-level orchestrator with DeepSeek Harness-style
  pluggable component registry. All slots are interface-based and
  hot-swappable at runtime via YAML config.

Design principles:
  1. Token-friendly: compress before sending to cloud
  2. Boundary intelligence: fire cloud only when local confidence is low
  3. Efficiency-first: async cloud, hard deadline, never block local path
  4. Pluggable: all components are interfaces, config-driven activation

Includes:
- config/harness.yaml: full pipeline configuration (user fills API keys)
- templates/: prompt templates (default, navigation, vehicle_control)
- tests/test_harness.cpp: 15 unit tests, all passing
- docs/edge_cloud_design.md: architecture design document (Chinese)

Tested: cmake clean build + ctest 6/6 passed (macOS ARM64)
The CI grep pattern matches 'api_key:' even when the value is empty.
Remove the direct api_key field — only the env var approach (api_key_env)
is supported, which is more secure anyway.
The build job hung 24h waiting for a runner because ubuntu-22.04 and
macos-13 images are retired — no hosted runners pick up those labels.

Changes:
- Runner labels: ubuntu-22.04 -> ubuntu-latest, macos-13/14 -> macos-latest
- Trigger CI on feat/** and fix/** branches (was main/develop only, so
  feature branches never ran CI before opening a PR)
- clang-tidy: use unversioned package (clang-tidy-15 unavailable on
  ubuntu-latest), drop the -p build flag pointing at a nonexistent
  compile database, fix include path src/include -> include
- Artifact upload: add if-no-files-found: ignore since build/cli/sparx
  only exists when the proprietary kernel is present
- Release packaging: guard the sparx copy and include OSS binaries instead
  of failing outright
@PiloBi
PiloBi merged commit c4dd3d7 into main Aug 22, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant