Skip to content

Roadmap: learn-once/compile-once execution model to reduce AI-in-the-loop cost for large regression suites #26

Description

@qtpsudhakar

I'm evaluating ARTEMIS for large-scale test automation and have a question about the intended long-term architecture.

ARTEMIS can currently execute end-to-end workflows using AI/agentic interaction. However, for large regression suites, having an AI model reason at every step can introduce latency, cost, and nondeterminism.

Is there (or is there planned) a workflow along these lines?

Learn once → generate/compile a reusable automation → execute cheaply/deterministically → invoke AI only when the learned interaction no longer matches the UI

For example:

  1. ARTEMIS explores a workflow such as: Login → Search → Product → Add to Cart → Checkout
  2. During exploration, it learns the semantic/UI targets and state transitions.
  3. The learned workflow is persisted as a reusable test artifact.
  4. Subsequent executions use the learned targets/path without invoking the LLM/VLM for every action.
  5. If a target/state cannot be resolved because the application UI has changed, ARTEMIS invokes the AI/VLM to understand the new UI and repair/update the workflow.
  6. The repaired workflow is validated and then reused for subsequent executions.

In other words, I'm interested in whether ARTEMIS is intended to become a learned/adaptive automation engine, rather than having the AI remain in the execution loop for every test run.

A few specific questions:

  • Does ARTEMIS currently persist knowledge learned during a successful workflow for later deterministic/reduced-AI execution?
  • Is there a concept of compiling/serializing an agentic workflow into a fast, reusable test?
  • Can AI/VLM invocation be triggered only when deterministic/local grounding fails?
  • Is self-healing/re-learning of previously discovered workflows part of the roadmap?
  • Is the planned lightweight Edge VLM intended partly to support this low-latency execution/recovery model?
  • If this isn't currently supported, is this an architectural direction the team is considering?

I'm particularly interested in how ARTEMIS is expected to scale to thousands of regression tests where AI-per-step latency would become significant.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions