Skip to content

Question about run-to-run consistency in skill evaluation #31

Description

@blak0p

I build tooling with a similar underlying concern in a different layer: a Git MCP server where an LLM never decides semantics directly — a deterministic layer classifies changes first, so the model can’t introduce inconsistency into decisions that matter.

Your train/test split + 3-run approach for skill evaluation caught my eye because it’s tackling the same root problem (LLM non-determinism) from the evaluation side instead. When two runs of the same skill produce different results on the same input, how does your system decide whether that’s the skill being poorly designed vs. just expected model noise? Trying to compare notes on how you draw that line.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions