Skip to content
#

evaluation-methodology

Here are 17 public repositories matching this topic...

Local Windows AI prompt enhancer (Tauri 2 + Rust). 100% offline: prompts, history, API key stay on disk (SQLCipher). Quantitative effect numbers are not published (v0.3.2 evaluation: 60 samples x 4 targets x 3 repeats + 48-unit blind human review found no quality advantage for enhancement; automated eval biases toward longer/structured outputs).

  • Updated Sep 6, 2026
  • Python

An evaluation methodology for context-augmentation experiments: a format-matched control that separates whether added context helps because of its information or its format. Pre-registered instrument, mutation-audited scorer, worked example in SVG reference resolution.

  • Updated Jul 31, 2026
  • Python
Validate-Before-Commit

Reproducibility artifact for 'Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection'. Drift alarms propose challengers; whether promotion helps, and which update policy looks best, depends on how the challenger was constructed and evidenced. Sealed results, preregistered protocols, claim audit.

  • Updated Sep 3, 2026
  • Python

Add this topic to your repo

To associate your repository with the evaluation-methodology topic, visit your repo's landing page and select "manage topics."

Learn more