琥珀式封存历史回放评测(AMBER):封进琥珀,重做当时的题 —— 真实事件回放、物理封存、预注册评分的方法规范
-
Updated
Sep 8, 2026
琥珀式封存历史回放评测(AMBER):封进琥珀,重做当时的题 —— 真实事件回放、物理封存、预注册评分的方法规范
Local Windows AI prompt enhancer (Tauri 2 + Rust). 100% offline: prompts, history, API key stay on disk (SQLCipher). Quantitative effect numbers are not published (v0.3.2 evaluation: 60 samples x 4 targets x 3 repeats + 48-unit blind human review found no quality advantage for enhancement; automated eval biases toward longer/structured outputs).
An evaluation methodology for context-augmentation experiments: a format-matched control that separates whether added context helps because of its information or its format. Pre-registered instrument, mutation-audited scorer, worked example in SVG reference resolution.
Diagnostic of physics literacy in frontier LLMs. Evaluates induction, formulation, and prediction inside unfamiliar physics frameworks via dual-judge inter-rater reliability with human-audit resolution.
MCDM using TOPSIS: A technique for multi-criteria decision making, utilizing the TOPSIS method to evaluate alternatives based on multiple criteria.
The instrument makes the winner: reconciling two contradictory 2026 clinical-AI evaluations (Real-POCQi vs Nature Medicine). A 2x2 instrument x rater decomposition shows the evaluation instrument (pairwise vs rubric), not the model, drives the disagreement.
Grounds simulated-user personas in sourced real evidence (HN/ProductHunt quotes), then has them live-evaluate the actual product — no invented personas, no guessed pain points.
Do models catch flawed ML results? Flawed reports with byte-identical matched controls.
Runtime-verifiable audit harness that falsifies coding-agent intervention claims — preregistered, oracle-checked, reports null.
JAMO — OCR-mediated Hangul rendering benchmark. Code companion to huggingface.co/datasets/Nasser4963/jamo-gold
Code, data, and paper for an evaluation-protocol-adjusted residual audit of public LLM benchmark scores — predicting score variation from public metadata and using the residuals as a review queue.
Reproducibility artifact for 'Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection'. Drift alarms propose challengers; whether promotion helps, and which update policy looks best, depends on how the challenger was constructed and evidenced. Sealed results, preregistered protocols, claim audit.
An audit of how wearable stress-detection performance is measured. Three public datasets, subject-independent evaluation, and what a personal baseline actually buys.
Measurement protocol for demographic bias in vision systems: frozen pools, paired designs, mandatory intervals, pre-registered criteria.
Paper artifacts for TASA, a theoretical framework for persistent, inspectable agent state.
An evidence-driven evaluation engine for AI systems in regulated financial workflows
This repository contains the data, model outputs, and analysis materials for EVAL4SD paper @ ACL: "Gender Associations in LLM-Mediated ADHD Self-Diagnosis" (Hansen & Kristensen-McLachlan, 2026).
To associate your repository with the evaluation-methodology topic, visit your repo's landing page and select "manage topics."