Skip to content

Migrate specsmith → spec-kit + AEE; core on AEE library; GDELT Web Ngrams default - #54

Merged
tbitcs merged 6 commits into
mainfrom
feat/spec-kit-aee-migration
Oct 5, 2026
Merged

tbitcs merged 6 commits into
mainfrom
feat/spec-kit-aee-migration

Conversation

@tbitcs

@tbitcs tbitcs commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Summary

Three-stage migration, executed per the repo's governance (read-first:
AGENTS.md, docs/governance/rules.md, LIFECYCLE.md) and ledgered
per stage in LEDGER.md (append-only, AI-disclosed).

Stage 1 — specsmith → spec-kit + AEE (2e603341)

  • specify init (v1.0.10, copilot integration) + aee & evaluator
    extensions installed from local clones as real files (verified: no
    gitlinks under .specify/).
  • .specify/memory/constitution.md preserves Glossa's real governance:
    citation/provenance, append-only LEDGER, foundation-check gate,
    public/private correspondence boundary, AI disclosure, falsifiable
    claims, spec-driven flow.
  • specs/001-glossa-lab-baseline/ documents the as-built platform
    (FastAPI + React, discovery engine, Evidence Graph, Indus status per
    README: 161 H+M readings, 90.96% coverage, preprint v4 DOI
    10.5281/zenodo.20414696).
  • specsmith retired as active tooling (skills, scaffold.yml, tracked
    rate-limit state removed); runtime persistence .specsmith/ →
    .glossa-state/ in fetchers/base.py + model_intelligence.py.
    Historical mentions in LEDGER/CHANGELOG/ledger-archive untouched.

Stage 2 — core wired to the AEE library (a915eca5)

  • Dependency: applied-epistemic-engineering>=1.0.4,<2.
  • backend/glossa_lab/aee_core.py adapts extracted-claims JSON onto
    AEE Claim/Evidence/ClaimGraph + ScoringEngine.
    falsification_condition maps natively to
    Claim.falsification_tests; untested → AEE DRAFT (original kept
    in Claim.metadata); status enters scoring through attached evidence
    strength (the engine takes no status input — documented in the
    docstring).
  • API, additive only: GET /api/v1/indus-evidence/claims/aee-scores
    and ?aee=true on GET /claims. Verified live: 31 claims, mean
    propagated score 0.485; default response shapes unchanged.

Stage 3 — GDELT guidance implemented (48c1ad8f)

  • New default GDELT source gdelt_ngrams (Web Ngrams dataset):
    ~5-min-ago request, bounded 15-min-heartbeat walk-back (24 marks),
    quadgram matching with tri/bi/unigram reduction (>4-word keywords via
    constituent windows), topic exclusions, TOC cross-reference,
    .glossa-state/ watermark.
  • DOC-API fetcher kept but opt-in only, docstring-noted as paused
    per GDELT's request during the Spanner migration.
  • specs/002-gdelt-ngrams-and-frontier-methods/ also records (not
    implemented): a future fully-cited daily briefing over the discovery
    corpus, and manuscript-method design notes for future seal/tablet
    vision work (physical metadata up front; discrete focused passes).

Sync note

Branch includes a clean merge of origin/main after Dependabot
#50–#53 landed (workflow actions v7, Pillow bump) — no conflicts; AEE
dependency and Dependabot changes coexist.

Test results

  • Full backend suite: 533 passed, 9 skipped, 0 failed
    (test_indus_evidence_api.py 25 passed run separately; the rest 508
    passed / 9 skipped).
  • New tests: test_aee_core.py 18 passed, test_gdelt_ngrams.py 11
    passed (synthetic fixtures, no network). Ruff clean on changed files.
  • backend/scripts/foundation_check.py cannot run off the Windows dev
    box (hardcoded C:\Users\trist\... path — pre-existing, untouched);
    noted in the stage-4 ledger entry.

Live smoke test (GDELT ngrams)

  • Endpoint/naming verified against the dataset's documented example
    pair (2026-06-30 20:16 UTC): 1,063,647 quadgrams, 1,977 TOC entries;
    end-to-end matching on real data works ("disease" → 16 items;
    "Indus script" → 0 in that single minute).
  • No current files exist: the 25-mark walk-back for today and spot
    marks over prior weeks all return 404 — the dataset appears not to be
    publishing at test time. The fetcher handles this correctly (bounded
    walk, no items, watermark untouched) and picks files up when
    publication resumes.

🤖 Executed by an AI agent (Muse Spark, via Muse) at the direction of
Tristen Pierson, per constitution §VI. Not merged — awaiting review.

tbitcs added 6 commits October 5, 2026 16:01
- specify init 1.0.10 (copilot/sh); aee + evaluator extensions vendored as real files
- constitution ratified at .specify/memory/constitution.md from existing governance
- specs/001-glossa-lab-baseline: as-built baseline spec/plan/tasks
- remove specsmith skills, scaffold.yml, tracked backend/.specsmith state
- runtime rate-limit state: .specsmith/ -> .glossa-state/ (base.py, model_intelligence.py), gitignored
- AGENTS.md + LIFECYCLE.md updated to spec-kit flow; LEDGER entry appended (history untouched)
@tbitcs
tbitcs merged commit 56a16a3 into main Oct 5, 2026
7 checks passed
@tbitcs
tbitcs deleted the feat/spec-kit-aee-migration branch October 5, 2026 19:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant