Repository navigation
[Product Gap] Add real-audio MIR accuracy acceptance benchmarks #770
Description
Activity
seonghobae commented
on Aug 9, 2026 CollaboratorAuthorMore actionsProgress update (2026-08-10): PR #828 now publishes exact head
8e34c0e563fafa08909307109b358cbbebf5537f(treea4e77ef4998ac7b30fea38aa02edb8a74b5c3489) for one bounded known-vocal slice of this issue.Delivered and reviewable:
- production YouTube intake → pinned creator-master identity → one composed global alignment → deterministic real htdemucs → zero-mean vocal SI-SDR improvement and semantic assignment margin;
- fail-closed full filename/size/SHA-256 model verification before restricted weights-only same-byte loading, with no runtime model retrieval;
- full ffmpeg+ffprobe path/hash/execute/sibling identity, bounded download/decode/cleanup, deterministic offline metric/alignment/integrity/failure tests, and an explicit authorized live opt-in;
- canonical PRD (11 requirements), TRD (13 requirements), three ADRs, root/discoverable Architecture, sequence/state/deployment/class UML, a logical run/evidence model, traceability, acceptance, operations, release, Security Notes, and [Product Gap] Add real-audio MIR accuracy acceptance benchmarks #770 doctoring;
- a shared cross-platform Python launcher for root checks and quickcheck rather than assuming either
python3orpythonexists.
Documentation fitness is now design-sufficient for this bounded known-vocal slice. A physical database ERD remains intentionally not applicable because fixtures, aligned windows, and separated stems are ephemeral and no benchmark database exists; the logical
BenchmarkRun/evidence relationships are documented instead.Exact-head publication evidence: all 9 hosted workflows completed successfully, CodeRabbit combined status is success, its exact-head command review reported no actionable issue, and all 29 review threads are resolved. Formal independent exact-head approval is still absent, so the PR is not merge-ready.
The only live production-path attempt remains historical negative evidence: archive/member/master/model integrity passed, then YouTube intake failed closed with HTTP 502 after 65.49 seconds before separation. It produced no identity correlation or SI-SDR score and is not a live pass. Creator-master-only calibration (+1.752 dB SI-SDRi, +7.631 dB assignment margin) is not YouTube validation.
This still does not close #770. Remaining work includes formal content/platform authorization; model rights/delivery and approved-pickle risk decisions; authorized exact-candidate YouTube calibration with drift ownership; accepted schema-v1 evidence persistence/access/TTL/deletion controls and an emitter; every advertised OS/architecture pass or verified fallback; recoverable failure UX; multi-fixture/four-stem coverage; the other MIR families; machine-readable/accessibility reports; and qualifying independent review.
- added a commit that references this issue
on Aug 17, 2026 - addedarea: apiAPI, protocol, event, or external contractAPI, protocol, event, or external contractarea: authAuthentication, authorization, identity, or tenant isolationAuthentication, authorization, identity, or tenant isolationarea: ci-cdCI, GitHub Actions, checks, release, or supply chainCI, GitHub Actions, checks, release, or supply chainarea: securitySecurity boundary, hardening, or vulnerability preventionSecurity boundary, hardening, or vulnerability preventionpriority: mediumNormal-priority or P2 workNormal-priority or P2 workscope: product-gapCustomer-visible product gapCustomer-visible product gapstatus: triagedOpen issue has an organization taxonomy assignmentOpen issue has an organization taxonomy assignmenttype: featureNew or expanded product capabilityNew or expanded product capability
on Aug 22, 2026 - added a commit that references this issue
on Aug 23, 2026 seonghobae commented
on Sep 13, 2026 CollaboratorAuthorMore actionsScientific acceptance hierarchy correction — 2026-09-13
Production/commercial MIR acceptance must not be satisfied by synthetic or procedurally generated audio, even when it is decoded through the real runtime. Generated WAV/FLAC fixtures remain useful for deterministic unit/regression coverage only. A release-grade scientific claim needs rights-cleared real recorded audio passed through the actual decode → MIR → rehearsal-insight boundary, with immutable content hash, source/license/authorization, annotation provenance, metric implementation/version, runtime/backend identity, uncertainty/claim boundary and reproducibility evidence.
Accordingly, the current issue's deterministic Tier-1 fixture lane is a test floor, not a substitute for the public/private real-audio acceptance lanes. Tempo acceptance continues to require the governed Acc1+Acc2 pair; beat F-measure, WCSR, SI-SDR and other task metrics remain claim-specific rather than interchangeable scoreboards.
PR #828 remains the canonical bounded implementation owner for the known-vocal slice; do not open a parallel MIR acceptance PR. Fresh comparison against protected
develop@314ddeae7b775a4957594b599358c8255617eb2eshows#828@d50a578739e62156839b8633a7ff118dd70f6fbfis diverged, 36 commits ahead and 2 behind, with merge base749511c3ad4000090048718f685c6bee6b3d2c25. Keep it Draft until the rights-cleared real-audio lane and its claim evidence are valid; when it moves, repair by ordinary/non-force reconciliation with current develop, preserving the existing metric/security/provenance work rather than closing, force-rebasing or replacing it.Acceptance remains RED until at least one authorized real-audio corpus/candidate produces reproducible exact-head measurements through the production boundary and the evidence supports the specific rehearsal claim. A synthetic-only GREEN is explicitly insufficient.
Buyer-visible gap
BandScope has strong unit, contract, parity, and security gates, but the protected merge path does not yet prove that an actual audio file traversing the production intake and analysis pipeline yields the expected musical result. A buyer cannot distinguish “the code executes” from “the rehearsal guidance is measurably accurate.”
This issue establishes a reproducible, rights-safe real-audio acceptance layer for harmony, tempo/beat, structure, source separation, role range, and rehearsal cue accuracy. The test must exercise decoded PCM audio through the same public analysis boundary used by the desktop product; mocked feature matrices alone are not acceptance evidence.
Product outcome
Every release candidate publishes an accuracy manifest answering:
Test tiers
Tier 1 — deterministic redistributable PCM fixtures
Generate or check in tiny, license-clean WAV/FLAC fixtures with immutable SHA-256 manifests. These must be real decoded waveforms, not direct chroma/onset arrays. Tier 1 is a deterministic unit/regression floor; it is not sufficient commercial/scientific acceptance evidence by itself.
Required cases:
The fixture generator, waveform parameters, annotations, checksums, and expected metric ranges must be versioned. Generation should be deterministic and independent of network access.
Tier 2 — redistributable public corpus slice
Use only audio whose redistribution and automated evaluation rights are documented. Record the license, source URL/DOI, exact file hash, annotation provenance, split, and any transformation. Do not commit copyrighted commercial recordings or annotations that require unavailable audio.
Tier 3 — private commercial-readiness benchmark
Run a scheduled/manual benchmark against a separately licensed private corpus. Store only aggregate metrics, bounded error exemplars, configuration hashes, and provenance-safe artifacts in GitHub. The workflow must fail closed when the corpus credential or manifest is absent and must never substitute synthetic evidence while claiming private-corpus success.
Metrics and acceptance contracts
Harmony
Beat and tempo
mir_eval.beat.f_measuresemantics with the documented 0.07 s (±70 ms) window; any alternative tolerance must be separately named and must not replace the default silently.mir_evalimplementation and Raffel et al. (2014) remain the executable metric authority for this repository.Structure
Source separation
museval-compatible SDR as a separately labelled compatibility metric. Do not substitute BSS Eval v4 SDR for SI-SDR or collapse the two into one score.Rehearsal-specific outputs
Regression policy
CPU/GPU and implementation boundary
Security and rights notes
Required repository changes
docs/doctoring/real-audio-accuracy-acceptance.mdwith metric definitions, claim boundaries, licenses, operational interpretation, rollback policy, and APA 7th references.AGENTS.md,ARCHITECTURE.md,CHANGELOG.md, and release acceptance documentation.Merge and release gate
Do not claim release readiness from unit tests alone after this acceptance layer exists. Commercial/scientific acceptance requires Tier 2 and/or separately licensed Tier 3 rights-cleared real recordings traversing the production decode → MIR → rehearsal-insight boundary; Tier 1 synthetic/deterministic fixtures cannot substitute for that evidence. Merge the implementation only after exact-head repository CI, central coverage, security/supply-chain checks, realistic audio acceptance, CPU/GPU parity where configured, current-head automated review, zero unresolved threads, and qualifying independent approval all succeed without bypass.
References — APA 7th
Chiu, C.-Y., Liu, L., Weiß, C., & Müller, M. (2025). Cross-modal approaches to beat tracking: A case study on Chopin Mazurkas. Transactions of the International Society for Music Information Retrieval, 8(1), 55–69. https://doi.org/10.5334/tismir.238
Le Roux, J., Wisdom, S., Erdogan, H., & Hershey, J. R. (2019). SDR—Half-baked or well done? In 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2019.8683855
Odekerken, D., Koops, H. V., & Volk, A. (2021). Improving audio chord estimation by alignment and integration of crowd-sourced symbolic music. Transactions of the International Society for Music Information Retrieval, 4(1), 141–155. https://doi.org/10.5334/tismir.81
Raffel, C., McFee, B., Humphrey, E. J., Salamon, J., Nieto, O., Liang, D., & Ellis, D. P. W. (2014). MIR_EVAL: A transparent implementation of common MIR metrics. In Proceedings of the 15th International Society for Music Information Retrieval Conference (pp. 367–372).
Schreiber, H., & Müller, M. (2020). Music tempo estimation: Are we done yet? Transactions of the International Society for Music Information Retrieval, 3(1), 111–125. https://doi.org/10.5334/tismir.43
Stöter, F.-R., Liutkus, A., & Ito, N. (2018). The 2018 signal separation evaluation campaign. In Latent Variable Analysis and Signal Separation (pp. 293–305). https://doi.org/10.1007/978-3-319-93764-9_28
Claim boundary
Passing this issue’s benchmark will support specific, versioned accuracy claims on the registered fixtures and corpora. It will not establish universal musical correctness, genre/culture invariance, or perceptual superiority without separate representative data and human validation.