feat(manual): add the manual test management API - #127
Merged
Merged
Conversation
Phase 1 of manual test case management: QAs can author, search, clone and delete manual test cases per team, with every version of a case preserved. A test case is split into a mutable head (ManualTestCase, always the latest content) and an append-only version collection (ManualTestCaseVersion). The version documents are only ever inserted - nothing updates or deletes them - which is what lets a later execution bind to the exact content it was run against. Storing a version *number* alone would not be enough: resolving it back to content requires that content to still exist somewhere unchanged. The frozen copy is deliberately complete rather than steps-only. An execution has to render what the tester actually saw, which includes the title, preconditions, priority and custom field values, so a partial snapshot would show a mix of old steps and current everything-else. Only content changes burn a version. A status transition (DRAFT -> ACTIVE) is workflow state, not content, and an edit that changes nothing is a no-op - both leave the version number where it is. Steps keep their _id so a step result can bind to a specific step of a specific version, which a later reorder cannot disturb. Versions are written before the head is saved, so a crash between the two leaves an orphan version rather than a head pointing at a version that does not exist. An orphan is harmless; a dangling pointer would make the case unreadable at that version. Permissions follow the existing split: team access to author and clone, team-lead access to delete. The `search` parameter is escaped before being interpolated into a $regex, per validation-utils. Bumps the version to 3.0.0 - this is the first phase of a substantial new feature area rather than an increment on the existing one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds the eight phase 1 operations, a "Manual Test Case" tag, and seven schemas covering the test case, its steps, the list response and the frozen version shape. The schemas follow the existing Base/Stored pairing: ManualTestCase is the request shape, StoredManualTestCase adds _id, version and the audit fields. StoredManualStep exists separately because a step's _id is meaningful to callers - it is preserved into every frozen version, so a recorded step result stays bound to the right step when a later version reorders or removes steps. The version endpoints document the immutability guarantee explicitly, since that is the part a client has to understand to use them correctly: a version response renders the content an execution was run against and is unaffected by later edits. The `unversioned` flag is documented as the one case where that guarantee does not hold, for cases predating version tracking. The line count of the diff overstates the change - appending to `paths` reindented the trailing top-level keys. No existing leaf value was removed or altered, verified by comparing the parsed documents. Note the spec does not pass swagger-cli validation, but did not before this change either: the top-level `definitions` key is not valid in OpenAPI 3.0 and two settings paths reference it. Left alone rather than fixed here. The paths added by this commit validate cleanly in isolation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 2 of manual test case management. Admins configure extra fields per team; test cases carry values for them, validated and coerced on every save. Values are checked against the team's definitions rather than stored blind. Unknown keys are rejected instead of dropped, so a typo in a client surfaces immediately rather than silently discarding what the author typed. Every problem in a submission is collected and reported together - a QA filling in six fields should see all six mistakes at once, not fix one 400 per save. Coercion is deliberate rather than strict typing: values arrive from JSON bodies and form inputs, so a number field legitimately receives "42" and a boolean receives "true". What is not accepted is anything ambiguous - "" for a number, an unparseable date - because storing those silently would surface much later as a broken render. Required fields are enforced only for a status other than DRAFT, so authoring stays incremental but a published case carries what it must. Publishing without restating the custom fields re-validates what is already stored, otherwise a draft could be promoted straight past the check. Custom fields are validated *before* the content comparison that decides whether to burn a version. Comparing raw request values against coerced stored ones would see an unchanged date re-sent as a string as a change, and burn a version for nothing. A field is archived rather than deleted once its values appear on any test case or in any frozen version - the definition is what supplies the label, type and options a historical value renders by, so removing it would silently lose part of what a tester saw. For the same reason the storage key can never change, and neither the type nor an in-use option can be removed while cases depend on them. Archived fields stay valid for existing values but are exempt from the required check and never receive defaults. Frozen versions now record the definitions for the fields the case actually holds values for, so a later relabel or archive leaves history intact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The spec failed `swagger-cli validate` with "#/ must NOT have additional properties": it carried a top-level `definitions` block, which is Swagger 2.0 syntax and is not a permitted key in OpenAPI 3.0. Four `$ref`s pointed into it, so tooling that resolves references strictly could not resolve them either. Moves AuthProvider and AuthProviderPublic into components/schemas, where OpenAPI 3 expects them, and repoints the four references. Neither schema collided with an existing name, and both are otherwise unchanged - comparing the parsed documents with the ref targets normalised shows them identical, so the only semantic change is where the two schemas live and what the references point at. The line count of the diff is almost entirely reindentation from the block moving one level deeper; `git diff -w` shows only the four ref changes. The spec now validates clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 3 of manual test case management. A team lead builds a library of re-usable step sequences; authors include one in a test case and it expands into the steps a tester follows. Expansion happens when a frozen version is written, not when a run reads one. That placement is the whole design. A version stores the expanded steps as literal content, so editing a shared step used by twelve cases cannot rewrite what twelve historical executions show. Had versions stored the reference unexpanded, every one of them would silently re-render with the new content. The consequence is that editing a shared step must cascade: each referencing case gets its version incremented and a fresh frozen version written, re-expanded against the new content. Without that the live case would show new steps while its latest version showed old, and the two would disagree about what "current" means. The head keeps its placeholder throughout, so it carries on tracking the shared step. Only a change to the steps cascades. Renaming or re-describing a shared step alters no test case content and must not version forty cases for nothing. Cascade failures are collected rather than thrown. The shared step edit has already been saved by that point, and reporting the whole request as failed would invite a retry that versions every case a second time. The response carries the count and the failures; the usage endpoint lets an author see the blast radius first. An expanded step keeps the shared step's own _id, so a recorded step result binds to a stable identity. A shared step cannot include another - expansion stays a single pass with no cycle risk - and a reference to a missing or another team's shared step is rejected at write time rather than silently expanding to nothing. Also fixes normaliseValue treating an empty array as a value: a stored step carries `attachments: []` from the schema default while a client re-sending that step omits the key, so re-sending unchanged steps compared unequal and burned a version for no change. This affected PUT /manual-test-case too, not just the shared step cascade. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 4 of manual test case management. A QA uploads an image against a test case or shared step and references it from a step's expected result as ``, which the UI resolves to /attachment/<id>/file. Attachments are their own model rather than an extension of Screenshot. That model requires a build and carries phash, platformId, view and baseline semantics an authoring image has no use for; making `build` optional to fit these in would weaken an invariant every screenshot query and the whole image-engine depends on. An attachment referenced by a frozen version is never deleted. The file is the only copy of the image that version's tester saw, so removing it would silently strip part of a historical execution. Taking the image off a step therefore drops it from the head only; deleting the whole test case removes the versions first, which does release the attachments. A permitted delete also pulls the reference out of any head step so nothing renders broken. The upload path is modelled on multer-config-screenshots: mime allowlist, generated filenames, and the same startsWith(ROOT + path.sep) containment check, because multer writes the file before express-validator ever runs. For the same reason a rejected upload has to clean up after itself - the file is already on disk by the time the team check fails, and multer created the directory before the owner was even confirmed to exist, so both are tidied away. Without that, every denied upload would leave something behind. Thumbnails are files rather than base64 on the document, unlike Screenshot: a case with thirty attachments would otherwise carry thirty blobs in one response, and the authoring UI lists them constantly. The thumbnail endpoint falls back to the original, so an image jimp cannot read is still usable. Step attachment references are validated on write - a dangling id renders as a broken image mid-test, and a cross-team one would expose another team's screenshots inside a case. Adds attachments/ to the ignore files, a volume to the Dockerfile, docker-compose and the k8s pod, and the collection and indexes to mongo-init. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… fields Phase 5 of manual test case management. Every mutation to a manual test case, shared step or custom field definition writes an audit entry naming the actor, the action and the field-level changes. History is deliberately separate from the version collection rather than derived from it. A version is the content a tester saw and is what an execution binds to; a history entry says who changed what and why. Two things the version collection cannot express make that distinction load-bearing: a DRAFT -> ACTIVE transition burns no version but is exactly the sort of change an auditor asks about, and a shared step edit re-versions every referencing case without anyone editing those cases - so each gets a SHARED_STEP_UPDATE entry naming the shared step as the cause, rather than an unexplained bump. Changes are diffed per field, and the two composite fields are broken down further: a steps change reports `steps[2]` with the before and after text, and a custom field reports `customFields.owner`. Reporting either as one blob would technically be accurate and practically useless. Steps are summarised to their text for display. The full subdocument carries _ids and attachment arrays that make a diff unreadable, and the reader wants to know what the tester was asked to do. recordChange is never awaited. The change it describes has already been saved and committed by the time it runs, so failing the user's request because the audit line could not be written would be strictly worse than losing the line - a test verifies the save still succeeds when the history write throws. Tracked paths are explicit per entity type rather than "every key that differs": timestamps, version counters and audit fields change on every save and would bury the changes a reader came for. changedBy is populated to _id and username only. The user document carries apiTokens, which must never leave the server through an audit endpoint, and a test asserts the response carries nothing else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…model
Phase 6 of manual test case management. A QA plans a run, executes each case
against the version it was bound to, and the results appear on the existing
dashboards alongside automated ones.
Manual results are written as ordinary TestExecution documents against a Build
tagged executionType 'manual', through the existing
buildMetricsUtils.addExecutionToBuild. Routing them that way rather than
recomputing the build here means the suite grouping, the result map and the
optimistic-concurrency retry loop are shared: manual and automated results
cannot drift, because there is only one implementation.
executionType defaults to 'automated' on both Build and TestExecution, so every
existing document reads back unchanged with no migration. What those documents
lack is the field on disk, which stops the new { team, executionType, start }
index covering them - scripts/backfill-execution-type.js sets it explicitly and
is safe to re-run.
Three decisions settled during implementation, and the plan updated to match:
BLOCKED maps to SKIPPED on the execution, not ERROR as the plan first proposed.
A blocked test was never executed: nothing was verified and no defect was
found. Recording it as ERROR would inflate the failure count on every dashboard
and alert that counts errors, and conflate "could not run" with "ran and
errored". The richer state stays on the run document, which is not constrained
by the closed executionStates enum the metrics layer switches on.
A run carries one component, matching Build.component; a case from another
component is rejected rather than quietly making the component metrics wrong.
A DEPRECATED case cannot be added to a run - deprecating is how a team retires
a test, and letting one in silently would defeat the status.
Version binding is fixed when a case is added, so editing a case mid-run does
not shift the steps under the tester. Picking up an edit is explicit: rebind
moves a NOT_RUN case onto the latest version, and is refused with a 409 once a
result exists, because re-pointing a recorded execution at content it was never
run against is the silent history rewrite this whole design prevents.
Platform capture reuses app/models/platform.js verbatim, so a manual run
records the same platform and device detail as an automated one and the
platform aggregations need no new fields.
Re-recording a result updates the existing execution rather than adding a
second, keeping the build's totals equal to the number of cases in the run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 7 of manual test case management. Manual runs already appear on the dashboards, having been written as builds and executions in phase 6; this adds the ability to tell the two apart. Every API change is an optional narrowing filter. Omitting executionType returns exactly what it returned before manual runs existed, which is what the rest of the suite passing unchanged verifies - the filter narrows, it never reshapes. The phase metrics response gains a per-period executionTypeBreakdown so the UI can stack automated against manual without a second request. A test asserts the two always sum to result.TOTAL: a breakdown that failed to account for every execution would make a stacked chart silently under-report rather than error. Executions and builds written before manual runs existed have no executionType on disk. Both aggregations default it to 'automated' with $ifNull rather than reading the field directly, so historical data lands in the bucket it belongs to instead of falling out of the breakdown or being labelled "unknown". For Prometheus, angles_builds_by_team gains an execution_type label and two new families expose the split directly. Adding a label changes a series' identity, so this is a breaking change for any query pinning the full label set - though aggregating queries, including the bundled dashboard's sum by (team), are unaffected. Documented under a "Changed in 3.0" note, and the alternative (leaving the series alone and exposing the split only globally) was rejected because it would not answer "which team is doing manual testing?". The execution metrics aggregation now sums by label pair before emitting, since two id-groups can collapse onto one label set - Prometheus rejects a scrape containing duplicate series outright, failing the whole scrape rather than the offending metric. Output validated with dockerised promtool (exit 0). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The plan had both out of scope on the grounds that manual test management has no programmatic caller. That reasoning still holds for the manual surface, but it missed that phases 6 and 7 added fields to the *automated* API those clients already report against - executionType on builds and executions, attachments on steps, and the executionType filter on the list endpoints. Neither client is broken by those additions today: Gson ignores unknown JSON fields, and the Python client returns raw dicts from requests.py rather than deserializing into its dataclasses at all. The gap is that a caller fetching a build cannot read back whether it was automated or manual. Splits the phase by what each client actually does. The JavaScript client backs angles-ui and needs the full manual test management surface across five new request classes. The Java and Python clients get the automated-reporting fields only. Both are explicitly barred from writing executionType. That value is what distinguishes a manual run from an automated one, and a reporting client setting it to "manual" would put a build on the dashboard with no manual run to explain it - it stays server-assigned. Also drops the swagger item: every phase has documented its own endpoints as it landed rather than deferring to this one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Covers the concepts and why executions bind to an immutable version document, authoring and shared step inclusion, custom field configuration and what is safe to change later, runs and the BLOCKED -> SKIPPED mapping, attachment immutability and the 409 on a referenced attachment, change history, the permissions matrix, the automated/manual filter, deployment (attachment volume plus the executionType backfill), and client library scope. Every route in the reference section was cross-checked against the registered routes, and every enum value, the 10MB limit and the custom field key regex against the models. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… chart stacking Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A test case that included a shared step could not be saved twice. The second
edit failed with a duplicate key error on manualtestcaseversions and kept
failing, leaving the case permanently unsaveable.
Two bugs compounded:
1. `action` was required on every step of the head, but a shared step
inclusion has no action of its own - the shared step's contents are
expanded in its place, and expandSteps even supplies a placeholder for an
unresolved reference. The version document stores those expanded steps and
validated fine; the head stored the placeholder and did not.
2. The version was written before the head was saved, so that failure left an
orphan version at the new number. Because { testCase, version } is unique,
the next edit retried the same number and collided - permanently.
Fixes:
- `action` is now required only for a literal step, in both the schema and the
route validator, matching what the feature already assumed.
- The head is validated before the version is written, so an invalid update
writes nothing at all.
- reconcileVersion advances a head that has fallen behind its versions, so a
case already in the broken state recovers on its next save instead of
needing manual repair.
- The same ordering fix is applied to the shared step cascade, which had it too.
- A mongoose ValidationError now maps to 422 rather than surfacing as a 500 -
it is bad input, not a server fault.
test/manual-version-recovery.tests.js covers both the inclusion case and
recovery from an already-wedged case: 6 of its 9 assertions fail without these
changes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A step validation failure reported "Invalid value" twice and named no step,
so an author with several steps could not tell which one was wrong.
withMessage() only applies to the validator immediately before it, so the
single trailing message covered isLength() while exists() and isString() fell
back to the default - and both fired, hence the duplication.
Each check now carries its own message and names the step ("Step 2 requires
an action, unless it includes a shared step"), with bail() so only the first
failure per step is reported.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds folders for manual test cases, nesting to any depth, so a team can file cases by component, feature or whatever structure suits them. Model: - ManualFolder stores a parent reference plus a materialised path of ancestor ids. The parent alone would make "everything under this folder" a recursive walk - one query per level - which the tree view, the delete guard and the subtree filter all need on every request. The path makes it a single indexed query, at the cost of rewriting a subtree's paths on a move. Moves are rare. - ManualTestCase.folder, deliberately outside CONTENT_FIELDS: filing is organisation, not content, so a move burns no version and never reaches a frozen version. An execution renders what was tested, not where the case has since been filed. The move is still audited, as a new MOVE history action. Routes: create, tree, get, rename/move, delete, plus a bulk move for cases. GET /manual-test-case gains `folder` (or "none" for unfiled) and `includeSubFolders`. Guards: - A folder cannot be moved inside itself or its own subtree, which would strand the branch where no query starting from a root could reach it. - Deleting is refused with a 409 while anything is inside, naming the counts. Re-parenting or cascading would rearrange or destroy work in one click, and a test case carries execution history. - Sibling names are unique; the partial index plus an explicit null default is what stops roots escaping the constraint. Also fixes create() to validate the head before freezing a version, matching the ordering already corrected in update(). 31 new tests; suite now 531 passing. Swagger documents all six endpoints. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A test case could not be saved once its steps had been expanded from a shared
step. Saving returned 422 on the first expanded step.
Two shapes carry no action of their own, and only one was exempt:
sharedStep - an inclusion the author just added. Already handled.
sharedStepRef - a step that came *from* a shared step, echoed back by a
client that read the case. Its action is whatever the shared
step said, which may legitimately be empty.
Both the route validator and the schema's conditional `required` now exempt
either shape, so a client can round-trip a case it has read without having to
strip or invent step content.
Adds two regression tests for the echoed-back shape, including an expanded
step whose source had no action.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sending the wrong shape for a reference field - an attachment document where an id was expected - surfaced as "Cast to ObjectId failed ..." with a 500, reading as a server fault rather than bad input. Mapped to 422, naming the offending path, so the message says which field is wrong instead of only that a cast failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The list endpoint sorted purely by recency, which interleaves folders. The UI groups a page into folder sections, so a recency-ordered page scatters one folder's cases across several pages and repeats the same heading on each of them. Sorting by folder first returns whole folders together; recency still orders the cases within a folder. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Manual test case management becomes an optional feature, toggled by an admin from the settings page. It is enabled by default, so an existing instance behaves exactly as it did before. The toggle covers test cases, test runs, shared steps and folders as one unit, because a run is a run *of* cases and the parts are not useful separately. Turning it off makes all four route groups respond 404 for every method, so the data is genuinely unreachable rather than only hidden in the UI. Nothing is deleted - re-enabling restores everything. Feature toggles get their own singleton document rather than joining the auth settings, so the two save independently and turning a feature off never rewrites the provider list. The persisted value is mirrored onto an in-memory config, which lets the route guard read it synchronously and makes a change take effect without a restart. The toggles are also reported by /auth/config, the one config endpoint every client already fetches before login, so the UI can decide which navigation to render. 404 rather than 403: a disabled feature does not exist on this instance, and "forbidden" would imply the caller could be granted access to it. ANGLES_MANUAL_TESTING_ENABLED seeds the initial value so a deployment can ship with the feature already off. It is a seed rather than an override, following ANGLES_ADMIN_PASSWORD: once the settings document exists the database is authoritative, so a restart never silently reverts a change an admin made in the UI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the manual test management API. This is the backend half of
AnglesHQ/angles-ui#100 — that PR calls these endpoints, so this one needs
to land first or the UI has nothing to talk to.
What's here
Test cases with immutable versioning. Editing a case writes a new version
rather than mutating the old one, so a result recorded months ago still shows
the steps the tester actually followed. Each version also freezes a snapshot
of the custom field definitions as they stood at the time, so a field that is
later relabelled, retyped or archived still renders historical executions
correctly.
Folders — a nested tree to organise cases, with moves and per-folder
counts. The list endpoint sorts by folder before recency so a page of results
arrives as whole folders rather than an interleaved slice.
Shared steps — re-usable step sequences with forward-cascading versions: a
change propagates to cases that include them, without rewriting history.
Manual test runs — execute a set of cases and record per-step results.
Results are written into a manual-tagged build, joining the existing automated
model, so dashboards and metrics need no separate reporting path. Both gain an
automated/manual execution-type filter.
Attachments — image uploads against manual steps.
Custom fields — admin-configurable definitions for manual test cases.
Change history — records who changed what on cases, shared steps and
custom fields.
Feature toggle — the whole area sits behind
manualTestingEnabled,settable by an admin and surfaced through
/auth/config.New models and routes
Models:
manual-test-case,manual-test-case-version,manual-folder,shared-step,manual-test-run,manual-change-history,custom-field-definition,attachment,feature-settings.Routes:
manual-test-case,manual-folder,shared-step,manual-test-run,custom-field,attachment.All documented in swagger.
Fixes along the way
covered by tests (
manual-version-recovery.tests.js).components/schemasin swagger.Testing
10 new test files, ~4,000 lines. The full suite is 549 passing locally
against the dockerised Mongo the CI workflow uses.
Notes for the reviewer
manualTestingEnableddefaults to on. Merging this enables manualtesting for every existing install rather than leaving it dark. That is
deliberate — the model comments note that a feature which already shipped
should not need a config change to keep working — but it is worth a
conscious nod, since the alternative is a dark launch.
pre-date this branch and are untouched here.
🤖 Generated with Claude Code