Skip to content

feat(ans-verify): accept an FQDN and resolve its ANS badge - #131

Open
sparkmastergrape wants to merge 3 commits into
agentnameservice:mainfrom
sparkmastergrape:pr/verify-fqdn-badge
Open

sparkmastergrape wants to merge 3 commits into
agentnameservice:mainfrom
sparkmastergrape:pr/verify-fqdn-badge

Conversation

@sparkmastergrape

Copy link
Copy Markdown
Contributor

Closes #120 — the FQDN entry point split out of #99, and the "happy to do it next as its own PR" from #112.

What. ans-verify gains -fqdn. It resolves TXT at _ans-badge.<fqdn>, requires v=ans-badge1, and reads the agent id out of the url= value's /v1/agents/<id> path, which is the shape BadgeRecord writes. That id then enters the existing receipt and proof flow with nothing else altered. -agent and the positional agent id keep working unchanged; passing both is an error. -check-metadata and -metadata-timeout apply on the FQDN path too.

This changes how a registration is located, not the evidence required to verify it. No SVCB endpoint discovery, no agent-card fetch, no reachable-but-unregistered outcome.

The log the badge names. The badge is published by the same operator as the agent, so it does not get to move verification to a log of its choosing. An explicit -url wins, and a badge that disagrees is reported (⚠ badge names X; verifying against -url Y instead). The badge's log is adopted only when -url was left at its default, where the operator expressed no preference and the localhost:18081 default would be useless. Say the word if you would rather it never adopt the badge's log and always require an explicit -url; that is a one-line change.

Failures, all exit 1 with a specific message. Lookup error, NXDOMAIN, no ans-badge1 record at the name, a name publishing two different badges (which registration to verify is ambiguous), or a url= that is not a /v1/agents/<agentId> path. That last one also rejects the endpoint-URL form BadgeRecord falls back to when a deployment configures no TL base URL: it names a reachable agent, not a log, so there is no ANS evidence to check. The agent id must match the UUID shape the RA issues before it is interpolated into a fetch path, reusing the walker's agentIDPattern for the same reason it exists there.

-dns host:port aims the lookup at one resolver, which the local ans-dns dev server needs. A failed lookup names that resolver: Go's own resolver error reports the server from the system configuration even when Dial sent the query elsewhere, which reads as though -dns was ignored.

Tests. badge_test.go: the canonical record, five tolerated forms (trailing slash, no spaces, a log under a path prefix, http for local, uppercase keys), nine rejected ones, the badge owner name actually queried, unrelated TXT at the same name ignored, the same badge published twice accepted, seven resolve failures, the three log-selection branches, flagWasSet, and normalizeResolverAddr. Two of them use a real nameserver: one serving the badge over UDP through newTXTLookup, one pointed at a dead port to check the error names the resolver.

Gate. gofmt clean, go vet ok, golangci-lint (v2.11.4) 0 issues, go test -race clean, every new function 87–100% covered. Also smoke-tested end to end with the real binaries: ans-dns serve publishing the badge, then ans-verify -fqdn resolving it and proceeding into Step 1; plus the -url disagreement, the no-badge, the both-flags and the no-target paths.

Assisted-by: Claude Code (claude-opus-5)

`_ans-badge.<fqdn>` carries the log and the agent id of every ANS
registration, and nothing consumed it (agentnameservice#99). ans-verify gains `-fqdn`:
it resolves the TXT record at that name, requires `v=ans-badge1`, and
reads the agent id out of the `url=` value's `/v1/agents/<id>` path,
which is the shape BadgeRecord writes. The id then enters the existing
receipt and proof flow with nothing else altered. `-agent` and the
positional agent id keep working unchanged, and `-check-metadata` /
`-metadata-timeout` apply on the FQDN path too.

This changes how a registration is located, not the evidence required to
verify it. There is no SVCB endpoint discovery, no agent-card fetch, and
no reachable-but-unregistered outcome. The endpoint-URL form BadgeRecord
falls back to when a deployment configures no TL base URL is rejected
rather than followed: it names a reachable agent, not a log, so there is
no ANS evidence to check.

On the log the badge names: the badge is published by the same operator
as the agent, so an explicit `-url` wins and a badge that disagrees with
it is reported. The badge's log is adopted only when `-url` was left at
its default, where the operator expressed no preference and the
localhost default would be useless.

Every failure exits nonzero with a specific message: a lookup error,
NXDOMAIN, no ans-badge1 record at the name, a name publishing two
different badges, or a `url=` that is not a `/v1/agents/<agentId>` path.
The agent id must match the UUID shape the RA issues before it is
interpolated into a fetch path, for the same reason the tile walker
checks it. `-dns host:port` aims the lookup at one resolver, which the
local ans-dns dev server needs; a failed lookup names that resolver,
since the Go resolver's own error reports the system one regardless.

Refs agentnameservice#99
Closes agentnameservice#120

Signed-off-by: J. DiMare <jdimare@pm.me>

@csnitker-godaddy csnitker-godaddy left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The FQDN-only scope matches #120. Two issues need correction before merge: the signed registration is not checked against the requested hostname, and the lookup can fall through DNS search domains even for an explicitly absolute input.

The focused race tests, go vet, and formatting checks pass. I also exercised the built CLI with signed receipt/status fixtures and loopback DNS/HTTP servers. Existing explicit and positional agent IDs, wrong-key rejection, metadata-mismatch rejection, the metadata opt-out, and the ordinary no-badge failure behaved as expected. Both incorrect-hostname cases still returned success. The separate comment about DNS-selected log trust is nonblocking for this diagnostic receipt verifier.

Comment thread cmd/ans-verify/badge.go
baseURL := b.TLBaseURL
switch {
case !in.URLExplicit:
fmt.Fprintf(out, " log: %s (named by the badge; -url not set)\n", baseURL)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Nonblocking] Explain the trust limits of automatic log discovery

The existing CLI fetches /root-keys from its configured log unless -pubkey is supplied. This branch additionally lets the target's DNS select that log. That is reasonable for receipt inspection, but success establishes cryptographic consistency with the discovered log's advertised keys; it does not establish that the log is independently trusted. Please make that limitation explicit in the output/docs and point users who want to select their verification authority independently to -url or a pinned -pubkey.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and the distinction is worth stating in the tool rather than assumed. Adopting the badge's log means the keys come from the same operator as the agent, so a pass proves the receipt is consistent with that log and nothing about whether the log is one the reader trusts.

2365215 adds it to the output in the branch where the badge's log is adopted:

    log:     https://tl.example.com (named by the badge; -url not set)
    note:    keys come from this log, so success proves consistency with it,
             not that the log is independently trusted - pass -url or -pubkey to choose.

f23a2db carries the same point into the README's -fqdn section, pointing readers who want to choose their own verification authority at -url or -pubkey. That commit also corrects the paragraph's opening claim that -fqdn only locates the registration, which stopped being true with the binding fix above.

Comment thread cmd/ans-verify/main.go
ConfiguredURL: baseURL,
URLExplicit: urlExplicit,
Out: os.Stdout,
})

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Bind the verified registration back to the requested FQDN

After this resolution the FQDN is never checked against the signed receipt. Even with an explicit -url and independently pinned -pubkey, I could point _ans-badge.unregistered.example at a valid registration for registered.example; both the receipt and fresh status token named registered.example, but the CLI printed VERIFIED and exited 0. An unregistered hostname can therefore borrow another registration's proof without breaking a signature. Carry the normalized requested hostname into the verification step and reject when the signed registration's primary host/ANSName does not agree.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed, and it is the defect I should have caught: the badge step resolved a name to an agentId and then dropped the name. Everything after that verified a registration without ever asking whose it was.

Fixed in 2365215. After Step 4 succeeds, -fqdn mode extracts the signed payload and binds it. The host is read from agent.host, falling back to the host segment of ansName parsed with domain.ParseAnsName rather than a split, since the version segment holds dots and a naive slice would compare the wrong name. Comparison is on a normalized form: trailing root dot removed, lowercased.

A mismatch is fatal rather than a warning. In -fqdn mode the question is whether this host is registered, and a receipt for another host is a no, not a qualified yes. An event that names no host at all is also a refusal rather than a skipped check, so a payload the tool cannot bind cannot reach VERIFIED.

Your point that pinning -url and -pubkey does not help is the part worth keeping in the code, and it is in the comment on bindRequestedFQDN: the forged link is the badge, not the log, so choosing the verification authority does not close it.

bind_test.go covers your exact case (unregistered.example pointing at a valid registered.example registration), the accept path including a rooted and a mixed-case request, sibling names in both directions so a suffix comparison cannot creep back in, the ansName fallback including the mis-sliced-version case, and four shapes that name no host. bindRequestedFQDN returns its error rather than exiting so the rejection is assertable.

Comment thread cmd/ans-verify/badge.go
return badge{}, errors.New("empty FQDN")
}
owner := badgeOwnerPrefix + name
txts, err := lookup(ctx, owner)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Query the badge owner as an absolute DNS name

The trailing dot is stripped from the supplied FQDN, but the lookup name is not rooted again. On systems with a DNS search list, even -fqdn absolute.example. can fall back to _ans-badge.absolute.example.<search-domain>.. I reproduced this with no TXT at _ans-badge.absolute.example. and a badge only at the search-qualified name: the CLI reported that it found the original owner's badge and exited 0. Append the DNS root dot to the lookup owner after normalization so another namespace cannot satisfy a missing FQDN badge.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed. Normalizing the input by trimming the dot and then never restoring it left the query relative, which is exactly the fall-through you reproduced.

Fixed in 2365215: the owner is re-rooted after normalization, so the lookup is _ans-badge.<fqdn>. A registration is anchored to one absolute name, and a missing badge should be a missing badge rather than an invitation to the search list.

TestResolveBadge_QueriesTheBadgeOwnerName now asserts the queried name and badge.Owner carry the trailing dot, so the regression is pinned on the string that actually goes to the resolver.

Verifying against keys fetched from the log the badge names proves the receipt is consistent with that log, not that the log is independently trusted. Say so in the -fqdn section, and point readers who want to choose their verification authority at -url or -pubkey. Also corrects the paragraph's opening claim: -fqdn now binds the signed registration to the requested host rather than only locating it.

Signed-off-by: J. DiMare <jdimare@pm.me>
@sparkmastergrape

Copy link
Copy Markdown
Contributor Author

Both corrected, and the nonblocking one too. Thanks for the reproductions — the hostname substitution was the useful kind of finding, since every signature in that scenario is genuine.

[P1] Bind the verified registration back to the requested FQDN — 2365215. The requested host is now carried into the verification step and checked against the signed event, immediately after the signature check and before anything reports success. It reads agent.host, falling back to the host segment of ansName parsed through domain.ParseAnsName rather than split by hand: the version segment carries dots, so splitting on the first label mis-slices v1.0.0 and compares the wrong name. The comparison normalizes case and the root dot. An event that names no host at all is a refusal rather than a pass — if the tool cannot bind what was signed, it cannot answer the question -fqdn asks.

Your case now exits nonzero and names both hosts, so the substitution is visible rather than just a failure.

[P2] Query the badge owner as an absolute DNS name — same commit. resolveBadge re-roots the owner after normalizing, so the lookup is absolute and the search list cannot satisfy it. Worth flagging that TestResolveBadge_QueriesTheBadgeOwnerName asserted the relative name, so it was pinning the defect; it is updated with a note on why the trailing dot is there.

[Nonblocking] Trust limits of automatic log discovery — f23a2db, plus an output line in 2365215. When the badge's log is adopted the step now prints that the keys come from that log, so a pass proves consistency with it and not independent trust, and points at -url or -pubkey. The README's -fqdn paragraph says the same at length; its opening claim needed correcting anyway, since -fqdn no longer only locates the registration.

New tests in cmd/ans-verify/bind_test.go cover the substitution case, case and trailing-dot normalization, sibling and parent names (so a suffix comparison cannot creep back in), the ansName fallback including a mis-sliced version, and the no-host refusal. make check is green.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Triage

Development

Successfully merging this pull request may close these issues.

ans-verify: accept an FQDN and resolve its ANS badge

2 participants