Skip to content

Exploratory: JMH benchmark for a DynamicLatch guarded-decode pattern - #12672

Draft
dougqh wants to merge 4 commits into
masterfrom
experiment/base64-guard-benchmark
Draft

dougqh wants to merge 4 commits into
masterfrom
experiment/base64-guard-benchmark

Conversation

@dougqh

@dougqh dougqh commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

What Does This Do

Benchmark-only. Reworks the exploratory FunctionsBase64Benchmark around an abstract DynamicLatch sketch that is local to the benchmark (no production code changes).

The earlier version compared the hand-written Breaker against a generic ParseHandler plus GenericBreakerCtor / GenericBreakerParam (strategy in a constructor field versus passed at call time). Those two variants are replaced by one abstract class whose subclass is the strategy:

  • handle(input): the operation done optimistically (for a parser, the parse), which may throw for bad input. Named handle, which matched the hook on Latch and ClassLatch when this was written; those now use apply, and this sketch will be aligned when DynamicLatch is revisited.
  • isKnownToFail(input): a cheap, correct pre-check, true only if handle would definitely fail; used while engaged.
  • stacklessFailure(input): the failure to throw while engaged, with no stack trace.

It has two public flavors, matching the decode / decodeOrNull split in #12671:

  • get(input): the failure flows to the caller. While engaged it throws the stackless stand-in instead of paying for a stack trace.
  • tryGetOrNull(input): the failure becomes null, and while engaged no exception is built at all.

State is a plain countdown with hysteresis (engage on a failure of the declared type, disengage after 20 consecutive successes), as in #12671. Latches are static final fields of a named final subclass, one per arm. The hand-written Breaker stays as the specialized baseline, and the unguarded and pre-check arms are unchanged. The Mix arms drop the generic variants and gain mixLatchFlow and mixLatchConvert. A @Setup self-check fails fast if the sketch does not behave as the benchmark assumes.

Motivation

Follow-up to #12671 (Kafka header Base64 guard). Enough independent exception and parsing-cost cases exist (APMLP-1760, APMLP-1769, APMLP-1772, APMLP-1783) to consider a shared primitive instead of repeating one-off copies. The design is tracked in APMLP-1884 (input-rate guard) and APMLP-1894 (the Latch family: Latch in #12670, ClassLatch in #12702). DynamicLatch is the self-resetting member of that family; Breaker is reserved for dependency health.

The constructor-versus-call-time question from the first version is resolved structurally: with an abstract class there is no separate strategy to wire. What remains to measure is the same question as APMLP-1893 (a static final subclass versus a passed function).

Additional Notes

  • Benchmark results. One run on Zulu 17.0.7 (HotSpot), MacBook M1, single thread, 5 forks, on a laptop with normal background activity (load about 4 to 7); the full tables are in the benchmark's Javadoc. JDK 8 and x86 are not measured. Error margins are below 4%.

    ns/op valid input invalid input
    unguarded (status quo) 30.5 905.7
    always pre-check 71.3 2.75
    hand-written Breaker, converting 31.2 2.73
    hand-written Breaker, throwing 31.4 11.6
    DynamicLatch.tryGetOrNull 31.1 2.78
    DynamicLatch.get 31.2 11.5

    Mix, ns/op, one invalid input in every N:

    N unguarded always pre-check Breaker Breaker throwing latch.get latch.tryGetOrNull
    2 467.1 36.6 47.7 49.9 51.9 47.3
    10 118.4 64.9 71.0 71.3 72.1 70.0
    100 41.9 70.0 46.1 46.4 47.9 45.9
    1,000 41.6 70.3 44.7 45.7 45.1 45.6
    10,000 34.3 70.2 36.2 35.3 39.6 35.5

    The sketch matches the hand-written breaker: within 0.2 ns in the single-input arms and within about 4 ns in the mixed ones (the largest gap is flow-through at one in 10,000, 39.6 against 35.3 ns). While engaged, the converting flavor costs about 2.8 ns where the status quo costs about 906 ns, and the flow-through flavor costs about 11.5 ns, the price of building a stackless exception. Always pre-checking more than doubles the cost of valid input (71 ns against 30 ns), which is what the adaptive form avoids.
    It is a tradeoff, not a free win. With one invalid input in 100 or rarer, the guard costs 1 to 4 ns over doing nothing and is about 24 to 34 ns cheaper than always pre-checking. With a high rate (one in 2 or one in 10), always pre-checking is faster than the adaptive guard (36.6 ns against about 47 to 52 ns at one in 2, and 65 ns against about 70 to 72 ns at one in 10). I did not investigate why.

  • Not a proposal to merge the abstraction. DynamicLatch cannot go into main until Latch and ClassLatch have merged, and it needs at least two real users. This may be folded into Fix repeated exception cost from malformed Kafka header Base64 decoding #12671 later.

  • Stackless failure gotchas (recorded in the benchmark's Javadoc and APMLP-1884): FastFailBase64Exception is a new instance per failure, never shared. A shared instance would also have to disable suppression (otherwise addSuppressed accumulates on it forever) and forbid initCause. IllegalArgumentException has no constructor that turns the stack trace off, so the stand-in overrides fillInStackTrace.

  • Declared failure type. The converting flavor must not swallow unrelated runtime exceptions, so the sketch takes a Class<X> and converts only that type. Whether that should be a class token or a predicate is an open question in APMLP-1884.

  • The branch is 24 commits behind master; I have not merged it.

Contributor Checklist

Jira ticket: APMLP-1884

🤖 Generated with Claude Code

Follow-up to the Kafka header Base64-decode guard fix. Compares the
hand-specialized GuardedBase64Decode/Breaker shape against a generic
ParseHandler + Breaker abstraction (construction-time vs. call-time
strategy composition) to see what devirtualization costs, if any, a
reusable version of this pattern would carry. Motivated by APMLP-1513's
children (APMLP-1760, APMLP-1769, APMLP-1772, APMLP-1783), which show
enough independent exception/parsing-cost cases to justify considering
a shared primitive rather than one-off copies.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@dougqh dougqh added comp: core Tracer core tag: no release notes Changes to exclude from release notes tag: experimental Experimental changes tag: ai generated Largely based on code generated by an AI or LLM labels Sep 28, 2026
@dd-octo-sts

dd-octo-sts Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.91 s 14.68 s [+0.7%; +2.5%] (maybe worse)
startup:insecure-bank:tracing:Agent 13.74 s 13.72 s [-0.8%; +1.1%] (no difference)
startup:petclinic:appsec:Agent 17.01 s 16.85 s [+0.1%; +1.8%] (maybe worse)
startup:petclinic:iast:Agent 16.93 s 17.07 s [-1.6%; +0.0%] (no difference)
startup:petclinic:profiling:Agent 16.58 s 16.82 s [-2.4%; -0.5%] (maybe better)
startup:petclinic:sca:Agent 16.93 s 16.74 s [+0.1%; +2.2%] (maybe worse)
startup:petclinic:tracing:Agent 15.87 s 15.82 s [-5.6%; +6.2%] (unstable)

Commit: e3158a50 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

Replace the ParseHandler plus GenericBreakerCtor/GenericBreakerParam
pair with one abstract DynamicLatch sketch, local to the benchmark, whose
subclass is the strategy: an optimistic parse, a cheap correct pre-check
and a stackless failure. That removes the constructor-versus-call-time
question, since there is no separate strategy object.

The latch has two flavors: get lets the failure flow to the caller
(throwing a stackless stand-in while engaged), and tryGetOrNull converts
it to null and builds no exception at all while engaged. Latches are
static final fields of a named final subclass. The hand-written Breaker
stays as the specialized baseline, and a setup self-check fails fast if
the sketch misbehaves.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@dougqh dougqh changed the title Exploratory: JMH benchmark for a generic guarded-decode/breaker pattern Exploratory: JMH benchmark for a DynamicLatch guarded-decode pattern Sep 30, 2026
Rename the optimistic hook parse to handle, so it reads the same across
the Latch family, and the pre-check isDefinitelyInvalid to isKnownToFail,
which is not specific to parsers. Naming only.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@datadog-prod-us1-5

This comment has been minimized.

Record a five-fork run on Zulu 17 (M1) in the benchmark's Javadoc: the
single-input arms and the mixed-input arms at five invalid rates. The
sketch matches the hand-written breaker; the adaptive form is a
tradeoff against always pre-checking at high invalid rates.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: core Tracer core tag: ai generated Largely based on code generated by an AI or LLM tag: experimental Experimental changes tag: no release notes Changes to exclude from release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant