Skip to content

Fix Scala firstCompletedOf continuation cleanup - #12681

Open
amarziali wants to merge 3 commits into
masterfrom
andrea.marziali/scala-flake
Open

amarziali wants to merge 3 commits into
masterfrom
andrea.marziali/scala-flake

Conversation

@amarziali

Copy link
Copy Markdown
Contributor

What Does This Do

Fixes a continuation leak in Scala 2.13.17+ when Future.firstCompletedOf unregisters callbacks from futures that did not complete first.

The instrumentation releases a callback’s continuation only when Scala’s callback-removal CAS succeeds. If completion has already dispatched the callback, ownership remains with the callback until it runs.

flowchart LR
  A[Capture callback continuation] --> B{Which operation wins?}
  B -->|Unregister CAS| C[Release continuation]
  B -->|Callback dispatch| D[Callback run consumes continuation]
Loading

Motivation

GitLab job 2087926766 intermittently timed out in:

ScalaInstrumentationTest > scala first completed future

Diagnostics identified an unresolved continuation captured by Promise$Transformation. Scala 2.13.17 introduced callback unregistration for firstCompletedOf, adding a terminal lifecycle path that bypassed the existing execution cleanup.

Unconditional cleanup on unregisterCallback would be unsafe because completion can commit the callback to execution before unregistration observes the promise state. The successful removal CAS is the point where ownership is transferred safely.

Motivation

Additional Notes

Contributor Checklist

Jira ticket: [PROJ-IDENT]

@amarziali
amarziali requested review from a team as code owners September 29, 2026 10:36
@amarziali amarziali added the type: bug fix Bug fix label Sep 29, 2026
@amarziali
amarziali removed the request for review from a team September 29, 2026 10:36
@amarziali amarziali added the tag: flaky test Flaky tests label Sep 29, 2026
@amarziali
amarziali requested a review from dougqh September 29, 2026 10:36
@amarziali amarziali added inst: scala Scala instrumentation tag: ai generated Largely based on code generated by an AI or LLM labels Sep 29, 2026
@amarziali
amarziali requested review from vandonr and removed request for a team September 29, 2026 10:36
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T10:40:01.886277Z a06df8d PR opened
🔒 Security Review ✅ Completed 2026-09-29T10:39:57.770830Z a06df8d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@amarziali amarziali added the tag: override groovy enforcement Override the "Enforce Groovy Migration" check label Sep 29, 2026

@datadog-prod-us1-5 datadog-prod-us1-5 Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bits Code Review: PASS

More details

The Scala 2.13.17+ path releases the continuation only after a successful state-changing removal CAS; when callback dispatch wins, the callback retains cleanup ownership.

Was this helpful? React 👍 or 👎

Open Bits AI session

🤖 Bits Code Review · Commit a06df8d · @DataDog review to ask questions

@datadog-prod-us1-5

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.09 s 13.98 s [-0.1%; +1.7%] (no difference)
startup:insecure-bank:tracing:Agent 12.96 s 13.01 s [-1.1%; +0.3%] (no difference)
startup:petclinic:appsec:Agent 16.99 s 16.81 s [+0.2%; +2.0%] (maybe worse)
startup:petclinic:iast:Agent 16.97 s 16.97 s [-1.0%; +1.1%] (no difference)
startup:petclinic:profiling:Agent 16.59 s 16.95 s [-3.4%; -0.9%] (maybe better)
startup:petclinic:sca:Agent 16.97 s 16.14 s [+1.0%; +9.4%] (maybe worse)
startup:petclinic:tracing:Agent 16.27 s 16.16 s [-0.1%; +1.5%] (no difference)

Commit: 9dd16ff7 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@amarziali
amarziali force-pushed the andrea.marziali/scala-flake branch from 1b00f2e to c9c419d Compare September 29, 2026 11:46
@amarziali
amarziali force-pushed the andrea.marziali/scala-flake branch from c9c419d to 188f30f Compare September 29, 2026 12:45

@vandonr vandonr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean, I'm approving to unblock things, but really I don't understand half of what's happening in this code...
@dougqh if you have more Scala knowledge, review carefully, make no mistake 😅

@dougqh

dougqh commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

(Comment drafted by Claude on behalf of @dougqh.)

I checked the fix against the 2.13.17 source, and the mechanism is correct: unregisterCallback has exactly the two CAS sites this rewrites, and slot 1 is t throughout. My hesitation is risk versus reward, not correctness.

Reward. As far as I can tell, this fixes a flaky test timeout, and the leak looks like a delayed trace flush rather than lost or incorrect data. firstCompletedOf is also a fairly niche API. Could you spell out the production impact? For example, how long until the trace flushes, and does pending-trace memory build up in the meantime?

Risk. This is ASM surgery on a private method of the Scala stdlib. It depends on the call-site count (< 2) and the callback being in local slot 1, and neither is checked by muzzle. A Scala 2.13.x change could silently reintroduce the leak, or worse, close the wrong callback's continuation. It also adds a muzzle directive and a suite that runs twice.

Two questions:

  1. Is this driven by Enforce continuation diagnostics in suite fixtures #12528's strict continuation diagnostics? If so, would exempting this scenario there, or a lighter approach, do the job?
  2. If we keep the rewrite, could we add the loud-failure guards from the inline comments, plus a test with three or more futures?

I'm not blocking. I'd just like us to decide deliberately that the reward justifies this kind of instrumentation.

@amarziali

Copy link
Copy Markdown
Contributor Author

Thanks for the thourough review @dougqh

I’ve added independent guards for the exact CAS count and preservation of the callback argument, plus mutation tests for violations.

The remaining tradeoff is explicit: an unsupported bytecode shape rejects the whole Datadog transformation for DefaultPromise, including its other advice. DefaultPromise instrumentation also, is opt-in.

Two questions:

  1. Is this driven by Enforce continuation diagnostics in suite fixtures #12528's strict continuation diagnostics? If so, would exempting this scenario there, or a lighter approach, do the job?

The diagnostic is there to highlight this kind of lifecycle misses. Now we can ignore it if it's too complex to maintain. The diagnostic allows to override the check by providing a reason.

  1. If we keep the rewrite, could we add the loud-failure guards from the inline comments, plus a test with three or more futures?

Yes I already did this in the PR as suggested.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

inst: scala Scala instrumentation tag: ai generated Largely based on code generated by an AI or LLM tag: flaky test Flaky tests tag: override groovy enforcement Override the "Enforce Groovy Migration" check type: bug fix Bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants