Skip to content

ci: adjust prof asan to avoid nightly (and other build refactoring) - #4241

Merged
morrisonlevi merged 3 commits into
masterfrom
levi/prof-rustc-bootstrap-instead-of-nightly
Sep 28, 2026
Merged

morrisonlevi merged 3 commits into
masterfrom
levi/prof-rustc-bootstrap-instead-of-nightly

Conversation

@morrisonlevi

@morrisonlevi morrisonlevi commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Description

The main goal here is to avoid +nightly toolchains so we can drop it from our images, which will save a substantial amount of bytes to pull on every job. Cleans up other build and CI stuff while I'm at it, such as refactoring the compile_profiler Make target and selecting features and some stale docs.

I tried separating out the "independent" changes to docs but the changes are kind of tied together, since this is small enough, I think we can just review it instead of trying to untangle it (not worth the effort IMO).

Reviewer checklist

  • Test coverage seems ok.
  • Appropriate labels assigned.

@datadog-prod-us1-5

This comment has been minimized.

@pr-commenter

pr-commenter Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Benchmarks [ profiler ]

Benchmark execution time: 2026-09-28 15:10:34

Comparing candidate commit 33488c7 in PR branch levi/prof-rustc-bootstrap-instead-of-nightly with baseline commit 8e9b2ff in branch levi/claude-docs-staleness.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 27 metrics, 9 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:php-profiler-exceptions-control

  • unstable cpu_system_time [+0.854ms; +8.003ms] or [+2.119%; +19.868%]
  • unstable cpu_user_time [-8.436ms; -1.911ms] or [-12.991%; -2.943%]

scenario:php-profiler-exceptions-with-profiler

  • unstable cpu_system_time [-5191.099µs; +6033.099µs] or [-11.486%; +13.349%]
  • unstable cpu_user_time [-8.676ms; +5.142ms] or [-11.304%; +6.700%]

scenario:php-profiler-exceptions-with-profiler-and-timeline

  • unstable cpu_system_time [+0.522ms; +6.640ms] or [+1.128%; +14.346%]

scenario:php-profiler-timeline-memory-control

  • unstable cpu_system_time [-6.141ms; +2.564ms] or [-13.784%; +5.755%]

scenario:php-profiler-timeline-memory-with-profiler

  • unstable cpu_system_time [-25.708ms; +35.212ms] or [-5.931%; +8.124%]
  • unstable cpu_usage_percentage [-1.581%; +10.066%]

scenario:php-profiler-timeline-memory-with-profiler-and-timeline

  • unstable cpu_system_time [-4.244ms; +41.957ms] or [-1.035%; +10.229%]

@pr-commenter

pr-commenter Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Benchmarks [ tracer ]

Benchmark execution time: 2026-09-28 15:47:59

Comparing candidate commit 33488c7 in PR branch levi/prof-rustc-bootstrap-instead-of-nightly with baseline commit 8e9b2ff in branch master.

📊 Benchmarking dashboard

Found 1 performance improvements and 2 performance regressions! Performance is the same for 190 metrics, 1 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:HookBench/benchWithoutHook

  • 🟥 execution_time [+4.088µs; +6.875µs] or [+5.003%; +8.413%]

scenario:MessagePackSerializationBench/benchMessagePackSerialization-opcache

  • 🟩 execution_time [-4.893µs; -2.847µs] or [-4.247%; -2.472%]

scenario:SamplingRuleMatchingBench/benchRegexMatching2

  • 🟥 execution_time [+55.093ns; +115.907ns] or [+3.797%; +7.988%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:LaravelBench/benchLaravelDdprof-opcache

  • unstable execution_time [-538.847µs; +867.367µs] or [-4.147%; +6.675%]

@morrisonlevi
morrisonlevi force-pushed the levi/prof-rustc-bootstrap-instead-of-nightly branch from 18f3f57 to 7b6172b Compare September 28, 2026 14:07
- Image tags bookworm-6 -> bookworm-11 and clang versions in the CI docs.
- The profiler is built from the root datadog-php crate; there is no
  profiling/Cargo.toml, and the root rust-toolchain.toml is the pin.
- Profiler CI job table/matrix: add UBSAN, drop the stale ASAN row name.
- buildPortableLibdatadogPhp uses stable Rust (compile_rust.sh sets
  RUSTC_BOOTSTRAP=1), not nightly.
@morrisonlevi
morrisonlevi force-pushed the levi/prof-rustc-bootstrap-instead-of-nightly branch from 7b6172b to 33488c7 Compare September 28, 2026 14:39
@morrisonlevi
morrisonlevi changed the base branch from master to levi/claude-docs-staleness September 28, 2026 14:46
@morrisonlevi
morrisonlevi changed the base branch from levi/claude-docs-staleness to master September 28, 2026 15:31
@morrisonlevi
morrisonlevi marked this pull request as ready for review September 28, 2026 15:43
@morrisonlevi
morrisonlevi requested review from a team as code owners September 28, 2026 15:43
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T15:49:58.550885Z 33488c7 Draft marked ready
🔒 Security Review ✅ Completed 2026-09-28T15:48:01.986251Z 33488c7 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@realFlowControl realFlowControl left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we can remove the "nightly" references completely, yes?

Comment thread .claude/ci/building-locally.md
Comment thread .claude/ci/github-actions-profiler.md Outdated
Co-authored-by: Florian Engelhardt <florian.engelhardt@datadoghq.com>
@morrisonlevi
morrisonlevi merged commit a4b6a50 into master Sep 28, 2026
32 of 35 checks passed
@morrisonlevi
morrisonlevi deleted the levi/prof-rustc-bootstrap-instead-of-nightly branch September 28, 2026 15:53
@github-actions github-actions Bot added profiling Relates to the Continuous Profiler tracing area:asm labels Sep 28, 2026
@github-actions github-actions Bot added this to the 1.26.0 milestone Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:asm profiling Relates to the Continuous Profiler tracing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants