Skip to content

fix(collector): keep collecting temperatures without CPU power status - #3767

Open
SaiPisey2 wants to merge 2 commits into
prometheus:masterfrom
SaiPisey2:fix/thermal-darwin-no-cpu-power-status
Open

fix(collector): keep collecting temperatures without CPU power status#3767
SaiPisey2 wants to merge 2 commits into
prometheus:masterfrom
SaiPisey2:fix/thermal-darwin-no-cpu-power-status

Conversation

@SaiPisey2

Copy link
Copy Markdown

Addresses #2906.

Apple Silicon does not implement IOPMCopyCPUPowerStatus, so fetchCPUPowerStatus gets back kIOReturnNotFound. Update returned that error straight away, which meant it never reached updateTemperatures, so no temperature metrics were collected at all. The error also isn't ErrNoData, so it was logged at error level on every scrape.

Measured on an M5 Pro (darwin/arm64) against master:

fetchCPUPowerStatus()      -> status=map[] err=no CPU power status has been recorded
Update()                   -> err=no CPU power status has been recorded ; metrics emitted=0
IsNoDataError(err)         -> false
updateTemperatures() alone -> err=<nil> ; temperature metrics available=52

52 usable temperature sensors on that machine, none of them exported, because an expected and unrelated condition aborted the collector.

Change

Treat kIOReturnNotFound as "this system does not report CPU power status" rather than a failure: skip the three CPU power metrics, log at debug level, and continue on to the temperature sensors. Any other non-success return code is still returned as an error, exactly as before, so systems that do report CPU power status are unaffected.

After the change, on the same machine, Update() returns no error and emits all 52 temperature metrics.

Scope

This does not make the CPU power metrics appear on Apple Silicon and does not resolve #2218 — the underlying API provides no data there, as @rexagod already established on #2906. It only stops their absence from suppressing the temperature metrics. I've left Closes #2906 out of the commit deliberately, since whether that issue is fully answered by this is a maintainer call.

Testing

Added collector/thermal_darwin_test.go, which fails on master:

--- FAIL: TestThermalUpdateWithoutCPUPowerStatus
    thermal_darwin_test.go:44: Update returned errNoCPUPowerStatus; a system
    without CPU power status must still collect temperatures

and passes with the change. Also verified locally:

  • go build ./...
  • go vet ./collector/
  • gofmt -l clean on both touched files
  • go test ./collector/ full package
  • go build -tags notherm ./collector/

@nicolastakashi

Copy link
Copy Markdown

/workflow-approve

@nicolastakashi

Copy link
Copy Markdown

The Darwin/macOS e2e job is failing on a fixture mismatch, node_scrape_collector_success{collector="thermal"} changed from 0 to 1. Can you regenerate the golden output on a Darwin/arm64 box and push it?

./end-to-end-test.sh -u

Then commit the updated collector/fixtures/e2e-output-darwin.txt.

Apple Silicon does not implement IOPMCopyCPUPowerStatus, so
fetchCPUPowerStatus returns kIOReturnNotFound there. Update returned that
error straight away, which aborted the collector before updateTemperatures
ran, so no temperature metrics were collected at all. The error was also
not ErrNoData, so it was logged at error level on every scrape.

On an M5 Pro, Update emitted 0 metrics and failed, while updateTemperatures
on its own returned 52 temperature metrics.

Treat kIOReturnNotFound as a system that does not report CPU power status:
skip the three CPU power metrics, log at debug level and carry on to the
temperature sensors. Any other non-success return code is still returned as
an error. Systems that do report CPU power status are unaffected.

This does not make the CPU power metrics available on Apple Silicon, since
the underlying API provides no data. It only stops their absence from
suppressing the temperature metrics.

Adds a regression test covering the case.

Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
Collecting temperatures on Apple Silicon surfaced a second problem that
was previously unreachable, because Update returned before
updateTemperatures ever ran.

A sensor is identified only by its IOHID "Product" name, and several
services report the same name. On an M5 Pro, 76 temperature-capable
services carry only 25 distinct names, and the services expose nothing
that tells them apart: RegistryID, UniqueID and SerialNumber are all
unset, and two services can share both Product and LocationID. Emitting
the same label set twice makes the registry reject those samples and log
an error on every scrape, so only the first reading for a name is
reported now and the rest are counted in a debug message.

Also update collector/fixtures/e2e-output-darwin.txt for
node_scrape_collector_success{collector="thermal"}, which is now 1
because the collector no longer fails.

Two fixes to end-to-end-test.sh so that regenerating that fixture on a
Darwin host does the right thing:

  - The Darwin fixture path was built with "${fixture_metrics::-4}". A
    negative substring length needs bash 4.2 and macOS still ships bash
    3.2, where the expansion fails and fixture_metrics keeps pointing at
    the Linux file. Running ./end-to-end-test.sh -u on macOS therefore
    overwrote collector/fixtures/e2e-output.txt with Darwin output.
    "${fixture_metrics%.txt}" is equivalent and portable.

  - node_thermal_temperature_celsius is per-machine, both in its sensor
    names and its values, so it is stripped along with the other
    non-deterministic metrics. The CI runner reports no sensors at all,
    but a developer running the suite on real hardware would otherwise
    see the whole sensor list as a diff.

Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
@SaiPisey2
SaiPisey2 force-pushed the fix/thermal-darwin-no-cpu-power-status branch from 8d1e633 to 907bf2f Compare August 6, 2026 10:22
@SaiPisey2

Copy link
Copy Markdown
Author

Thanks for the review. Pushed, but regenerating the fixture turned up two things worth flagging rather than just committing the output.

The fixture change itself

I did not commit a regenerated file. Running ./end-to-end-test.sh -u on real hardware produced a 469-line diff, because this machine reports actual sensors while the CI runner reports none, and the values are live temperatures. So the fixture now carries only the single line the CI diff asked for:

-node_scrape_collector_success{collector="thermal"} 0
+node_scrape_collector_success{collector="thermal"} 1

./end-to-end-test.sh -u overwrites the Linux fixture on macOS

Worth knowing before anyone else follows the same instruction. The Darwin fixture path is built with ${fixture_metrics::-4}, and a negative substring length needs bash 4.2. macOS still ships bash 3.2, where that expansion fails, fixture_metrics keeps pointing at e2e-output.txt, and the update copies Darwin output over the Linux fixture. My first run rewrote 5330 lines of collector/fixtures/e2e-output.txt before I noticed. ${fixture_metrics%.txt} is equivalent and works on both, so I've changed it.

I also added node_thermal_temperature_celsius to non_deterministic_metrics. It is per-machine in both its sensor names and its values, so without it anyone running the suite on real hardware sees the whole sensor list as a diff. It makes no difference in CI, which reports no sensors.

(Unrelated and left alone: the sed -i /pattern/d calls in that loop need an extension argument on BSD sed, so the non-deterministic stripping aborts on stock macOS. It behaves the same on master, so it is not something this PR introduces.)

A real bug the fixture regeneration exposed

This is the part I would not have caught otherwise. Once temperatures are actually collected, the collector emits duplicate label sets:

emitted=52  uniqueSensorNames=17  duplicateSeries=35

and the scrape logs 35 error(s) occurred: ... was collected before with the same name and label values every time, with those samples dropped.

A sensor is identified only by its IOHID Product name, and several services report the same one. Probing the services directly, 76 temperature-capable services carry just 25 distinct Product values, and there is nothing to tell them apart: RegistryID, UniqueID, SerialNumber and Manufacturer are all unset, VendorID/ProductID/PrimaryUsage are constant across every service, and two services can share both Product and LocationID (PMU tdie2 at 1414541922 appears twice). There are also genuinely distinct sensors sharing a name — two gas gauge battery services with different LocationIDs.

Since there is no property that makes them addressable, I report the first reading per name and count the rest in a debug line. The endpoint now returns 200 with 17 unique series and no gather errors. TestThermalTemperaturesAreUnique covers it and fails without the change.

CI never sees this, since the runner has no sensors, but every Apple Silicon machine would have.

Verified locally: go build ./..., go vet ./collector/, gofmt clean, go test ./collector/ including 20 repeats of the thermal tests, and a -tags notherm build.

Happy to split the two end-to-end-test.sh changes into their own PR if you would rather keep this one to the collector.

@nicolastakashi

Copy link
Copy Markdown

Thanks for the fixes. Let's split the end-to-end-test.sh changes into a separate PR.

On the duplicate sensors: dropping readings loses data. #3646 hit the same colliding-label problem for hwmon and fixed it by disambiguating the label instead (suffix it, falling back to something always unique when the first suffix still collides). Take a look there for the approach.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SIGTRAP: trace trap on M1

2 participants