Hi @BETA15, I'm the one who replied on Reddit with the long messages about my own experience and testing.
@DTS Engineer, I can corroborate and extend BETA15's findings from a different chassis and OS train, with mechanism-level instrumentation. MacBook Pro 16-inch, M5 Max (40-core GPU), 128 GB, on AC (original 140W power block and MagSafe cable). Reproduced across macOS 27.0 Developer Beta 3 (26A5378n) and Beta 4 (26A5388g). Filed as FB23754032 (currently showing 10+ similar reports), with three follow-ups and full raw telemetry attached.
Why my data point matters for this thread specifically: the 16-inch M5 Max was independently measured as completely stable under sustained GPU load at launch (Notebookcheck: stable in Automatic mode, no throttling). Whatever both of us are now measuring is therefore not a chassis limitation.
What I measured, using a native Metal/MLX video-inference pipeline (not a windowed or iOS-compatibility benchmark) with two independent telemetry channels recorded simultaneously (sudo powermetrics at 1 Hz + the SMC/IOReport channel):
Back-to-back identical runs degrade monotonically: 513 s -> 642 s -> 627 s on Beta 4 (497 -> 565 -> 570 s on Beta 3). Output hashes are bit-identical across all runs, so the computation is constant; only the regulator state changes.
During the degraded runs, the GPU sits at 100% utilization while collapsed to 840 MHz / 16 W, with fans held at ~2,900 rpm of 5,800 available and the die pinned at 68.7 C. Deepest live sample: 534 MHz / 8.6 W. A fourth run launched after ~50 min of idle (die verified at ~50 C beforehand) came back fully healthy: 457 s, ~1,230 MHz / 30 W, fans allowed to reach ~4,000 rpm, die allowed to rise to 78 C. Same binary, same input, bit-identical output. The accumulated state fully resets with idle time.
3DMark Steel Nomad Stress Test (native macOS app), 20 loops, run twice ~30 min apart with a single variable changed - fan policy. High Power was selected in both runs. Fans on an auto curve: monotonic decline on all 20 loops (41.2 -> 32.7 fps, 79.3% stability) while the GPU die temperature FELL from 79 C to 72 C and system power decayed from ~103 W to ~93 W. Fans manually forced to 100%: 91.6% stability, sustained 41.2 fps at a stable ~73 C / ~114 W, with loop times improving monotonically from loop 2 onward. Converted to scores, my fans-auto run (3982 first loop / ~3460 sustained) matches BETA15's 14-inch results (3983 / 3432) nearly digit-for-digit.
Performance falling together with temperature, at power levels far below the launch-review equilibria, with fan headroom unused, is the inverse of thermal throttling. The consistent mechanism across all our data: the regulator reduces clock first, at die temperatures below any fan curve's ramp thresholds, instead of using the available cooling capacity - and the constraint tightens with sustained load (integrator-like) and relaxes with
idle time.
Note that the fans-auto decline above happened with High Power selected: the mode whose sole documented function is to raise the fan-speed ceiling presided over a run where the fans never ramped and the clock was progressively reduced instead.
On the High Power Mode engagement question discussed above: accepting the definition given in this thread - engagement is heat-triggered, not work-triggered - our data shows why "High Power Mode: No" is permanent on these machines under this policy. The clock-reduction loop suppresses the die temperature below any plausible engagement threshold before it can be reached. Under the current regulation, the documented engagement condition for High Power Mode is unreachable by construction. The status readout is not cosmetic; it is the observable of the clock-before-fans priority order.
One change worth noting between Beta 3 and Beta 4: on 26A5378n, powermetrics reported a constant 1620 MHz / 100% residency / Nominal pressure throughout degraded episodes (verified against 3,082 and 3,629 raw samples), while the IOReport channel showed the true collapsed frequencies. On 26A5388g, powermetrics now reports the real frequency during collapse (543 MHz, "1620 MHz: 0%") and agrees with IOReport. The observability defect appears fixed; the regulation itself is unchanged or slightly worse.
Happy to provide the dual-channel telemetry, the .3dmark-result files, or a sysdiagnose captured during a degraded episode - everything is already attached to FB23754032, and I would welcome that report being related to the FBs referenced in this thread.