[Tooling] (DeviceBench): verify the two campaign devices' pin profiles and add a pin helper that works on them

- bench.sh's pin_freqs writes /proc/ppm and /proc/gpufreq, which exist on neither 35d0befa
  (SM8750) nor 3B159D009VZ00000 (MT6993 dropped both for /proc/gpufreqv2); the guard the
  PROFILE_VERIFIED key exists for was about to be defeated on the device its own comment
  guessed was safe
- pin_device.sh pins big/little/GPU through the nodes each device actually has (Xiaomi:
  policy6 1958400 / policy0 1555200 / kgsl pwrlevel 0; Oppo: policy4 2000000 + policy7
  2000000 / policy0 1600000 / gpufreqv2 fix_target_opp_index 0), checks only against its own
  pins because ColorOS moves stock maxima by itself, and refuses to call a zero-key read a verdict
- both profiles are PROFILE_VERIFIED=1 on 30 s pinned windows with zero drift; the evidence and
  the two things not verified (Oppo GPU held at DVFS-request level only, no full bench window)
  are in docs/Disaggregated/devices/pin-verification-2026-09-07.md
This commit is contained in:
2026-09-07 23:43:11 -04:00
parent 7a2e256133
commit 2d690754dd
5 changed files with 1003 additions and 72 deletions
@@ -0,0 +1,417 @@
# P2 frequency-pin profiles: what was read, what was written, what was proven
Date: 2026-09-07. Devices: `35d0befa` (Xiaomi 24129PN74C, Snapdragon 8 Elite / SM8750, Adreno
830v2) and `3B159D009VZ00000` (Oppo PLG110 / ColorOS, MediaTek MT6993 "Dimensity 9500", Mali).
Both rooted via Magisk (`su` returns uid 0 in context `u:r:magisk:s0`). Host: Windows, Git Bash,
`MSYS_NO_PATHCONV=1` on every adb invocation.
Outputs, all in this directory:
| file | what it is |
|---|---|
| `xiaomi-adreno830.env` | completed profile, `PROFILE_VERIFIED=1` |
| `oppo-mali.env` | completed profile, `PROFILE_VERIFIED=1` |
| `pin_device.sh` | standalone `<serial> pin\|unpin\|check` helper, pure adb + su |
| `REPORT.md` | this file |
Both devices were left **unpinned**, confirmed by `pin_device.sh <serial> check` at the end
(transcripts at the bottom). Every node written was restored.
---
## Headline: bench.sh cannot pin either of these devices
This is the finding that matters most, because it inverts an assumption written into the
existing profiles.
The `pin_freqs()` / `unpin_freqs()` pair in bench.sh writes
`/proc/ppm/policy/hard_userlimit_{min,max}_cpu_freq` and `/proc/gpufreq/gpufreq_opp_freq`.
On these devices:
| path | `35d0befa` (Qualcomm) | `3B159D009VZ00000` (MediaTek) |
|---|---|---|
| `/proc/ppm/policy/` | absent | **absent** |
| `/proc/gpufreq/` | absent | **absent** (superseded by `/proc/gpufreqv2/`) |
The Qualcomm case was already anticipated - `xiaomi-adreno830.env` said so. The MediaTek case
was **not**: the old `oppo-mali.env` reasoned "this part is a MediaTek SoC, so unlike the Adreno
profile the harness's existing /proc/ppm + /proc/gpufreq pin path is probably the right one" and
set `PIN_STYLE=ppm`. It is not right. MT6993 is new enough to have dropped both legacy
interfaces. So the exact failure mode the `PROFILE_VERIFIED` guard was written to prevent - a
root `echo` redirect into a missing /proc path exiting 0, and the run then reporting itself as
pinned - was waiting on the device the guard's own comment guessed was safe.
Consequence for P2: run bench.sh / session.sh with **`--no-pin`**, drive the pin with
`pin_device.sh`, and use `pin_device.sh check` in place of the `big_cur`/`little_cur`/`gpu_cur_khz`
integrity fields. `PROFILE_VERIFIED=1` in the two profiles certifies *the nodes and pins in those
files*, not that bench.sh can drive them; both files say so in a header block. Teaching bench.sh
a `PIN_STYLE` switch is the real fix and is not done here.
---
## Device 1 - `35d0befa`, Xiaomi 24129PN74C / SM8750 / Adreno 830v2
### CPU: which policy is "big", which is "little"
`ro.board.platform=sun`, `ro.soc.model=SM8750`. Two policies, and **no true little cluster** -
SM8750 is a 2+6 part:
```
policy0 cpus 0 1 2 3 4 5 gov=walt cpuinfo 384000..3532800 (6x performance)
avail: 384000 556800 748800 960000 1152000 1363200 1555200 1785600 1996800 2227200
2400000 2745600 2918400 3072000 3321600 3532800
policy6 cpus 6 7 gov=walt cpuinfo 1017600..4320000 (2x prime)
avail: 1017600 1209600 1401600 1689600 1958400 2246400 2438400 2649600 2841600 3072000
3283200 3513600 3801600 4089600 4204800 4320000
```
So **big = policy6, little = policy0**, and both protocol targets are *exact* members of
`scaling_available_frequencies` - no rounding:
| role | policy | target | chosen | rounding |
|---|---|---|---|---|
| big | policy6 | 1.96 GHz | **1958400** | exact |
| little | policy0 | 1.55 GHz | **1555200** | exact |
### CPU: how to pin, and does it hold
Method: `scaling_min_freq = scaling_max_freq = target`, **stock `walt` governor left alone**. No
performance or userspace governor is needed - once min == max the governor has no room, and not
switching it means nothing extra to restore. (`migov walt conservative powersave performance
schedutil` are all offered.)
Writes returned exit 0 and took effect immediately. Held for the whole window under load
(8 concurrent `yes > /dev/null`), sampled every 5 s:
```
t=1 p0_cur=1555200 p0_min=1555200 p0_max=1555200 p6_cur=1958400 p6_min=1958400 p6_max=1958400 gpuclk=1100000000 cpuss-0-0=47100
t=2 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=49400
t=3 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=49800
t=4 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=50200
t=5 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=51300
t=6 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=52100
t=7 p0_cur=1555200 (min/max same) p6_cur=1958400 (min/max same) gpuclk=1100000000 cpuss-0-0=52100
```
7 samples over 30 s, **zero drift** on either cluster or the GPU, while the SoC heated 47 -> 52 C.
Restore sequence (verified back to the stock snapshot):
```
policy6: scaling_min_freq=1017600 scaling_max_freq=2841600
policy0: scaling_min_freq=556800 scaling_max_freq=2745600
```
`pin_device.sh` writes min -> `cpuinfo_min_freq` first, then max, then min, so the order works
regardless of where stock sits relative to the target (writing min above the current max is
clamped by cpufreq).
### GPU: kgsl pwrlevels
```
gpu_model=Adreno830v2 num_pwrlevels=14 default_pwrlevel=12
gpu_available_frequencies: 1100000000 1050000000 967000000 900000000 832000000 734000000
660000000 607000000 525000000 443000000 389000000 342000000
222000000 160000000
stock: min_pwrlevel=12 max_pwrlevel=0 gpuclk=222000000
devfreq/governor=msm-adreno-tz devfreq/min_freq=160000000 devfreq/max_freq=1050000000
```
Pin: `echo 0 > /sys/class/kgsl/kgsl-3d0/min_pwrlevel` and the same to `max_pwrlevel`, collapsing
the range onto level 0. `gpuclk` read `1100000000` immediately and for the whole 30 s window.
Worth noting: **the devfreq route cannot reach the top step.** `devfreq/max_freq` is 1050000000,
one OPP below level 0, so a `devfreq/min_freq` + `max_freq` pin tops out at 1050 MHz. The
pwrlevel pin unlocks 1100 MHz. Restore: `min_pwrlevel=12`, `max_pwrlevel=0`.
**Unit trap.** bench.sh reports `gpu_cur_khz` from `$GPU_CURFREQ_NODE` via `awk` taking the last
field. The kgsl `gpuclk` node is in **Hz**; the MediaTek `current_freqency` node is in **kHz**.
So on this device the field bench.sh calls `gpu_cur_khz` is actually Hz. The profile carries
`GPU_PIN_HZ=1100000000` and `GPU_CURFREQ_UNIT=hz` alongside a schema-compatible
`GPU_PIN_KHZ=1100000`; compare against the former.
**GPU util node corrected.** The draft profile had `GPU_UTIL_NODE=/sys/class/kgsl/kgsl-3d0/gpubusy`.
Reading `gpubusy` prints `263821 1008410` - busy and total *cycles* - and bench.sh takes the
first field, so it would have logged raw cycle counts as a busy percentage. Changed to
`gpu_busy_percentage`, which prints `26 %` and whose first field is what bench.sh means.
### Thermal
84 zones; type strings are unique, so the match-first-zone-by-type loop in bench.sh is
unambiguous. Chosen: **`cpuss-0-0`** (`thermal_zone13`), the CPU-subsystem sensor for the 6-core
cluster that carries most of the load, and the hottest of the candidates under the load test:
| zone | idle | end of 30 s load |
|---|---|---|
| `cpuss-0-0` (13) | 34700 | 52100 |
| `cpuss-1-0` (20) | - | 49000 |
| `quiet_therm` (71) | 29835 | 32668 (skin sensor, not SoC) |
| `gpuss-0` (23) | 31200 | - |
The 40 C gate is reachable: it idles at 34.7 C, and was observed between 36.6 and 40.1 C across
the session.
---
## Device 2 - `3B159D009VZ00000`, Oppo PLG110 / MT6993 / Mali
### CPU: three policies, and the one that would have silently ruined the pin
`ro.soc.model=MT6993`. MT6993 is a **4+3+1** part:
```
policy0 cpus 0 1 2 3 gov=sugov_ext cpuinfo 300000..2700000 stock max 2100000
policy4 cpus 4 5 6 gov=sugov_ext cpuinfo 300000..3500000 stock max 3500000
policy7 cpu 7 gov=sugov_ext cpuinfo 300000..4210000 stock max 3200000
```
Decision: **big = policy4, little = policy0**, *and policy7 is pinned to the big target as well*.
policy4 rather than policy7 is the structural analogue of the 2-core policy6 on the Xiaomi - a
multi-core cluster the render and worker threads can share, rather than a single core they would
contend for. But policy7 cannot be left alone: an unpinned 4.21 GHz core defeats the entire pin
the moment the scheduler lands a hot thread on it, and nothing in the result JSON would show it.
Pinned, the device reads as 4 little + 4 big.
Neither target is an exact OPP here - both are the **nearest available step**:
| role | policy | target | candidates | chosen | offset |
|---|---|---|---|---|---|
| big | policy4 | 1958400 | 1900000 (-58400), 2000000 (+41600) | **2000000** | +2.1% |
| little | policy0 | 1555200 | 1500000 (-55200), 1600000 (+44800) | **1600000** | +2.9% |
| extra | policy7 | 1958400 | 2000000 is an exact member | **2000000** | +2.1% |
Carry that 2-3% offset when comparing absolute per-thread CPU cost against the Xiaomi.
### CPU: how to pin, and does it hold
Same method (min = max = target, stock `sugov_ext` left alone). All six writes returned 0.
Held under the same 8-way load, sampled every 5 s, in `cur/min-max` form:
```
t=1 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=52833
t=2 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=56859
t=3 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=58063
t=4 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=59661
t=5 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=60353
t=6 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=61354
t=7 p0=1600000/1600000-1600000 p4=2000000/2000000-2000000 p7=2000000/2000000-2000000 gpu="0 1716000" soc_max=62025
```
7 samples over 30 s, **zero drift** on all three policies, SoC 52.8 -> 62.0 C. The ColorOS
performance daemons did not contend for the cpufreq nodes during the window.
**Read permissions.** `scaling_governor`, `scaling_min_freq` and `scaling_max_freq` are `0660
system:system` here. A plain `adb shell cat .../scaling_governor` answers `Permission denied`,
which a careless parser reads as an empty governor rather than a failure. Every read goes through
su. Note `scaling_cur_freq` and `scaling_available_frequencies` *are* world-readable, which makes
the trap worse - a script can look like it is working.
### GPU: gpufreqv2, not legacy gpufreq
57 working OPPs, index 0 fastest:
```
/proc/gpufreqv2/gpu_working_opp_table
[00] freq: 1716000 [01] 1690000 [02] 1664000 ... [56] 390000 (kHz)
```
Pin: `echo 0 > /proc/gpufreqv2/fix_target_opp_index`; restore: `echo -1`. Evidence:
```
before : [GPUFREQ-DEBUG] fix GPU/STACK OPP index is disabled
after 0 : [GPUFREQ-DEBUG] fix GPU/STACK OPP index: 0/0
/sys/kernel/ged/hal/current_freqency = "0 1716000" (index 0, 1716000 kHz)
after -1: [GPUFREQ-DEBUG] fix GPU/STACK OPP index is disabled
```
**Deliberately not the ged `custom_boost_gpu_freq` / `custom_upbound_gpu_freq` pair**, the usual
MTK route, for two reasons both visible in the nodes:
1. On this kernel they take an **OPP index, not kHz**. Stock reads `boost=56`, `upbound=0` - a
floor at the slowest OPP and a ceiling at the fastest. Porting the kHz-valued
`GPU_PIN_KHZ=902000` from odinlite.env into them would write garbage.
2. Reading them prints a request log showing
`/odm/bin/hw/vendor.oplus.hardware.urcc-service` (pid 1534) writing the same nodes. A pin
there is one voter among several. `fix_target_opp_index` has no such contention.
### Thermal
87 zones, unique types. Chosen: **`soc_max`** (`thermal_zone14`) - the SoC-wide max aggregate,
which is what "SoC temperature" means for a gate, and the strictest candidate.
Cooldown series after the load test (screen awake, idle), 15 s apart:
```
soc_max 39094 37177 36806 35745 36954 36312
soc-top0 36361 35082 34641 34226 34079 34031
ap_ntc 36414 35034 34342 33965 33746 33566
```
It settles to **36-37 C**, so `THERMAL_START_MAX_MC=40000` is reachable - but only just, and it
was still reading 42.4 C several minutes after ordinary use. Expect this gate to actually wait,
unlike the one on the Xiaomi. Under load it went 52.8 -> 62.0 C in 30 s.
---
## Every node read or written
Read-only (both devices): `/sys/devices/system/cpu/cpufreq/policy*/{affected_cpus,
scaling_governor,scaling_available_governors,scaling_available_frequencies,cpuinfo_min_freq,
cpuinfo_max_freq,scaling_cur_freq}`, `/sys/class/thermal/thermal_zone*/{type,temp}`,
`getprop`, `id`.
Read-only, Xiaomi: `/sys/class/kgsl/kgsl-3d0/{gpu_model,num_pwrlevels,default_pwrlevel,
thermal_pwrlevel,max_gpuclk,gpuclk,clock_mhz,max_clock_mhz,min_clock_mhz,
gpu_available_frequencies,freq_table_mhz,gpubusy,gpu_busy_percentage,throttling,temp}`,
`/sys/class/kgsl/kgsl-3d0/devfreq/{name,governor,available_governors,available_frequencies,
cur_freq,min_freq,max_freq,target_freq,gpu_load,mod_percent}`, `/sys/class/devfreq/`.
Read-only, Oppo: `/proc/gpufreqv2/{gpufreq_status,gpu_working_opp_table,fix_target_opp_index}`,
`/sys/kernel/ged/hal/{current_freqency,custom_boost_gpu_freq,custom_upbound_gpu_freq,
gpu_utilization,total_gpu_freq_level_count}`.
**Written** (all restored):
| device | node | stock | pinned to | restored to |
|---|---|---|---|---|
| Xiaomi | `policy6/scaling_min_freq` | 1017600 | 1958400 | 1017600 |
| Xiaomi | `policy6/scaling_max_freq` | 2841600 | 1958400 | 2841600 |
| Xiaomi | `policy0/scaling_min_freq` | 556800 | 1555200 | 556800 |
| Xiaomi | `policy0/scaling_max_freq` | 2745600 | 1555200 | 2745600 |
| Xiaomi | `kgsl-3d0/min_pwrlevel` | 12 | 0 | 12 |
| Xiaomi | `kgsl-3d0/max_pwrlevel` | 0 | 0 | 0 |
| Oppo | `policy0/scaling_min_freq` | 300000 | 1600000 | 300000 |
| Oppo | `policy0/scaling_max_freq` | 2100000 | 1600000 | 2100000 |
| Oppo | `policy4/scaling_min_freq` | 300000 | 2000000 | 300000 |
| Oppo | `policy4/scaling_max_freq` | 3500000 | 2000000 | 3500000, then 3200000 by the vendor daemon - see below |
| Oppo | `policy7/scaling_min_freq` | 300000 | 2000000 | 300000 |
| Oppo | `policy7/scaling_max_freq` | 3200000 | 2000000 | 3200000 |
| Oppo | `/proc/gpufreqv2/fix_target_opp_index` | disabled | 0 | -1 (disabled) |
No governor was ever written on either device; both stayed on their vendor governor (`walt`,
`sugov_ext`) throughout. Nothing else was written - no `settings put`, no thermal-engine
tampering, no app installs.
### The one restore that does not settle at its stock value
On the Oppo, `policy4/scaling_max_freq` was restored to its snapshotted 3500000 and a ColorOS
daemon lowered it to **3200000** by itself within seconds. This is correct behaviour - the range
is handed back and the vendor takes over - but it means **stock maxima on this device are not
constants**.
It caught a real bug in the first version of `pin_device.sh check`, which classified each node as
at-pin / at-stock / neither and reported DRIFT for anything else: it declared DRIFT on a device
that had been released perfectly. The classifier now asserts only against *our* pins (pinned
means `min == max == pin`; anything else is not pinned, with a `min == max` at some other value
flagged as an external clamp, which is equally fatal to comparability). The hardcoded stock
values still drive the restore path, where handing the range back is all they have to do.
---
## `pin_device.sh`
Usage: `pin_device.sh <serial> pin|unpin|check`. Pure adb + su; no profile sourcing, no bench.sh.
The device table (policies, pins, stock values, GPU style, thermal zone type) is embedded so it
has no inputs to get wrong. All three actions print the live big/little/GPU frequencies,
governors and gate temperature, so the output of `pin` and of `unpin` is itself the before/after
evidence.
Exit codes: **0** PINNED (every pinned node at its pin and every live frequency equal to it),
**1** DRIFT (partly pinned, externally clamped, or a live frequency has left its pin - discard any
overlapping run), **2** UNPINNED (no node is at a pin this script set). UNPINNED is non-zero on
purpose, so that `pin_device.sh X check && measure` cannot silently measure an unpinned device.
Three things it does that are load-bearing:
- **Write order** is min -> `cpuinfo_min_freq`, then max -> target, then min -> target. Setting
min above the current max is clamped by cpufreq, so a naive two-write pin half-applies
depending on where stock sits relative to the target.
- **Stock values are hardcoded, not sampled at pin time.** A restore that read "stock" off an
already-pinned device would make the pin permanent.
- **The device-side read emits raw node contents and parses locally.** The read is delivered as a
single-quoted `su -c` argument, so a single quote inside it (an awk program, a sed expression)
closes that quoting and the whole read returns nothing. That bug was hit during this work: the
Oppo `check` printed `<unread>` for every field and reported DRIFT on a correctly pinned
device. A zero-key read is now a hard error (exit 66), not a verdict.
### Final `check` on both devices, unpinned state
```
$ pin_device.sh 35d0befa check
== check ==
device : Xiaomi 24129PN74C / SM8750 / Adreno 830v2 (35d0befa)
big policy6 gov=walt cur=2841600 min=1017600 max=2841600 (pin 1958400 / stock 1017600-2841600)
little policy0 gov=walt cur=2400000 min=556800 max=2745600 (pin 1555200 / stock 556800-2745600)
gpu kgsl-3d0 pwrlevel=0..12 freq=222000000 Hz busy=7% (pin lvl 0 = 1100000000 Hz / stock lvl 0..12)
thermal cpuss-0-0 38200 mC = 38.2 C (/sys/class/thermal/thermal_zone13)
VERDICT: UNPINNED - none of the 3 nodes is at a pin this script set; the vendor
governors have their range back and nothing this script writes is in effect.
rc=2
```
```
$ pin_device.sh 3B159D009VZ00000 check
== check ==
device : Oppo PLG110 / MT6993 / Mali (gpufreqv2) (3B159D009VZ00000)
big policy4 gov=sugov_ext cur=300000 min=300000 max=3200000 (pin 2000000 / stock 300000-3500000)
little policy0 gov=sugov_ext cur=2100000 min=300000 max=2100000 (pin 1600000 / stock 300000-2100000)
extra policy7 gov=sugov_ext cur=2750000 min=300000 max=3200000 (pin 2000000 / stock 300000-3200000)
gpu gpufreqv2 fix_opp=off freq=1508000 kHz busy=0% (pin idx 0 = 1716000 kHz / stock off)
thermal soc_max 45292 mC = 45.3 C (/sys/class/thermal/thermal_zone14)
VERDICT: UNPINNED - none of the 4 nodes is at a pin this script set; the vendor
governors have their range back and nothing this script writes is in effect.
note: big(policy4) unpinned, range 300000-3200000 (recorded stock 300000-3500000, which vendor daemons move)
rc=2
```
Both report their stock governors (`walt`, `sugov_ext`), the GPU back on DVFS
(pwrlevel 0..12 / `fix_opp=off`), and cpufreq ranges wide open. Nothing written is still in
effect.
For the record, the pinned half of the same cycle on each device:
```
$ pin_device.sh 35d0befa pin (after pin)
big policy6 gov=walt cur=1958400 min=1958400 max=1958400
little policy0 gov=walt cur=1555200 min=1555200 max=1555200
gpu kgsl-3d0 pwrlevel=0..0 freq=1100000000 Hz
VERDICT: PINNED - all 3 pinned nodes at their pins, live frequencies match. rc=0
$ pin_device.sh 3B159D009VZ00000 pin (after pin)
big policy4 gov=sugov_ext cur=2000000 min=2000000 max=2000000
little policy0 gov=sugov_ext cur=1600000 min=1600000 max=1600000
extra policy7 gov=sugov_ext cur=2000000 min=2000000 max=2000000
gpu gpufreqv2 fix_opp=0 freq=1716000 kHz
VERDICT: PINNED - all 4 pinned nodes at their pins, live frequencies match. rc=0
```
---
## What is NOT verified
1. **bench.sh cannot drive these pins.** Not fixed here. The profiles are verified; the harness
integration is not. Use `--no-pin` plus `pin_device.sh` until a `PIN_STYLE` switch exists.
2. **The Oppo GPU pin was confirmed at the DVFS-request level, not under graphics load.** The
pinned window was a CPU load, so the GPU was power-collapsed throughout (`gpufreq_status`
showed `PowerCount: 0`). `current_freqency` reported OPP index 0 / 1716000 kHz for the whole
window - which is exactly the field bench.sh samples as `gpu_cur_khz`, so the check bench.sh
would perform does pass - but nobody has watched this GPU actually clock 1716 MHz while
rendering. Related readout quirk: while powered down, `/proc/gpufreqv2/gpufreq_status` keeps
printing `Freq: 26000` for the fixed OPP entry. That is a parked rail, not the pin failing;
`current_freqency` is the node to believe.
3. **Pin durability over a full bench window is untested.** The verification window is 30 s of
synthetic CPU load, which is what the README defines. Neither device was held pinned through a
180 s warmup plus a 30-sample measurement with the game running, where a vendor game-boost
daemon has far more reason to intervene than it did here. Running `pin_device.sh check` after
each run is what closes that gap; treat `rc=1` as "discard the run".
4. **Thermal gate behaviour under the real protocol is untested.** 40 C is reachable on both, but
on the Oppo it idles only 3-4 C below the gate, so back-to-back paired runs will spend real
time waiting. Nobody has measured how long.
5. **No reboot test.** All pins were applied and released within one session. Nothing here
persists across a reboot by construction (every node written is volatile sysfs or procfs), but
that was not demonstrated.
6. **The `walt` governor and the Xiaomi thermal daemon were never provoked.** 30 s at 52 C is
well short of the thermal-limit regime where `thermal_pwrlevel` (read 0 throughout) would
start capping the GPU.
+10
View File
@@ -82,3 +82,13 @@ them: the pre/post pair must show the same scene, or the run is invalid.
- **F3 off** for standard numbers (the F3 debug overlay multiplies per-draw
overhead and skews backends differently).
- The FPS overlay itself must be ON (it is what produces the FCLFPS lines).
## Pinning on the two MGPipe campaign devices (2026-09-07)
`bench.sh`'s `pin_freqs()` writes `/proc/ppm/policy/*` and `/proc/gpufreq/gpufreq_opp_freq`; **neither
path exists on `35d0befa` (SM8750) nor on `3B159D009VZ00000` (MT6993, which dropped both legacy
interfaces for `/proc/gpufreqv2/`)**. For those two devices run `bench.sh --no-pin` and pin with
`tools/device_bench/pin_device.sh <serial> pin|unpin|check` (pure adb + su; exit 0 PINNED, 1 DRIFT,
2 UNPINNED, so `check && measure` cannot measure unpinned). The profiles under `devices/` carry
`PROFILE_VERIFIED=1` for the nodes and pins named in the file, not for `bench.sh`'s ability to drive
them; the verification evidence is `docs/Disaggregated/devices/pin-verification-2026-09-07.md`.
+125 -35
View File
@@ -1,24 +1,134 @@
# Device profile: Oppo / ColorOS, Mali GPU, adb serial 3B159D009VZ00000.
# Device profile: Oppo PLG110 / ColorOS, MediaTek MT6993 (Dimensity 9500),
# Mali (gpufreqv2 + ged), adb serial 3B159D009VZ00000.
#
# The second of the two devices the disaggregation campaign is measured on (the other is
# devices/xiaomi-adreno830.env). Same purpose: keep the pinning and thermal protocol in the
# repository rather than in one operator's shell history.
#
# ============================ NOT YET DEVICE-VERIFIED ============================
# PROFILE_VERIFIED=0, and bench.sh / session.sh refuse to run against it - or against a profile
# that omits the key - unless --allow-unverified-profile is passed. (profile.sh only records a
# simpleperf profile and pins nothing, so it carries no such guard.) This part is a MediaTek SoC, so unlike the Adreno
# profile the harness's existing /proc/ppm + /proc/gpufreq pin path is probably the right one -
# but "probably" is exactly the state a measurement profile must not ship in. The cluster
# indices, the available OPPs, the top GPU OPP and the thermal zone TYPE all differ between
# MediaTek generations, and odinlite.env's values are for an MT6877, not for this device.
# ============================== DEVICE-VERIFIED ==============================
# Read off the device on 2026-09-07 and confirmed against one pinned window: pins written,
# 8 concurrent busy loops as load, scaling_cur_freq for all three policies and the ged GPU
# frequency sampled 7 times over 30 s, every sample equal to the pin, then unpinned and every
# node confirmed back at its stock value. Evidence is in ../REPORT.md.
#
# To promote it: read the four TODO fields off the device
# (`cat /sys/devices/system/cpu/cpufreq/policy*/scaling_available_frequencies`,
# `cat /proc/gpufreq/gpufreq_opp_dump`, `for tz in /sys/class/thermal/thermal_zone*; do
# echo "$tz $(cat $tz/type)"; done`), run one pinned window, check big_cur/little_cur/gpu_cur_khz
# in the result JSON against the pins, then set PROFILE_VERIFIED=1.
# READ THIS BEFORE USING IT WITH bench.sh / session.sh
# ----------------------------------------------------
# The old draft of this profile said PIN_STYLE=ppm on the reasoning that "this is a MediaTek
# SoC, so the harness's existing /proc/ppm + /proc/gpufreq pin path is probably the right one".
# It is not. On MT6993:
# /proc/ppm/ does not exist
# /proc/gpufreq/ does not exist (superseded by /proc/gpufreqv2/)
# So bench.sh's pin_freqs() writes four ppm lines and one gpufreq line into nothing, exits 0,
# and reports a run it believes was pinned - the failure the PROFILE_VERIFIED guard describes,
# on the device the guard's own comment guessed was safe. PROFILE_VERIFIED=1 below certifies
# THE NODES AND THE PINS IN THIS FILE, not that bench.sh can drive them.
#
# Until bench.sh grows a PIN_STYLE switch, drive the pin out of band:
# tools/device_bench/pin_device.sh 3B159D009VZ00000 pin
# tools/device_bench/bench.sh --device <this> --backend X --no-pin
# tools/device_bench/pin_device.sh 3B159D009VZ00000 check # non-zero on drift
# tools/device_bench/pin_device.sh 3B159D009VZ00000 unpin
# =============================================================================
PROFILE_VERIFIED=1
PIN_STYLE=mtk-gpufreqv2
DEVICE_SERIAL=3B159D009VZ00000
# --- CPU -------------------------------------------------------------------------------------
# MT6993 is a 4+3+1 part, THREE policies - which is the trap on this device:
# policy0 = cpus 0-3, 300000..2700000 kHz <- the protocol's "little"
# policy4 = cpus 4-6, 300000..3500000 kHz <- the protocol's "big"
# policy7 = cpu 7, 300000..4210000 kHz <- must ALSO be pinned, see below
# policy4 is chosen as "big" over the single-core policy7 because it is the structural analogue
# of the Xiaomi's 2-core policy6: a multi-core cluster the game's render and worker threads can
# actually share. But policy7 cannot simply be left alone - an unpinned 4.21 GHz core silently
# defeats the whole pin the moment the scheduler puts a hot thread on it, so the protocol pins
# it to the SAME target as policy4 and the device then behaves as 4 little + 4 big.
#
# Neither campaign target is an exact OPP here; both are the NEAREST available step:
# big 1958400 -> policy4 has 2000000 (D +41600, +2.1%) vs 1900000 (D -58400) -> 2000000
# little 1555200 -> policy0 has 1600000 (D +44800, +2.9%) vs 1500000 (D -55200) -> 1600000
# policy7 offers 2000000 exactly, so the big pin transfers to it unrounded.
# Record this 2-3% offset when comparing absolute per-thread CPU cost against the Xiaomi.
#
# Pin method: scaling_min_freq = scaling_max_freq = target, stock `sugov_ext` governor left in
# place. Verified to hold for 30 s under load on all three policies with no drift; the ColorOS
# performance daemons did not contend for the cpufreq nodes during the window. The nodes are
# 0660 system:system, so every read AND write has to go through su - a plain `adb shell cat
# scaling_governor` returns "Permission denied" and an unwary script reads that as an empty
# governor.
CPU_BIG_POLICY=policy4
CPU_BIG_FREQ=2000000
CPU_LITTLE_POLICY=policy0
CPU_LITTLE_FREQ=1600000
# The third policy. Pinned to CPU_BIG_FREQ; not reported as big_cur, but if it is not pinned
# the run is not pinned.
CPU_EXTRA_POLICY=policy7
CPU_EXTRA_FREQ=2000000
# Stock values sampled 2026-09-07 (screen on, idle), for the restore path. Note that policy0's
# and policy7's stock maxima (2100000 / 3200000) are BELOW cpuinfo_max_freq (2700000 / 4210000):
# ColorOS is already holding a limit there, so restore to these, not to cpuinfo_max_freq.
# Governor is `sugov_ext` on all three and is never written.
#
# THESE ARE NOT CONSTANTS. Observed directly: an unpin restored policy4 to 300000-3500000 and
# within seconds a ColorOS daemon had lowered its max to 3200000 by itself. So they are good
# enough to hand the range back to the governor with (which is all the restore path needs), but
# a verification that asserts "the device is released" by exact-matching them will report a
# false failure. pin_device.sh therefore asserts only against OUR pins, never against these.
CPU_BIG_STOCK_MIN=300000
CPU_BIG_STOCK_MAX=3500000
CPU_LITTLE_STOCK_MIN=300000
CPU_LITTLE_STOCK_MAX=2100000
CPU_EXTRA_STOCK_MIN=300000
CPU_EXTRA_STOCK_MAX=3200000
# --- GPU -------------------------------------------------------------------------------------
# 57 working OPPs (indices 0..56), index 0 fastest:
# [00] 1716000 kHz ... [56] 390000 kHz
# Pin at the top OPP with the gpufreqv2 hard override:
# echo 0 > /proc/gpufreqv2/fix_target_opp_index # pin (fixes GPU and STACK to index 0)
# echo -1 > /proc/gpufreqv2/fix_target_opp_index # restore DVFS
# Confirmed: after the write the node reads "fix GPU/STACK OPP index: 0/0" and
# /sys/kernel/ged/hal/current_freqency reads "0 1716000"; after `echo -1` it reads
# "fix GPU/STACK OPP index is disabled".
#
# Deliberately NOT the ged custom_boost_gpu_freq / custom_upbound_gpu_freq pair, even though it
# is the usual MTK route. Two reasons, both visible in the node: (a) on this kernel they take an
# OPP INDEX, not a kHz value (stock reads boost=56, upbound=0 - a floor at the slowest OPP and a
# ceiling at the fastest), so an operator porting odinlite's kHz-valued GPU_PIN would write
# garbage; (b) `cat` on them prints a request log showing /odm/bin/hw/vendor.oplus.hardware.urcc-
# service writing the same nodes, so a pin there is one voter among several and can be revised
# by the vendor daemon. fix_target_opp_index has no such contention.
GPU_PIN_OPP_INDEX=0
GPU_PIN_KHZ=1716000
GPU_CURFREQ_UNIT=khz
GPU_STOCK_BOOST_INDEX=56
GPU_STOCK_UPBOUND_INDEX=0
# ged nodes are unchanged from odinlite and parse the same way: gpu_utilization prints
# "7 0 100" ($1 = busy%), current_freqency prints "<opp_index> <freq_khz>" ($NF = kHz).
GPU_UTIL_NODE=/sys/kernel/ged/hal/gpu_utilization
GPU_CURFREQ_NODE=/sys/kernel/ged/hal/current_freqency
# CAVEAT on the GPU verification: the pinned window was a CPU load, so the GPU was power-
# collapsed throughout (gpufreq_status showed PowerCount: 0). The pin was therefore confirmed at
# the DVFS-request level - ged reported OPP index 0 / 1716000 kHz for the whole window - and not
# by observing the GPU actually clocking 1716 MHz under graphics load. Note also that while
# powered down, /proc/gpufreqv2/gpufreq_status keeps printing "Freq: 26000" for the fixed OPP
# entry; that is a parked-rail readout, not the pin failing. current_freqency is the node to
# believe, and it is the one bench.sh already samples.
# --- Thermal ---------------------------------------------------------------------------------
# 87 thermal zones. `soc_max` (thermal_zone14) is the SoC-wide max aggregate - the right meaning
# for a "SoC temperature" gate, and the strictest of the candidates. Type strings are unique, so
# bench.sh's match-first-zone-by-type loop resolves it unambiguously.
# 40 C is reachable but only just: it settles to 36-37 C at rest with the screen awake, and it
# was still reading 42.4 C several minutes after ordinary use. Expect the gate to actually wait.
# Under the load test it went 52.8 -> 62.0 C in 30 s.
THERMAL_ZONE_TYPE=soc_max
THERMAL_START_MAX_MC=40000
# ColorOS traps that belong with this device, and cost a run each when forgotten:
# * the first install of a not-yet-installed package blocks on
# com.oplus.appdetail InstallGuideActivity until "continue install" is tapped
@@ -26,24 +136,4 @@
# * a foreign-signed APK has to be uninstalled before a rebuild will install;
# * pass MSYS_NO_PATHCONV=1 on every adb invocation from Git Bash, or a /data/... argument is
# rewritten into a Windows path.
# =================================================================================
PROFILE_VERIFIED=0
PIN_STYLE=ppm
DEVICE_SERIAL=3B159D009VZ00000
# Campaign protocol constants (perf-test-protocol): big 1.96 GHz, little 1.55 GHz, GPU at its
# top OPP, 40 C start gate. As above, the kHz values are the protocol's targets and the nearest
# actual OPP has to be confirmed on the device.
CPU_BIG_POLICY=TODO_VERIFY_ON_DEVICE
CPU_BIG_FREQ=1958000
CPU_LITTLE_POLICY=TODO_VERIFY_ON_DEVICE
CPU_LITTLE_FREQ=1550000
# MediaTek legacy gpufreq, same node family as odinlite. The top OPP is device-specific.
GPU_PIN_KHZ=
GPU_UTIL_NODE=/sys/kernel/ged/hal/gpu_utilization
GPU_CURFREQ_NODE=/sys/kernel/ged/hal/current_freqency
THERMAL_ZONE_TYPE=TODO_VERIFY_ON_DEVICE
THERMAL_START_MAX_MC=40000
# * every cpufreq read needs su (0660 system:system), unlike the Xiaomi where reads are open.
+85 -37
View File
@@ -1,55 +1,103 @@
# Device profile: Xiaomi, Snapdragon 8 Elite (Adreno 830), adb serial 35d0befa.
# Device profile: Xiaomi 24129PN74C, Snapdragon 8 Elite (SM8750, "sun"),
# Adreno 830v2, adb serial 35d0befa.
#
# One of the two devices the disaggregation campaign is measured on (the other is
# devices/oppo-mali.env). It exists so that the pinning and thermal protocol the campaign
# actually runs is written down in the repository instead of living in one operator's shell
# history, and so that a `--device` argument names something reviewable.
#
# ============================ NOT YET DEVICE-VERIFIED ============================
# PROFILE_VERIFIED=0 below, and bench.sh / session.sh REFUSE to run against a profile that says
# so - or that omits the key - unless --allow-unverified-profile is passed. (profile.sh is not in
# that list: it records a simpleperf profile and pins nothing, so it has nothing to pin wrongly.) Two of the values here are
# protocol constants that are known (the campaign pins big 1.96 GHz / little 1.55 GHz and gates
# at 40 C), but the sysfs node names and the exact available OPPs are NOT: this is a Qualcomm
# part and the harness was written against MediaTek, where the pin goes through
# /proc/ppm/policy/hard_userlimit_* and the GPU through /proc/gpufreq/gpufreq_opp_freq. Neither
# path exists on this SoC - Adreno pins through /sys/class/kgsl/kgsl-3d0/devfreq/{min,max}_freq
# and its cpufreq policies are not policy6/policy0.
# ============================== DEVICE-VERIFIED ==============================
# Read off the device on 2026-09-07 and confirmed against one pinned window: pins written,
# 8 concurrent busy loops as load, scaling_cur_freq / gpuclk sampled 7 times over 30 s, every
# sample equal to the pin, then unpinned and every node confirmed back at its stock value.
# Evidence is in ../REPORT.md.
#
# A profile that quietly wrote MediaTek paths on this device would be the worst outcome
# available: `su -c 'echo ... > /proc/ppm/...'` fails silently, bench.sh would report a run it
# believes was pinned, and the pin-integrity fields it samples at window end would be the only
# clue. So the unknown fields are left EMPTY and marked, rather than guessed, and the refusal is
# the mechanism that keeps them from being used before somebody has read them off the device.
# READ THIS BEFORE USING IT WITH bench.sh / session.sh
# ----------------------------------------------------
# bench.sh's pin_freqs()/unpin_freqs() are MediaTek-only: they write /proc/ppm/policy/* and
# /proc/gpufreq/gpufreq_opp_freq. NEITHER PATH EXISTS ON THIS DEVICE, and `su -c 'echo ... >
# /proc/...'` against a missing path fails without a non-zero exit - which is the exact silent
# failure the PROFILE_VERIFIED guard was built to prevent. PROFILE_VERIFIED=1 below certifies
# THE NODES AND THE PINS IN THIS FILE, not that bench.sh can drive them.
#
# To promote this profile: fill in the four TODO fields from the device
# (`cat /sys/devices/system/cpu/cpufreq/policy*/scaling_available_frequencies`,
# `ls /sys/class/kgsl/kgsl-3d0/devfreq/`, `for tz in /sys/class/thermal/thermal_zone*; do
# echo "$tz $(cat $tz/type)"; done`), teach bench.sh the Qualcomm pin path, run one pinned
# window, check big_cur/little_cur/gpu_cur_khz in the result JSON against the pins, and only
# then set PROFILE_VERIFIED=1 in the same commit as the bench.sh change.
# =================================================================================
PROFILE_VERIFIED=0
# So until bench.sh grows a PIN_STYLE switch, drive the pin out of band:
# tools/device_bench/pin_device.sh 35d0befa pin # before the run
# tools/device_bench/bench.sh --device <this> --backend X --no-pin
# tools/device_bench/pin_device.sh 35d0befa check # AFTER, exits non-zero on drift
# tools/device_bench/pin_device.sh 35d0befa unpin
# `--no-pin` is what keeps bench.sh from writing MediaTek paths into the void and labelling the
# result pinned; pin_device.sh check is what replaces the pin-integrity fields it would have
# sampled. (bench.sh's read_temp and its GPU-busy sampling ARE portable - they only read
# /sys/class/thermal and $GPU_UTIL_NODE - so those fields stay meaningful.)
# =============================================================================
PROFILE_VERIFIED=1
PIN_STYLE=qualcomm-kgsl
DEVICE_SERIAL=35d0befa
# Campaign protocol constants (perf-test-protocol): big 1.96 GHz, little 1.55 GHz, GPU at its
# top OPP, and a 40 C start gate. The kHz values are the protocol's targets; the nearest actual
# OPP has to be read off the device before they are used, because a cpufreq write that names a
# frequency the policy does not offer is rounded silently.
CPU_BIG_POLICY=TODO_VERIFY_ON_DEVICE
# --- CPU -------------------------------------------------------------------------------------
# SM8750 is a 2+6 part with NO true little cluster:
# policy0 = cpus 0-5, 384000..3532800 kHz (6x performance) <- the protocol's "little"
# policy6 = cpus 6-7, 1017600..4320000 kHz (2x prime) <- the protocol's "big"
# The campaign's targets (big 1.96 GHz / little 1.55 GHz) are both EXACT members of
# scaling_available_frequencies here, so no rounding is involved:
# policy6 avail: ... 1689600 [1958400] 2246400 ...
# policy0 avail: ... 1363200 [1555200] 1785600 ...
# Pin method: scaling_min_freq = scaling_max_freq = target, leaving the stock `walt` governor
# in place. Verified to hold: 30 s under load, zero drift, and the walt governor has no say
# once min == max. A `performance` governor is offered but is not needed and would have to be
# restored by name, so the min/max clamp is the lighter touch.
# Restore: write the stock min/max back (see CPU_*_STOCK_* below). Order matters on policy0 -
# its stock min (556800) is BELOW its pinned value, so lower the max first, then the min, or
# the min write is clamped against the still-pinned max.
CPU_BIG_POLICY=policy6
CPU_BIG_FREQ=1958400
CPU_LITTLE_POLICY=TODO_VERIFY_ON_DEVICE
CPU_LITTLE_POLICY=policy0
CPU_LITTLE_FREQ=1555200
# Adreno pins through the kgsl devfreq knobs, not /proc/gpufreq. Left empty deliberately: see
# the block above.
GPU_PIN_KHZ=
GPU_UTIL_NODE=/sys/class/kgsl/kgsl-3d0/gpubusy
# Stock values sampled 2026-09-07, for the restore path. Governor is `walt` on both policies
# and is never written, so it needs no restore.
CPU_BIG_STOCK_MIN=1017600
CPU_BIG_STOCK_MAX=2841600
CPU_LITTLE_STOCK_MIN=556800
CPU_LITTLE_STOCK_MAX=2745600
# --- GPU -------------------------------------------------------------------------------------
# Adreno pins through the kgsl pwrlevel knobs, not /proc/gpufreq. 14 pwrlevels, index 0 fastest:
# [0] 1100 MHz ... [12] 222 MHz [13] 160 MHz
# Stock range is min_pwrlevel=12 / max_pwrlevel=0, and the devfreq governor (msm-adreno-tz)
# holds its own max_freq at 1050000000 - i.e. stock DVFS never reaches level 0. Writing
# echo 0 > /sys/class/kgsl/kgsl-3d0/min_pwrlevel
# echo 0 > /sys/class/kgsl/kgsl-3d0/max_pwrlevel
# collapses the range onto level 0 and DOES unlock the 1100 MHz step: gpuclk read 1100000000
# immediately and for the whole 30 s window. Prefer this over the devfreq min_freq/max_freq
# pair, which cannot express the top step.
GPU_PIN_PWRLEVEL=0
GPU_PIN_KHZ=1100000
# UNIT TRAP, do not delete: bench.sh reports `gpu_cur_khz` from $GPU_CURFREQ_NODE via
# `awk '{print $NF}'`. On MediaTek that node yields kHz; kgsl's gpuclk yields Hz. So on this
# device the field bench.sh calls gpu_cur_khz is actually Hz, and the value to compare it
# against is GPU_PIN_HZ, not GPU_PIN_KHZ. GPU_PIN_KHZ is carried only to keep the schema
# identical across profiles.
GPU_PIN_HZ=1100000000
GPU_CURFREQ_UNIT=hz
GPU_STOCK_MIN_PWRLEVEL=12
GPU_STOCK_MAX_PWRLEVEL=0
# NOT gpubusy. `cat gpubusy` prints "<busy_cycles> <total_cycles>" and bench.sh takes $1, which
# would record raw busy cycles as a percentage. gpu_busy_percentage prints "26 %", whose $1 is
# the percentage bench.sh means.
GPU_UTIL_NODE=/sys/class/kgsl/kgsl-3d0/gpu_busy_percentage
GPU_CURFREQ_NODE=/sys/class/kgsl/kgsl-3d0/gpuclk
# Thermal gate: 40 C, the campaign's threshold. The zone TYPE differs per SoC and bench.sh
# matches on it by name, so it has to be read off the device.
THERMAL_ZONE_TYPE=TODO_VERIFY_ON_DEVICE
# --- Thermal ---------------------------------------------------------------------------------
# 84 thermal zones; `cpuss-0-0` (thermal_zone13) is the CPU-subsystem sensor for the 6-core
# cluster that carries most of the load, and was the hottest of the candidates under the load
# test (47.1 -> 52.1 C while gpuss/quiet_therm stayed near 30 C). Type strings are unique on
# this device, so bench.sh's match-first-zone-by-type loop resolves it unambiguously.
# 40 C is reachable at rest: it idles at 34.7 C.
THERMAL_ZONE_TYPE=cpuss-0-0
THERMAL_START_MAX_MC=40000
# Xiaomi has no fan; bench.sh's `settings put global fan_mode 3` is an Odin Lite line and is a
# harmless no-op here.
+366
View File
@@ -0,0 +1,366 @@
#!/usr/bin/env bash
# pin_device.sh <serial> pin|unpin|check
#
# Frequency pinning for the two disaggregation-campaign devices, standalone: pure adb + su, no
# sourcing of a device profile, no dependency on bench.sh. It exists because bench.sh's
# pin_freqs()/unpin_freqs() are MediaTek-legacy-only (/proc/ppm + /proc/gpufreq) and NEITHER of
# these two devices has those paths - on both of them bench.sh's pin writes into nothing and
# exits 0, which is precisely the silent "the run looks pinned but is not" failure the
# PROFILE_VERIFIED guard was built to catch. So: pin here, run bench.sh with --no-pin, and use
# `check` in place of the pin-integrity fields bench.sh would otherwise have sampled.
#
# ./pin_device.sh 35d0befa pin
# ./bench.sh --device devices/xiaomi-adreno830.env --backend magma --no-pin
# ./pin_device.sh 35d0befa check || echo "PIN DRIFTED - discard this run"
# ./pin_device.sh 35d0befa unpin
#
# All three actions print the live big/little/GPU frequencies, the governors and the gate
# temperature, so the output of `pin` and of `unpin` is itself the before/after evidence.
#
# check exit codes (deliberately three-valued: "no pin at all" and "a pin that slipped" are
# different facts, and collapsing them to one non-zero would hide which one happened):
# 0 PINNED - every pinned node is at its pin AND every live frequency equals it
# 1 DRIFT - the device is partly pinned, or a live frequency has left its pin. The
# dangerous state: a run overlapping this is not comparable. Discard it.
# 2 UNPINNED- no node is at a pin this script set; the vendor governors have their range
# back. That is the correct state to leave a device in, and it is still non-zero
# so that `pin_device.sh X check && measure` cannot silently measure unpinned.
# Note this is asserted against OUR pins, not against the stock range: ColorOS
# moves policy4's max on its own within seconds of a release, so an exact-stock
# comparison reported DRIFT on a correctly unpinned device.
#
# Values were read off each device on 2026-09-07 and confirmed against a pinned window; see
# ../REPORT.md and the matching *.env profiles. Stock values are hardcoded rather than sampled
# at pin time on purpose - a restore that reads "stock" from an already-pinned device would
# make the pin permanent, which is how a device silently stays clamped across a reboot-less
# week of runs.
set -u -o pipefail
# Git Bash: stop MSYS rewriting /sys/... and /proc/... arguments into C:/Program Files/...
export MSYS_NO_PATHCONV=1 MSYS2_ARG_CONV_EXCL='*'
SERIAL=${1:-}
ACTION=${2:-}
case "$ACTION" in
pin|unpin|check) ;;
*) echo "usage: $0 <serial> pin|unpin|check" >&2; exit 64 ;;
esac
CPUFREQ=/sys/devices/system/cpu/cpufreq
# ---------------------------------------------------------------------------------------------
# Device table. Everything device-specific lives here; the actions below are generic.
# POLICIES - "label:policy:pinned_khz:stock_min:stock_max" per cpufreq policy to pin.
# Every policy that can run a hot thread must be listed, not just the two
# the protocol reports: on MT6993 an unpinned policy7 at 4.21 GHz defeats
# the entire pin the moment the scheduler lands a thread on it.
# THERMAL_TYPE - matched against /sys/class/thermal/thermal_zone*/type by name, never by
# zone number: numbering is not stable across boots.
# ---------------------------------------------------------------------------------------------
case "$SERIAL" in
35d0befa) # Xiaomi 24129PN74C, Snapdragon 8 Elite (SM8750), Adreno 830v2
DEV_NAME="Xiaomi 24129PN74C / SM8750 / Adreno 830v2"
GPU_STYLE=kgsl
POLICIES="big:policy6:1958400:1017600:2841600 little:policy0:1555200:556800:2745600"
THERMAL_TYPE=cpuss-0-0
KGSL=/sys/class/kgsl/kgsl-3d0
GPU_PIN_LEVEL=0 # pwrlevel 0 = 1100 MHz, above the stock devfreq ceiling of 1050
GPU_STOCK_MIN_LEVEL=12
GPU_STOCK_MAX_LEVEL=0
GPU_PINNED_FREQ=1100000000 # gpuclk is in Hz on kgsl, NOT kHz
GPU_FREQ_UNIT=Hz
;;
3B159D009VZ00000) # Oppo PLG110 / ColorOS, MediaTek MT6993 (Dimensity 9500), Mali
DEV_NAME="Oppo PLG110 / MT6993 / Mali (gpufreqv2)"
GPU_STYLE=gpufreqv2
# 2000000 / 1600000 are the nearest available OPPs to the protocol's 1958400 / 1555200
# (+2.1% / +2.9%); neither target is an exact step on this part.
POLICIES="big:policy4:2000000:300000:3500000 little:policy0:1600000:300000:2100000 extra:policy7:2000000:300000:3200000"
THERMAL_TYPE=soc_max
GPU_FIX_NODE=/proc/gpufreqv2/fix_target_opp_index
GPU_PIN_LEVEL=0 # OPP index 0 = 1716000 kHz, the top working OPP
GPU_PINNED_FREQ=1716000
GPU_FREQ_UNIT=kHz
;;
*)
echo "$0: unknown serial '$SERIAL'." >&2
echo "Known: 35d0befa (Xiaomi/Adreno830), 3B159D009VZ00000 (Oppo/Mali)." >&2
echo "Refusing to guess: the pin path differs per SoC and a wrong one fails silently." >&2
exit 64 ;;
esac
A="adb -s $SERIAL"
# Quote the whole su invocation for the DEVICE shell, or the redirect runs unprivileged.
su_() { $A shell "su -c '$*'" 2>&1 | tr -d '\r'; }
$A get-state >/dev/null 2>&1 || { echo "$0: device $SERIAL is not connected" >&2; exit 65; }
[ "$(su_ 'id -u')" = "0" ] || { echo "$0: no root on $SERIAL (su failed)" >&2; exit 65; }
# --- read live state in ONE device round trip -------------------------------------------------
# Every read goes through su: on the Oppo the cpufreq nodes are 0660 system:system and a plain
# `adb shell cat scaling_governor` answers "Permission denied", which a careless parser reads
# as an empty governor rather than as a failure.
read_state() {
local script="" spec label pol
for spec in $POLICIES; do
IFS=: read -r label pol _ _ _ <<<"$spec"
script="$script echo \"${label}_gov=\$(cat $CPUFREQ/$pol/scaling_governor)\";"
script="$script echo \"${label}_min=\$(cat $CPUFREQ/$pol/scaling_min_freq)\";"
script="$script echo \"${label}_max=\$(cat $CPUFREQ/$pol/scaling_max_freq)\";"
script="$script echo \"${label}_cur=\$(cat $CPUFREQ/$pol/scaling_cur_freq)\";"
done
script="$script for tz in /sys/class/thermal/thermal_zone*; do"
script="$script if [ \"\$(cat \$tz/type 2>/dev/null)\" = \"$THERMAL_TYPE\" ]; then"
script="$script echo \"temp_mc=\$(cat \$tz/temp)\"; echo \"temp_zone=\$tz\"; break; fi; done;"
# Emit the GPU nodes RAW and pick them apart locally. The device-side script is delivered as
# `su -c '<script>'`, so a single quote anywhere inside it (an awk program, a sed expression)
# closes that quoting and the whole read silently returns nothing - which `check` then reports
# as every node being "neither pin nor stock", i.e. a false DRIFT on a perfectly pinned device.
# Parsing on this side keeps the device-side text quote-free.
if [ "$GPU_STYLE" = kgsl ]; then
script="$script echo \"gpu_minlvl=\$(cat $KGSL/min_pwrlevel)\";"
script="$script echo \"gpu_maxlvl=\$(cat $KGSL/max_pwrlevel)\";"
script="$script echo \"gpu_freqraw=\$(cat $KGSL/gpuclk)\";"
script="$script echo \"gpu_busyraw=\$(cat $KGSL/gpu_busy_percentage)\";"
else
script="$script echo \"gpu_fixraw=\$(cat $GPU_FIX_NODE | head -1)\";"
script="$script echo \"gpu_freqraw=\$(cat /sys/kernel/ged/hal/current_freqency | head -1)\";"
script="$script echo \"gpu_busyraw=\$(cat /sys/kernel/ged/hal/gpu_utilization | head -1)\";"
fi
su_ "$script"
}
# gpuclk -> "222000000" ($NF)
# current_freqency -> "<opp_index> <freq_khz>" ($NF)
# gpu_busy_percentage -> "26 %" / gpu_utilization -> "7 0 100" ($1)
# fix_target_opp_index -> "[GPUFREQ-DEBUG] fix GPU/STACK OPP index: 0/0" when pinned
# -> "[GPUFREQ-DEBUG] fix GPU/STACK OPP index is disabled" when not
derive_gpu() {
ST[gpu_freq]=$(echo "${ST[gpu_freqraw]:-}" | awk '{print $NF}')
ST[gpu_busy]=$(echo "${ST[gpu_busyraw]:-}" | awk '{print $1}')
[ "$GPU_STYLE" = kgsl ] && return 0
case "${ST[gpu_fixraw]:-}" in
*"is disabled"*) ST[gpu_fix]=off ;;
*"index: "*) ST[gpu_fix]=$(echo "${ST[gpu_fixraw]}" | sed -e 's/.*index: //' -e 's|/.*||') ;;
*) ST[gpu_fix]="<unread>" ;;
esac
}
declare -A ST
load_state() {
local line k v
ST=()
while IFS= read -r line; do
k=${line%%=*}; v=${line#*=}
case "$line" in *=*) ST[$k]=$v ;; esac
done < <(read_state)
# A read that came back empty is a BROKEN READ, not an unpinned device and not a drifted one.
# Without this the classifier sees every node as "neither pin nor stock" and prints DRIFT,
# which reads as "your pin slipped" when the truth is "this script could not see the device".
# Bail loudly instead: a pin/unpin whose verification cannot run must not look like a verdict.
if [ "${#ST[@]}" -eq 0 ]; then
echo "$0: could not read any state from $SERIAL (su read returned nothing)." >&2
echo " The device-side read is delivered as su -c '<script>'; check that nothing in it" >&2
echo " contains a single quote, and that su still works: adb -s $SERIAL shell su -c id" >&2
exit 66
fi
derive_gpu
}
field() { echo "${ST[$1]:-<unread>}"; }
print_state() {
local spec label pol pin smin smax
echo " device : $DEV_NAME ($SERIAL)"
for spec in $POLICIES; do
IFS=: read -r label pol pin smin smax <<<"$spec"
printf " %-7s %-8s gov=%-10s cur=%-9s min=%-9s max=%-9s (pin %s / stock %s-%s)\n" \
"$label" "$pol" "$(field ${label}_gov)" "$(field ${label}_cur)" \
"$(field ${label}_min)" "$(field ${label}_max)" "$pin" "$smin" "$smax"
done
if [ "$GPU_STYLE" = kgsl ]; then
printf " %-7s %-8s pwrlevel=%s..%s freq=%s %s busy=%s%% (pin lvl %s = %s %s / stock lvl %s..%s)\n" \
gpu kgsl-3d0 "$(field gpu_maxlvl)" "$(field gpu_minlvl)" "$(field gpu_freq)" "$GPU_FREQ_UNIT" \
"$(field gpu_busy)" "$GPU_PIN_LEVEL" "$GPU_PINNED_FREQ" "$GPU_FREQ_UNIT" \
"$GPU_STOCK_MAX_LEVEL" "$GPU_STOCK_MIN_LEVEL"
else
printf " %-7s %-8s fix_opp=%-4s freq=%s %s busy=%s%% (pin idx %s = %s %s / stock off)\n" \
gpu gpufreqv2 "$(field gpu_fix)" "$(field gpu_freq)" "$GPU_FREQ_UNIT" "$(field gpu_busy)" \
"$GPU_PIN_LEVEL" "$GPU_PINNED_FREQ" "$GPU_FREQ_UNIT"
fi
printf " %-7s %-8s %s mC = %s C (%s)\n" thermal "$THERMAL_TYPE" \
"$(field temp_mc)" "$(awk -v t="$(field temp_mc)" 'BEGIN{if(t+0==0){print "?"}else{printf "%.1f", t/1000}}')" \
"$(field temp_zone)"
}
# --- actions ----------------------------------------------------------------------------------
# Write order is min -> floor, then max -> target, then min -> target. Setting min above the
# current max (or max below the current min) is clamped by cpufreq, so a naive two-write pin
# succeeds on one device and silently half-applies on another depending on where stock sits
# relative to the target. Dropping min to the policy floor first makes the order stock-agnostic.
do_pin() {
local spec label pol pin smin smax script=""
for spec in $POLICIES; do
IFS=: read -r label pol pin smin smax <<<"$spec"
script="$script echo \$(cat $CPUFREQ/$pol/cpuinfo_min_freq) > $CPUFREQ/$pol/scaling_min_freq;"
script="$script echo $pin > $CPUFREQ/$pol/scaling_max_freq;"
script="$script echo $pin > $CPUFREQ/$pol/scaling_min_freq;"
done
su_ "$script" >/dev/null
if [ "$GPU_STYLE" = kgsl ]; then
# Collapse the pwrlevel range onto level 0. This also unlocks the top step: the devfreq
# governor's own max_freq sits one step below it, so a devfreq min_freq/max_freq pin cannot
# reach 1100 MHz at all.
su_ "echo $GPU_PIN_LEVEL > $KGSL/min_pwrlevel; echo $GPU_PIN_LEVEL > $KGSL/max_pwrlevel" >/dev/null
else
su_ "echo $GPU_PIN_LEVEL > $GPU_FIX_NODE" >/dev/null
fi
}
do_unpin() {
local spec label pol pin smin smax script=""
for spec in $POLICIES; do
IFS=: read -r label pol pin smin smax <<<"$spec"
script="$script echo \$(cat $CPUFREQ/$pol/cpuinfo_min_freq) > $CPUFREQ/$pol/scaling_min_freq;"
script="$script echo $smax > $CPUFREQ/$pol/scaling_max_freq;"
script="$script echo $smin > $CPUFREQ/$pol/scaling_min_freq;"
done
su_ "$script" >/dev/null
if [ "$GPU_STYLE" = kgsl ]; then
su_ "echo $GPU_STOCK_MIN_LEVEL > $KGSL/min_pwrlevel; echo $GPU_STOCK_MAX_LEVEL > $KGSL/max_pwrlevel" >/dev/null
else
su_ "echo -1 > $GPU_FIX_NODE" >/dev/null
fi
}
# Classify every node this script writes as at-pin or not-at-pin.
#
# NOT "at-pin / at-stock / neither". An earlier version compared against the hardcoded stock
# range and called anything else DRIFT, and ColorOS broke it within seconds: after an unpin
# restored policy4 to 300000-3500000, the Oppo performance daemon lowered the max to 3200000 on
# its own, and the next check reported DRIFT on a correctly released device. Stock maxima are
# daemon-managed and are NOT constants. So the only thing asserted here is our own pin - which
# is the question that actually matters ("is this run pinned?"). The hardcoded stock values
# still drive the restore path, where handing the range back to the governor is all they have
# to do.
do_check() {
local spec label pol pin smin smax n_pin=0 n_stock=0 n_other=0 drift="" notes=""
for spec in $POLICIES; do
IFS=: read -r label pol pin smin smax <<<"$spec"
local mn mx cu
mn=$(field ${label}_min); mx=$(field ${label}_max); cu=$(field ${label}_cur)
if [ "$mn" = "$pin" ] && [ "$mx" = "$pin" ]; then
n_pin=$((n_pin+1))
# min==max leaves the governor no room, so a cur that is not the pin means something
# outside cpufreq (thermal engine, vendor limiter) is overriding it.
[ "$cu" = "$pin" ] || drift="$drift ${label}(${pol}) pinned to $pin but scaling_cur_freq=$cu;"
elif [ "$mn" = "$mx" ]; then
# Clamped, but not by us. Someone else (game mode, thermal engine) is holding this policy
# at a fixed frequency, which is just as fatal to comparability as a missing pin.
n_other=$((n_other+1))
drift="$drift ${label}(${pol}) is clamped at $mn by something other than this script (our pin is $pin);"
else
n_stock=$((n_stock+1))
[ "$mn" = "$smin" ] && [ "$mx" = "$smax" ] || \
notes="$notes ${label}(${pol}) unpinned, range $mn-$mx (recorded stock $smin-$smax, which vendor daemons move);"
fi
done
if [ "$GPU_STYLE" = kgsl ]; then
if [ "$(field gpu_minlvl)" = "$GPU_PIN_LEVEL" ] && [ "$(field gpu_maxlvl)" = "$GPU_PIN_LEVEL" ]; then
n_pin=$((n_pin+1))
[ "$(field gpu_freq)" = "$GPU_PINNED_FREQ" ] || \
drift="$drift gpu pinned to pwrlevel $GPU_PIN_LEVEL but gpuclk=$(field gpu_freq) (expected $GPU_PINNED_FREQ);"
elif [ "$(field gpu_minlvl)" = "$GPU_STOCK_MIN_LEVEL" ] && [ "$(field gpu_maxlvl)" = "$GPU_STOCK_MAX_LEVEL" ]; then
n_stock=$((n_stock+1))
else
n_other=$((n_other+1))
drift="$drift gpu pwrlevel range $(field gpu_maxlvl)..$(field gpu_minlvl) is neither pin nor stock;"
fi
else
if [ "$(field gpu_fix)" = "$GPU_PIN_LEVEL" ]; then
n_pin=$((n_pin+1))
# The GPU parks its rail when idle and then reports a fallback frequency; only treat a
# mismatch as drift while the GPU is actually doing something.
if [ "$(field gpu_freq)" != "$GPU_PINNED_FREQ" ]; then
if [ "$(field gpu_busy)" = "0" ]; then
notes="$notes gpu fixed at OPP $GPU_PIN_LEVEL, freq reads $(field gpu_freq) with busy=0 - rail parked, not drift;"
else
drift="$drift gpu fixed at OPP $GPU_PIN_LEVEL but current_freqency=$(field gpu_freq) (expected $GPU_PINNED_FREQ) at busy=$(field gpu_busy)%;"
fi
fi
elif [ "$(field gpu_fix)" = "off" ]; then
n_stock=$((n_stock+1))
else
n_other=$((n_other+1))
drift="$drift gpu fix_target_opp_index='$(field gpu_fix)' is neither $GPU_PIN_LEVEL nor off;"
fi
fi
# Notes are commentary (a parked GPU rail is not a slipped pin) and are collected in their own
# variable rather than tagged inside $drift and filtered back out: the glob that would strip
# them is greedy, so it could swallow a real drift message that happened to follow one.
show() { [ -n "${1// /}" ] && echo "$1" | tr ';' '\n' | sed -e "s|^ *| $2|" -e "/^ *$2 *\$/d"; return 0; }
echo
if [ -n "${drift// /}" ]; then
echo " VERDICT: DRIFT"
show "$drift" "- "; show "$notes" "note: "
return 1
fi
if [ "$n_other" -gt 0 ] || { [ "$n_pin" -gt 0 ] && [ "$n_stock" -gt 0 ]; }; then
echo " VERDICT: DRIFT (partially pinned: $n_pin at pin, $n_stock at stock, $n_other neither)"
show "$notes" "note: "
return 1
fi
if [ "$n_pin" -gt 0 ] && [ "$n_stock" -eq 0 ]; then
echo " VERDICT: PINNED - all $n_pin pinned nodes at their pins, live frequencies match."
show "$notes" "note: "
return 0
fi
echo " VERDICT: UNPINNED - none of the $n_stock nodes is at a pin this script set; the vendor"
echo " governors have their range back and nothing this script writes is in effect."
show "$notes" "note: "
return 2
}
# --- main -------------------------------------------------------------------------------------
case "$ACTION" in
check)
load_state
echo "== check =="
print_state
do_check
exit $?
;;
pin)
load_state
echo "== before pin =="
print_state
do_pin
sleep 1
load_state
echo
echo "== after pin =="
print_state
do_check
rc=$?
[ "$rc" = 0 ] || echo " (pin did not take - do not measure against this)" >&2
exit $rc
;;
unpin)
load_state
echo "== before unpin =="
print_state
do_unpin
sleep 1
load_state
echo
echo "== after unpin =="
print_state
do_check
rc=$?
# 2 (UNPINNED) is success for this action.
[ "$rc" = 2 ] && exit 0
echo " (unpin did not fully restore - the device is still clamped)" >&2
exit 1
;;
esac