Skip to content

vulkan: 在骁龙 8 Gen 3(Adreno 750,安卓)上跑通并调优无协作矩阵档 - #22

Merged
sdcb merged 12 commits into
feature/2.0from
vulkan-android-8gen3
Sep 29, 2026
Merged

sdcb merged 12 commits into
feature/2.0from
vulkan-android-8gen3

Conversation

@sdcb

@sdcb sdcb commented Sep 29, 2026

Copy link
Copy Markdown
Owner

概要

让 OcrBackend.Vulkan 在安卓手机(真我 GT5 Pro,骁龙 8 Gen 3 / Adreno 750,Android 16)上跑通 DET / CLS / REC,精度与同一台手机的 CPU、同一份 dataset/ 一致,并针对这颗 GPU 做了有实测收益的调优。三档端到端都稳定快于本机 CPU,所以 Auto 在“无协作矩阵、compute subgroup 最小 64”的设备上改走 GPU。

完整报告(探针原文、二分过程、所有表格、放弃的尝试)在 docs/vulkan-8gen3.md。

桌面 sg16 / sg32 / sg32l 的 shader、spv 和路由都没有改;全量重编后所有未改动的 .spv 逐字节一致。3080 Ti 上与 feature/2.0 交替复测,Vulkan 三档逐图 100/100、速度在噪声内。

设备能力(--caps 探针,走与正式管线相同的 VkDevice.NewPipeline)

项 值
设备 / 加载库 Adreno (TM) 750,Vulkan 1.3.128,libvulkan.so
subgroup 默认 64,sgRange 64–128;请求 64 → 实际 64,请求 128 → 实际 128;16 / 32 超出范围,按规范未请求
协作矩阵 没有 VK_KHR_cooperative_matrix
共享内存 32 KB
绑定限制 maxStorageBufferRange 128 MB,minStorageBufferOffsetAlignment 64 B
队列 / 时间戳 选中仅计算队列 family 1;timestampValidBits 48,周期 52 ns,可用
峰值 纯寄存器 fp32 FMA 2.09 TFLOPS(f16vec4 只有 0.91,不翻倍)

因此走 5.3 的第 3 条路由(无协作矩阵):不创建任何 conv1x1_cm* / convk_cm* 管线。锁不住 8 lane,进新增的 _ncWide 子档。

改动(12 个提交,一个改动一个提交)

加载与宿主

  • c0473cf vulkan: load libvulkan.so on Android — 在现有唯一的 resolver 里,libvulkan.so.1 失败后再试 libvulkan.so;失败信息列出试过的文件名。没有新建第二个 resolver。
  • 98da6e2 test: add an Android bench host — test/Sdcb.SimdPaddleOCR.AndroidBench(net10.0-android,CoreCLR,引用库的 net10.0,IsPackable=false)。run.ps1 把“编译 → adb install → 推模型和数据集 → 启动 → 等 .done → 拉回结果/logcat”收成一条命令。模式:--caps、--e2e(与 Tests 相同的选项和 JSON 字段,cmp.ps1 / steady.ps1 / --summarize 直接可用)、--detprof / --recprof、--detcmp / --reccmp、--conc,以及两个二分工具 --layers(把 GPU 图截断到节点 k 并读回该节点输出)和 --gemm(单个 GEMM 对 fp64 参考)。GpuBench 链接了其中不依赖安卓的文件,方便在桌面取对照值。没有加进 .slnx 和 CI(需要 android workload)。
  • b865f97、b294c2a 宿主工具:GEMM 变体扫描、FMA 峰值探针、gprof.ps1、CPU/Vulkan 交替对比 ab.ps1,以及两个启动竞态修复。

正确性(原样上机三档全错:DET 概率图几乎全零,会话却报告 GPU 正常)

  • 981b6fe direct-load gemm_nc where no narrow subgroup can be pinned — Adreno 编译器会算错 gemm_nc 的寄存器预取(下一块 A/B 先读进 f16vec4 数组)。单测 6400×32×64 有 188871/204800 个输出错;换 fma 形式、换 select 都无效,去掉预取就完全正确。gemm_nc.comp 新增 -DDIRECT,编成 gemm_nc_d / gemm_nc_ds;只在锁不住 8 lane 时使用。UHD 770 仍用原 gemm_nc / gemm_nc_s(spv 不变)。
  • 04f38bc pad fp16 constant uploads to whole 16 bytes — 标量常量原来只上传 2 字节,elem4 按 uint 读 b[0]。桌面驱动不查越界;Adreno 按绑定范围查,越界读返回 0,CLS 头里乘常量那步整个变 0(tiny CLS 306/1022,连带 REC CER 57%)。ConstF16 / VecF16 补齐到 8 个 half 的整数倍,只追加 0,所有内核读到的值不变。
  • afb0f38 keep the im2col matrix within one binding's maxStorageBufferRange — medium DET 960×960 的 7×7 卷积 im2col 有 180 MB,超出 128 MB 后读成 0,大图丢框(CER 13.2%)。放不进一个绑定时改走 convk_dot;仍会超范围(slab 或无法走直接卷积的 im2col)时,在运行前抛 NotSupportedException,会话回退 CPU,而不是返回半张零图。桌面范围是 GB 级,不触发。

调优(每项都做了同一时段交替 A/B,数据在提交说明里)

  • e228638 fp16 shared tiles with BK 8 in the direct-load gemm_nc — 11 个最热 GEMM 形状合计 109.5 → 76.2 ms(0.44 → 0.64 TFLOPS,峰值的 31%)。暂存值本来就是 fp16,每个输出仍按 K 递增逐个 fma,结果逐位不变(端到端 100/100);非 SIMPLE 尾部单独用 fp32 暂存,累加器不会先被截成 fp16。
  • 8961dc2 two-phase spatial mean at any size on wide no-coopmat parts — _ncWide 下小图也走 reduce_hw4;medium REC 8×480 269.7 → 243.3 ms。
  • 9156482 direct kxk conv at every K on wide no-coopmat parts — _ncWide 下 kxk 稠密卷积默认一律走直接卷积(SIMD_OCR_CONVD_KMAX 仍可覆盖);medium DET 960 304.0 → 299.3,640×960 216.9 → 206.2。

Auto

  • a318575 Auto takes the no-coopmat tier where subgroups cannot go below 64 — GpuBackend.UsesGpu 多了 dev.SubgroupMin >= 64 这一条(IsGpuBackend、REC 的 16 行批次跟着变)。只按能力判定,没有写死骁龙 / Adreno。UHD 770(能锁 8 lane)和其它最小 subgroup < 64 的无矩阵设备,Auto 仍走 CPU。SIMD_OCR_BACKEND=cpu 仍可强制 CPU。

文档

  • 23e53e2 docs/vulkan-8gen3.md;docs/perf.md 加链接;中英文 README 的支持范围表说明了各系统的 Vulkan 加载库名、安卓宿主,以及桌面路径未变。

手机上的实测

端到端(--workers 4 --warmup 1,n=99,CPU / Vulkan 同一时段交替,median ms/图)

模型 本机 CPU Vulkan 对 CPU
tiny 218.8 / 202.6 / 276.4 156.1 / 170.4 / 154.9 / 151.7 约 1.3×
small 867.1 / 785.0 / 811.7 251.3 / 247.3 / 246.2 / 255.1 约 3.3×
medium 3073.8 / 3294.1 / 3368.9 1466.4 / 1458.1 / 1467.0 约 2.2×

热状态 2–3(medium 全程 3,机身明显发烫)。发热是最大的误差来源:同一版本 medium Vulkan 在热状态 2 时是 833 ms,热状态 3 时是 1460 ms。工作集峰值 medium CPU 约 1.60 GB,Vulkan 约 1.09 GB。

精度(相对同一台手机的 CPU)

模型 CPU exact_lines / CER Vulkan exact_lines / CER 逐图 hash/detected/texts 框数一致
tiny 742 / 2.37% 743 / 2.40% 78/100 99/100
small 950 / 0.41% 949 / 0.42% 90/100 100/100
medium 1004 / 0.14% 1003 / 0.15% 97/100 100/100

手机 CPU 三档与 Windows 基线完全一致;CLS 三档满分。

纯 GPU(gprof.ps1,逐 dispatch 时间戳取最小,ms)

用例 修正确性后、调优前 调优后
medium DET 960×960 380.5 298.6
medium DET 640×960 270.1 205.2
medium REC 8×480 358.9 242.3
medium REC 1×320 40.0 34.4
small DET 960×960 64.5 52.7
small REC 8×480 72.6 54.5
tiny DET 960×960 27.9 23.7
tiny REC 8×480 15.9 12.3

试过但没留下

  • 9×9 depthwise 改走分块 conv_dw4t:更慢(medium DET 303.7 → 317.0)。
  • GEMM 不用共享内存:慢 2–4 倍,全展开 8×8 版本又被编译器算错。BK 4 / BK 32 / 32×32 块、f16vec4 共享内存读:无收益。
  • arena slab 与 SE 计数器按 64 B 对齐、描述符 range 截到 128 MB:在这个驱动上输出逐像素不变,没有收益,所以没留(按规范仍属无效用法,文档里记为遗留项)。
  • dot → GEMM 门槛(2¹²–2²⁰)和 StreamCuts 切分比例重扫后保持原值。

希望 reviewer 重点看的地方

  1. 共享代码对桌面的影响:GpuGraph 里会在桌面执行的只有两处——fp16 常量补齐(ConstF16 / VecF16 只追加 0),以及按 maxStorageBufferRange 的门槛和运行前保护(桌面范围是 GB 级,理论上不触发)。请确认没有别的路径会被 _ncWide 或新门槛误伤。
  2. _ncWide = !sg8 的范围:内核选择(gemm_nc_d*、reduce、直接卷积)挂在“无矩阵且锁不住 8 lane”上,比 Auto 的条件(SubgroupMin >= 64)宽。也就是说:无协作矩阵、最小 subgroup 16/32 的设备(例如不带协作矩阵驱动的 NVIDIA、Mali)在显式 Vulkan 下也会用这套内核,但它们没测过。
  3. Auto 规则:SubgroupMin >= 64 是否足够保守?目前命中的已知设备只有 Adreno;只支持 wave64 的老 AMD GCN(无协作矩阵)理论上也会命中,但没测过。
  4. gemm_nc.comp 的 DIRECT 分支:共享内存改成 float16_t、尾部用单独的 fp32 数组 St;不带 DIRECT 的构建宏展开后与原文一致(spv 逐字节验证过)。
  5. 运行前回退:两处新的 NotSupportedException 都发生在 BuildPlan 里,也就是任何 dispatch 之前。
  6. 安卓宿主:run.ps1 的启动/等待逻辑、global.json 固定 10.0.x SDK、DebugType=embedded(绕开 CoreCLR 安卓打包的 MSB4096)、UseMonoRuntime=false(CoreCLR 才有 AdvSimd)。

测试

  • dotnet build src/Sdcb.SimdPaddleOCR -c Release(net10.0 与 netstandard2.0 都过)
  • dotnet test test/Sdcb.SimdPaddleOCR.UnitTests -c Release 96/96;-p:UseNs20Library=true 96/96
  • 安卓宿主单独编译、安装、运行(.NET SDK 10.0.400 + android workload 36.1,Android platform 36,JDK 21)
  • 手机:三档 CPU / Vulkan 交替 3 轮(上表),精度与本机 CPU 对齐
  • 手机:--conc tiny 4/8 线程、small 4、medium 4、Auto tiny 4,mismatches=0,日志无 fallback;medium 4 线程工作集峰值 980 MB,未被系统杀
  • 手机:每项调优后端到端与上一版逐图 100/100(GEMM、reduce、直接卷积)
  • 手机:--backend auto usesGpu=True,与 --backend vulkan 100/100;Auto + SIMD_OCR_BACKEND=cpu 与 CPU 100/100
  • 全量 build-shaders.bat 后 git status 干净(所有 spv 逐字节一致),se_fused_dbg.spv 已删、未提交
  • 3080 Ti(驱动 581.80)与 feature/2.0 交替 2 轮:--engine vulkan median tiny 18.0/19.9 → 18.5/20.7、small 26.7/30.3 → 29.2/28.5、medium 36.5/37.6 → 37.3/37.3;精度 741 / 950 / 1006 不变;三档逐图 100/100;--engine sharp tiny / small 100/100
  • B580、880M、UHD 770:本次无法测。它们都不满足 SubgroupMin >= 64,路由不变,只受 fp16 常量补齐这一处共享改动影响,建议有机会复测一轮 --engine vulkan

复现

cd test/Sdcb.SimdPaddleOCR.AndroidBench
.\run.ps1 -Build -Install -Push -Bench "--caps"
.\ab.ps1 -tag x -models tiny,small,medium -rounds 3
.\gprof.ps1 -tag x

sdcb added 12 commits September 28, 2026 22:07
Android exposes the Vulkan loader only under the unversioned NDK name
(/system/lib64/libvulkan.so); the resolver tried libvulkan.so.1 alone, so
VkDevice.Create failed and every session fell back to the CPU. It now tries
libvulkan.so.1, then libvulkan.so, in the one existing resolver, and a miss
names every file it tried instead of the P/Invoke's "vulkan-1".

Realme GT5 Pro (Snapdragon 8 Gen 3, Android 16): the loader resolves to
libvulkan.so and VkDevice.Create picks "Adreno (TM) 750", Vulkan 1.3.128.
Windows / desktop Linux resolve the same file names as before.
test/Sdcb.SimdPaddleOCR.AndroidBench is a net10.0-android app (CoreCLR,
references the library's net10.0 build: AdvSimd kernels + Vulkan) driven
over adb by run.ps1: build, install, push bench-out/models and dataset/
into the app's external files dir, launch with args/env extras, wait for
<run>.done, pull files/out into bench-out/android and keep logcat when the
process dies. Models are read from files, nothing is packaged.

Modes: --caps (device, sgRange, gl_SubgroupSize under required size
0/16/32/64/128 through VkDevice.NewPipeline, coopmat list, limits, heaps,
queues, timestamp check), --e2e (the Tests harness's options and JSON rows,
so cmp.ps1 / steady.ps1 / --summarize work), --detprof / --recprof,
--detcmp / --reccmp (CPU vs Vulkan on the same input), --conc, and two
bisection aids: --layers (truncate the GPU graph after node k and read that
node's output back) and --gemm (one GEMM dispatch vs an fp64 reference).
GpuBench links the layer probe, GEMM check and CPU/GPU compare so the same
numbers can be taken on a desktop GPU.

Not in the .slnx or CI: it needs the android workload.
On Adreno 750 (no cooperative matrix, compute subgroups 64-128, so the
no-coopmat tier with the driver-chosen wave64) gemm_nc returned garbage:
tiny DET's probability map was all zeros while the session reported the
GPU alive. The layer probe put the first divergence from the 3080 Ti at
the first conv routed to gemm_nc_s, and --gemm reproduced it alone
(6400x32x64: 188871 of 204800 outputs wrong; 1024x128x256: 126660 of
131072). Variants: plain fma vs mul+add and if-guarded vs select loads
change nothing; dropping the register prefetch of the next A/B tile, or
shrinking the tile to 32x32 with 4x4 accumulators, makes it exact. The
driver's compiler mishandles the f16vec4 prefetch arrays carried across
the K loop next to 64 fp32 accumulators.

gemm_nc.comp gains -DDIRECT (each tile goes from global straight into
shared memory, same fp32 accumulation order in K) built as gemm_nc_d /
gemm_nc_ds. The tier takes them whenever it cannot pin an 8-lane subgroup;
where it can (UHD 770) gemm_nc / gemm_nc_s with requiredSubgroupSize=8
stay as they were. gemm_nc.spv and gemm_nc_s.spv rebuild byte-identical.

GT5 Pro, --gemm vs an fp64 reference: 0 wrong everywhere, and faster than
the broken form (1024x128x256 0.254 -> 0.211 ms, 6400x32x64 0.231 ->
0.191 ms). tiny DET vs the phone CPU: box counts equal on 10/10 images,
map maxAbs 0.05-0.15 (3080 Ti vs its CPU: 0.03-0.18).
A scalar constant (numel 1) was uploaded as a 2-byte buffer, and elem4
reads it as b[0] of a uint array. Desktop drivers do not bounds-check that
load, Adreno does (per binding range) and returns 0 for it: the CLS head's
Mul by a constant came out as 0, tiny CLS got 306 of 1022 right and the
wrongly rotated lines took REC to CER 57%.

ConstF16 and VecF16 now round their element count up to a multiple of 8.
Only zeros are appended; every kernel reads the same values as before.

GT5 Pro, tiny, --backend vulkan: cls 306 -> 1022/1022, exact_lines
216 -> 743 / 1036, CER 57.5% -> 2.40% (phone CPU: 742, 2.37%). No
fallback in the log.
…ange

Adreno 750 caps a storage binding at 128 MB. medium DET's 7x7 conv at
960x960 input (M 57600, K 1568) builds a 180 MB im2col matrix, and every
load past 128 MB from its binding start read 0: the layer probe put the
first divergence from the 3080 Ti at that conv (meanAbs -25%), and medium
lost lines on large images (818 of 1036 classified, CER 13.2%).

A dense kxk conv whose im2col matrix would not fit one binding now takes
the direct convk_dot path instead (no K cap there). As a guard, a plan that
would still bind more than maxStorageBufferRange (an arena slab, or an
im2col the direct path cannot take) throws before anything runs, so the
session serves that shape from the CPU instead of returning a partly zero
map. Desktop GPUs report multi-GB ranges; none of this triggers there.

GT5 Pro medium, --backend vulkan: DET boxes vs phone CPU equal in count on
6/6 images (map maxAbs 0.03-0.07, was up to 1.0); end to end cls
1035/1035, exact_lines 1003 / 1036, CER 0.15% (Windows CPU baseline 1004,
0.14%). Clamping the descriptor ranges themselves changed nothing on this
driver, so that part is not included.
…cript

--gemmsweep checks each .spv variant on an odd shape against fp64, then
times it on the eleven hottest GEMM shapes of medium/small DET and REC,
variants interleaved per round so clock / thermal drift hits all of them
alike; --peak measures fp32 / fp16 FMA throughput from registers.
gprof.ps1 is the phone twin of GpuBench's gprof.ps1 (same case list,
per-dispatch timestamps, min over reps). run.ps1 no longer reports a
finished run as dead when the app exits right after writing .done.
On Adreno 750 the direct-load GEMM ran at ~0.44 TFLOPS against a
measured 2.0 TFLOPS fp32 FMA peak. Sweeping the kernel on the eleven
hottest shapes (sum of min GPU ms, variants interleaved in one run):

  64x64 BK16 fp32 tiles (was)        109.5
  BK 8                                83.1
  128x128, 8x4 per thread, 512 thr    84.1
  BK 8 + fp16 tiles (kept)            76.2   0.64 TFLOPS
  same with 128x64 / 64x128 blocks    74.6 / 75.0 (noise; 64x128 loses
                                       half a block on N=64)
  BK 4 / BK 32 / 32x32 4x4            83.2 / 200.6 / 148.1
  no shared memory, per-thread f16vec4 rows: 2-4x slower, and the fully
  unrolled 8x8 form is miscomputed by the driver like the prefetch was

The DIRECT build now keeps the A/B tiles as fp16 in shared memory (they
come from fp16 memory, so this is exact) with BK 8; the staged tail gets
its own fp32 array so accumulators are never narrowed before bias,
residual and activation. Each output still takes one fma per k in
increasing K, so results are bit-identical: e2e tiny and medium match
the previous build 100/100 (hash, detected, texts). gemm_nc.spv and
gemm_nc_s.spv (UHD 770, 8-lane pin) rebuild byte-identical.

GT5 Pro, same-period pure GPU (gprof, min ms, two alternating rounds,
both rounds identical to 0.1 ms): medium DET 960 381 -> 311, 640x960
270 -> 221, medium REC 8x480 359 -> 270, 1x320 40.1 -> 34.7, small DET
64.5 -> 52.7, small REC 72.7 -> 54.6, tiny DET 28.0 -> 23.7, tiny REC
15.9 -> 12.3.
…arts

After the GEMM, medium REC's third hotspot was reduce_hw (29 of 270 ms):
below 4096 pixels the spatial mean ran one channel per workgroup, which
on Adreno reads 2 bytes at a C*2 stride (the c512 means of the 8x480
batch took 7.6 ms each). The lite sg32 parts already take the two-phase
reduce_hw4 / reduce_hw4b pair at any size for the same reason; the
no-coopmat tier now does too when it cannot pin a narrow subgroup
(_ncWide). UHD 770 (8-lane pin) keeps the old gate.

GT5 Pro, same-period pure GPU (min ms, two alternating rounds, identical
to 0.1 ms): medium REC 8x480 269.7 -> 243.3, medium DET 960 324.4 ->
317.0, 640x960 229.8 -> 226.1; small / tiny unchanged. e2e tiny, small,
medium match the previous build 100/100 (hash, detected, texts).

Tried and dropped: the shared-tile conv_dw4t for 9x9 depthwise (sg32
takes it up to 9) is slower here, medium DET 303.7 -> 317.0, small DET
52.7 -> 55.3.
Dense kxk convs switch from the direct convk_dot to im2col + GEMM once
Kp > 1024 (SIMD_OCR_CONVD_KMAX); on UHD 770 lowering that cap lost. On
Adreno 750 the direct kernel keeps up with the tuned GEMM (57600x2304x64
at ~0.56 TFLOPS) and skips writing and re-reading the im2col matrix, so
the wide no-coopmat tier now defaults to direct at every K (the env var
still overrides). Every dense conv in these models has Kp <= 2304.

GT5 Pro, same-period pure GPU (min ms, two rounds each) with the cap at
256 / 512 / 1024 / 2304: medium DET 960 315.2 / 312.3 / 304.0 / 299.3,
small DET 64.6 / 60.7 / 52.8 / 52.8, tiny DET 32.8 / 29.9 / 23.7 / 23.7;
REC unchanged. With this build: medium DET 960 299.7, 640x960 216.9 ->
206.2. e2e medium matches the previous build 100/100.

The dot -> GEMM cut for K <= 128 pointwise convs was re-swept too (M*cout
> 2^12 / 2^14 / 2^16 / 2^18 / 2^20): 2^12..2^16 within 0.3 ms, 2^18 and
up cost tiny DET 23.7 -> 26.5 ms. It stays at 2^16.
…s before a launch

ab.ps1 runs each model round by round with the backends alternating and a
cool-down in between, and prints median / last-50 median, peak working
set, thermal status and accuracy per run. run.ps1 now force-stops the
app and waits for the process to go away before launching: an intent
delivered to an instance that is still exiting was swallowed and the run
reported as dead.
…w 64

Auto skipped the no-coopmat tier everywhere because it trails the CPU on
UHD 770. On Adreno 750 (no cooperative matrix, compute subgroups 64-128)
the tuned tier beats the phone's own CPU on all three models, in every
alternating round. GpuBackend.UsesGpu (and with it IsGpuBackend, i.e. the
16-line REC batches) now also says yes for Auto when the device reports
SubgroupMin >= 64. UHD 770 (pins 8 lanes) and every other no-coopmat
device keep Auto on the CPU; coopmat devices are unchanged.
SIMD_OCR_BACKEND=cpu still forces the CPU.

GT5 Pro, alternating rounds (median ms, cpu vs vulkan; thermal status
2-3, medium at 3 throughout):
  tiny    219 / 203 / 276   vs  156 / 170 / 155 / 152
  small   867 / 785 / 812   vs  251 / 247 / 246 / 255
  medium  3074 / 3294 / 3369  vs  1466 / 1458 / 1467
Accuracy cpu vs vulkan: 742 / 2.37% vs 743 / 2.40%, 950 / 0.41% vs
949 / 0.42%, 1004 / 0.14% vs 1003 / 0.15%.

--backend auto: usesGpu=True, 100/100 identical to --backend vulkan;
auto + SIMD_OCR_BACKEND=cpu: 100/100 identical to --backend cpu.
docs/vulkan-8gen3.md: the phone and driver, the --caps probe (sgRange
64-128, no cooperative matrix, 32 KB shared memory, 128 MB binding range,
timestamps usable, 2.09 TFLOPS fp32 FMA), the no-coopmat route it takes,
the three Adreno-only correctness bugs and how they were bisected,
alternating CPU / Vulkan end to end on all three models with thermal
status, accuracy and per-image parity, pure GPU before / after tuning and
the hotspots, what was tuned, re-swept, dropped and left, the Auto rule,
and the 3080 Ti regression. perf.md links it; the README rows now name
the Vulkan loader per OS and the Android host, and say the desktop path
is unchanged.
@sdcb

sdcb commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

在 Arc B580(Ryzen 9 5950X,Windows,驱动 32.0.101.8331)上复测:没有性能回退,输出逐图一致。

方法

  • 基线:feature/2.0 6e0dc36(本 PR 的 merge-base)。本 PR:23e53e2。两个 worktree 分别 Release 编译。
  • --workers 4 --benchmark-kind simd --warmup 1,同一 dataset/ 100 张,n=99。
  • Vulkan 三档各 3 轮,base / PR 交替,每轮换先后顺序;sharp 三档各 1 轮。
  • B580 的 coopmat 是 8×16×16,走 sg16 cm。_nocm 和 _ncWide 都不开,CoopGemm 为真,Auto 不受新规则影响。

端到端(median ms/图)

模型 base Vulkan r1 / r2 / r3 PR Vulkan r1 / r2 / r3 base sharp PR sharp
tiny 19.6 / 21.8 / 18.1 19.9 / 20.4 / 20.1 49.5 48.9
small 26.5 / 29.0 / 29.3 30.4 / 28.7 / 28.4 198.8 200.6
medium 54.7 / 66.4 / 62.5 54.0 / 56.2 / 54.6 538.8 545.2

两边的最小值基本重合。base medium 的 r2 / r3 受机器干扰(那一轮 det 36.8 ms),与本 PR 无关。

正确率、内存、逐图

  • exact_lines / CER 两边相同:Vulkan 740 / 2.36%、950 / 0.40%、1003 / 0.14%;sharp 742 / 2.37%、950 / 0.41%、1004 / 0.14%。
  • WS peak 两边相同:Vulkan 约 581 / 647 / 937 MB。
  • base 与 PR 的 r1 逐图比较 hash / detected / texts:tiny、small、medium 的 Vulkan 和 sharp 全部 100/100 一致。这说明 fp16 常量补齐不改变结果。
  • backend=vulkan 没有回退 CPU。这也间接说明 Vk.cs 里新加的 MaxStorageBufferRange 偏移(p + 4 + 28)读出的值正确,否则每个 slab 都会在 BuildPlan 里抛错。

代码审阅

桌面上会执行到的新代码:ConstF16 / VecF16 补齐、maxStorageBufferRange 门槛和运行前抛错、UsesGpu 的 SubgroupMin >= 64、resolver 候选列表。在 B580 上都不改变路由和结果。

不阻塞合并,留作记录:

  1. resolver 加载失败时改为直接抛 DllNotFoundException。原来返回 IntPtr.Zero,运行时还会按 vulkan-1 这个名字做一遍默认探测;现在这一步没了。TryLoad("vulkan-1.dll") 已经覆盖同样的搜索路径,TryGetDevice 也会吞掉异常,实际影响很小。
  2. _ncWide = !sg8 比 Auto 的条件宽:最小 subgroup 是 16 或 32、又没有 coopmat 的设备,在显式选 Vulkan 时也会走 gemm_nc_d 和 Adreno 专用的路由,这些设备没有测过(PR 正文已列出)。
  3. 只按 slab 检查 maxStorageBufferRange,没有截断描述符 range。按规范仍属无效用法,文档里已记为遗留项。桌面上不会触发。

3080 Ti / 880M / UHD 770 不在本次测试范围内。

@sdcb

sdcb commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

880M 复测(本机 AMD Radeon 880M)

对照 feature/2.0 6e0dc36 与本 PR 23e53e2。--workers 4 --engine vulkan,同一时段交替,先基线再 PR,三轮。中位 ms/图:

模型 feature/2.0 本 PR
tiny 27.5 / 28.5 / 28.1 27.4 / 27.3 / 27.8
small 46.4 / 45.7 / 44.5 46.7 / 45.3 / 45.1
medium 118.6 / 118.9 / 120.8 120.3 / 120.7 / 127.3

去掉 warmup 后的逐图配对中位差(PR − 基线):tiny −0.29 / −0.25 / +0.06 ms,small −0.52 / −0.24 / +0.53 ms,medium −0.05 / −0.83 / +3.11 ms。medium 第三轮那 +3 ms 是整轮最后一次,且 PR 总排在基线后面。倒序再跑一轮 medium(先 PR 再基线):118.8 对 118.9,配对差 +0.55 ms。

三轮精度一致:tiny 740 / CER 2.39%,small 951 / 0.39%,medium 1003 / 0.15%。含倒序那一轮,hash / detected / texts 全部 100/100。

没有看到回退。880M 有 16×16×16 协作矩阵且能把 subgroup 钉到 32,不会走进这次的无矩阵档。

@sdcb
sdcb merged commit b5c6ef2 into feature/2.0 Sep 29, 2026
66 checks passed
@sdcb
sdcb deleted the vulkan-android-8gen3 branch September 29, 2026 02:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant