Repository navigation
vulkan: 在骁龙 8 Gen 3(Adreno 750,安卓)上跑通并调优无协作矩阵档 - #22
Conversation
Android exposes the Vulkan loader only under the unversioned NDK name (/system/lib64/libvulkan.so); the resolver tried libvulkan.so.1 alone, so VkDevice.Create failed and every session fell back to the CPU. It now tries libvulkan.so.1, then libvulkan.so, in the one existing resolver, and a miss names every file it tried instead of the P/Invoke's "vulkan-1". Realme GT5 Pro (Snapdragon 8 Gen 3, Android 16): the loader resolves to libvulkan.so and VkDevice.Create picks "Adreno (TM) 750", Vulkan 1.3.128. Windows / desktop Linux resolve the same file names as before.
test/Sdcb.SimdPaddleOCR.AndroidBench is a net10.0-android app (CoreCLR, references the library's net10.0 build: AdvSimd kernels + Vulkan) driven over adb by run.ps1: build, install, push bench-out/models and dataset/ into the app's external files dir, launch with args/env extras, wait for <run>.done, pull files/out into bench-out/android and keep logcat when the process dies. Models are read from files, nothing is packaged. Modes: --caps (device, sgRange, gl_SubgroupSize under required size 0/16/32/64/128 through VkDevice.NewPipeline, coopmat list, limits, heaps, queues, timestamp check), --e2e (the Tests harness's options and JSON rows, so cmp.ps1 / steady.ps1 / --summarize work), --detprof / --recprof, --detcmp / --reccmp (CPU vs Vulkan on the same input), --conc, and two bisection aids: --layers (truncate the GPU graph after node k and read that node's output back) and --gemm (one GEMM dispatch vs an fp64 reference). GpuBench links the layer probe, GEMM check and CPU/GPU compare so the same numbers can be taken on a desktop GPU. Not in the .slnx or CI: it needs the android workload.
On Adreno 750 (no cooperative matrix, compute subgroups 64-128, so the no-coopmat tier with the driver-chosen wave64) gemm_nc returned garbage: tiny DET's probability map was all zeros while the session reported the GPU alive. The layer probe put the first divergence from the 3080 Ti at the first conv routed to gemm_nc_s, and --gemm reproduced it alone (6400x32x64: 188871 of 204800 outputs wrong; 1024x128x256: 126660 of 131072). Variants: plain fma vs mul+add and if-guarded vs select loads change nothing; dropping the register prefetch of the next A/B tile, or shrinking the tile to 32x32 with 4x4 accumulators, makes it exact. The driver's compiler mishandles the f16vec4 prefetch arrays carried across the K loop next to 64 fp32 accumulators. gemm_nc.comp gains -DDIRECT (each tile goes from global straight into shared memory, same fp32 accumulation order in K) built as gemm_nc_d / gemm_nc_ds. The tier takes them whenever it cannot pin an 8-lane subgroup; where it can (UHD 770) gemm_nc / gemm_nc_s with requiredSubgroupSize=8 stay as they were. gemm_nc.spv and gemm_nc_s.spv rebuild byte-identical. GT5 Pro, --gemm vs an fp64 reference: 0 wrong everywhere, and faster than the broken form (1024x128x256 0.254 -> 0.211 ms, 6400x32x64 0.231 -> 0.191 ms). tiny DET vs the phone CPU: box counts equal on 10/10 images, map maxAbs 0.05-0.15 (3080 Ti vs its CPU: 0.03-0.18).
A scalar constant (numel 1) was uploaded as a 2-byte buffer, and elem4 reads it as b[0] of a uint array. Desktop drivers do not bounds-check that load, Adreno does (per binding range) and returns 0 for it: the CLS head's Mul by a constant came out as 0, tiny CLS got 306 of 1022 right and the wrongly rotated lines took REC to CER 57%. ConstF16 and VecF16 now round their element count up to a multiple of 8. Only zeros are appended; every kernel reads the same values as before. GT5 Pro, tiny, --backend vulkan: cls 306 -> 1022/1022, exact_lines 216 -> 743 / 1036, CER 57.5% -> 2.40% (phone CPU: 742, 2.37%). No fallback in the log.
…ange Adreno 750 caps a storage binding at 128 MB. medium DET's 7x7 conv at 960x960 input (M 57600, K 1568) builds a 180 MB im2col matrix, and every load past 128 MB from its binding start read 0: the layer probe put the first divergence from the 3080 Ti at that conv (meanAbs -25%), and medium lost lines on large images (818 of 1036 classified, CER 13.2%). A dense kxk conv whose im2col matrix would not fit one binding now takes the direct convk_dot path instead (no K cap there). As a guard, a plan that would still bind more than maxStorageBufferRange (an arena slab, or an im2col the direct path cannot take) throws before anything runs, so the session serves that shape from the CPU instead of returning a partly zero map. Desktop GPUs report multi-GB ranges; none of this triggers there. GT5 Pro medium, --backend vulkan: DET boxes vs phone CPU equal in count on 6/6 images (map maxAbs 0.03-0.07, was up to 1.0); end to end cls 1035/1035, exact_lines 1003 / 1036, CER 0.15% (Windows CPU baseline 1004, 0.14%). Clamping the descriptor ranges themselves changed nothing on this driver, so that part is not included.
…cript --gemmsweep checks each .spv variant on an odd shape against fp64, then times it on the eleven hottest GEMM shapes of medium/small DET and REC, variants interleaved per round so clock / thermal drift hits all of them alike; --peak measures fp32 / fp16 FMA throughput from registers. gprof.ps1 is the phone twin of GpuBench's gprof.ps1 (same case list, per-dispatch timestamps, min over reps). run.ps1 no longer reports a finished run as dead when the app exits right after writing .done.
On Adreno 750 the direct-load GEMM ran at ~0.44 TFLOPS against a
measured 2.0 TFLOPS fp32 FMA peak. Sweeping the kernel on the eleven
hottest shapes (sum of min GPU ms, variants interleaved in one run):
64x64 BK16 fp32 tiles (was) 109.5
BK 8 83.1
128x128, 8x4 per thread, 512 thr 84.1
BK 8 + fp16 tiles (kept) 76.2 0.64 TFLOPS
same with 128x64 / 64x128 blocks 74.6 / 75.0 (noise; 64x128 loses
half a block on N=64)
BK 4 / BK 32 / 32x32 4x4 83.2 / 200.6 / 148.1
no shared memory, per-thread f16vec4 rows: 2-4x slower, and the fully
unrolled 8x8 form is miscomputed by the driver like the prefetch was
The DIRECT build now keeps the A/B tiles as fp16 in shared memory (they
come from fp16 memory, so this is exact) with BK 8; the staged tail gets
its own fp32 array so accumulators are never narrowed before bias,
residual and activation. Each output still takes one fma per k in
increasing K, so results are bit-identical: e2e tiny and medium match
the previous build 100/100 (hash, detected, texts). gemm_nc.spv and
gemm_nc_s.spv (UHD 770, 8-lane pin) rebuild byte-identical.
GT5 Pro, same-period pure GPU (gprof, min ms, two alternating rounds,
both rounds identical to 0.1 ms): medium DET 960 381 -> 311, 640x960
270 -> 221, medium REC 8x480 359 -> 270, 1x320 40.1 -> 34.7, small DET
64.5 -> 52.7, small REC 72.7 -> 54.6, tiny DET 28.0 -> 23.7, tiny REC
15.9 -> 12.3.
…arts After the GEMM, medium REC's third hotspot was reduce_hw (29 of 270 ms): below 4096 pixels the spatial mean ran one channel per workgroup, which on Adreno reads 2 bytes at a C*2 stride (the c512 means of the 8x480 batch took 7.6 ms each). The lite sg32 parts already take the two-phase reduce_hw4 / reduce_hw4b pair at any size for the same reason; the no-coopmat tier now does too when it cannot pin a narrow subgroup (_ncWide). UHD 770 (8-lane pin) keeps the old gate. GT5 Pro, same-period pure GPU (min ms, two alternating rounds, identical to 0.1 ms): medium REC 8x480 269.7 -> 243.3, medium DET 960 324.4 -> 317.0, 640x960 229.8 -> 226.1; small / tiny unchanged. e2e tiny, small, medium match the previous build 100/100 (hash, detected, texts). Tried and dropped: the shared-tile conv_dw4t for 9x9 depthwise (sg32 takes it up to 9) is slower here, medium DET 303.7 -> 317.0, small DET 52.7 -> 55.3.
Dense kxk convs switch from the direct convk_dot to im2col + GEMM once Kp > 1024 (SIMD_OCR_CONVD_KMAX); on UHD 770 lowering that cap lost. On Adreno 750 the direct kernel keeps up with the tuned GEMM (57600x2304x64 at ~0.56 TFLOPS) and skips writing and re-reading the im2col matrix, so the wide no-coopmat tier now defaults to direct at every K (the env var still overrides). Every dense conv in these models has Kp <= 2304. GT5 Pro, same-period pure GPU (min ms, two rounds each) with the cap at 256 / 512 / 1024 / 2304: medium DET 960 315.2 / 312.3 / 304.0 / 299.3, small DET 64.6 / 60.7 / 52.8 / 52.8, tiny DET 32.8 / 29.9 / 23.7 / 23.7; REC unchanged. With this build: medium DET 960 299.7, 640x960 216.9 -> 206.2. e2e medium matches the previous build 100/100. The dot -> GEMM cut for K <= 128 pointwise convs was re-swept too (M*cout > 2^12 / 2^14 / 2^16 / 2^18 / 2^20): 2^12..2^16 within 0.3 ms, 2^18 and up cost tiny DET 23.7 -> 26.5 ms. It stays at 2^16.
…s before a launch ab.ps1 runs each model round by round with the backends alternating and a cool-down in between, and prints median / last-50 median, peak working set, thermal status and accuracy per run. run.ps1 now force-stops the app and waits for the process to go away before launching: an intent delivered to an instance that is still exiting was swallowed and the run reported as dead.
…w 64 Auto skipped the no-coopmat tier everywhere because it trails the CPU on UHD 770. On Adreno 750 (no cooperative matrix, compute subgroups 64-128) the tuned tier beats the phone's own CPU on all three models, in every alternating round. GpuBackend.UsesGpu (and with it IsGpuBackend, i.e. the 16-line REC batches) now also says yes for Auto when the device reports SubgroupMin >= 64. UHD 770 (pins 8 lanes) and every other no-coopmat device keep Auto on the CPU; coopmat devices are unchanged. SIMD_OCR_BACKEND=cpu still forces the CPU. GT5 Pro, alternating rounds (median ms, cpu vs vulkan; thermal status 2-3, medium at 3 throughout): tiny 219 / 203 / 276 vs 156 / 170 / 155 / 152 small 867 / 785 / 812 vs 251 / 247 / 246 / 255 medium 3074 / 3294 / 3369 vs 1466 / 1458 / 1467 Accuracy cpu vs vulkan: 742 / 2.37% vs 743 / 2.40%, 950 / 0.41% vs 949 / 0.42%, 1004 / 0.14% vs 1003 / 0.15%. --backend auto: usesGpu=True, 100/100 identical to --backend vulkan; auto + SIMD_OCR_BACKEND=cpu: 100/100 identical to --backend cpu.
docs/vulkan-8gen3.md: the phone and driver, the --caps probe (sgRange 64-128, no cooperative matrix, 32 KB shared memory, 128 MB binding range, timestamps usable, 2.09 TFLOPS fp32 FMA), the no-coopmat route it takes, the three Adreno-only correctness bugs and how they were bisected, alternating CPU / Vulkan end to end on all three models with thermal status, accuracy and per-image parity, pure GPU before / after tuning and the hotspots, what was tuned, re-swept, dropped and left, the Auto rule, and the 3080 Ti regression. perf.md links it; the README rows now name the Vulkan loader per OS and the Android host, and say the desktop path is unchanged.
|
在 Arc B580(Ryzen 9 5950X,Windows,驱动 32.0.101.8331)上复测:没有性能回退,输出逐图一致。 方法
端到端(median ms/图)
两边的最小值基本重合。base medium 的 r2 / r3 受机器干扰(那一轮 det 36.8 ms),与本 PR 无关。 正确率、内存、逐图
代码审阅桌面上会执行到的新代码: 不阻塞合并,留作记录:
3080 Ti / 880M / UHD 770 不在本次测试范围内。 |
880M 复测(本机 AMD Radeon 880M)对照
去掉 warmup 后的逐图配对中位差(PR − 基线):tiny −0.29 / −0.25 / +0.06 ms,small −0.52 / −0.24 / +0.53 ms,medium −0.05 / −0.83 / +3.11 ms。medium 第三轮那 +3 ms 是整轮最后一次,且 PR 总排在基线后面。倒序再跑一轮 medium(先 PR 再基线):118.8 对 118.9,配对差 +0.55 ms。 三轮精度一致:tiny 740 / CER 2.39%,small 951 / 0.39%,medium 1003 / 0.15%。含倒序那一轮,hash / detected / texts 全部 100/100。 没有看到回退。880M 有 16×16×16 协作矩阵且能把 subgroup 钉到 32,不会走进这次的无矩阵档。 |
概要
让
OcrBackend.Vulkan在安卓手机(真我 GT5 Pro,骁龙 8 Gen 3 / Adreno 750,Android 16)上跑通 DET / CLS / REC,精度与同一台手机的 CPU、同一份dataset/一致,并针对这颗 GPU 做了有实测收益的调优。三档端到端都稳定快于本机 CPU,所以Auto在“无协作矩阵、compute subgroup 最小 64”的设备上改走 GPU。完整报告(探针原文、二分过程、所有表格、放弃的尝试)在
docs/vulkan-8gen3.md。桌面 sg16 / sg32 / sg32l 的 shader、spv 和路由都没有改;全量重编后所有未改动的
.spv逐字节一致。3080 Ti 上与feature/2.0交替复测,Vulkan 三档逐图 100/100、速度在噪声内。设备能力(
--caps探针,走与正式管线相同的VkDevice.NewPipeline)Adreno (TM) 750,Vulkan 1.3.128,libvulkan.sosgRange64–128;请求 64 → 实际 64,请求 128 → 实际 128;16 / 32 超出范围,按规范未请求VK_KHR_cooperative_matrixmaxStorageBufferRange128 MB,minStorageBufferOffsetAlignment64 BtimestampValidBits48,周期 52 ns,可用因此走 5.3 的第 3 条路由(无协作矩阵):不创建任何
conv1x1_cm*/convk_cm*管线。锁不住 8 lane,进新增的_ncWide子档。改动(12 个提交,一个改动一个提交)
加载与宿主
c0473cfvulkan: load libvulkan.so on Android — 在现有唯一的 resolver 里,libvulkan.so.1失败后再试libvulkan.so;失败信息列出试过的文件名。没有新建第二个 resolver。98da6e2test: add an Android bench host —test/Sdcb.SimdPaddleOCR.AndroidBench(net10.0-android,CoreCLR,引用库的 net10.0,IsPackable=false)。run.ps1把“编译 → adb install → 推模型和数据集 → 启动 → 等.done→ 拉回结果/logcat”收成一条命令。模式:--caps、--e2e(与 Tests 相同的选项和 JSON 字段,cmp.ps1/steady.ps1/--summarize直接可用)、--detprof/--recprof、--detcmp/--reccmp、--conc,以及两个二分工具--layers(把 GPU 图截断到节点 k 并读回该节点输出)和--gemm(单个 GEMM 对 fp64 参考)。GpuBench 链接了其中不依赖安卓的文件,方便在桌面取对照值。没有加进.slnx和 CI(需要 android workload)。b865f97、b294c2a宿主工具:GEMM 变体扫描、FMA 峰值探针、gprof.ps1、CPU/Vulkan 交替对比ab.ps1,以及两个启动竞态修复。正确性(原样上机三档全错:DET 概率图几乎全零,会话却报告 GPU 正常)
981b6fedirect-load gemm_nc where no narrow subgroup can be pinned — Adreno 编译器会算错gemm_nc的寄存器预取(下一块 A/B 先读进f16vec4数组)。单测 6400×32×64 有 188871/204800 个输出错;换 fma 形式、换 select 都无效,去掉预取就完全正确。gemm_nc.comp新增-DDIRECT,编成gemm_nc_d/gemm_nc_ds;只在锁不住 8 lane 时使用。UHD 770 仍用原gemm_nc/gemm_nc_s(spv 不变)。04f38bcpad fp16 constant uploads to whole 16 bytes — 标量常量原来只上传 2 字节,elem4按uint读b[0]。桌面驱动不查越界;Adreno 按绑定范围查,越界读返回 0,CLS 头里乘常量那步整个变 0(tiny CLS 306/1022,连带 REC CER 57%)。ConstF16/VecF16补齐到 8 个 half 的整数倍,只追加 0,所有内核读到的值不变。afb0f38keep the im2col matrix within one binding's maxStorageBufferRange — medium DET 960×960 的 7×7 卷积 im2col 有 180 MB,超出 128 MB 后读成 0,大图丢框(CER 13.2%)。放不进一个绑定时改走convk_dot;仍会超范围(slab 或无法走直接卷积的 im2col)时,在运行前抛NotSupportedException,会话回退 CPU,而不是返回半张零图。桌面范围是 GB 级,不触发。调优(每项都做了同一时段交替 A/B,数据在提交说明里)
e228638fp16 shared tiles with BK 8 in the direct-load gemm_nc — 11 个最热 GEMM 形状合计 109.5 → 76.2 ms(0.44 → 0.64 TFLOPS,峰值的 31%)。暂存值本来就是 fp16,每个输出仍按 K 递增逐个 fma,结果逐位不变(端到端 100/100);非 SIMPLE 尾部单独用 fp32 暂存,累加器不会先被截成 fp16。8961dc2two-phase spatial mean at any size on wide no-coopmat parts —_ncWide下小图也走reduce_hw4;medium REC 8×480 269.7 → 243.3 ms。9156482direct kxk conv at every K on wide no-coopmat parts —_ncWide下 kxk 稠密卷积默认一律走直接卷积(SIMD_OCR_CONVD_KMAX仍可覆盖);medium DET 960 304.0 → 299.3,640×960 216.9 → 206.2。Auto
a318575Auto takes the no-coopmat tier where subgroups cannot go below 64 —GpuBackend.UsesGpu多了dev.SubgroupMin >= 64这一条(IsGpuBackend、REC 的 16 行批次跟着变)。只按能力判定,没有写死骁龙 / Adreno。UHD 770(能锁 8 lane)和其它最小 subgroup < 64 的无矩阵设备,Auto仍走 CPU。SIMD_OCR_BACKEND=cpu仍可强制 CPU。文档
23e53e2docs/vulkan-8gen3.md;docs/perf.md加链接;中英文 README 的支持范围表说明了各系统的 Vulkan 加载库名、安卓宿主,以及桌面路径未变。手机上的实测
端到端(
--workers 4 --warmup 1,n=99,CPU / Vulkan 同一时段交替,median ms/图)热状态 2–3(medium 全程 3,机身明显发烫)。发热是最大的误差来源:同一版本 medium Vulkan 在热状态 2 时是 833 ms,热状态 3 时是 1460 ms。工作集峰值 medium CPU 约 1.60 GB,Vulkan 约 1.09 GB。
精度(相对同一台手机的 CPU)
手机 CPU 三档与 Windows 基线完全一致;CLS 三档满分。
纯 GPU(
gprof.ps1,逐 dispatch 时间戳取最小,ms)试过但没留下
conv_dw4t:更慢(medium DET 303.7 → 317.0)。StreamCuts切分比例重扫后保持原值。希望 reviewer 重点看的地方
GpuGraph里会在桌面执行的只有两处——fp16 常量补齐(ConstF16/VecF16只追加 0),以及按maxStorageBufferRange的门槛和运行前保护(桌面范围是 GB 级,理论上不触发)。请确认没有别的路径会被_ncWide或新门槛误伤。_ncWide = !sg8的范围:内核选择(gemm_nc_d*、reduce、直接卷积)挂在“无矩阵且锁不住 8 lane”上,比Auto的条件(SubgroupMin >= 64)宽。也就是说:无协作矩阵、最小 subgroup 16/32 的设备(例如不带协作矩阵驱动的 NVIDIA、Mali)在显式 Vulkan 下也会用这套内核,但它们没测过。Auto规则:SubgroupMin >= 64是否足够保守?目前命中的已知设备只有 Adreno;只支持 wave64 的老 AMD GCN(无协作矩阵)理论上也会命中,但没测过。gemm_nc.comp的DIRECT分支:共享内存改成float16_t、尾部用单独的 fp32 数组St;不带DIRECT的构建宏展开后与原文一致(spv 逐字节验证过)。NotSupportedException都发生在BuildPlan里,也就是任何 dispatch 之前。run.ps1的启动/等待逻辑、global.json固定 10.0.x SDK、DebugType=embedded(绕开 CoreCLR 安卓打包的 MSB4096)、UseMonoRuntime=false(CoreCLR 才有 AdvSimd)。测试
dotnet build src/Sdcb.SimdPaddleOCR -c Release(net10.0 与 netstandard2.0 都过)dotnet test test/Sdcb.SimdPaddleOCR.UnitTests -c Release96/96;-p:UseNs20Library=true96/96--conctiny 4/8 线程、small 4、medium 4、Autotiny 4,mismatches=0,日志无 fallback;medium 4 线程工作集峰值 980 MB,未被系统杀--backend autousesGpu=True,与--backend vulkan100/100;Auto+SIMD_OCR_BACKEND=cpu与 CPU 100/100build-shaders.bat后git status干净(所有 spv 逐字节一致),se_fused_dbg.spv已删、未提交feature/2.0交替 2 轮:--engine vulkanmedian tiny 18.0/19.9 → 18.5/20.7、small 26.7/30.3 → 29.2/28.5、medium 36.5/37.6 → 37.3/37.3;精度 741 / 950 / 1006 不变;三档逐图 100/100;--engine sharptiny / small 100/100SubgroupMin >= 64,路由不变,只受 fp16 常量补齐这一处共享改动影响,建议有机会复测一轮--engine vulkan复现