Repository navigation
Metal: pure-C# macOS GPU backend (DET/CLS/REC full-graph schedule) - #21
Conversation
sdcb
left a comment
There was a problem hiding this comment.
Review 总结
整体结构和 Vulkan seam 对齐得很好:纯托管 MetalSchedule 加 grow-only 共享 arena,所有 GPU 失败路径都会永久回落 CPU,SE 拆成两阶段也写清了原因。下面按严重程度列出问题,具体位置见行内评论。
需要修(正确性)
mm_ic_sg隐式 GEMM 分支丢了 residual / dual-write 绑定(MetalGraph.cs两处icconv)。flags仍带 bit1(res)和 bit3(dual),kernel 也会照做,但 buffer(3)、buffer(5) 绑的是arena+0:residual 会读到错误数据,o2 会写坏 arena 偏移 0 上的张量。第二个分支还没有 gatescalarBias == 0,而scalarBias = 8u正好和 mm kernel 的 dual-write bit3 冲突。当前三档模型的行数一致,可能只是没走到这些组合,属于潜伏 bug。- Auto 在所有 macOS 上都选 Metal,包括 Intel Mac(AMD / Intel GPU)。GEMM kernel 写死了 32-lane simdgroup(
tid >> 5),而 Intel Mac 上既没测过性能也没测过正确性。Vulkan 那边对不达标的设备会让 Auto 留在 CPU。建议 Auto 只在 Apple Silicon(SupportsApple7+,且ThreadExecutionWidth == 32)上选 Metal,其余机器要显式指定OcrBackend.Metal。
应该修(资源 / 性能)
MetalGraphModel.Release只释放 buffer,MtlLibrary和约 40 个 PSO 从不释放。model 每次被重新 Acquire 都会泄漏一整套。另外每个 model(det/cls/rec)都要从源码现编一遍整个 MSL library,建议像VkDevice.GetPipeline那样把 library 和 PSO 缓存到MtlDevice上。- Runner 用的是 serial compute encoder,本身就保证 dispatch 按序执行、内存可见,所以每个 rec 后面的
memoryBarrierWithScope:是多余的 objc_msgSend。更关键的是,interleave(多 unit 按 level 交织)在 serial encoder 下拿不到任何并发。建议 interleave 模式改用BeginCommands(concurrent: true),每个 level 放一次 barrier,然后实测。paravirt 上 dispatch 地板约 45µs,这里可能有明显收益。
可维护性
- 把类名归一化后,
MetalGraph.cs和GpuGraph.cs约 88% 的行相同(2438 行对 2375 行,只有约 290 行不同)。MetalSession.cs和GpuSession.cs几乎逐行一样,ResolveCtcProjection算上InferenceSession已经是第三份拷贝。这个 PR 不要求合并,但建议开 issue 跟进:把 BuildPlan 抽成后端无关的 planner,后端只提供 pipeline 选择和Emit;session 抽成公共基类。否则以后每修一个融合 bug 都得同步改两份。 - 有一批死代码和过时注释(sg32/_lite 标志、SupportsApple*、concurrent 重载、intConstants、调试 reader、se_fused 注释等),见行内评论。
文档和验证
docs/metal-m4.md里的--dtr、--mprobe、--mgemm都不在仓库中,结论没法复现。CER 0.67% 对 0.14% 和“文本 3 图各差 1 字符”对不上,请确认口径。“单元测试在非 macOS 上 SKIP”不成立,这个 PR 没有新增任何 Metal 测试。- 建议加一个 macOS arm64 CI smoke:CI 的
macos-26runner 就是 paravirt GPU,用显式的OcrBackend.Metal对几张图,跟 CPU 比较文本,并断言 stderr 里没有[metal] ... fallback。现在 Auto 回落 CPU 只打一行 stderr,CI 是静默全绿的。
… device-level MSL lib/PSO cache, concurrent interleave encoder, GpuSessionBase dedupe - mm_ic_* emit sites bound residual/dual tensors as (arena,0)/(arena,0,2) while flags could carry bit1/bit3 — bind the real slots, and gate the second site on scalarBias==0 (the scalar bias channel is repurposed as the dual-write flag in the mm kernels). - Auto no longer selects Metal on Intel Macs: Auto now requires Arm64+Apple7; explicit OcrBackend.Metal still allowed elsewhere via UsesGpu (probe validates simdgroup matrix live). - Move the embedded MSL library + ~40 PSOs to MtlDevice (OcrLibrary / OcrPipe): one compile per process instead of per MetalGraphModel, released on device Dispose. - Interleave now records through a concurrent compute encoder with one buffer barrier per level — under a serial encoder the per-rec barrier was a no-op and the level overlap unreachable. Serial path drops the dead per-rec barrier. - Extract GpuSessionBase + IOcrGraphRunner: GpuSession/MetalSession (identical 324-line bodies) become thin constructors over the shared session logic; GpuUnitsReady/MetalUnitsReady unify on CtcUnitsReady. - ProbeSimdMatrix now verifies PSO geometry (execWidth=32, maxTg>=256) and compares a full 64x64 tile against a reference GEMM. - Comment/doc nits: correct DWA/NOIC direction comments, mm_sg_dd gate comment, se two-phase explanation (relaxed atomics are universal MSL, not paravirt-specific), mm_ic pad-k semantics, MtlDataType UInt=30 fix (was Int2), MTLCreateSystemDefaultDevice +1 rule, doc CER metric and tool references.
还需要改的上一轮的绑定、Auto 门控、PSO 泄漏和 concurrent encoder 在 1.
|
|
第二轮意见全部采纳( 1. 精度口径已用同一轮 JSON 重报(docs/metal-m4.md §精度):同一轮跑出三口径——逐图 2. CI smoke 加强 + 矛盾句删除:osx-arm64 Metal smoke 在 3. 性能表已在 review-HEAD 重测(同一轮、同机):tiny 68.8→39.3(1.75×)、small 218.4→54.6(4.00×)、medium 693.0→107.0(6.47×)。注:medium 相对上轮 178.4ms 的额外收益正来自 review 修订的 concurrent encoder + device 级 MSL lib/PSO cache(rec_graph 112.6→68.9ms)。 4. |
…dule) - ObjC/Metal P/Invoke layer (no native deps), MetalBackend probe + auto-select on macOS, MetalSession/MetalRunner/MetalGraph mirroring the Vulkan seam: compiled managed MetalSchedule (shape-keyed LRU, zero-cost per run) + grow-only shared device arena. - ~40 MSL compute kernels embedded as resources; GEMM family incl. simdgroup MMA 64x64 (mm_sg_dd direct-load, mm_ic_sg implicit-GEMM conv), dot-product conv kernels, SE squeeze-excite two-phase split (paravirt lacks acq_rel atomics). - OcrBackend.Metal (additive); session-create fallback to CPU permanent. - GpuBench PipelineBench: System.Drawing -> ImageSharp (cross-platform). - docs/metal-m4.md: 3-tier baselines (tiny 1.9x / small 4.7x / medium 5.2x vs sharp, 4w), correctness gates (lines identical, --conc 0/72, --dtr bitwise clean), tuning log and paravirt caps.
… device-level MSL lib/PSO cache, concurrent interleave encoder, GpuSessionBase dedupe - mm_ic_* emit sites bound residual/dual tensors as (arena,0)/(arena,0,2) while flags could carry bit1/bit3 — bind the real slots, and gate the second site on scalarBias==0 (the scalar bias channel is repurposed as the dual-write flag in the mm kernels). - Auto no longer selects Metal on Intel Macs: Auto now requires Arm64+Apple7; explicit OcrBackend.Metal still allowed elsewhere via UsesGpu (probe validates simdgroup matrix live). - Move the embedded MSL library + ~40 PSOs to MtlDevice (OcrLibrary / OcrPipe): one compile per process instead of per MetalGraphModel, released on device Dispose. - Interleave now records through a concurrent compute encoder with one buffer barrier per level — under a serial encoder the per-rec barrier was a no-op and the level overlap unreachable. Serial path drops the dead per-rec barrier. - Extract GpuSessionBase + IOcrGraphRunner: GpuSession/MetalSession (identical 324-line bodies) become thin constructors over the shared session logic; GpuUnitsReady/MetalUnitsReady unify on CtcUnitsReady. - ProbeSimdMatrix now verifies PSO geometry (execWidth=32, maxTg>=256) and compares a full 64x64 tile against a reference GEMM. - Comment/doc nits: correct DWA/NOIC direction comments, mm_sg_dd gate comment, se two-phase explanation (relaxed atomics are universal MSL, not paravirt-specific), mm_ic pad-k semantics, MtlDataType UInt=30 fix (was Int2), MTLCreateSystemDefaultDevice +1 rule, doc CER metric and tool references.
…x OcrBackend.Auto comment
第三轮 review(HEAD
|
| 意见 | 状态 |
|---|---|
| 精度口径矛盾 | 已改。exact_lines、CER 和相对 CPU 的逐图文本差来自同一轮。tiny 最差的 img-078 差 38 个字符,和「CLS 1 图翻转」对得上。 |
CI 只比 detected |
test.yml 不改(维护者决定)。「无 CI Metal 依赖」一句已删,但换上去的描述多写了 texts 断言,见第 1 条。 |
| 性能表不是最终二进制 | 已在 review 修订后的 HEAD 重测。 |
OcrBackend.Auto 注释 |
已改。 |
3. 共享代码对 Vulkan / CPU 的影响(Arc B580 实测)
GpuSession移到GpuSessionBase:新旧代码归一化后逐行比对,只有三处不同:类型名、ObjectDisposedException的名字、日志前缀改成参数(Vulkan 仍是[gpu])。逻辑是原样搬移。- Windows 选路:
Auto仍走 Vulkan;--engine metal回到 CPU(backend=metal,结果与 sharp 一致)。 - 构建与测试:lib 的 net10 / netstandard2.0、Tests、GpuBench(改用 ImageSharp 后)都能编;
dotnet test test/Sdcb.SimdPaddleOCR.UnitTests -c Release96/96。
A/B:Ryzen 9 5950X + Arc B580,基线 b5c6ef2(合入 #22 后),本 PR 1bf0350。两个 worktree 分别编译,--workers 4 --benchmark-kind simd --warmup 1,同一 dataset/ 100 张,n=99,Vulkan 交替跑并逐轮换先后。median ms/图:
| 模型 | base | PR | 结论 |
|---|---|---|---|
| tiny | 24.0 / 19.1 / 18.4 | 20.8 / 17.0 / 19.8 | 持平 |
| small | 31.0 / 29.8 / 30.3 | 28.7 / 28.3 / 28.9 | 持平 |
| medium(7 轮) | 最小 54.4,各轮中位 56.6 | 最小 54.0,各轮中位 56.3 | 持平 |
- medium 前两轮 PR 偏慢(70.5 / 69.5),补跑 4 轮后都在 54–57,属于机器噪声。
- sharp 三档持平:49.5 / 49.6、193.5 / 199.1、538.3 / 535.1。
--engine autotiny 19.0 ms,走 GPU。- 正确率两边相同:Vulkan 740 / 950 / 1003,sharp 742 / 950 / 1004。WS peak 两边相同。
- base 与 PR 逐图比较
hash/detected/texts:Vulkan 和 sharp 三档全部 100/100 一致。
4. 小问题(不阻塞)
GpuRunner.cs 删掉了 GpuUnitsReady 委托,但它上面那段 /// <summary>Streamed-run callback… 注释还留着,现在紧贴在 GpuDetGraph 的类注释前,变成两段 summary。委托已经移到 IBatchedCtcSession.cs 的 CtcUnitsReady,这段注释可以删掉。
|
第三轮意见已采纳( |
Summary
Pure-C# Metal GPU backend for macOS, mirroring the Vulkan seam from #18: a compiled managed
MetalSchedule(shape-keyed LRU, no Metal objects — zero-cost to rebuild per run) over a grow-only shared device arena, with ~40 MSL compute kernels embedded as resources.OcrBackend.Metalis additive;Autopicks Metal on macOS when a device probes, and every GPU failure path falls back to the CPU interpreter permanently for that call.Architecture follows the plan doc:
MetalBackend(device probe + selection rules incl.SIMD_OCR_BACKENDenv) →MetalSession(IOcrSession+IBatchedCtcSession) →MetalGraph(BuildPlan emit: conv/matmul/SE/pool/elemwise/xform dispatch list) →MetalRunner(per-run CB record + one commit + fence; CTC head stays CPU-side bit-exact).Measured on this dev VM (Apple M4 paravirt GPU, 8 vCPU, .NET 10, 99-image
dataset/, 4w, median ms/img):Correctness gates: detected line counts identical to CPU on all 3 models;
--conc4-thread ×72 runs 0 mismatch;--dtr228-node det graph bitwise-deterministic across 8 runs; 3 images differ by fp16-threshold-edge pixels (same known-limitation class as Vulkan). Full numbers, caps, and per-optimization tuning log:docs/metal-m4.md.Key measured details a reviewer should know:
mm_sg_dd64×64 simdgroup-MMA tile loads operands directly from device memory (no threadgroup staging) — gatedK%16==0; edge-tile overread is safe by construction (arena ≥1MB tail slack, W padded /128, epilogue masksmm<p.M/nn<p.N).mm_ic_sgis an implicit-GEMM conv: gathers tap-major during A staging so kxk convs never materialize im2col (det medium 960² 86→80ms, ~1.28GB less traffic).memory_order_relaxed, so cross-threadgroup ticket handshakes are impossible — SE is two-phasese_part+ buffer barrier +se_join.SIMD_OCR_DWA(flat depthwise loses to staging),SIMD_OCR_IC32(BK=32 mixed), plus the standardNO*debug gates.PipelineBenchswitchedSystem.Drawing.Bitmap→ ImageSharp so the bench runs on macOS/Linux.Also fixes pre-existing nullable warnings; unit tests 95/96 net10.0 (only failure is the pre-existing AVX-only
Depthwise9PackedMatchesReferenceon ARM64, unrelated).Link to Devin session: https://app.devin.ai/sessions/680b2404b8304ce1979919406191fa47
Open in Devin Desktop: https://app.devin.ai/desktop/session/680b2404b8304ce1979919406191fa47?variant=devin
Requested by: @sdcb