Skip to content

feat(operators): NE16 requantized conv benchmark generator - #23

Merged
runwangdl merged 1 commit into
pulp-platform:develfrom
runwangdl:feat/ne16-conv-benchmarks
Jul 30, 2026
Merged

runwangdl merged 1 commit into
pulp-platform:develfrom
runwangdl:feat/ne16-conv-benchmarks

Conversation

@runwangdl

Copy link
Copy Markdown
Collaborator

NE16 has three native convolution modes (CONV3x3, CONV3x3_DW, CONV1x1) and Deeploy ships a functional kernel test for each, but they are all small enough that dispatch overhead dominates -- the existing Dense_2D_RQ is 16x16x8x8, about 13% NE16 utilisation. There is no way to measure NE16 throughput from them.

Add an operator generator that emits one Conv -> RequantShift test per mode, sized on the NE16 granularity (c_in % 16, c_out % 32, 8-bit weights) so the MAC array stays fed:

dense 3x3 64 -> 64 @32x32 37.8 MMAC
dw 3x3 128ch @32x32 1.2 MMAC
pw 1x1 256 -> 128 @16x16 8.4 MMAC

python3 scripts/gen_ne16_bench_tests.py --deeploy

Deeploy's integer path uses RequantShift, which ONNX Runtime cannot execute, so run_inference is overridden with an integer NumPy reference instead of the base class's ORT golden. That reference reproduces the three shipped fixtures (Dense_2D_RQ, DW_2D_RQ, PW_2D_RQ/Regular_RQ) bit-exactly, including the rounding term, so it is checked against the semantics Deeploy actually implements rather than against my reading of them.

Two properties the generated fixtures need, both learned from the existing Dense_2D_RQ_NE16Bench fixture, which has neither:

  • no Conv bias. After the requant merge the node inputs are mapped positionally to [data_in, weight, mul, add], so a third Conv input displaces mul/add and the NE16 parser rejects the node.
  • unsigned activations. NE16's CONFIG0 has no signed-input flag, so a signed input silently prevents the offload and the op falls back to the cluster.

mul/add are fitted per channel from the real accumulator distribution rather than drawn at random. Random requant parameters clipped 74-94% of the outputs to the int8 bounds, which makes the fixture insensitive to arithmetic errors; the fitted version saturates 0.1%.

Verified end to end on the gap9-ne16 branch: all three generate, build, and emit ne16_nnx_dispatch in the generated Network.c (dense and dw need --enable-3x3).

NE16 has three native convolution modes (CONV3x3, CONV3x3_DW, CONV1x1) and
Deeploy ships a functional kernel test for each, but they are all small enough
that dispatch overhead dominates -- the existing Dense_2D_RQ is 16x16x8x8, about
13% NE16 utilisation. There is no way to measure NE16 throughput from them.

Add an operator generator that emits one Conv -> RequantShift test per mode,
sized on the NE16 granularity (c_in % 16, c_out % 32, 8-bit weights) so the MAC
array stays fed:

  dense  3x3  64 -> 64   @32x32  37.8 MMAC
  dw     3x3  128ch      @32x32   1.2 MMAC
  pw     1x1  256 -> 128 @16x16   8.4 MMAC

  python3 scripts/gen_ne16_bench_tests.py --deeploy <deeploy-checkout>

Deeploy's integer path uses RequantShift, which ONNX Runtime cannot execute, so
run_inference is overridden with an integer NumPy reference instead of the base
class's ORT golden. That reference reproduces the three shipped fixtures
(Dense_2D_RQ, DW_2D_RQ, PW_2D_RQ/Regular_RQ) bit-exactly, including the rounding
term, so it is checked against the semantics Deeploy actually implements rather
than against my reading of them.

Two properties the generated fixtures need, both learned from the existing
Dense_2D_RQ_NE16Bench fixture, which has neither:

  - no Conv bias. After the requant merge the node inputs are mapped positionally
    to [data_in, weight, mul, add], so a third Conv input displaces mul/add and
    the NE16 parser rejects the node.
  - unsigned activations. NE16's CONFIG0 has no signed-input flag, so a signed
    input silently prevents the offload and the op falls back to the cluster.

mul/add are fitted per channel from the real accumulator distribution rather than
drawn at random. Random requant parameters clipped 74-94% of the outputs to the
int8 bounds, which makes the fixture insensitive to arithmetic errors; the fitted
version saturates 0.1%.

Verified end to end on the gap9-ne16 branch: all three generate, build, and emit
ne16_nnx_dispatch in the generated Network.c (dense and dw need --enable-3x3).
@runwangdl
runwangdl requested a review from Victor-Jung as a code owner July 30, 2026 01:42
@runwangdl
runwangdl merged commit 4f4b5d8 into pulp-platform:devel Jul 30, 2026
11 checks passed
@runwangdl
runwangdl deleted the feat/ne16-conv-benchmarks branch July 30, 2026 01:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant