Repository navigation
feat(operators): NE16 requantized conv benchmark generator - #23
Merged
runwangdl merged 1 commit intoJul 30, 2026
Merged
Conversation
NE16 has three native convolution modes (CONV3x3, CONV3x3_DW, CONV1x1) and Deeploy ships a functional kernel test for each, but they are all small enough that dispatch overhead dominates -- the existing Dense_2D_RQ is 16x16x8x8, about 13% NE16 utilisation. There is no way to measure NE16 throughput from them. Add an operator generator that emits one Conv -> RequantShift test per mode, sized on the NE16 granularity (c_in % 16, c_out % 32, 8-bit weights) so the MAC array stays fed: dense 3x3 64 -> 64 @32x32 37.8 MMAC dw 3x3 128ch @32x32 1.2 MMAC pw 1x1 256 -> 128 @16x16 8.4 MMAC python3 scripts/gen_ne16_bench_tests.py --deeploy <deeploy-checkout> Deeploy's integer path uses RequantShift, which ONNX Runtime cannot execute, so run_inference is overridden with an integer NumPy reference instead of the base class's ORT golden. That reference reproduces the three shipped fixtures (Dense_2D_RQ, DW_2D_RQ, PW_2D_RQ/Regular_RQ) bit-exactly, including the rounding term, so it is checked against the semantics Deeploy actually implements rather than against my reading of them. Two properties the generated fixtures need, both learned from the existing Dense_2D_RQ_NE16Bench fixture, which has neither: - no Conv bias. After the requant merge the node inputs are mapped positionally to [data_in, weight, mul, add], so a third Conv input displaces mul/add and the NE16 parser rejects the node. - unsigned activations. NE16's CONFIG0 has no signed-input flag, so a signed input silently prevents the offload and the op falls back to the cluster. mul/add are fitted per channel from the real accumulator distribution rather than drawn at random. Random requant parameters clipped 74-94% of the outputs to the int8 bounds, which makes the fixture insensitive to arithmetic errors; the fitted version saturates 0.1%. Verified end to end on the gap9-ne16 branch: all three generate, build, and emit ne16_nnx_dispatch in the generated Network.c (dense and dw need --enable-3x3).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
NE16 has three native convolution modes (CONV3x3, CONV3x3_DW, CONV1x1) and Deeploy ships a functional kernel test for each, but they are all small enough that dispatch overhead dominates -- the existing Dense_2D_RQ is 16x16x8x8, about 13% NE16 utilisation. There is no way to measure NE16 throughput from them.
Add an operator generator that emits one Conv -> RequantShift test per mode, sized on the NE16 granularity (c_in % 16, c_out % 32, 8-bit weights) so the MAC array stays fed:
dense 3x3 64 -> 64 @32x32 37.8 MMAC
dw 3x3 128ch @32x32 1.2 MMAC
pw 1x1 256 -> 128 @16x16 8.4 MMAC
python3 scripts/gen_ne16_bench_tests.py --deeploy
Deeploy's integer path uses RequantShift, which ONNX Runtime cannot execute, so run_inference is overridden with an integer NumPy reference instead of the base class's ORT golden. That reference reproduces the three shipped fixtures (Dense_2D_RQ, DW_2D_RQ, PW_2D_RQ/Regular_RQ) bit-exactly, including the rounding term, so it is checked against the semantics Deeploy actually implements rather than against my reading of them.
Two properties the generated fixtures need, both learned from the existing Dense_2D_RQ_NE16Bench fixture, which has neither:
mul/add are fitted per channel from the real accumulator distribution rather than drawn at random. Random requant parameters clipped 74-94% of the outputs to the int8 bounds, which makes the fixture insensitive to arithmetic errors; the fitted version saturates 0.1%.
Verified end to end on the gap9-ne16 branch: all three generate, build, and emit ne16_nnx_dispatch in the generated Network.c (dense and dw need --enable-3x3).