PR #272 review follow-up for parquet_writer.go: #272 (comment)
The three alternating Linux ARM64 ingestion-only repetitions against #270 reproduced a roughly 15% durable-throughput regression. Optimize the measured typed writer/copy/allocation path while retaining format 3, typed values, nanosecond UTC columns, complete trace ordering, durable acknowledgements, and conservative admission bounds.
Median baseline -> DuckDB 2: 484,312 -> 411,796 acknowledged rows/s; 3,826 -> 4,123 Go allocated bytes/row; 1.316 -> 1.478 client+server CPU cores; 36.54 -> 42.09ms export p95. CPU cores rose about 12%; CPU time per acknowledged row rose about 32%, since throughput also declined. All six runs had zero export errors and passed storage verification. The fixture is 16 shared-transport clients, 1,000 rows/export, equal signals, 4 CPUs/6GiB, ten-second intervals, in-process clients/server, no dashboard queries. It is not the harness mixed-p8 workload. ValueBytes was about 2% cumulative CPU in the first current profile; its per-key charge is accounting, not allocated bytes.
Validation: repeat baseline/current alternating runs at least three times with CPU/post-load heap profiles, then repeat realistic mixed-p8 and cooldown RSS measurements. Preserve deep typed/nested/binary round trips, shredded residual types, compaction and crash/corruption checks. Profiles are generated with FANOUT_BENCH_PROFILE_DIR and the transportbench build tag; include inspectable reports/artifacts with the final measurements.
Implementation is authorized in PR #272; this issue tracks the fix and its validation.
PR #272 review follow-up for
parquet_writer.go: #272 (comment)The three alternating Linux ARM64 ingestion-only repetitions against #270 reproduced a roughly 15% durable-throughput regression. Optimize the measured typed writer/copy/allocation path while retaining format 3, typed values, nanosecond UTC columns, complete trace ordering, durable acknowledgements, and conservative admission bounds.
Median baseline -> DuckDB 2: 484,312 -> 411,796 acknowledged rows/s; 3,826 -> 4,123 Go allocated bytes/row; 1.316 -> 1.478 client+server CPU cores; 36.54 -> 42.09ms export p95. CPU cores rose about 12%; CPU time per acknowledged row rose about 32%, since throughput also declined. All six runs had zero export errors and passed storage verification. The fixture is 16 shared-transport clients, 1,000 rows/export, equal signals, 4 CPUs/6GiB, ten-second intervals, in-process clients/server, no dashboard queries. It is not the harness mixed-p8 workload. ValueBytes was about 2% cumulative CPU in the first current profile; its per-key charge is accounting, not allocated bytes.
Validation: repeat baseline/current alternating runs at least three times with CPU/post-load heap profiles, then repeat realistic mixed-p8 and cooldown RSS measurements. Preserve deep typed/nested/binary round trips, shredded residual types, compaction and crash/corruption checks. Profiles are generated with FANOUT_BENCH_PROFILE_DIR and the transportbench build tag; include inspectable reports/artifacts with the final measurements.
Implementation is authorized in PR #272; this issue tracks the fix and its validation.