Skip to content

Add ALP Encoding blog - #195

Merged
alamb merged 67 commits into
apache:productionfrom
sdf-jkl:alp-blog
Sep 22, 2026
Merged

alamb merged 67 commits into
apache:productionfrom
sdf-jkl:alp-blog

Conversation

@sdf-jkl

@sdf-jkl sdf-jkl commented Jul 8, 2026 •

Copy link
Copy Markdown
Member

@sdf-jkl

sdf-jkl commented Jul 8, 2026

Copy link
Copy Markdown
Member Author

@alamb still need to add some charts, but the text is ready for review.

cc. @devanbenz

@alamb

alamb commented Jul 10, 2026

Copy link
Copy Markdown
Collaborator

Amazing -- thank you @sdf-jkl

@alamb

alamb commented Jul 20, 2026

Copy link
Copy Markdown
Collaborator

BTW there is a vote thread to add ALP here: https://lists.apache.org/thread/hgmd58wrv9yoopcrf61m1bg211l65tbt

its-happening

@alamb alamb left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @sdf-jkl -- this is a great start. I am sorry it has taken so long to get you feedback on this.

I have left quite a few comments. Let me know if they make sense.

My core feedback is that I recommend we reframe this piece to focus and explain what was added to Parquet, rather than ALP more broadly. I am thinking that this blog will likely the first page that someone will find when they search for ALP in parquet (or that an agent will find / summarize for them). Ideally it has all the info about when to choose ALP (what data / systems are good candidates) and then a technical overview of how it works.

Finally, I would love to be a coauthor on this post with you if you are willing

Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
@sdf-jkl

sdf-jkl commented Aug 2, 2026

Copy link
Copy Markdown
Member Author

Finally, I would love to be a coauthor on this post with you if you are willing

100% 🔥

@alamb

alamb commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator

Ok great -- thank you.

I plan to focus on getting the code / PR ready in rust first and then come back to this one. If you haven't had a chance to incorporate the feedback by then, no worries, I can take a direct pass; That is probably sometime next week (hopefully)

@alamb

This comment was marked as outdated.

@sdf-jkl

sdf-jkl commented Aug 7, 2026

Copy link
Copy Markdown
Member Author

This is a branch with the blog bench - https://github.com/sdf-jkl/arrow-rs/tree/alp-benchmarks

Example output on my machine:

Details

Benchmark environment

Environment Value
UTC timestamp 2026-08-07T22:10:51Z
Commit 0ffde54952cc6fc6ac0fd25a739b4a43022a770b
Worktree clean
CPU AMD Ryzen AI 9 HX PRO 470 w/ Radeon 890M
Architecture x86_64
SIMD ISA AVX-512F, AVX2, AVX
Logical CPUs 24
OS and kernel Linux 6.19.10-300.fc44.x86_64
CPU governor powersave
Rust rustc 1.96.1 (31fca3adb 2026-06-26)
LLVM 22.1.2
Cargo cargo 1.96.1 (356927216 2026-06-26)
RUSTFLAGS -C target-cpu=native
Dataset archive SHA-256 1070817918b9e2b2cc7003995927bd04fe7b942045383913d3f40437eda29831

Parquet compression results

Dataset Parquet choice Compression (GB/s) Decompression (GB/s) Compressed size (bits/value)
arade4 PLAIN 74.479 74.591 64.01
arade4 PLAIN + ZSTD 0.653 1.640 37.39
arade4 ALP 2.662 34.503 24.99
basel_temp_f PLAIN 67.349 74.978 64.01
basel_temp_f PLAIN + ZSTD 0.461 1.655 23.07
basel_temp_f ALP 1.362 22.070 29.23
basel_wind_f PLAIN 61.719 74.922 64.01
basel_wind_f PLAIN + ZSTD 0.583 1.668 18.53
basel_wind_f ALP 2.444 26.226 29.87
bird_migration_f PLAIN 62.842 106.453 64.01
bird_migration_f PLAIN + ZSTD 0.633 1.774 23.49
bird_migration_f ALP 2.657 25.455 20.24
bitcoin_f PLAIN 83.667 113.938 64.07
bitcoin_f PLAIN + ZSTD 0.570 1.640 50.01
bitcoin_f ALP 1.789 29.517 27.18
bitcoin_transactions_f PLAIN 64.196 74.489 64.01
bitcoin_transactions_f PLAIN + ZSTD 1.088 2.000 47.96
bitcoin_transactions_f ALP 2.331 20.391 41.27
city_temperature_f PLAIN 75.474 76.966 64.01
city_temperature_f PLAIN + ZSTD 0.568 1.372 17.67
city_temperature_f ALP 2.788 34.158 10.80
cms1 PLAIN 76.093 30.520 64.01
cms1 PLAIN + ZSTD 0.666 1.526 26.84
cms1 ALP 1.411 14.770 35.19
cms25 PLAIN 71.629 73.437 64.01
cms25 PLAIN + ZSTD 0.834 1.841 58.11
cms25 ALP 2.202 24.216 41.17
cms9 PLAIN 76.616 77.507 64.01
cms9 PLAIN + ZSTD 0.719 1.471 11.71
cms9 ALP 2.803 33.547 12.16
food_prices PLAIN 68.581 74.763 64.01
food_prices PLAIN + ZSTD 0.580 1.353 18.13
food_prices ALP 1.154 20.379 23.20
gov10 PLAIN 75.242 76.877 64.01
gov10 PLAIN + ZSTD 0.518 1.260 29.12
gov10 ALP 1.783 26.573 29.88
gov26 PLAIN 75.922 76.636 64.01
gov26 PLAIN + ZSTD 12.506 25.474 0.20
gov26 ALP 2.107 94.626 1.40
gov30 PLAIN 75.812 76.574 64.01
gov30 PLAIN + ZSTD 2.203 5.277 4.52
gov30 ALP 1.205 39.760 17.88
gov31 PLAIN 64.361 65.519 64.01
gov31 PLAIN + ZSTD 3.542 8.255 1.65
gov31 ALP 2.551 40.989 6.77
gov40 PLAIN 57.109 59.596 64.01
gov40 PLAIN + ZSTD 8.097 14.910 0.43
gov40 ALP 2.710 60.151 2.59
medicare1 PLAIN 60.313 58.146 64.01
medicare1 PLAIN + ZSTD 0.519 1.364 31.68
medicare1 ALP 1.218 14.356 40.46
medicare9 PLAIN 62.343 63.324 64.01
medicare9 PLAIN + ZSTD 0.644 1.314 11.86
medicare9 ALP 2.497 27.809 12.82
neon_air_pressure PLAIN 73.548 74.262 64.01
neon_air_pressure PLAIN + ZSTD 0.805 2.029 11.85
neon_air_pressure ALP 2.692 34.409 16.48
neon_bio_temp_c PLAIN 74.851 75.700 64.01
neon_bio_temp_c PLAIN + ZSTD 0.563 1.559 16.84
neon_bio_temp_c ALP 2.776 33.096 10.81
neon_dew_point_temp PLAIN 73.178 73.636 64.01
neon_dew_point_temp PLAIN + ZSTD 0.478 1.651 23.73
neon_dew_point_temp ALP 2.728 30.639 13.63
neon_pm10_dust PLAIN 47.279 70.301 64.01
neon_pm10_dust PLAIN + ZSTD 0.848 1.663 7.79
neon_pm10_dust ALP 1.794 34.353 8.41
neon_wind_dir PLAIN 73.187 74.370 64.01
neon_wind_dir PLAIN + ZSTD 0.493 1.486 24.41
neon_wind_dir ALP 2.689 46.726 15.94
nyc29 PLAIN 72.889 70.456 64.01
nyc29 PLAIN + ZSTD 0.615 1.483 24.67
nyc29 ALP 2.434 24.030 40.43
poi_lat PLAIN 72.573 18.951 64.01
poi_lat PLAIN + ZSTD 0.682 1.533 57.78
poi_lat ALP 1.522 10.882 88.19
poi_lon PLAIN 73.593 19.404 64.01
poi_lon PLAIN + ZSTD 0.854 1.774 60.44
poi_lon ALP 1.713 16.118 79.12
ssd_hdd_benchmarks_f PLAIN 81.743 113.024 64.02
ssd_hdd_benchmarks_f PLAIN + ZSTD 0.800 1.772 12.98
ssd_hdd_benchmarks_f ALP 2.662 34.116 16.04
stocks_de PLAIN 73.346 74.265 64.01
stocks_de PLAIN + ZSTD 0.684 1.662 10.07
stocks_de ALP 1.496 33.555 11.20
stocks_uk PLAIN 74.759 76.045 64.01
stocks_uk PLAIN + ZSTD 0.673 1.491 11.29
stocks_uk ALP 0.936 35.561 12.75
stocks_usa_c PLAIN 71.756 73.168 64.01
stocks_usa_c PLAIN + ZSTD 0.717 1.585 8.24
stocks_usa_c ALP 2.714 35.625 7.95
ALL AVG. PLAIN 70.548 71.427 64.01
ALL AVG. PLAIN + ZSTD 1.453 3.183 22.75
ALL AVG. ALP 2.128 31.954 24.27

GB/s is decimal billions of uncompressed input bytes processed per second; higher is better. Compressed size includes Parquet data-page headers but excludes the file footer. Speed processes every value in pages of up to 131072 values and excludes file I/O. PLAIN + ZSTD includes both stages: PLAIN encoding plus ZSTD compression, and ZSTD decompression plus PLAIN decoding. Short pages are repeated for timing stability and normalized to one page.

Random access

Time to decode 100 deterministic, uniformly distributed rows from city_temperature_f (lower is better). Each lookup starts from the encoded page.

Parquet choice 100 random rows (µs)
PLAIN 2.817
PLAIN + ZSTD 75898.583
ALP 10.100

PLAIN and ALP reset the page decoder, skip to the selected row, and decode one value. PLAIN + ZSTD additionally decompresses the complete target page for every independent lookup. Encoded pages are already in memory; file I/O and page lookup are excluded.

30 datasets. Arithmetic mean: PLAIN 64.01, PLAIN + ZSTD 22.75, ALP 24.27 bits/value.
Median ALP: 17.88 bits/value. ALP is 0.28x the size of PLAIN and 1.25x the size of PLAIN + ZSTD by geometric mean.
ALP is smaller than PLAIN + ZSTD on 10/30 datasets.
Arithmetic mean compression/decompression speed in GB/s: PLAIN 70.548/71.427, PLAIN + ZSTD 1.453/3.183, ALP 2.128/31.954.

@alamb

alamb commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator

the output looks great

@alamb

alamb commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator

My hope is in the next day or two to take a pass on this doc and try to incorporate feedback

@sdf-jkl

sdf-jkl commented Aug 12, 2026

Copy link
Copy Markdown
Member Author

I'll address the existing feedback myself, would be down for another review 😃

@alamb

alamb commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator

I'll address the existing feedback myself, would be down for another review 😃

As in you plan to address it and update this PR?

Also I think @prtkgaur was working on a blog post as well and he may wish to help out.

@sdf-jkl

sdf-jkl commented Aug 12, 2026

Copy link
Copy Markdown
Member Author

I'll address the existing feedback myself, would be down for another review 😃

As in you plan to address it and update this PR?

I meant I'll address your existing comment on this PR.

Also I think @prtkgaur was working on a blog post as well and he may wish to help out.

🔥

Comment thread content/en/blog/features/alp_encoding.md Outdated
@alamb

alamb commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

I am working on some diagrams and charts for this post

@sdf-jkl

sdf-jkl commented Aug 14, 2026

Copy link
Copy Markdown
Member Author

Lmk if you're working on the rest of your suggestions too. I was working on the prose earlier and don't want our work to conflict.

@alamb

alamb commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

Lmk if you're working on the rest of your suggestions too. I was working on the prose earlier and don't want our work to conflict.

I was not going to do any pose editing likewise because I didn't want our work to conflict

I'll work on diagrams and charts instead and then when you have something you can push it up -- that way we won't conlflict

@alamb

This comment was marked as outdated.

@alamb

alamb commented Aug 14, 2026 •

Copy link
Copy Markdown
Collaborator

Here are some charts https://docs.google.com/spreadsheets/d/1ZSo8Rx3bPzCGkRn4VNivivZzuiIVUfBFy7QNVXdLfhM/edit

Screenshot 2026-08-14 at 5 34 46 PM Screenshot 2026-08-14 at 5 35 00 PM

I haven't really dug into the numbers yet or tried yet to reproduce the results (aka is gov26 really so much better than ALP, or is there something that )

Screenshot 2026-08-14 at 5 36 00 PM

But I think this gives some ideas

For the actual post, I think we can pick a few representative data sets that reflect the pattern

@sdf-jkl

sdf-jkl commented Aug 14, 2026

Copy link
Copy Markdown
Member Author

I also wanted to add byte stream split + zstd to the bench as it's another way to store floats

@sdf-jkl

sdf-jkl commented Aug 14, 2026

Copy link
Copy Markdown
Member Author

Sorry for taking a while. I can post changes more frequently so we're on the same page.

@alamb

alamb commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

I also wanted to add byte stream split + zstd to the bench as it's another way to store floats

Sure -- I also want to be careful not to treat this as an academic paper (really explore the ins/outs). I think we need enough data to back up the key points, but not a full on treatment. Comparing to other parquet encoding implementations is a good idea

@alamb

alamb commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator

Here is the result of running the benchmark on a gcp mchine

Details

Benchmark environment

Environment Value
UTC timestamp 2026-08-15T09:38:17Z
Commit 0ffde54952cc6fc6ac0fd25a739b4a43022a770b
Worktree dirty
CPU Intel(R) Xeon(R) CPU @ 3.10GHz
Architecture x86_64
SIMD ISA AVX-512F, AVX2, AVX
Logical CPUs 8
OS and kernel Linux 6.17.0-1022-gcp
CPU governor unavailable
Rust rustc 1.96.1 (31fca3adb 2026-06-26)
LLVM 22.1.2
Cargo cargo 1.96.1 (356927216 2026-06-26)
RUSTFLAGS -C target-cpu=native
Dataset archive SHA-256 1070817918b9e2b2cc7003995927bd04fe7b942045383913d3f40437eda29831

Parquet compression results

Dataset Parquet choice Compression (GB/s) Decompression (GB/s) Compressed size (bits/value)
arade4 PLAIN 13.428 4.618 64.01
arade4 PLAIN + ZSTD 0.304 0.768 37.39
arade4 ALP 1.146 4.222 24.99
basel_temp_f PLAIN 14.190 15.634 64.01
basel_temp_f PLAIN + ZSTD 0.261 0.703 23.07
basel_temp_f ALP 0.660 9.281 29.23
basel_wind_f PLAIN 15.096 15.333 64.01
basel_wind_f PLAIN + ZSTD 0.322 0.807 18.53
basel_wind_f ALP 1.066 10.160 29.87
bird_migration_f PLAIN 15.472 22.338 64.01
bird_migration_f PLAIN + ZSTD 0.295 0.816 23.49
bird_migration_f ALP 1.158 5.688 20.24
bitcoin_f PLAIN 19.500 21.789 64.07
bitcoin_f PLAIN + ZSTD 0.265 0.909 50.01
bitcoin_f ALP 0.825 8.965 27.18
bitcoin_transactions_f PLAIN 6.660 14.798 64.01
bitcoin_transactions_f PLAIN + ZSTD 0.520 1.147 47.96
bitcoin_transactions_f ALP 0.958 7.396 41.27
city_temperature_f PLAIN 13.562 4.369 64.01
city_temperature_f PLAIN + ZSTD 0.302 0.561 17.67
city_temperature_f ALP 1.233 4.532 10.80
cms1 PLAIN 13.507 6.339 64.01
cms1 PLAIN + ZSTD 0.353 0.804 26.84
cms1 ALP 0.683 4.894 35.19
cms25 PLAIN 13.829 12.360 64.01
cms25 PLAIN + ZSTD 0.407 1.100 58.11
cms25 ALP 1.020 8.744 41.17
cms9 PLAIN 13.907 12.828 64.01
cms9 PLAIN + ZSTD 0.354 0.702 11.71
cms9 ALP 1.224 11.271 12.16
food_prices PLAIN 13.603 8.478 64.01
food_prices PLAIN + ZSTD 0.314 0.693 18.13
food_prices ALP 0.587 7.385 23.20
gov10 PLAIN 13.788 13.031 64.01
gov10 PLAIN + ZSTD 0.275 0.699 29.12
gov10 ALP 0.966 9.393 29.88
gov26 PLAIN 13.832 13.086 64.01
gov26 PLAIN + ZSTD 3.281 5.765 0.20
gov26 ALP 0.924 18.473 1.40
gov30 PLAIN 13.828 13.085 64.01
gov30 PLAIN + ZSTD 1.112 2.440 4.52
gov30 ALP 0.601 11.786 17.88
gov31 PLAIN 13.809 13.074 64.01
gov31 PLAIN + ZSTD 1.756 3.387 1.65
gov31 ALP 1.250 13.075 6.77
gov40 PLAIN 13.789 13.057 64.01
gov40 PLAIN + ZSTD 2.971 5.021 0.43
gov40 ALP 1.346 16.493 2.59
medicare1 PLAIN 13.721 12.113 64.01
medicare1 PLAIN + ZSTD 0.299 0.823 31.68
medicare1 ALP 0.650 6.870 40.46
medicare9 PLAIN 13.787 13.137 64.01
medicare9 PLAIN + ZSTD 0.351 0.707 11.86
medicare9 ALP 1.208 10.952 12.82
neon_air_pressure PLAIN 13.822 13.118 64.01
neon_air_pressure PLAIN + ZSTD 0.452 1.069 11.85
neon_air_pressure ALP 1.199 10.680 16.48
neon_bio_temp_c PLAIN 13.783 13.082 64.01
neon_bio_temp_c PLAIN + ZSTD 0.314 0.763 16.84
neon_bio_temp_c ALP 1.237 11.511 10.81
neon_dew_point_temp PLAIN 13.855 12.947 64.01
neon_dew_point_temp PLAIN + ZSTD 0.266 0.799 23.73
neon_dew_point_temp ALP 1.213 10.294 13.63
neon_pm10_dust PLAIN 13.550 14.790 64.01
neon_pm10_dust PLAIN + ZSTD 0.471 0.896 7.79
neon_pm10_dust ALP 0.834 12.082 8.41
neon_wind_dir PLAIN 13.774 13.098 64.01
neon_wind_dir PLAIN + ZSTD 0.272 0.743 24.41
neon_wind_dir ALP 1.208 12.560 15.94
nyc29 PLAIN 13.751 12.170 64.01
nyc29 PLAIN + ZSTD 0.343 0.847 24.67
nyc29 ALP 1.061 8.720 40.43
poi_lat PLAIN 13.366 5.087 64.01
poi_lat PLAIN + ZSTD 0.329 0.889 57.78
poi_lat ALP 0.623 3.158 88.19
poi_lon PLAIN 13.681 5.958 64.01
poi_lon PLAIN + ZSTD 0.414 0.961 60.44
poi_lon ALP 0.702 3.836 79.12
ssd_hdd_benchmarks_f PLAIN 17.871 22.035 64.02
ssd_hdd_benchmarks_f PLAIN + ZSTD 0.334 0.684 12.98
ssd_hdd_benchmarks_f ALP 1.205 11.831 16.04
stocks_de PLAIN 13.730 12.883 64.01
stocks_de PLAIN + ZSTD 0.381 0.846 10.07
stocks_de ALP 0.790 11.214 11.20
stocks_uk PLAIN 13.792 13.182 64.01
stocks_uk PLAIN + ZSTD 0.364 0.762 11.29
stocks_uk ALP 0.546 11.074 12.75
stocks_usa_c PLAIN 13.793 13.056 64.01
stocks_usa_c PLAIN + ZSTD 0.410 0.854 8.24
stocks_usa_c ALP 1.250 11.793 7.95
ALL AVG. PLAIN 13.936 12.696 64.01
ALL AVG. PLAIN + ZSTD 0.603 1.266 22.75
ALL AVG. ALP 0.979 9.611 24.27

GB/s is decimal billions of uncompressed input bytes processed per second; higher is better. Compressed size includes Parquet data-page headers but excludes the file footer. Speed processes every value in pages of up to 131072 values and excludes file I/O. PLAIN + ZSTD includes both stages: PLAIN encoding plus ZSTD compression, and ZSTD decompression plus PLAIN decoding. Short pages are repeated for timing stability and normalized to one page.

Random access

Time to decode 100 deterministic, uniformly distributed rows from city_temperature_f (lower is better). Each lookup starts from the encoded page.

Parquet choice 100 random rows (µs)
PLAIN 4.044
PLAIN + ZSTD 145706.934
ALP 16.891

PLAIN and ALP reset the page decoder, skip to the selected row, and decode one value. PLAIN + ZSTD additionally decompresses the complete target page for every independent lookup. Encoded pages are already in memory; file I/O and page lookup are excluded.

30 datasets. Arithmetic mean: PLAIN 64.01, PLAIN + ZSTD 22.75, ALP 24.27 bits/value.
Median ALP: 17.88 bits/value. ALP is 0.28x the size of PLAIN and 1.25x the size of PLAIN + ZSTD by geometric mean.
ALP is smaller than PLAIN + ZSTD on 10/30 datasets.
Arithmetic mean compression/decompression speed in GB/s: PLAIN 13.936/12.696, PLAIN + ZSTD 0.603/1.266, ALP 0.979/9.611.

@JigaoLuo

Copy link
Copy Markdown

Thanks again to all for the great work.

Question+discussion to this: #195 (comment). I see this comment, but I'm not sure if there's a plan to keep such text in this blog.

The meta-question from me is whether we should treat this blog as a positioning. By positioning, I mean showing that Parquet with ALP is just as good as the newer file formats (I won't name-drop here). The positioning should give the reader a clear takeaway on Parquet: future Parquet with ALP is the file format to go with.

  • A small experiment, like a benchmark comparing Parquet with the new file formats, could also be helpful for that.

@alamb

alamb commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

Thanks again to all for the great work.

👍

The meta-question from me is whether we should treat this blog as a positioning. By positioning, I mean showing that Parquet with ALP is just as good as the newer file formats (I won't name-drop here). The positioning should give the reader a clear takeaway on Parquet: future Parquet with ALP is the file format to go with.

  • A small experiment, like a benchmark comparing Parquet with the new file formats, could also be helpful for that.

I don't think it is appropriate for Parquet to try and position itself against other formats. In my mind:

  1. There is room for many formats for different use cases
  2. Parquet / ASF should take the high ground (and purposely not try to treat this as a marketing exercise)

I would love to hear other opinions as well

@alamb

alamb commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator

Update here:

Once we merge these two PRs Parquet 2.14 will be available. Then I will refresh this PR (and probably update the adoption section a bit) and get it pubished

@alamb

alamb commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator

Ok, now we have a few more PRs to publish 2.14 format changes to the website. Once those are in, I will polish up the links and get this one out for a final round of review

@alamb

alamb commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

I am adding some final touchups here;

  1. An acknowledgments section, including a list of those who commented on the spec and mailing list
  2. Updated the implementation status to reflect the parquet-format 2.14 release and the Rust parquet crate release
  3. Removed some unecessary headings to make the document more concise

I will re-update the preview site and update the screen shots and then send it out to the mailing list one more time. I hope to merge early next week (though we still need a committer to approve)

@sdf-jkl

sdf-jkl commented Sep 18, 2026

Copy link
Copy Markdown
Member Author

We can rerun ALP encoding benchmarks now that bit-packing is SIMD on arrow-rs 🤓

@alamb

alamb commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

We can rerun ALP encoding benchmarks now that bit-packing is SIMD on arrow-rs 🤓

We can indeed -- though I think the numbers in the current post stand for themselves already quite well.

It would be interesting to see how much the bit packing improvements (apache/arrow-rs#10432 for anyone else who is curious) improve things

@alamb

alamb commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

BTW sorry for obsessing over this blog -- I have spent way more time than I would like to admit on it, but I am quite proud of the result

@prtkgaur prtkgaur left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took another pass.


## Why ALP?

Encoding floating-point data is a complicated engineering problem due to the nature of floating-point values. They do not exactly represent most real values. This leads to rounding errors that prevent using existing lightweight encodings like Delta and Frame of Reference (FOR).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Most lines wrap around 80-100 characters. This one doesn't.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f75a4f4

<p/>
</div>

Note that these numbers are for the pre-release Rust implementation of ALP, and

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would be add a link the the performance numbers that we got for c++ impl too?
https://docs.google.com/spreadsheets/d/1NmCg0WZKeZUc6vNXXD8M3GIyNqF_H3goj6mVbT8at7A/?

(to show cross language performance improvement?)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in f75a4f4

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe exclude PLAIN from the table's conditional formatting? RN it looks like everything (including ALP) is bad compared to it. Instead we want to show that ALP performs better than every other encoding.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is a comment for @prtkgaur for the spreadsheet. I agree that it was somewhat hard to fully grok the numbers there quickly without study

@vinoo-ganesh-kepler vinoo-ganesh-kepler left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@vinooganesh vinooganesh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!(from my non-work account)

@alamb

alamb commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

I plan to publish this tomorrow Sep 22, 2026

@sdf-jkl sdf-jkl left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @alamb, did another pass with fresh eyes.

Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated

<div class="row g-3 td-max-width-on-larger-screens">
<div class="col-12 col-md-6">
<img src="/blog/alp/avg_compression_ratio.png" alt="Average compression ratio benchmark" class="img-fluid">

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's no point in comparing the compression ratio between different machines. The algorithm stays the same

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

that is technically true. Though it currently has a nice symmetry -- if we changed it then one of the charts will be visually different than the others which might cause confusion 🤔

<p/>
</div>

Note that these numbers are for the pre-release Rust implementation of ALP, and

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No longer pre-release numbers

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't regenerate the numbers with the released version 🤔 Though I don't think we made any changes, I would want to regenerate the numbers if we claimed it was with the released version

<p/>
</div>

Note that these numbers are for the pre-release Rust implementation of ALP, and

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe exclude PLAIN from the table's conditional formatting? RN it looks like everything (including ALP) is bad compared to it. Instead we want to show that ALP performs better than every other encoding.

Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
Comment thread content/en/blog/features/alp_encoding.md Outdated
value = encoded × 10<sup>f</sup> × 10<sup>-e</sup>
</pre>

This calculation uses floating-point arithmetic, which rounds to the nearest

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the first mention of floating-point arithmetic in the blog. There's another link to IEEE 754, but we might add a link to a wiki page explaining the range of representable values?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like I would like to assume readers of this post will have a basic familiarity with floating point (or can look it up themselves if they have questions). If they don't understand floating point is not exact I think there is a lot of this blog that wont make sense.

alamb and others added 2 commits September 21, 2026 15:06
Co-authored-by: Kosta Tarasov <33369833+sdf-jkl@users.noreply.github.com>
@alamb alamb changed the title Add ALP support blog Add ALP Encoding blog Sep 22, 2026
@alamb
alamb merged commit 1b223ca into apache:production Sep 22, 2026
1 check passed
@alamb

alamb commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

Thank you everyone for your comments and reviews. This has been a long time in the making

@alamb
alamb deleted the alp-blog branch September 22, 2026 10:41
@alamb

alamb commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

Blog is published: https://parquet.apache.org/blog/2026/09/22/alp-adaptive-lossless-floating-point-encoding-in-apache-parquet/

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Write a Parquet Blog on new encoding Adaptive Lossless Floating Point (ALP)

8 participants