Pure-Rust Zstandard (zstd) compression and decompression. All 22 standard levels plus the negative ultra-fast range, streaming, dictionaries, no_std support and a WebAssembly build — with plain cargo: no cmake, no system zstd, no FFI.
- Production-grade decoder — complete RFC 8878 implementation: dictionary-backed streams, raw / RLE / compressed blocks, the full frame format, optional content checksums, runtime-dispatched SIMD kernels (SSE2 / BMI2 / AVX2 / NEON, opt-in AVX-512).
- Full-range encoder — every C-zstd level (
-131072..=22) produces valid frames decodable by this crate and by upstream C zstd; named presets, per-knob parameter overrides, long-distance matching, streaming viastd::io::Write. - Dictionaries end to end — compress and decompress with the same dictionary format C zstd consumes; reusable parsed handles; pure-Rust COVER / FastCOVER training behind the
dict-builderfeature. - Wire-compatible both ways — frames interoperate with C zstd in either direction; interop is enforced in CI against the reference implementation.
no_stdready — the decoder builds with--no-default-featuresfor embedded and sandboxed targets.- WebAssembly / npm — the same codec as an npm package with automatic SIMD selection; no native addons, no postinstall scripts.
- Continuously benchmarked — a public dashboard tracks speed and ratio against C zstd on every merge.
cargo add structured-zstduse structured_zstd::encoding::{compress_to_vec, CompressionLevel};
let compressed = compress_to_vec(&b"hello world"[..], CompressionLevel::from_level(7));For no_std builds disable the default features:
cargo add structured-zstd --no-default-featuresRelease notes for every version live in zstd/CHANGELOG.md (maintained by release-plz).
The tool ships from this same crate, so one name covers both uses:
cargo add structured-zstd # the library
cargo install structured-zstd # the `structured-zstd` binaryIt carries no dependencies of its own — argument parsing, progress display and
error reporting are written against std — so depending on the library pulls
in nothing extra.
The binary speaks the upstream zstd
command line: levels (-1..-19, --ultra for -20..-22, --fast[=N]),
-d, -c, -o, -t, -l, -D, --train, -b, and the usual
-f/-k/--rm file handling.
Flags that only steer how the work is done (-T, -B, --adapt,
--[no-]progress, …) are accepted and ignored — their values are still
validated, so a typo is an error rather than silence.
--target-compressed-block-size does take effect: it bounds what goes into a
block, so blocks flush sooner. --long means --long=27, as upstream
documents, and is capped there: a larger window would produce frames this
build's decoder refuses. It needs level 16 or above, where long-distance
matching actually runs — below that it is refused rather than accepted as a
wider window and nothing else. A window is never declared larger than the
source can fill, so a small file compressed with --long does not ask its
decoders to reserve 128 MiB.
Flags that would change the result are refused instead: --format= for
anything but zstd, --patch-from, --rsyncable, --no-check,
--[no-]compress-literals, and the not-yet-implemented --pass-through /
--exclude-compressed. -M is treated as the safety promise it is: on the
runs that decode, a limit covering the 128 MiB window, the decoder's buffers
and the -D dictionary is kept and a tighter one is refused rather than
ignored. Compressing, listing and training allocate no decoder, so the flag is
accepted there and describes nothing, as upstream has it.
--train and --train-fastcover both train with FastCOVER, the algorithm
upstream also defaults to. --train-cover and --train-legacy name algorithms
this build does not have, so they are refused rather than quietly served by
FastCOVER. -D takes either a dictionary produced by --train or any file at
all, which is then used as raw content the way upstream does — such a
dictionary has no ID, so the same bytes must be supplied when decoding.
It is deliberately not installed as zstd, so it never shadows the system
tool. It does dispatch on the name it is invoked under, so linking it as the
familiar names works:
ln -s "$(command -v structured-zstd)" ~/.local/bin/unzstd # defaults to -d
ln -s "$(command -v structured-zstd)" ~/.local/bin/zstdcat # defaults to -d -cDistributions should register it with their alternatives mechanism rather than
overwriting /usr/bin/zstd, e.g.
update-alternatives --install /usr/bin/zstd zstd /usr/bin/structured-zstd 100use structured_zstd::encoding::{compress, compress_to_vec, CompressionLevel};
let data: &[u8] = b"hello world";
// Named level
let compressed = compress_to_vec(data, CompressionLevel::Fastest);
// Numeric level (C zstd compatible: 0 = default, 1-22, negative for ultra-fast)
let compressed = compress_to_vec(data, CompressionLevel::from_level(7));use structured_zstd::encoding::{CompressionLevel, StreamingEncoder};
use std::io::Write;
let mut out = Vec::new();
let mut encoder = StreamingEncoder::new(&mut out, CompressionLevel::Fastest);
encoder.write_all(b"hello ")?;
encoder.write_all(b"world")?;
encoder.finish()?;
# Ok::<(), std::io::Error>(())- Named presets:
Fastest(≈1),Default(≈3),Better(≈7),Best(≈13) - Frame Content Size:
FrameCompressorwrites FCS automatically;StreamingEncoderrequiresset_pledged_content_size()before the first write - Content checksums: opt-in via
set_content_checksum(true)
Override individual compression knobs (the drop-in equivalent of C zstd's
ZSTD_CCtx_setParameter). Every knob left unset inherits the base level's
default, so a parameter set that overrides nothing reproduces plain
level-based compression. Long-distance matching is off at every level preset
and is activated only here; it also needs the (default-on) ldm feature.
Without it the builder still accepts enable_long_distance_matching(true) and
the frame is still valid, but no long-distance matches are produced:
use structured_zstd::encoding::{
compress_with_parameters, CompressionLevel, CompressionParameters, Strategy,
};
let data: &[u8] = b"hello world";
let params = CompressionParameters::builder(CompressionLevel::Level(19))
.window_log(22)
.strategy(Strategy::Btultra2)
.enable_long_distance_matching(true)
.build()
.expect("parameters within bounds");
let compressed = compress_with_parameters(data, ¶ms);Each parameter's valid range is queryable via CParameter::bounds() (the
analogue of ZSTD_cParam_getBounds); the builder validates every set knob.
use structured_zstd::decoding::StreamingDecoder;
use structured_zstd::io::Read;
let compressed_data: Vec<u8> = vec![];
let mut source: &[u8] = &compressed_data;
let mut decoder = StreamingDecoder::new(&mut source).unwrap();
let mut result = Vec::new();
decoder.read_to_end(&mut result).unwrap();use structured_zstd::decoding::{DictionaryHandle, FrameDecoder, StreamingDecoder};
use structured_zstd::io::Read;
let compressed: Vec<u8> = vec![];
let dict_bytes: Vec<u8> = vec![];
let mut output = vec![0u8; 1024];
// Parse dictionary once, then reuse handle.
let handle = DictionaryHandle::decode_dict(&dict_bytes).unwrap();
let mut decoder = FrameDecoder::new();
let _written = decoder
.decode_all_with_dict_handle(compressed.as_slice(), &mut output, &handle)
.unwrap();
// Compatibility path: pass raw dictionary bytes directly.
let mut decoder = FrameDecoder::new();
let _written = decoder
.decode_all_with_dict_bytes(compressed.as_slice(), &mut output, &dict_bytes)
.unwrap();
// Streaming helpers exist for both handle- and bytes-based paths.
let mut source: &[u8] = &compressed;
let mut stream = StreamingDecoder::new_with_dictionary_handle(&mut source, &handle).unwrap();
let mut sink = Vec::new();
stream.read_to_end(&mut sink).unwrap();Compression takes the same dictionary format through
FrameCompressor::set_dictionary_from_bytes / EncoderDictionary::from_bytes
(one parse, reusable across frames).
Behind the dict-builder feature, the dictionary module trains dictionaries
in pure Rust:
- COVER (
create_raw_dict_from_source) and FastCOVER (create_fastcover_raw_dict_from_source) raw dictionaries finalize_raw_dictto produce the full zstd dictionary formatcreate_fastcover_dict_from_sourcefor train + finalize in one call
| Feature | Default | What it enables |
|---|---|---|
std |
✅ | Runtime CPU detection, std::io adapters |
hash |
✅ | XXH64 content checksums |
ldm |
✅ | Long-distance matching (implies hash: LDM hashes each window with XXH64) |
kernel-sse, kernel-bmi2, kernel-avx2 |
✅ | x86 SIMD kernels (kernel-sse covers both the SSE2 and SSE4.2 tiers) |
kernel-neon, kernel-sve |
✅ | aarch64 SIMD kernels |
kernel-simd128 |
✅ | WebAssembly SIMD kernel (needs -C target-feature=+simd128) |
kernel-vbmi2 |
❌ | AVX-512 decode kernel (see note below) |
kernel-scalar |
✅ | Marker for the always-compiled scalar fallback |
dict-builder |
❌ | Pure-Rust COVER / FastCOVER dictionary training |
lsm |
❌ | Storage-format extensions |
Each flag gates its tier wherever that tier exists. kernel-sse,
kernel-bmi2, kernel-avx2, kernel-neon and kernel-simd128 cover both
the decoder and the encoder; kernel-vbmi2 and kernel-sve are decoder-only
(the encoder has no AVX-512 or SVE tier), and kernel-scalar gates nothing,
since the scalar path is the mandatory fallback and always compiled. So
--no-default-features (optionally with --features kernel-scalar) compiles
every per-tier dispatch and all explicit SIMD intrinsics out of the crate.
On x86 and aarch64 with std, the tier is picked at runtime from CPU
detection; on no_std it comes from the target's target_feature set at
compile time. x86 has two 128-bit tiers under kernel-sse: SSE4.2 when
available, otherwise a plain-SSE2 tier, so pre-SSE4.2 CPUs still get vector
match compares instead of dropping to scalar.
WebAssembly is compile-time only, with or without std: wasm has no
runtime feature detection, so both the decoder kernels and the encoder
fastpath additionally require target_feature = "simd128". Building for
wasm32 with default features and no extra flags therefore stays scalar —
pass -C target-feature=+simd128 to get the SIMD tier. (The npm package
sidesteps this by shipping separately compiled scalar and +simd128
payloads and picking one at load time.)
In every case these features control only the crate's own explicit SIMD; the compiler's autovectorizer is unaffected.
Why AVX-512 is off by default
On AVX-512 hosts the kernel-vbmi2 tier measures slower than kernel-avx2
for this decode workload: AVX-512's license-based frequency downclocking
stalls the surrounding bursty, memory-bound code and the heavier kernel never
amortizes. By default runtime dispatch is therefore capped at AVX2, and
AVX-512 hosts use the (faster) AVX2 tier. Opt in with
--features kernel-vbmi2 for a sustained AVX-512 workload that genuinely
benefits.
- Per-merge benchmarks publish to a public dashboard: structured-world.github.io/structured-zstd/dev/bench — speed and ratio against upstream C zstd over time.
- The CI matrix covers
x86_64-linux-gnu,i686-linux-gnuandx86_64-musl, with per-target / stage / scenario / level filtering on the dashboard. - A dedicated section tracks the WebAssembly build (
simd128+ scalar) against the most popular npm wasm zstd,@bokuweb/zstd-wasm. - Methodology in BENCHMARKS.md: small payloads, entropy extremes, a
100 MiBlarge-stream scenario, repository corpus fixtures, optional local Silesia corpora.
Internal: compression strategy backends
| Level range | Strategy | Backend |
|---|---|---|
| 1-2 | Fast |
Simple matcher |
| 3-4 | Dfast |
Dfast two-tier hash |
| 5-12 | Greedy / Lazy / Lazy2 |
Row lazy parse (lazy_depth=0/1/2): row match-finder above a 2^14 window, hash chain at or below it |
| 13-15 | Btlazy2 |
Row lazy parse over the lazily-sorted binary tree |
| 16-17 | BtOpt |
HashChain candidates + btopt price parser |
| 18 | BtUltra |
HashChain candidates + btultra price parser |
| 19-22 | BtUltra2 |
HashChain candidates + btultra2 dual-profile parse |
The level → strategy column matches upstream zstd ZSTD_defaultCParameters[0] at zstd/lib/compress/clevels.h:25-50 (srcSize > 256 KiB tier); smaller sources shift the row per upstream's size tiers. The whole greedy..btlazy2 band runs upstream's ZSTD_compressBlock_lazy_generic parse on the Row backend over the three upstream match finders (rows / hash chain per ZSTD_resolveRowMatchFinderMode, lazily-sorted binary tree for btlazy2).
JavaScript / TypeScript consumers can use the codec from npm — no native addons, no build step:
npm install @structured-world/structured-zstdimport { compress, decompress } from "@structured-world/structured-zstd";
const framed = await compress(new TextEncoder().encode("hello"), 19);
const plain = await decompress(framed);The package ships two WebAssembly payloads — one built with the simd128
SIMD tier, one scalar — and selects the fast one at runtime from the host
engine's capabilities. Pure ESM, strict TypeScript types. Frames interoperate
with native zstd. Source lives in
zstd-wasm/;
see the
package README.
Behind the lsm feature (default off), the crate adds building blocks for
storage-format authors:
- Skippable frames — a typed
SkippableFrameAPI (structured_zstd::skippable) for interleaving application metadata with zstd data. - Block-subset partial decode —
FrameDecoder::decode_blocks_partialdecodes only the inner blocks covering a requested range (skipping the trailing ones) and preserves the clean prefix on a corrupt block. - Block-to-byte-range lookup —
FrameEmitInfo::decompressed_byte_range(block_index)maps a block to its decompressed byte range, so a range query can locate which blocks cover a target byte window. - Resumable decoding — request a
ResumeState(cross-block entropy tables + repcode history + next-block coordinates) from a partial decode, then feed it back to continue from a later block, even across a dropped decoder. The state does not carry the match window: the resuming call also supplies the tail of the already-decompressed output (the lastmin(window_size, resume_offset)bytes) viaResumeInput::window_prime.
[dependencies]
structured-zstd = { version = "0", features = ["lsm"] }The ecosystem registry of allocated skippable-frame magic variants and the allocation policy live in docs/SKIPPABLE_MAGIC_ALLOCATIONS.md.
Maintained fork of KillingSpark/zstd-rs (ruzstd) by Dmitry Prudnikov. We sync periodically with upstream but maintain an independent development trajectory focused on the CoordiNode database engine's per-label dictionary needs.
Apache License 2.0. Contributions will be published under the same Apache 2.0 license.