Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Preserve raw benchmark snapshots across platform-specific newline settings.
/bench/data/**/*.json -text
/bench/data/**/*.log -text
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
/target
bench/server/target/

# Byte-compiled / optimized / DLL files
__pycache__/
Expand Down Expand Up @@ -86,6 +87,11 @@ dist

keylog.txt
*.json
!/bench/data/**/*.json
!/bench/data/**/*-stdout.log
!/bench/data/**/*-stderr.log
*.csv
docs/site
docs/source/benchmark.md
docs/source/assets/benchmark/
*.pdb
2 changes: 1 addition & 1 deletion .readthedocs.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ build:
- python -m pip install -r docs/requirements.txt
build:
html:
- python -m zensical build -f docs/mkdocs.yml
- python docs/build.py
post_build:
- mkdir -p $READTHEDOCS_OUTPUT/html/
- cp --recursive docs/site/* $READTHEDOCS_OUTPUT/html/
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ Due to the complexity of TLS encryption and the widespread adoption of HTTP/2, b

2. **Device Emulation**

**TLS** and **HTTP/2** fingerprints are often identical across various browser models because these underlying protocols evolve slower than browser release cycles. **100+ browser device emulation [profiles](https://python.wreq.org/en/latest/getting-started/introduction/#behavior)** are maintained in **wreq**.
**TLS** and **HTTP/2** fingerprints are often identical across various browser models because these underlying protocols evolve slower than browser release cycles. **100+ browser device emulation [profiles](https://python.wreq.org/en/latest/api/emulation/#wreq.emulation.Profile)** are maintained in **wreq**.

## Building

Expand Down Expand Up @@ -107,7 +107,7 @@ maturin build --release

## Benchmark

Outperforms `requests`, `httpx`, `aiohttp`, and `curl_cffi` according to our [benchmark](https://github.com/0x676e67/wreq-python/tree/main/bench) suite driven by [pyperf](https://github.com/psf/pyperf), though results are for reference only as they vary by environment.
Our [benchmark suite](https://github.com/0x676e67/wreq-python/tree/main/bench) compares Python HTTP clients over HTTPS HTTP/1.1 and HTTP/2. It tests complete and streamed uploads, reading every response to the end. The results include client versions and machine details so you can check the conditions behind each measurement. Performance varies by workload and environment.

## Services

Expand Down
179 changes: 150 additions & 29 deletions bench/README.md
Original file line number Diff line number Diff line change
@@ -1,43 +1,164 @@
# Benchmark
# HTTPS benchmarks

benchmark between wreq and other python http clients
This suite compares Python HTTP clients over HTTPS HTTP/1.1 and HTTP/2. It
follows the workload model of the Rust wreq benchmark: Full or Stream request
bodies, with every echoed response streamed to EOF. Async and blocking results
are reported separately.

Sync clients
------
## Workload

- curl_cffi
- requests
- niquests
- pycurl
- httpx
- wreq
- ry
- Async: one reused client per case, with 10, 50, 100, or 150 closed-loop workers.
- Blocking: one reused client per logical worker, using a persistent thread
pool. Submit a whole worker batch, not one executor task per request.
- Seven payload and upload-chunk cases matching the Rust benchmark, listed below.
- A local Rust TLS 1.3 echo server, restricted to the selected protocol by ALPN.
- Async clients: wreq default runtime, wreq single worker, pyreqwest
single-threaded and multithreaded runtimes, ry, httpx, aiohttp, niquests, and
curl_cffi.
- Blocking clients: wreq, ry, requests, httpx, niquests, curl_cffi, and pycurl.
- Three rounds, each with one warmup and one timed batch of 300 requests.
Case order is shuffled with a recorded seed.
- Status, negotiated HTTP version, and total response length are checked for
every request. Unsupported protocol or upload combinations are reported as
N/A, not downgraded or silently substituted. TLS verification is disabled
only for the local self-signed
benchmark certificate. Proxy environment variables are not used.

Async clients
------
Client construction is outside the timer. Timed batches include uploads, TLS
and HTTP processing, and response iteration; they are not isolated TLS IO
measurements. Untimed requests establish reusable connections. Client variants
run in separate processes; async clients use standard asyncio rather than
uvloop. Blocking results include thread-pool scheduling, and their independent
worker clients have separate connection pools.

- curl_cffi
- httpx
- niquests
- aiohttp
- wreq
- ry
| Upload / echo payload | Stream upload chunk |
| --- | --- |
| 1 KiB | 1 KiB |
| 10 KiB | 10 KiB |
| 64 KiB | 16 KiB |
| 128 KiB | 32 KiB |
| 1 MiB | 64 KiB |
| 2 MiB | 128 KiB |
| 4 MiB | 256 KiB |

Target
------
Both Full and Stream uploads are tested for every payload and concurrency.
The complete matrix contains 1,680 supported cells across 16 client variants;
each cell has three measured rounds. The Rust benchmark uses Criterion and
600 requests per iteration; this suite uses fixed batches of 300 requests.
Custom payload sizes use upload chunks of at most 64 KiB.

## Run locally

All the clients run with session/client enabled.
Use CPython 3.14 and a Rust toolchain compatible with the manifests. BoringSSL
requires CMake, Clang, and the usual native build tools.

## Run benchmark
```bash
uv venv --python 3.14
uv pip install -r bench/requirements.txt
uv run --no-sync maturin develop --release --uv --locked
cargo build --release --locked --manifest-path bench/server/Cargo.toml
uv run --no-sync python bench/run.py \
--server bench/server/target/release/wreq-benchmark-server
```

On Windows, the server executable ends in `.exe`. If Cargo uses a custom target
directory, pass the actual executable path to `--server`.

`bench/run.py` passes workload options to the measurement runner. Before it
starts, it prints the number of cases and batches. It saves stdout/stderr logs
beside the raw JSON and generates an English `.report.md` once validation passes.

The default suite has 1,680 supported cases, 5,040 timed batches and the same
number of warm-up batches. Allow several hours for a full run. The setup commands
above prepare the dependencies and native binaries; the runner won't install or
build them for you. Publishing the results to the docs is optional.

The saved WSL measurements built wreq with `--features jemalloc` in addition
to the release flags above. Check a run's build records for its allocator and
toolchain. The setup command above uses the platform's default allocator.

For an integration check, use the same clients and protocols with fewer
requests. This smoke run checks behavior; it isn't enough to rank performance:

```bash
# Install project + benchmark dependencies
pip install -e .[bench]
uv run --no-sync python bench/run.py \
--server bench/server/target/release/wreq-benchmark-server \
--sizes 10240,1048576 --concurrency 2 --requests 4 \
--rounds 1 --warmup 0 --samples 1 --output bench/data/smoke/RUN.json
uv run --no-sync python -m pytest bench/test_*.py
```

## Results and interpretation

The JSON records the source SHA and dirty state, Python and package versions,
CPU model, operating system, machine architecture, logical CPU count, runtime
configurations, native module hashes, server hash, and individual timings.
Results are written atomically only after all supported combinations in the
requested matrix pass validation.

# Start server
python server.py
Requests/s is total completed requests divided by total elapsed seconds, not an
arithmetic mean of rates. MB/s counts the response payload only, in decimal MB;
the timer includes both upload and download. Responses are streamed, never
collected into a complete body. Native streaming interfaces are used where
available; adapters requiring a read size use documented 64 KiB reads.

# Run benchmark suite
python benchmark.py
Other programs on the machine compete for CPU and I/O. Repeat measurements and
check their environments before treating a difference as stable. A client that
wins here may perform differently in your application. These tests also don't
measure browser-emulation compatibility.

Compare two snapshots with identical recorded workloads and compatible metadata:

```bash
uv run --no-sync python bench/compare.py \
--before bench/data/BEFORE.json --after bench/data/AFTER.json \
--output bench/data/COMPARISON.md
```

The English report shows each case's RPS change and variation between rounds.
That variation isn't a confidence interval. Running one snapshot after another
also doesn't isolate TLS I/O improvements; check the build and harness records
before attributing a change. Omit `--output` to print UTF-8 Markdown to stdout.
File exports refuse to overwrite any existing path, including the input JSON.

## Stored data and documentation

We run the benchmarks locally, outside GitHub Actions, and keep JSON snapshots
in [`bench/data`](data/) on `main`. Without `--output`, the runner creates a
timestamped filename containing the measured source SHA. If you supply an output
path, it must be new. Completed runs never overwrite an earlier snapshot.

Keep every completed snapshot. After reviewing a full measurement, select it
with:

```bash
uv run --no-sync python bench/run.py --input bench/data/RUN.json --publish
```

`--input` reads a saved run and generates its report, or reuses the report if its
contents match. It never starts a measurement. With `--publish`, the runner also
freezes the candidate JSON and builds the docs. Only a successful build lets it
atomically select the original bytes as `bench/data/latest.json`. Historical
JSON, logs and reports stay untouched.

Smoke runs belong in `bench/data/smoke/` and cannot be published. Publication
requires the complete default matrix, at least 300 requests per batch and
three rounds, with at least one warm-up and timed sample per round. These checks
confirm coverage; you still need to review the quality of the measurements.

Use `--build-docs` instead of `--publish` to preview a complete recorded matrix
without changing `latest.json`. If your docs dependencies are in a separate
virtual environment, add `--docs-python PATH/TO/python`. You can also pass
`--publish` to a new measurement to run those steps after it finishes. The
runner doesn't commit, push or enable benchmarks in CI.

`python docs/build.py` reads and validates the checked-in `bench/data/latest.json`.
It generates responsive light/dark SVG charts and fills
`docs/templates/benchmark.md` with body-size controls and expandable tables.
The built site includes a frozen raw JSON download. The build won't fetch
measurements or start a benchmark, and missing or invalid data stops it.
Read the Docs uses this same entry point.

Use `python docs/build.py --data bench/data/RUN.json` to preview another run.
Each measurement shows its source revision and whether the checkout had
uncommitted changes. Saving data on `main` doesn't change which code was tested.
Loading
Loading