Problem
channel.Channel.WaitReady polls IsReadyToSend() with a new 5 ms timer on every unsuccessful check:
for {
if c.IsReadyToSend() {
return nil
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(5 * time.Millisecond):
}
}
On a low-RTT Link this imposes a throughput ceiling unrelated to the underlay. Once the Channel TX window fills, a delivery proof may free a slot in substantially less than 5 ms, but the blocked writer does not observe it until the polling timer fires.
With the Python-compatible fast-link window and current stream payload size:
WindowMaxFast = 48
Link MDU = 431 bytes
Stream payload = 431 - 6-byte Channel header - 2-byte stream header = 423 bytes
48 * 423 bytes / 5 ms = 4,060,800 bytes/s
That theoretical polling ceiling almost exactly matches the measured sustained throughput below.
Reproduction and measurements
Reticulum-Go v1.2.0, macOS arm64, Go 1.27.1, local Backbone/TCP -> Link -> Channel -> StreamDataMessage path, standard uncompressed stream messages:
| workload |
result |
| 1 MiB single stream |
275.9 ms / 3.62 MiB/s |
| 10 MiB single stream |
2.501 s / 4.00 MiB/s |
| 10 concurrent x 1 MiB |
1.743 s / 5.74 MiB/s aggregate |
| 512-byte request + 2048-byte response |
782 us per round trip |
The 10 MiB row converges almost exactly on the window * payload / polling interval ceiling.
A CPU profile of the RNS benchmark subtest also shows substantial scheduler sleep/timer overhead:
flat function
31.96% syscall.rawsyscalln
15.92% runtime.usleep
12.49% runtime.kevent
Relevant cumulative paths include Channel.Send, Link packet processing, explicit Channel proofs, and synchronous Backbone writes. The polling delay is not the only remaining cost, but it is the first artificial cap after removing unnecessary compression work.
Expected behavior
WaitReady(ctx) should wake promptly when Channel sendability can change:
- a delivery proof removes an envelope from the TX ring;
- timeout handling removes an envelope or changes the window;
- Channel or Link shutdown makes the outlet unusable;
- context cancellation occurs.
It should not add up to 5 ms of latency after the TX window has already gained capacity.
Proposed implementation
Use an event-driven readiness notification, for example a generation/channel signal or a carefully implemented condition variable:
- Check link state and
len(txRing) < window while holding the Channel lock.
- Capture the current readiness notification/generation under the same lock to avoid a lost wakeup.
- Wait for either notification or
ctx.Done().
- Signal/broadcast after delivery, terminal timeout/removal, window changes, and Channel close.
A replace-and-close notification channel is one possible Go-friendly shape because it supports select with context cancellation. The implementation should avoid blocking proof handlers and should not start one goroutine or allocate one timer per polling iteration.
Please add deterministic tests covering:
- a waiter wakes immediately after a delivery callback frees one slot;
- cancellation returns promptly;
- multiple waiters do not deadlock or lose notifications;
- close/link-not-ready wakes waiters with
ErrLinkNotReady;
- delivery between the readiness check and the wait cannot be lost;
- no timer or goroutine accumulation under sustained transfers.
A Channel/Buffer throughput benchmark with low RTT and incompressible data would make the regression visible.
Python RNS compatibility
This change requires no wire or protocol modification:
- keep the Python-compatible window constants, sequence handling, proofs and retransmission behavior;
- do not change
StreamDataMessage encoding;
- only change how the Go caller is notified that the existing TX window has room.
Python RNS exposes is_ready_to_send() and its raw writer reports no progress when the Channel is full; it does not define a wire-visible 5 ms polling interval. An event-driven Go WaitReady therefore preserves full interoperability with Python nodes.
Related issues
Problem
channel.Channel.WaitReadypollsIsReadyToSend()with a new 5 ms timer on every unsuccessful check:On a low-RTT Link this imposes a throughput ceiling unrelated to the underlay. Once the Channel TX window fills, a delivery proof may free a slot in substantially less than 5 ms, but the blocked writer does not observe it until the polling timer fires.
With the Python-compatible fast-link window and current stream payload size:
That theoretical polling ceiling almost exactly matches the measured sustained throughput below.
Reproduction and measurements
Reticulum-Go v1.2.0, macOS arm64, Go 1.27.1, local
Backbone/TCP -> Link -> Channel -> StreamDataMessagepath, standard uncompressed stream messages:The 10 MiB row converges almost exactly on the
window * payload / polling intervalceiling.A CPU profile of the RNS benchmark subtest also shows substantial scheduler sleep/timer overhead:
Relevant cumulative paths include
Channel.Send, Link packet processing, explicit Channel proofs, and synchronous Backbone writes. The polling delay is not the only remaining cost, but it is the first artificial cap after removing unnecessary compression work.Expected behavior
WaitReady(ctx)should wake promptly when Channel sendability can change:It should not add up to 5 ms of latency after the TX window has already gained capacity.
Proposed implementation
Use an event-driven readiness notification, for example a generation/channel signal or a carefully implemented condition variable:
len(txRing) < windowwhile holding the Channel lock.ctx.Done().A replace-and-close notification channel is one possible Go-friendly shape because it supports
selectwith context cancellation. The implementation should avoid blocking proof handlers and should not start one goroutine or allocate one timer per polling iteration.Please add deterministic tests covering:
ErrLinkNotReady;A Channel/Buffer throughput benchmark with low RTT and incompressible data would make the regression visible.
Python RNS compatibility
This change requires no wire or protocol modification:
StreamDataMessageencoding;Python RNS exposes
is_ready_to_send()and its raw writer reports no progress when the Channel is full; it does not define a wire-visible 5 ms polling interval. An event-driven GoWaitReadytherefore preserves full interoperability with Python nodes.Related issues
context.Background()/cancellation parity gap and Channel/Backbone correctness defects.