reticulum-go version
v1.2.0 (github.com/Quad4-Software/Reticulum-Go v1.2.0)
OS / platform
macOS arm64, Go 1.27.1, including go test -race. The affected code paths are not Darwin-specific: they are in the transport worker pool, Channel, Buffer, and Backbone hub.
Summary
Long-lived byte streams built as Backbone/TCP -> Link -> Channel -> Buffer can be corrupted or stall under burst traffic. Comparing v1.2.0 with Python RNS 1.5.4 shows two confirmed concurrency defects plus several Buffer API parity gaps that downstream stream adapters currently have to compensate for.
The two confirmed defects are:
- inbound Channel messages can be dispatched out of sequence by parallel transport workers;
- the Backbone hub can lose
evWrite interest when QueueSend races writeStream, stranding queued bytes until the Link goes stale.
Steps to reproduce
- Start two in-process Reticulum-Go transports connected by a Backbone/TCP client/server pair.
- Establish one Link, obtain its Channel, and attach a
buffer.RawChannelReader and buffer.RawChannelWriter using one stream ID.
- Send several hundred incompressible, marker-bearing chunks sized to the negotiated payload MDU while reading concurrently on the peer.
- Repeat under
go test -race and with the default transport worker pool.
- Separately, repeatedly call
QueueSend while the Backbone poller drains the same stream.
Observed failure modes include:
- delivered stream chunks appearing in a different order from their Channel sequence numbers;
- queued Backbone bytes remaining unsent after a stale
pollerMod(evRead) overwrites a concurrent evRead|evWrite arm;
- both Links becoming stale after roughly 10 seconds with queued traffic;
- setting
MaxPacketHandlers=1 avoiding the dispatch race but reducing packetQ capacity to one and losing burst traffic (RawChannelWriter.Write eventually returns link not ready; one run failed around byte 79,524).
I have deterministic/stress regression tests for the Channel ordering and Backbone progress paths and can submit them with focused patches.
Expected behavior
- Handler delivery for one Channel must remain strictly ordered by Channel sequence, regardless of how many transport packet workers process the underlying packets.
- If a Backbone stream has buffered bytes, write interest must remain armed until the buffer is drained.
- A reliable TCP underlay must not lose RNS frames merely because the application uses one logical stream under burst load.
RawChannelWriter should derive its payload limit from the live Channel MDU, as Python does.
Actual behavior
1. Parallel Channel dispatch
Transport drains inbound work into a multi-worker pool. Channel.HandleInbound protects RX-ring mutation, but releases the Channel mutex before invoking handlers. A worker can drain envelope N and be preempted before its handler runs; another worker can then drain and dispatch N+1 first. Both handlers append to the same RawChannelReader, corrupting byte-stream order.
Python RNS avoids this because its inbound priority queue has one inbound_job consumer, and Channel._receive performs contiguous RX-ring delivery serially.
A working fix is to serialize emplace + drain + dispatch per Channel. A dedicated serial dispatcher would be preferable to a global/single transport worker, since MaxPacketHandlers=1 also shrinks the packet queue and causes drops.
2. Backbone write-interest race
In v1.2.0, writeStream can observe an empty txBuf, release s.mu, and then call pollerMod(fd, evRead). A concurrent QueueSend can append a frame and arm evRead|evWrite in that gap; the stale disarm then wins and no further write event flushes the queued frame. The partial-write branch also leaves buffered bytes without explicitly re-arming evWrite.
The fix is to decide and apply poller interest while holding the same lock that protects txBuf/wantOut, and to explicitly retain evWrite after partial writes.
Related Buffer parity gaps
These are not required to reproduce the two races, but they affect the same stream use case:
- Python
RawChannelWriter computes _mdu = channel.mdu - StreamDataMessage.HEADER_LEN. Go uses the constant MaxDataLen = 457. With a negotiated Link MDU of 431, the Channel body limit is 425 and the stream payload limit is 423; the fixed 457-byte limit can therefore produce a Channel ErrTooBig.
- Python's raw reader uses
None for would-block and exposes ready callbacks. Go maps an empty, non-EOF read to (0, nil), which causes normal io.Reader consumers to spin and lets a read immediately after an asynchronous EOF observe no EOF.
- Go invokes ready callbacks synchronously while holding the raw reader mutex. A callback that reads or unregisters itself can deadlock; Python invokes listeners asynchronously.
- Go
RawChannelWriter.Write waits with context.Background(), so a caller cannot cancel blocked backpressure. Python's raw writer reports no progress when the Channel is not ready, and its close wait is bounded.
- Python Backbone couples queue pressure to ingress gating and has transmit-buffer high-water/stall handling. The Go transport sheds on worker overflow but does not feed that pressure back to the Backbone reader.
Python RNS comparison
Compared against Python RNS 1.5.4, commit 192898864c008b6287dd56781d89fccef0bb5f7a:
RNS/Transport.py: one inbound_job drains the priority queue sequentially;
RNS/Channel.py: contiguous envelopes are delivered serially;
RNS/Buffer.py: writer payload size is derived from channel.mdu, and reader readiness is represented explicitly;
RNS/Interfaces/BackboneInterface.py and TransmitBuffer.py: a single epoll loop rechecks sendable data before disabling EPOLLOUT, with ingress and egress pressure controls.
I am happy to split the Buffer/API parity items into follow-up issues if that is easier to review. The Channel ordering and Backbone write-interest fixes are independently testable and should be safe as separate commits.
reticulum-go version
v1.2.0 (
github.com/Quad4-Software/Reticulum-Go v1.2.0)OS / platform
macOS arm64, Go 1.27.1, including
go test -race. The affected code paths are not Darwin-specific: they are in the transport worker pool, Channel, Buffer, and Backbone hub.Summary
Long-lived byte streams built as
Backbone/TCP -> Link -> Channel -> Buffercan be corrupted or stall under burst traffic. Comparing v1.2.0 with Python RNS 1.5.4 shows two confirmed concurrency defects plus several Buffer API parity gaps that downstream stream adapters currently have to compensate for.The two confirmed defects are:
evWriteinterest whenQueueSendraceswriteStream, stranding queued bytes until the Link goes stale.Steps to reproduce
buffer.RawChannelReaderandbuffer.RawChannelWriterusing one stream ID.go test -raceand with the default transport worker pool.QueueSendwhile the Backbone poller drains the same stream.Observed failure modes include:
pollerMod(evRead)overwrites a concurrentevRead|evWritearm;MaxPacketHandlers=1avoiding the dispatch race but reducingpacketQcapacity to one and losing burst traffic (RawChannelWriter.Writeeventually returnslink not ready; one run failed around byte 79,524).I have deterministic/stress regression tests for the Channel ordering and Backbone progress paths and can submit them with focused patches.
Expected behavior
RawChannelWritershould derive its payload limit from the live Channel MDU, as Python does.Actual behavior
1. Parallel Channel dispatch
Transportdrains inbound work into a multi-worker pool.Channel.HandleInboundprotects RX-ring mutation, but releases the Channel mutex before invoking handlers. A worker can drain envelope N and be preempted before its handler runs; another worker can then drain and dispatch N+1 first. Both handlers append to the sameRawChannelReader, corrupting byte-stream order.Python RNS avoids this because its inbound priority queue has one
inbound_jobconsumer, andChannel._receiveperforms contiguous RX-ring delivery serially.A working fix is to serialize emplace + drain + dispatch per Channel. A dedicated serial dispatcher would be preferable to a global/single transport worker, since
MaxPacketHandlers=1also shrinks the packet queue and causes drops.2. Backbone write-interest race
In v1.2.0,
writeStreamcan observe an emptytxBuf, releases.mu, and then callpollerMod(fd, evRead). A concurrentQueueSendcan append a frame and armevRead|evWritein that gap; the stale disarm then wins and no further write event flushes the queued frame. The partial-write branch also leaves buffered bytes without explicitly re-armingevWrite.The fix is to decide and apply poller interest while holding the same lock that protects
txBuf/wantOut, and to explicitly retainevWriteafter partial writes.Related Buffer parity gaps
These are not required to reproduce the two races, but they affect the same stream use case:
RawChannelWritercomputes_mdu = channel.mdu - StreamDataMessage.HEADER_LEN. Go uses the constantMaxDataLen = 457. With a negotiated Link MDU of 431, the Channel body limit is 425 and the stream payload limit is 423; the fixed 457-byte limit can therefore produce a ChannelErrTooBig.Nonefor would-block and exposes ready callbacks. Go maps an empty, non-EOF read to(0, nil), which causes normalio.Readerconsumers to spin and lets a read immediately after an asynchronous EOF observe no EOF.RawChannelWriter.Writewaits withcontext.Background(), so a caller cannot cancel blocked backpressure. Python's raw writer reports no progress when the Channel is not ready, and its close wait is bounded.Python RNS comparison
Compared against Python RNS 1.5.4, commit
192898864c008b6287dd56781d89fccef0bb5f7a:RNS/Transport.py: oneinbound_jobdrains the priority queue sequentially;RNS/Channel.py: contiguous envelopes are delivered serially;RNS/Buffer.py: writer payload size is derived fromchannel.mdu, and reader readiness is represented explicitly;RNS/Interfaces/BackboneInterface.pyandTransmitBuffer.py: a single epoll loop rechecks sendable data before disablingEPOLLOUT, with ingress and egress pressure controls.I am happy to split the Buffer/API parity items into follow-up issues if that is easier to review. The Channel ordering and Backbone write-interest fixes are independently testable and should be safe as separate commits.