Skip to content

fix(examples): reconnect the otel exporters after a dropped connection - #91

Merged
nhobin219 merged 1 commit into
mainfrom
otel-reconnect
Oct 3, 2026
Merged

nhobin219 merged 1 commit into
mainfrom
otel-reconnect

Conversation

@nhobin219

Copy link
Copy Markdown
Owner

The OTel exporters each held one Publication for the life of the process. A broker restart therefore ended a service's telemetry for good, silently: every later batch failed against the closed connection, and OTel's SDK ignores a FAILURE result.

Change

  • common.Publisher replaces the common.publish(publication, loop, rows) helper. It holds the stream's URI:
    • it connects on the first batch;
    • a batch that fails on the connection drops it, so the next batch connects again;
    • there's no retry loop, because the SDK calls export every half second (logs, spans) or every second (metrics), and that is the retry;
    • a lock around connecting stops two services' processors, which share an exporter, from opening two connections at once.
  • The exporters take a Publisher: StreamLogExporter(publisher), StreamSpanExporter(publisher), StreamMetricExporter(publisher). services.exporting() builds one per stream and closes them at the end.
  • Still never raises. A failed batch is dropped and export returns FAILURE; nothing reaches the application.
  • One printed line when telemetry starts being dropped, and one when it's published again, so an outage is visible without a line per batch. It's printed rather than logged: with OTel's LoggingHandler on the root logger, a log record about a failing exporter would be exported through that same exporter. That's 8 lines.

Tests

  • New: TestThePublisher. It sends a batch, stops the broker, sends twice (both return False, nothing raises), restarts the broker on the same port and sends again.
    • The last send succeeds on a new connection.
    • Exactly two lines are printed: dropped, then published again.
    • The stored rows are the two sent while the broker was up.
  • Falsified: never reconnecting, printing per batch, and never printing recovery each fail the test.
  • Live run: a broker and services.py as separate processes, with the broker killed mid-run and restarted 5 s later. Each of the three streams printed one "dropped" line and one "published again" line. The producer kept taking orders throughout (36), with no tracebacks.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi

Each exporter held one publication for the life of the process, so a
broker restart silently ended a service's telemetry. Exporters now take
a common.Publisher holding the stream's URI: it connects on the first
batch, drops a connection that fails, and the next batch reconnects.
A failed batch is still dropped, never raised, and one line is printed
when telemetry starts being dropped and one when it recovers. Printed,
not logged, so a record about a failing exporter cannot loop back
through it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FsSDkeb5rVAxA1FSmKKfQi
@nhobin219
nhobin219 merged commit 7451c22 into main Oct 3, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant