Skip to content

DM-52067: Make Sasquatch dispatch robust to an unresponsive or broken proxy - #492

Open
mfisherlevine wants to merge 1 commit into
mainfrom
tickets/DM-52067
Open

mfisherlevine wants to merge 1 commit into
mainfrom
tickets/DM-52067

Conversation

@mfisherlevine

Copy link
Copy Markdown
Contributor

Rapid analysis is about to publish metric bundles from its worker pods via the SasquatchDatastore, where a Sasquatch outage must never fail the pipeline task doing the put. Three gaps closed:

  • No HTTP request had a timeout, so a proxy that accepted connections but never answered would block the put, and the task, indefinitely. The dispatcher now has a timeout (default 30 s) applied to every request, configurable via the datastore's "timeout" config entry.
  • The cluster id lookup parsed the JSON body before checking the status, so a 5xx (or a 2xx of the wrong shape) escaped as a bare KeyError or JSONDecodeError rather than SasquatchDispatchFailure; it now checks the status first and translates all such failures.
  • SasquatchDatastore.put only caught the two dispatch failures. Anything else propagated through the butler, rolled back the put and failed the task; it is now logged (with traceback) and swallowed, as publishing is best-effort by design.

Rapid analysis is about to publish metric bundles from its worker pods via
the SasquatchDatastore, where a Sasquatch outage must never fail the
pipeline task doing the put. Three gaps closed:

- No HTTP request had a timeout, so a proxy that accepted connections but
  never answered would block the put, and the task, indefinitely. The
  dispatcher now has a timeout (default 30 s) applied to every request,
  configurable via the datastore's "timeout" config entry.
- The cluster id lookup parsed the JSON body before checking the status,
  so a 5xx (or a 2xx of the wrong shape) escaped as a bare KeyError or
  JSONDecodeError rather than SasquatchDispatchFailure; it now checks the
  status first and translates all such failures.
- SasquatchDatastore.put only caught the two dispatch failures. Anything
  else propagated through the butler, rolled back the put and failed the
  task; it is now logged (with traceback) and swallowed, as publishing is
  best-effort by design.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant