Skip to content

[ServiceBus] pyamqp receiver-link settlement outcomes are unobservable — completes can be silently lost and redeliver at lock expiry #48143

Description

@dtucker-auro
  • Package Name: azure-servicebus
  • Package Version: 7.14.3 (pyamqp transport)
  • Operating System: Linux (AKS)
  • Python Version: 3.10

Describe the bug

In a file-ingestion pipeline settling batches of 400–4,000 messages over a peek-lock receiver, complete_message() returns success for every message, yet a fraction (observed 4–47% per batch across many runs) redelivers exactly one lock-duration after receive. Message IDs are deterministic and downstream writes are idempotent, so we can prove via ground-truth accounting that these are re-deliveries of messages the client believes it completed — 648 duplicate deliveries across a ~16,600-message campaign, each wave arriving at receive + lockDuration.

The mechanism appears to be acknowledged in the SDK's own source — _settle_message_with_retry in azure/servicebus/aio/_servicebus_receiver_async.py contains:

"pyamqp doesn't currently (and uamqp doesn't have the ability to) wait to receive disposition result returned from the service after settlement, so there's no way we could tell whether a disposition succeeds or not and there's no error condition info. (for uamqp, see issue: Azure/azure-uamqp-c#274)"

i.e., receiver-link settlement sends a settled=True disposition frame fire-and-forget, so a disposition that the service never processes is indistinguishable client-side from a successful completion. For contrast, the .NET SDK's AMQP layer awaits disposition outcomes, so this appears to be a Python-transport gap rather than a service limitation.

To Reproduce

We could not reproduce synthetically despite five controlled attempts (idle and loaded, standalone VM and in-cluster with matching image/identity/CPU-limits/network, 800-message batches, settle-then-close-immediately and settle-then-linger variants) — all settled cleanly. Production hits it consistently on large batches processed over multi-minute hold times. Happy to share timing profiles and environment details.

Expected behavior

Either settlement failures surface as exceptions/return values, or an option exists to await the disposition outcome.

Additional context / feature request

As a workaround we settle through the management link (com.microsoft:update-disposition), which is request/response and accepts lock-token arrays — batched acknowledged completion measures ~18–31 ms per 500-token call. This currently requires the private _settle_message_via_mgmt_link surface, which we'd love to stop depending on. Two suggestions, either of which would let us delete the workaround:

  1. An option for complete_message() (and friends) to await/verify the disposition outcome, matching .NET semantics.
  2. A public batched settlement API over the management link (the operation already accepts token arrays).

Metadata

Metadata

Assignees

No one assigned

    Labels

    ClientThis issue points to a problem in the data-plane of the library.Service AttentionWorkflow: This issue is responsible by Azure service team.Service Buscustomer-reportedIssues that are reported by GitHub users external to the Azure organization.needs-team-attentionWorkflow: This issue needs attention from Azure service team or SDK teamquestionThe issue doesn't require a change to the product in order to be resolved. Most issues start as that

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions