Skip to content

Prove namespace noisy-neighbor containment for shared Server deployments #91

Description

@rmcdaniel

Problem

Namespace authorization prevents cross-tenant data access, but a shared Durable Workflow Server also needs resource isolation. One namespace must not monopolize API capacity, task leasing, scheduler work, database growth, Redis state, or downstream service-call budgets until unrelated namespaces stop making progress.

Server already has queue/namespace lease and dispatch admission controls plus task-queue fairness primitives. This work must audit those controls as one system, close uncovered paths, and prove an adversarial namespace remains contained under realistic constrained capacity.

Required contract

  • Enforce server-owned per-namespace admission independently of caller-supplied queue, priority, or fairness metadata.
  • Preserve bounded progress for quiet namespaces while another namespace floods workflow starts, activities, timers, signals, queries, updates, schedules, child workflows, and Nexus/service calls.
  • Bound namespace-owned active leases, pending tasks, dispatch rate, API concurrency/rate, open workflows, history growth, timers/schedules, external payload usage, and other durable/cardinality growth that can exhaust shared infrastructure.
  • Return explicit retryable admission responses and operator-visible metrics identifying the namespace and exhausted budget; do not silently drop durable work.
  • Keep counters and leases correct through process restart, Redis interruption, expired leases, and retry/replay.
  • Define which limits are configurable defaults, per-namespace overrides, and hard server ceilings. Unsafe unlimited defaults must be explicit rather than accidental.

Adversarial evidence

Build a repeatable constrained-cell experiment with at least one noisy namespace and one control namespace. It must cover:

  • sustained workflow-start and activity-dispatch saturation;
  • same-queue and different-queue contention;
  • timer, signal, query, update, schedule, child-workflow, and Nexus/service-call storms;
  • large/replay-heavy histories and external payload pressure;
  • worker/server restart and Redis interruption during saturation;
  • database, Redis, memory, CPU, queue depth, rejection counts, and p50/p95 control-namespace latency.

The control namespace must continue starting and completing a documented minimum workload within a bounded latency envelope while the noisy namespace receives deterministic throttling. After pressure stops, both namespaces must recover without manual data repair.

Completion

  • Inventory every shared resource and admission path, with existing coverage and gaps.
  • Implement missing namespace-level limits/fair scheduling and tests.
  • Publish a reusable experiment in this repository, not provider-specific orchestration.
  • Record an exact Server/image tuple and adversarial results on this issue.
  • State the proven shared-server operating envelope and remaining non-isolated resources.
  • Only then recommend whether mutually untrusted namespaces can safely share one small Server deployment.

This is a Server product capability. Hosting products can consume the result, but provider topology and commercial plan decisions are out of scope here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    authority:githubGitHub is the authoritative lifecycle record for this workcompletion:evidence-verifiedAcceptance, fixed version, and required operational evidence are publicly verifiedkind:featureA public product capability or experience is requestedpriority:P1High-priority product or release riskrepo:serverOwned by the standalone server repositorystatus:doneDerived from the authoritative closed issue state

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions