Skip to content

Async extract/addMessage/commit endpoints return 200 OK even when the work fails — silent data loss for the Python integration #56

Description

@yueyue0w0

TL;DR

OpenMemoryController.extract, addMessage, and commit are fire-and-forget. They return 200 OK immediately and dispatch the real work onto a separate scheduler. Any failure in the dispatched work is logged on the server and lost everywhere else. The Python ingest.py retry spool only triggers on transport-level failures, so a server-side exception here results in silent, unrecoverable data loss from the client's perspective.

Details

This is a common pattern that can lead to subtle data loss when async endpoints are placed in front of a retry-aware client. Looking at the controller:

// memind-server/.../controller/openapi/OpenMemoryController.java
@PostMapping("/extract")
public Mono<ApiResult<Void>> extract(@Valid @RequestBody ExtractMemoryRequest request) {
    service.extractAsync(request);
    return Mono.just(ApiResult.ok());
}

extractAsync schedules onto Schedulers.boundedElastic (see OpenMemoryApplicationService.dispatchAsync, ~line 222) and any error is handled with a log.error(...) only. The HTTP response has already been sent before the scheduled task runs.

Now the client side, memind-integrations/claude-code/scripts/lib/client.py:

def is_success_envelope(response):
    return response.status_code == 200 and response.json().get("success") is True

A 200 OK with a successful envelope is the success criterion. Combined with ingest.py (lines 82–95), where the retry spool only enqueues on requests-level exceptions or non-success envelopes, this means:

  • DB unavailable mid-extract → server logs it → client thinks it succeeded → message is marked submitted in state.json → never retried.
  • Vector store down → same outcome.
  • Any extraction-pipeline bug that throws after the HTTP boundary → same outcome.

References

  • memind-server/src/main/java/com/openmemind/ai/memory/server/controller/openapi/OpenMemoryController.java:41–53
  • memind-server/src/main/java/com/openmemind/ai/memory/server/service/openapi/OpenMemoryApplicationService.java (dispatchAsync ~line 222)
  • memind-integrations/claude-code/scripts/ingest.py:82–95

Suggested next steps

Two reasonable paths:

  1. Synchronous endpoint variants for clients that need durability — /extract/sync returns the actual outcome. The async variant remains for callers who explicitly want fire-and-forget.
  2. Server-side durable queue: persist the request before returning 200, and have the async worker drain from that queue. Then 200 truly means "we have it." This is heavier but matches the user's intuition.

Either is fine; the status quo just shouldn't be the only option.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions