Skip to content

Retain: make semantic conversation chunking the default #5

Description

@Danielxu0208

Background

Retain currently chunks long JSON conversations by character count between individual turns. This can place a user turn and its response in different chunks. It also ignores topic boundaries, which can weaken the context available to fact extraction.

Current behavior

  • Long conversations use deterministic, size-based chunking.
  • The chunker does not identify topic transitions.
  • Retain does not persist a versioned semantic plan for safe reuse.
  • LongMemEval artifacts do not record the chunking policy that produced their memories.

Expected behavior

Use the Retain language model to select topic boundaries for long JSON conversations. The model must return boundary indices only. HMS must build every chunk from the original exchanges and must not let the model rewrite, omit, or reorder source values.

Semantic planning should be enabled by default. Short content, trusted pre-chunked input, and non-conversation content should avoid the extra provider call.

Acceptance criteria

  • Long role-and-content JSON conversations above the configured chunk size use semantic boundary planning by default.
  • The provider returns only ordered end-exchange indices that cover every exchange exactly once.
  • Materialization preserves all source values and keeps each user exchange intact. A single oversized exchange remains atomic and is marked as such.
  • Short content passes through unchanged. Trusted pre-chunked input bypasses planning. Non-conversation content uses deterministic chunking.
  • Operators can select fixed_fallback or raise when planning fails. fixed_fallback remains the default.
  • Retain persists versioned, text-free manifests with input hashes, policy fingerprints, and plan digests.
  • Compatible retries reuse the durable plan without another provider call.
  • Boundary drift triggers a conservative FULL plan, while invalid or tampered recovery data fails closed.
  • Unsupported provider Batch requests are rejected when semantic planning is active.
  • LongMemEval records the complete Retain chunking policy and rejects incompatible resume artifacts.
  • Focused unit, service, configuration, extraction-boundary, and release-integrity tests pass in CI.

Dependency

Depends on PR #6, which implements #4 and introduces the Retain ingestion pipeline and checkpoint contracts extended by this work.

Risks

  • Long conversations require an additional model call, which adds latency and cost.
  • Enabling the feature by default changes chunk identities and can require a conservative FULL refresh.
  • Provider failures can affect ingestion when the raise policy is selected.
  • Provider Batch extraction cannot safely use semantic planning until its checkpoint binds results to a semantic plan digest.

Owner

Owner: @Danielxu0208

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions