Background
Retain currently chunks long JSON conversations by character count between individual turns. This can place a user turn and its response in different chunks. It also ignores topic boundaries, which can weaken the context available to fact extraction.
Current behavior
- Long conversations use deterministic, size-based chunking.
- The chunker does not identify topic transitions.
- Retain does not persist a versioned semantic plan for safe reuse.
- LongMemEval artifacts do not record the chunking policy that produced their memories.
Expected behavior
Use the Retain language model to select topic boundaries for long JSON conversations. The model must return boundary indices only. HMS must build every chunk from the original exchanges and must not let the model rewrite, omit, or reorder source values.
Semantic planning should be enabled by default. Short content, trusted pre-chunked input, and non-conversation content should avoid the extra provider call.
Acceptance criteria
Dependency
Depends on PR #6, which implements #4 and introduces the Retain ingestion pipeline and checkpoint contracts extended by this work.
Risks
- Long conversations require an additional model call, which adds latency and cost.
- Enabling the feature by default changes chunk identities and can require a conservative FULL refresh.
- Provider failures can affect ingestion when the
raise policy is selected.
- Provider Batch extraction cannot safely use semantic planning until its checkpoint binds results to a semantic plan digest.
Owner
Owner: @Danielxu0208
Background
Retain currently chunks long JSON conversations by character count between individual turns. This can place a user turn and its response in different chunks. It also ignores topic boundaries, which can weaken the context available to fact extraction.
Current behavior
Expected behavior
Use the Retain language model to select topic boundaries for long JSON conversations. The model must return boundary indices only. HMS must build every chunk from the original exchanges and must not let the model rewrite, omit, or reorder source values.
Semantic planning should be enabled by default. Short content, trusted pre-chunked input, and non-conversation content should avoid the extra provider call.
Acceptance criteria
fixed_fallbackorraisewhen planning fails.fixed_fallbackremains the default.Dependency
Depends on PR #6, which implements #4 and introduces the Retain ingestion pipeline and checkpoint contracts extended by this work.
Risks
raisepolicy is selected.Owner
Owner: @Danielxu0208