Skip to content

feat: --max-prompt-tokens rejects over-long prompts before prefill - #205

Merged
solderzzc merged 2 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/memory-guard
Oct 4, 2026
Merged

solderzzc merged 2 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/memory-guard

Conversation

@CodeAndCanvas728

@CodeAndCanvas728 CodeAndCanvas728 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

What

--max-prompt-tokens N is an opt-in limit. A prompt longer than N tokens gets an OpenAI-shaped 400 before any prefill work:

{"error":{"message":"This model's maximum context length is N tokens. However, your messages resulted in M tokens.","type":"invalid_request_error","code":"context_length_exceeded"}}
  • It applies to /v1/chat/completions and /v1/completions, streaming or not.
  • The generation slot and stats are released, so the next request runs normally.
  • Agent clients key on context_length_exceeded to compact their history. Today the only signal they get is a multi-minute prefill.
  • Why it's a separate option: --ctx-size is documented as a sliding window, so rejecting at --ctx-size would change what that flag means. The default is no limit, so behaviour is unchanged unless the flag is set.

Why not a full memory-pressure guard

I set out to build a pre-prefill memory guard. On an older mlx-swift-lm, a 16 GB M2 running a 35B-A3B hybrid MoE (qwen3_5_moe, --stream-experts) was SIGKILLed mid-prefill by the VM compressor (compressor_exhausted): the VLM path materialised fp32 [heads, N, N] attention.

Re-measured on main (mlx-swift-lm b782, which has the windowed Qwen3.5 VLM prefill) with the same rig and flags (--stream-experts --ssd-prefetch --prefill-size 2048 --ctx-size 32768, auto-detected VLM):

Prompt tokens Prefill tok/s MEM_DEMAND Peak swap Result
12,689 (turn 1) 88.7 8.3 GB 10.0 GB ok
14,655 (turn 2; this one died on the old code) 83.8 8.7 GB 11.3 GB ok
17,269 (the old ceiling was ~17.2k) 77.3 8.7 GB 13.2 GB ok
23,623 73.0 8.8 GB 12.6 GB ok
31,493 64.6 9.0 GB 12.6 GB ok

There was no compressor_exhausted event, and memory demand stayed nearly flat as prompts grew. A demand-based preflight would never have triggered on this rig, so I left it out rather than ship a guessed threshold. It fits better alongside a multi-entry prompt cache, which is what would add memory that can grow.

Testing

  • swift test --filter SwiftLMTests: 210 tests, 0 failures. New PromptLengthGuardTests cover the boundary, the OpenAI shape and the CLI parse.
  • Live check against Ministral-8B-Instruct-2410-4bit with --max-prompt-tokens 50:
    • A 204-token prompt gets a 400 in 11 ms on chat, and also on streaming text completions.
    • Short prompts still return 200, and the slot is freed after a rejection.
  • Built with Xcode 27.0 (27A266a).

Documented in the README flags table. This repo has no CHANGELOG or version file.

🤖 Generated with Claude Code

An opt-in limit: a prompt longer than N tokens gets an OpenAI-shaped 400
(code context_length_exceeded) on /v1/chat/completions and /v1/completions,
streaming or not, before any prefill work. The slot and stats are released.
Agent clients recognise that code and compact their history instead of
waiting minutes on a prefill and retrying it.

It is a separate option because --ctx-size is documented as a sliding
window; rejecting at --ctx-size would change what that flag does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014J4GsqznbKNX8tcKzNRxEr
@solderzzc

Copy link
Copy Markdown
Member

Reviewed (with AI assistance, Claude Code). The count is taken after the chat template and image/audio expansion, the slot and stats are released on reject, both endpoints return the same OpenAI-shaped 400 for streaming and non-streaming, and the default leaves behaviour unchanged. No blocking issues.

Merging.

Keep both --max-prompt-tokens and --no-token-echo in README and ServerConfig.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@solderzzc
solderzzc merged commit 4fcea47 into SharpAI:main Oct 4, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants