Repository navigation
feat: --max-prompt-tokens rejects over-long prompts before prefill - #205
Merged
Merged
Conversation
An opt-in limit: a prompt longer than N tokens gets an OpenAI-shaped 400 (code context_length_exceeded) on /v1/chat/completions and /v1/completions, streaming or not, before any prefill work. The slot and stats are released. Agent clients recognise that code and compact their history instead of waiting minutes on a prefill and retrying it. It is a separate option because --ctx-size is documented as a sliding window; rejecting at --ctx-size would change what that flag does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014J4GsqznbKNX8tcKzNRxEr
Member
|
Reviewed (with AI assistance, Claude Code). The count is taken after the chat template and image/audio expansion, the slot and stats are released on reject, both endpoints return the same OpenAI-shaped 400 for streaming and non-streaming, and the default leaves behaviour unchanged. No blocking issues. Merging. |
Keep both --max-prompt-tokens and --no-token-echo in README and ServerConfig. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
--max-prompt-tokens Nis an opt-in limit. A prompt longer than N tokens gets an OpenAI-shaped 400 before any prefill work:{"error":{"message":"This model's maximum context length is N tokens. However, your messages resulted in M tokens.","type":"invalid_request_error","code":"context_length_exceeded"}}/v1/chat/completionsand/v1/completions, streaming or not.context_length_exceededto compact their history. Today the only signal they get is a multi-minute prefill.--ctx-sizeis documented as a sliding window, so rejecting at--ctx-sizewould change what that flag means. The default is no limit, so behaviour is unchanged unless the flag is set.Why not a full memory-pressure guard
I set out to build a pre-prefill memory guard. On an older mlx-swift-lm, a 16 GB M2 running a 35B-A3B hybrid MoE (
qwen3_5_moe,--stream-experts) was SIGKILLed mid-prefill by the VM compressor (compressor_exhausted): the VLM path materialised fp32[heads, N, N]attention.Re-measured on
main(mlx-swift-lm b782, which has the windowed Qwen3.5 VLM prefill) with the same rig and flags (--stream-experts --ssd-prefetch --prefill-size 2048 --ctx-size 32768, auto-detected VLM):There was no
compressor_exhaustedevent, and memory demand stayed nearly flat as prompts grew. A demand-based preflight would never have triggered on this rig, so I left it out rather than ship a guessed threshold. It fits better alongside a multi-entry prompt cache, which is what would add memory that can grow.Testing
swift test --filter SwiftLMTests: 210 tests, 0 failures. NewPromptLengthGuardTestscover the boundary, the OpenAI shape and the CLI parse.Ministral-8B-Instruct-2410-4bitwith--max-prompt-tokens 50:Documented in the README flags table. This repo has no CHANGELOG or version file.
🤖 Generated with Claude Code