概述
Together with the finish reason the API returns, they decide where an answer ends and tell you whether it ended naturally or was cut off. Getting them right prevents truncated JSON, half-finished answers and wasted spend.
深入探討
A language model generates one token at a time until something tells it to stop. There are three common reasons it stops. It produces its own end-of-turn token, meaning it decided the answer is complete. It reaches the maximum number of output tokens you allowed. Or it produces text that matches one of your stop sequences. Max tokens is often confused with the context window. The context window is the total number of tokens the model can handle at once, input plus output. Max tokens only limits the output of a single call. A model with a very large context window can still produce short responses if max tokens is set low, and input plus max tokens generally has to fit inside the context window. APIs report why generation ended. OpenAI's Chat Completions API returns a finish_reason such as 'stop' for a natural end or a matched stop sequence, 'length' when the output limit was hit, and 'tool_calls' when the model wants to call a tool. Anthropic's Messages API returns a stop_reason such as 'end_turn', 'max_tokens', 'stop_sequence' or 'tool_use'. Anthropic requires max_tokens on every request. For its reasoning models, OpenAI introduced max_completion_tokens, which counts hidden reasoning tokens as well as visible output. Stop sequences are useful for structured formats: ending after a closing tag, stopping before a speaker label, or cutting off after one list item. The matched text is normally not included in the returned output, so code should not expect to see it. The most common mistake is ignoring the finish reason. A truncated response can look plausible, especially prose, while structured output simply breaks. Another misconception is that a low max tokens makes the model write concisely. It does not; the model writes as it normally would and gets cut off. Ask for brevity in the prompt and use the limit as a safety cap.
戰略影響
成本與預算
多年來,架構決策決定著效能和營運成本。
更明確的決策
技術教育幫助團隊選擇正確的堆疊,而不僅僅是最新的堆疊。
品質管控
更好的工程選擇可以減少生產中的可靠性事故。
The Future of Max Tokens and Stop Sequences
Output limits have grown as models support longer generations, and structured output features that constrain responses to a schema reduce, but do not remove, the risk of truncation, since a long enough response can still hit the cap. Reasoning models have made output budgeting more important because hidden tokens now share the same limit. Parameter names and finish reason values differ between providers and have changed over time, so it is sensible to check current documentation and wrap these details in a small layer of your own code.
現實世界的實施
A summarizer sets max tokens to 150 and some summaries end mid-sentence; checking the finish reason shows 'length' for those calls, so the app raises the limit and asks for shorter summaries in the prompt instead.
A script that extracts data as JSON fails to parse about one response in fifty; each failure had hit the output limit, so the code now retries with a higher limit when the finish reason says the output was truncated.
A developer generating one line of dialogue at a time uses the stop sequence 'User:' so the model does not continue and write the user's next line too.
A team moving to a reasoning model finds answers coming back empty because hidden reasoning used up the output limit, so they raise the limit to leave room for both reasoning and the answer.
風險與防護欄
優化一項基準測試可以隱藏更廣泛的系統弱點。
基礎設施和維護成本常常被低估。
隨著系統變得更加複雜,安全性和可觀察性差距可能會擴大。
實施路線圖
在實施之前定義延遲、品質和成本目標。
在實際負載和資料條件下進行基準測試。
儀器監控錯誤、漂移和使用者影響。
在擴展之前準備回滾和事件回應路徑。
不斷探索
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Max Tokens and Stop Sequences quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
常見問題
What is Max Tokens and Stop Sequences?
Max tokens is the cap on how many tokens a model may generate in one response, and stop sequences are strings that end generation as soon as the model produces them. Together with the finish reason the API returns, they decide where an answer ends and tell you whether it ended naturally or was cut off. Getting them right prevents truncated JSON, half-finished answers and wasted spend.
What does max tokens limit?
Max tokens caps output for one call. The context window is the separate total limit covering input and output.
Your call returns a finish reason of 'length' in OpenAI's API. What happened?
'length' means generation stopped because the output limit was reached, so the response may be incomplete.
Which Anthropic stop_reason indicates the model decided its answer was complete?
'end_turn' means the model ended naturally. 'max_tokens' means it was cut off, and 'stop_sequence' means one of your strings matched.
Why does setting a low max tokens value not make answers concise?
The limit is a hard stop, not an instruction. Ask for brevity in the prompt and keep the limit as a safety cap.
Is the matched stop sequence normally included in the returned text?
Stop sequences end generation and are typically not included in the returned output, so your code should not rely on seeing them.
繼續學習
相關指南
為此主題精選的更多指南