概述
Together with the finish reason the API returns, they decide where an answer ends and tell you whether it ended naturally or was cut off. Getting them right prevents truncated JSON, half-finished answers and wasted spend.
深入探讨
A language model generates one token at a time until something tells it to stop. There are three common reasons it stops. It produces its own end-of-turn token, meaning it decided the answer is complete. It reaches the maximum number of output tokens you allowed. Or it produces text that matches one of your stop sequences. Max tokens is often confused with the context window. The context window is the total number of tokens the model can handle at once, input plus output. Max tokens only limits the output of a single call. A model with a very large context window can still produce short responses if max tokens is set low, and input plus max tokens generally has to fit inside the context window. APIs report why generation ended. OpenAI's Chat Completions API returns a finish_reason such as 'stop' for a natural end or a matched stop sequence, 'length' when the output limit was hit, and 'tool_calls' when the model wants to call a tool. Anthropic's Messages API returns a stop_reason such as 'end_turn', 'max_tokens', 'stop_sequence' or 'tool_use'. Anthropic requires max_tokens on every request. For its reasoning models, OpenAI introduced max_completion_tokens, which counts hidden reasoning tokens as well as visible output. Stop sequences are useful for structured formats: ending after a closing tag, stopping before a speaker label, or cutting off after one list item. The matched text is normally not included in the returned output, so code should not expect to see it. The most common mistake is ignoring the finish reason. A truncated response can look plausible, especially prose, while structured output simply breaks. Another misconception is that a low max tokens makes the model write concisely. It does not; the model writes as it normally would and gets cut off. Ask for brevity in the prompt and use the limit as a safety cap.
战略影响
成本与预算
多年来,架构决策决定着性能和运营成本。
更清晰的判决
技术教育帮助团队选择正确的堆栈,而不仅仅是最新的堆栈。
质量控制
更好的工程选择可以减少生产中的可靠性事故。
The Future of Max Tokens and Stop Sequences
Output limits have grown as models support longer generations, and structured output features that constrain responses to a schema reduce, but do not remove, the risk of truncation, since a long enough response can still hit the cap. Reasoning models have made output budgeting more important because hidden tokens now share the same limit. Parameter names and finish reason values differ between providers and have changed over time, so it is sensible to check current documentation and wrap these details in a small layer of your own code.
现实世界的实施
A summarizer sets max tokens to 150 and some summaries end mid-sentence; checking the finish reason shows 'length' for those calls, so the app raises the limit and asks for shorter summaries in the prompt instead.
A script that extracts data as JSON fails to parse about one response in fifty; each failure had hit the output limit, so the code now retries with a higher limit when the finish reason says the output was truncated.
A developer generating one line of dialogue at a time uses the stop sequence 'User:' so the model does not continue and write the user's next line too.
A team moving to a reasoning model finds answers coming back empty because hidden reasoning used up the output limit, so they raise the limit to leave room for both reasoning and the answer.
风险与防护栏
优化一项基准测试可以隐藏更广泛的系统弱点。
基础设施和维护成本常常被低估。
随着系统变得更加复杂,安全性和可观察性差距可能会扩大。
实施路线图
在实施之前定义延迟、质量和成本目标。
在实际负载和数据条件下进行基准测试。
仪器监控错误、漂移和用户影响。
在扩展之前准备回滚和事件响应路径。
不断探索
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Max Tokens and Stop Sequences quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
常见问题
What is Max Tokens and Stop Sequences?
Max tokens is the cap on how many tokens a model may generate in one response, and stop sequences are strings that end generation as soon as the model produces them. Together with the finish reason the API returns, they decide where an answer ends and tell you whether it ended naturally or was cut off. Getting them right prevents truncated JSON, half-finished answers and wasted spend.
What does max tokens limit?
Max tokens caps output for one call. The context window is the separate total limit covering input and output.
Your call returns a finish reason of 'length' in OpenAI's API. What happened?
'length' means generation stopped because the output limit was reached, so the response may be incomplete.
Which Anthropic stop_reason indicates the model decided its answer was complete?
'end_turn' means the model ended naturally. 'max_tokens' means it was cut off, and 'stop_sequence' means one of your strings matched.
Why does setting a low max tokens value not make answers concise?
The limit is a hard stop, not an instruction. Ask for brevity in the prompt and keep the limit as a safety cap.
Is the matched stop sequence normally included in the returned text?
Stop sequences end generation and are typically not included in the returned output, so your code should not rely on seeing them.
继续学习
相关指南
为此主题精选的更多指南