Hacker News
new
past
comments
ask
show
jobs
points
by
TomGarden
3 hours ago
|
comments
by
petu
3 hours ago
|
next
[-]
That's max output tokens per response limit, separate from context length
reply
by
simonw
3 hours ago
|
prev
|
next
[-]
It's the output token limit, which has been 128,000 for Claude models for quite a while note
reply
by
croemer
2 hours ago
|
parent
|
[-]
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
reply
by
simonw
2 hours ago
|
parent
|
[-]
I think this is a bug. I've not seen this problem from any of the other frontier models.
reply
by
NewJazz
1 hours ago
|
parent
|
next
[-]
I would also consider this a bug. I think ajy reasonable consumer would.
reply
by
Insanity
2 hours ago
|
parent
|
prev
|
[-]
Do other models put a hard cap on the output tokens it can generate?
reply
by
simonw
1 hours ago
|
parent
|
[-]
Yes, the OpenAI GPT-6 Astra limit is 128,000 as well:
https://developers.openai.com/api/docs/models/gpt-6-astra
Gemini 3.8 Flash is 65,536
https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
reply
by
3 hours ago
|
prev
|
[-]
deleted
reply