Horrible model .... Overthink eats all context tokens

#136
by kremerneil - opened

Huge expectation but this is overthinking pro max model which wastes token, time and electricity.
Does not honor the reasoning_effort variable as I tried medium and low.
Steer to other files in the directory that are not asked for, compare them with the file in questions and keep making permutation and combinations of all expected scenario of expected results and then fail because all context are eaten up and that's too when task are simple.
Please release this model so it does not let us down with these issues.

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

Have you given this template a try yet?

No, but I used the model default template and changed the reasoning_effort to medium and low.
If I have to use non-default template, does it means the Qwen own templates is terrible?
If yes, my original comment stands, it is complaint to Qwen as a user.

Qwen has had a chat-template problem since 3.5, on 3.6 the template was making the model do wrong tool calls, and now in 3.8 it seems to ignore the thinking budget constraints all together, they seem to be unable to get it right, people started used froggerics template back then, he released a new one for 3.8, might still be a little buggy.
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

No, but I used the model default template and changed the reasoning_effort to medium and low.
If I have to use non-default template, does it means the Qwen own templates is terrible?
If yes, my original comment stands, it is complaint to Qwen as a user.

Yes, it does. It has been relatively bad since 3.5, this template has fixed all issues for me,.from tool calling, some looping.

Here is the fix to your problem, you should try it.

man people will get frontier intelligence on their own hardware, and still find a way to complain

man people will get frontier intelligence on their own hardware, and still find a way to complain

That's genuine criticism, and @kremerneil does the right thing. They're not complaining, they're showing that they wish local "frontier intelligence" to be even better.

Such honesty is vital against stagnation, times more useful than a thousand of "I love ProductName". So, thanks to @kremerneil for speaking up.

How you gonna complain and give attitude after receiving something for free is beyond me … lol people are so fk ungrateful these days…

Go start on some reading material and stop complaining, be glad they continuously release models, as pretty much last org that drops consumer hardware able models, that are actually quite decent mind you.

(Looking at you Gemini 😂)

Since when criticism is a sign of being ungratefulness.... people who assume and start abusing others are the worst kind.
Nowhere I said to be ungrateful to Qwen but this attitude of not telling the truth about something is the worst kind of behavior to which I am not part of.
Criticism is a part of feedback which inform the one concerned that something need to be improved. If it does not work for the folk who want to use it with local hardware as intended, what is the point of it being free? It is not solving the problem, and mind you I have absolutely zero problem with Qwen3.6 which works as it should. It is only this horrible model Qwen3.8.

How you gonna complain and give attitude after receiving something for free is beyond me … lol people are so fk ungrateful these days…

Go start on some reading material and stop complaining, be glad they continuously release models, as pretty much last org that drops consumer hardware able models, that are actually quite decent mind you.

(Looking at you Gemini 😂)

The real problem isn't that it's free. It's that Qwen markets this model as "Frontier," but in reality it only works well with code, under very specific conditions. What’s more, saying it’s free is particularly stupid, you have to buy GPUs, mainly NVIDIA ones. Remind me, what’s NVIDIA’s stake in this?
For my part, I’ve spent a long time benchmarking Qwen 3.8, and I find that it excels at certain tasks but is generally poor. This is my personal take on my work tasks at every possible level. Why? Because it’s slow, glitches with a LOT of hallucination in the reasoning, sometimes uses way too many tokens, and a good result (when it finally comes) is magical, but it’s more magic than math. Its only advantage is that it’s stable in terms of reproducibility, but what does that even mean… that I have to code my own templates? or my own custom formats? Everyone’s raving about this model that requires some tinkering, and they’re calling it “Frontier”? Are you kidding me? Let’s be serious, if I were to release a product like that, all of you kick my ass because it's not a product, it's just showtime. Qwen release it just because of Muse Glimmer. And it worked, in the end, no one tested it. But to go so far as to criticize @kremerneil just because he's unhappy with an unfinished product... seriously??

@YoRandom Its a coding model buddy, generic models already exist and they suck at this particular niche that qwen is supplying. There is a reason qwen3.6 was hailed as the greatest of its time and it wasnt because it was capable of telling good stories.
Go get the fixed chat template, it was the same one used in 3.6, it was a problem then and its a problem now, these things happen and we deal with it as best we can, stop bitching about getting a product you payed nothing to get. Learn the quirks so it can do its job properly like you would with any other model and enjoy yourself.
Or go use muse glimmer if you want, doesnt matter.

Qwen release it just because of Muse Glimmer.

Are you sure it's not the other way round? Meta wanted to get some traction before the new Qwen model inevitably moved it down a tier, and it worked! Muse-Glimmer's caveman-like thinking and insanely fast speculative decoding definitely put it on my radar. I'll be looking forward to their next release.

If you are about capping this excellent model's intelligence by reducing its reasoning (which actually is a big part of intelligence).. then you probably do not need such model at all. Better use lighter model: faster, less electricity. 100% win scenario.
There are some people who appreciate that this model can actually solve really hard problems through overthinking. It was not possible for local consumer setups before.

If you are about capping this excellent model's intelligence by reducing its reasoning (which actually is a big part of intelligence).. then you probably do not need such model at all. Better use lighter model: faster, less electricity. 100% win scenario.

This model can't solve a simple stuff in reasonable token or time and keep on thinking, trying all permutation and combination instead of giving a result that Qwen3.6 can give. I gave it a refactoring job, simple one, it ate double the tokens to do the same task as Qwen3.6. Shouldn't a intelligent model can do quick thinking? How does an intelligent person do it?
It is slower because it has a disease called analysis paralysis, a habit of going on path of ...Hmm...Wait..But wait ...No, wait ... Is intelligence mere a play of permutation and combination with learned pattern?
It may be a specific use case model for those who can have massive amount of VRAM for context where it may shine but for simple task or refactoring, it is not special and in fact worse than Qwen3.6.

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
Have you given this template a try yet?

This template is better than the default template for sure, it uses less token but still much verbose than Qwen3.6. More serious issue is, it steers away from the given files and their pattern and many times ignore points from the prompt. Anyway, this is my finding as a user and it is not something Qwen may be interested in. Qwen does not release anything in this regard as it is not their purpose, it is for their own use in their own ecosystem of hardware and software. Not a complaint on it as it is free though a guidance to fit it to users like us will be appreciated.

@TTA21 That's not how Qwen communicates it, though. But let's move on, the point isn't whether it's "Chat" or "Code." What matters is providing realistic specifications at launch. I agree with you that people choose how they use it, but try to understand the frustration of anyone who doesn't just write code.

@bfkxnr We can't be certain, but that's what Alibaba does, releasing models right after a competitor, either a very large one to be compared against, or a direct answer. Alibaba needs to stay the leader in small models and hold the best portable model. We just don't know exactly why they need this so badly yet. And don't tell me it's some purely altruistic gift to the world...

I'm not saying Qwen is the only one doing this. Everyone does, Anthropic, OpenAI, all the Chinese labs. Even more so since the Frontier OpenSource release. I do think, however, that Meta released Glimmer and Qwen pushed 3.8 out very quickly to overshadow it. That's fair game, but given your reactions, it makes sense. Especially since Glimmer isn't code-focused at all, it's actually pretty bad at that, but it's excellent at prose, finance, and communication. And it has a completely different architecture. On a 5090, it provides 600K of context and runs at 3,500 tok/s once parallelized at 64 slot. (https://huggingface.co/meta-models/Muse-Glimmer-30B/discussions/56)

@NikiKrutan That's what I said... It's a brilliant model for code, but poor for everything else. That way of thinking is problematic for many users because, in many cases, it costs more than it's worth. You're right, it's fantastic at this scale in many ways, but again, look beyond code, and no. I think you're only looking at the problem from one angle. 3.8 is clearly designed to be used in xhigh mode, but then why would Qwen go through the trouble of customizing it and making reasoning effort configurable? Furthermore, yes, using a smaller model for specific tasks often offers greater value. Combining multiple calls with context, a well-prepared prompt, and available documentation takes just as much time, if not less, to solve a complex problem.

If you think that just throwing out a long prompt is the best use for it, then fine, but we're clearly not on the same page. Qwen3.8 is great as a coding assistant, but in the real world, in actual applications, it's more hit-or-miss.

That's the whole problem: it's easier to get results for code. It generates code templates or basic chat responses. But the moment you want a real-world application in the professional world, which isn't limited to code, guys, everything becomes more complicated, and often nonexistent.

This is where two visions of what an LLM should be clash. On one side, the ideal use case: code. On the other, the rest of the world. I think that before we even criticize a model, we should first ask ourselves, and state explicitly, what purpose we have in mind, so we can actually speak meaningfully about it. I’d add that LLMs are meant to save us time on the job. If we spend too much time proofreading and correcting, we might as well do it ourselves. If we have to check how our agents are performing in real time, we might as well do it ourselves. If we have to buy a 200K server to meet our needs properly, we might as well sell our kidneys right now and hand over our credit cards for APIs.

Also, you're all forgetting to mention reproducibility and stability. 3.8 has major flaws in its line of reasoning, even though it often gives a very good answer. But what does that really mean, that you can't reproduce the reasoning, and that you have to systematically verify everything the moment it's done? I know I'm repeating myself, but it all depends on how you use it.

Finally, I want to remind you that none of this is free. RAM prices have risen by 500% in 12 months. GPU prices and availability, the 4090 at $4000 and PRO600 have surpassed $15,000. We are the alpha testers in a race between the West and China to produce models that will later be embedded in everyday consumer products, not for coders, but for the general public. So no, it's anything but free. But if you'd rather ignore that, that's your call.

I've found that recommended by Qwen temperature setting of 1.0 is not good. 0.8 is much much better. But yeah, I do code design docs and agentic coding mostly. So this really depends on use case.

About price.. Well, I managed to work with 3.6-27B on 1x5060Ti 16 Gb VRAM with 32 Gb RAM with ~130K context (fully in VRAM). That is not kidney-selling setup. Slow? Yes! But cheaper than what I would spend on paid APIs (I use free though) doing the same job. Now with 3.8-27B most free options become inferior an paid ones win mostly by speed. That's again for my use case which differs from more casual use cases.

Since when criticism is a sign of being ungratefulness.... people who assume and start abusing others are the worst kind.
Nowhere I said to be ungrateful to Qwen but this attitude of not telling the truth about something is the worst kind of behavior to which I am not part of.
Criticism is a part of feedback which inform the one concerned that something need to be improved. If it does not work for the folk who want to use it with local hardware as intended, what is the point of it being free? It is not solving the problem, and mind you I have absolutely zero problem with Qwen3.6 which works as it should. It is only this horrible model Qwen3.8.

Haven’t seen you thank them for their work anywhere? so it doesn't seem like you’re too grateful. Its easy to complain innit
me on the other hand I have no issues. Besides Qwen might have their own reasons as to why to ship with the template they do. Id assume they also serve their own market first - swap out for froggeric and you’ll have less problems.
How about you send them $$, and then ask for a private tour and tutoring on how to use it - as I personally have no issues with it.

Since when criticism is a sign of being ungratefulness.... people who assume and start abusing others are the worst kind.
Nowhere I said to be ungrateful to Qwen but this attitude of not telling the truth about something is the worst kind of behavior to which I am not part of.
Criticism is a part of feedback which inform the one concerned that something need to be improved. If it does not work for the folk who want to use it with local hardware as intended, what is the point of it being free? It is not solving the problem, and mind you I have absolutely zero problem with Qwen3.6 which works as it should. It is only this horrible model Qwen3.8.

Haven’t seen you thank them for their work anywhere? so it doesn't seem like you’re too grateful. Its easy to complain innit
me on the other hand I have no issues. Besides Qwen might have their own reasons as to why to ship with the template they do. Id assume they also serve their own market first - swap out for froggeric and you’ll have less problems.
How about you send them $$, and then ask for a private tour and tutoring on how to use it - as I personally have no issues with it.

people who assume and start abusing others are the worst kind

Since when criticism is a sign of being ungratefulness.... people who assume and start abusing others are the worst kind.
Nowhere I said to be ungrateful to Qwen but this attitude of not telling the truth about something is the worst kind of behavior to which I am not part of.
Criticism is a part of feedback which inform the one concerned that something need to be improved. If it does not work for the folk who want to use it with local hardware as intended, what is the point of it being free? It is not solving the problem, and mind you I have absolutely zero problem with Qwen3.6 which works as it should. It is only this horrible model Qwen3.8.

Haven’t seen you thank them for their work anywhere? so it doesn't seem like you’re too grateful. Its easy to complain innit
me on the other hand I have no issues. Besides Qwen might have their own reasons as to why to ship with the template they do. Id assume they also serve their own market first - swap out for froggeric and you’ll have less problems.
How about you send them $$, and then ask for a private tour and tutoring on how to use it - as I personally have no issues with it.

people who assume and start abusing others are the worst kind

Id say cheap and unappreciative - in my book.

throw in: clueless, bitchy and entitled and you get @kremerneil

Since when criticism is a sign of being ungratefulness.... people who assume and start abusing others are the worst kind.
Nowhere I said to be ungrateful to Qwen but this attitude of not telling the truth about something is the worst kind of behavior to which I am not part of.
Criticism is a part of feedback which inform the one concerned that something need to be improved. If it does not work for the folk who want to use it with local hardware as intended, what is the point of it being free? It is not solving the problem, and mind you I have absolutely zero problem with Qwen3.6 which works as it should. It is only this horrible model Qwen3.8.

Haven’t seen you thank them for their work anywhere? so it doesn't seem like you’re too grateful. Its easy to complain innit
me on the other hand I have no issues. Besides Qwen might have their own reasons as to why to ship with the template they do. Id assume they also serve their own market first - swap out for froggeric and you’ll have less problems.
How about you send them $$, and then ask for a private tour and tutoring on how to use it - as I personally have no issues with it.

people who assume and start abusing others are the worst kind

Id say cheap and unappreciative - in my book.

throw in: clueless, bitchy and entitled and you get @kremerneil

Thanks for showing your real face.
You can go and cry. I already said, I am not for personal BS in the beginning because I know people like you don't read comments and if do, don't understand and keep assuming and abusing others.
So no BS needed from you or your kind, I had an issue and I said it and it wasn't a personal attack, it was for Qwen. So if you can't add something helpful like a kind soul in the second comment, then you can burn and cry.

Back to the topic.
This confirms that much of intelligence comes from reasoning for this model:
image
Tokens spent on full benchmark suite:
xhigh: 160M
medium: 75M
low: 43M
For the reference:
GLM 5.3 max: 170M
So actually 160M is ok.

This confirms that much of intelligence comes from reasoning for this model:
So actually 160M is ok.

Youtube channel named "Luke's Dev Lab" has done the comparisons of xhigh, medium and low in the video called "Qwen 3.8 27B Reasoning Levels Tested - Not What I Expected".
In short summary of his video:
For simpler task, low and medium reasoning efforts have similar tokens usage with medium marginally lower while xhigh is highest. Medium has lowest all-around, low having highest token usage in complicated task and xhigh have higher token if not the highest for complicated tasks.
Interestingly, the medium and xhigh coding task results are not dramatically better. In Qwen default chat templates, medium does not eject any prompt like it does for low and xhigh.

So questions to ponder:

  1. if the results are not 2-fold better, will it be okay to assume token generation consumption to be 2-fold?
  2. Does this model has inherent intelligence to decide how to accomplish a task or it rely on the outside direction via chat template to force it to reason in a particular direction or waypoints?
  3. If this model rely on the chart template to reason effectively for its intelligence (but inefficiently in token consumption), does it means we can also improve the intelligence of other models say Qwen3.6 27B via custom chat templates?
    The last point is important because at low reasoning efforts, not only the model performs worst but also consume highest tokens during reasoning due to hallucinations, overthinking etc.

Chat template is important for sure. But how much important? I don't know. It can be benched, but such benches are "expensive" to be done at home (I am not going to let it generate millions of tokens on my local setup just for benchmarking templates). I use froggeric template and it works fine with Kilo Code. I haven't tried original chat template (well, for some reason original Qwen's templates usually are far from good).

Another point to consider: model and KV-cache quantization. BF16/f16 and let's say IQ2_XS/q4_0 are very far from each other. I use my "Q5_K_S" with q8_0/q5_1 (K/V) and I am impressed with results in agentic coding. I can't say that reasoning is much more than other models (API) that give same results. But API are fast mostly and local highly depends on hardware. So I often prefer API if decent free tier is available.

Following up with measured numbers rather than more heat, since several of these claims are testable and we ran the eval.

Where this comes from: for our serving stack, Qwen and Muse Glimmer are the only two models we consider worth running at all. So this is not a Glimmer-vs-Qwen thing, and it is not ingratitude, it is the opposite. Qwen's intelligence is climbing fast and we want to run it. The problem is that the architecture and serving ergonomics are not stabilizing at the same pace as the intelligence, and at production scale that gap is what makes it unusable for us right now, not the raw capability.

Setup, so it is reproducible: Qwen3.8-27B (NVFP4) on vLLM, FinanceReasoning Hard (Program-of-Thought), 238 items, concurrency 4, 8192-token generation ceiling. We looked hard at the effort dial. Three things:

1. "reasoning_effort does nothing" is partly real, partly a harness trap. The dial is enable_thinking (a hard <|think_on|>/<|think_off|> token) plus reasoning_effort (xhigh default, medium, low) delivered as an injected instruction phrase, not a hard cap, so control is genuinely soft. On top of that we hit three ways to think you set low while actually running xhigh: (a) our config initially carried Glimmer's reasoning_strength key, which Qwen silently ignores since it only honors enable_thinking + reasoning_effort, so we stayed at the xhigh default; (b) a "none" mode that passed no kwarg at all, also defaulting to xhigh; (c) vLLM 0.27 renamed reasoning_content to reasoning, so a parser keyed on the old field silently drops the trace. Since the default is xhigh, any of these lands you there. Concrete check: GET /v1/models, confirm the served name, and count thinking tokens per response before concluding the dial is inert.

2. The "30K for 3K" framing did not hold at low in our numbers. Median thinking was ~650 tokens/problem (mean ~1,287, under Glimmer's ~1,426 on the same eval). The runaway is real but it is a tail: roughly 3% of items (6-8 of 238) ran to the 8K ceiling, p95 ~3.5K. That tail was invariant to server config, so it reads as a model property, not a serving bug. A tail is more tractable than a bloated mean.

3. More reasoning is not free quality on every task family. On fact-grounded checks we measured thinking mode increasing fabrication rather than reducing it (roughly 7% to 17% on a related Qwen model), including brand new fabrications the non-thinking pass never made. So for extraction and grounding work, low or no thinking can be the correct call on quality grounds, not only token grounds.

One caveat on any quality comparison: NVFP4 is a lossy quant and PoT is code-heavy, where we have seen roughly -5 HumanEval versus GGUF Q4_K_M elsewhere, so part of any gap can be the quant rather than the model.

Net, and I mean this as a model we want to deploy: the capability is there and rising. What is not there yet is a stable, hard-guaranteed effort dial and a serving story that holds steady across versions. Until that stabilizes, the intelligence gains are hard to actually bank at production scale.

Sign up or log in to comment