Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I had to fix this on 35B A3B -- I have a proxy that just shuts it down if it gets to 2K thinking tokens and injects something like "We have thought enough, let's begin working." and it almost always finishes the turn then. It rarely needs more than 2K thinking tokens and if it does there is always next turn. I would need to see what 27B is actually doing, but these smaller Qwen models seem prone to this.


Unfortunately in xhigh reasoning effort it will burn through 2K tokens before it has even finished its bullet point overview. It really is intense and obsessive. You might need ten times more!

Your strategy would likely help in medium reasoning effort (because there it gets caught up in the very typical Qwen looping).

Not seen looping in the “low” reasoning effort mode.


I have been using Muse Glimmer for a few days instead of A3B. It gets the job done quicker than A3B despite being several times slower.


Yes — I just found out that you can set reasoning level in the prompt — like with Qwen 3.8 27B it is actually really pretty solid at "Reasoning level: low".

Ten to thirteen tokens per second on my M1 Max (might be some room to improve this) but it indeed solved as fast as the Qwen 35B. 40 seconds faster on one of my tests that involves three steps.

This is very striking.


>I had to fix this on 35B A3B -- I have a proxy that just shuts it down if it gets to 2K thinking tokens and injects something like "We have thought enough, let's begin working." and it almost always finishes the turn then.

that is amazing, thanks for sharing.


Could you share the proxy and config ? This sounds good.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: