After waiting 49 minutes and 16 seconds while the model was still thinking, I'm simply giving up...

#92
by MrDevolver - opened

Hello,

After waiting 49 minutes and 16 seconds while the model was still thinking about the following prompt

Create a character from retro side scroller game, html single file.

I'm simply giving up with conclusion that this model is not for me.

I hope you guys find the model more useful than I did.

That's all I wanted to say here, thanks for your time reading this.

Have a nice day.

hopefully the community makes a version that doesn't overthink like thinking cap on 3.8 27B or something

Lower reasoning effort, default is xhigh

Same here, I'm switching to 3.6 27B

i really like it though! glm 5.2 or gpt sol like king of thinking

I asked for information on triangular numbers and it took over 40 mins to think about it generating 40k of thought tokens?? clearly a problem!

there is a new paremeter!!

"-rea", "on",
"--reasoning-budget 256",

rea budget control the model think tokens

if it's sitting in think for 50 min that's the loop, not the gpu. i have it at 0.156s first token / 140 tok/s on one 6000, reasoning effort is a knob.
https://inference.tiyuvta.ai/app

You can change the thinking effort. It's a trade-off of speed and quality. Consider using a harness that'll execute smaller thinking tasks

Ha-ha, you gave up too early, 2h 12min 2 timeout errors from qwen code, 2 retry attempts, still waiting for 3D aquarium in single HTML 😃

Strix Halo 128 GB VRAM

Sign up or log in to comment