After waiting 49 minutes and 16 seconds while the model was still thinking, I'm simply giving up...

#92
by MrDevolver - opened

Hello,

After waiting 49 minutes and 16 seconds while the model was still thinking about the following prompt

Create a character from retro side scroller game, html single file.

I'm simply giving up with conclusion that this model is not for me.

I hope you guys find the model more useful than I did.

That's all I wanted to say here, thanks for your time reading this.

Have a nice day.

hopefully the community makes a version that doesn't overthink like thinking cap on 3.8 27B or something

Lower reasoning effort, default is xhigh

Same here, I'm switching to 3.6 27B

i really like it though! glm 5.2 or gpt sol like king of thinking

I asked for information on triangular numbers and it took over 40 mins to think about it generating 40k of thought tokens?? clearly a problem!

there is a new paremeter!!

"-rea", "on",
"--reasoning-budget 256",

rea budget control the model think tokens

if it's sitting in think for 50 min that's the loop, not the gpu. i have it at 0.156s first token / 140 tok/s on one 6000, reasoning effort is a knob.
https://inference.tiyuvta.ai/app

You can change the thinking effort. It's a trade-off of speed and quality. Consider using a harness that'll execute smaller thinking tasks

Ha-ha, you gave up too early, 2h 12min 2 timeout errors from qwen code, 2 retry attempts, still waiting for 3D aquarium in single HTML 😃

Strix Halo 128 GB VRAM

I think I found a better approach! Add these flags to llama.cpp, and use Pi Agent works like a charm:

-rea on --reasoning-preserve
--reasoning-effort medium --reasoning-budget 32768
--reasoning-budget-message "\n\n[Reasoning budget exhausted. Stop thinking and provide the final answer immediately.]" `

Ok, i'm giving up too, 4+ hours still thinking about 3D aquarium in single HTML. 🐢
I have never seen so slow models. Not really usable for local coding.

I run it on 4x GPUs and on xhigh reasoning it thinks for maybe 10-20 minutes then starts working.

On medium reasoning it thinks for 3-5 minutes

On low reasoning it thinks for about 30 seconds.

These are the only three reasoning levels you can set. By default it is set to xhigh and most harnesses don’t let you change the default.

I have yet to see any meaningful difference in output quality between low and xhigh but it must be there?

You can also turn thinking off entirely. This is good for stuff like tab completions or very small edits. With thinking off entirely then it starts generating instantly.

Sign up or log in to comment

Free AI Image Generator No sign-up. Instant results. Open Now