are we there yet?

#1
by pirola - opened

are we there yet?

https://huggingface.co/Akahsizrr/Mini-Whale-1-12B HERE WE GO! with a few issues though, speed is not the best will update it tomorrow, and mlx, vllm and all the other ones arent availalbe yet, will add them tomorrow.
might build my own version of llama and mlx just for my architecture ahahah.

new better checkpoint coming tomorrow too, im all in on this model for the next week

Nice! Sir, what is the reason You chose QWEN3-4B and not LFM2.5-2.6B ?

I started this model before fuse-1, I don't think LFM released their 2.6B model when i started working on this

bro that's too slow.

What's the advantage that you reached?

its very slow but if i continue to work on it it has the potential to basically be deepseek v4 flash level intelligence, local, and at over 50 tps if i continue to lock in.

i get it at 10 tps on my rtx 3060. 2 days ago it was at 0.06 tps. this is real progression day after day.

Today I finished 1500 steps of training on the model, making the deepseek weights better inside the model too. 1-2 more models like this and i'll be able to provide you guys with local FRONTIER

that's a strong statement that can only be sustained through benchmarks

Sign up or log in to comment