RUN on CPU + RAM

#121
by Anngo554 - opened

Guys do we have any similar models optimized for cpus not gpus? model is already quantized have no idea how to retrain it for cpus, what do you think what approach is the best

does sci-fi level consumer SIMD processors count? if so then yes, otherwise no. current off the shelf cpus do not have the capabilities to compute a causal model at this size at a satiable rate.

I know It's gonna be slow on simd but at least it will work on relatively cheap hardware

life is short and time is precious. please don't do it to yourself.

You might get close to 5 tok/s on a 4 epyc cpu cluster with 32 memory channels with 32 ddr5 dims of course, which would still be atleast 10,000$

Well, yep, but problem is cpus don't have native MXFP4, even epycs
All said, model will waste more compute on transforming to floats then compute itself, I guess complete redesign is needed to use floats efficiently on llms this size

By far the cheapest and most reasonable way to run Kimi K3 locally is a cluster of 16 DGX Sparks networked using either 100G or 200G DAC cables depending on what networking switch is chosen. This set up costs about $100k currently, can be purchased through retail channels, and uses multiple times less power in watts than a minimal system using data center GPUs, which of course can't even be bought retail. Using CPU or offloading the weights onto disk or system memory is going to be too slow to even be worth the time to explore. $100k is the minimum entry price to run this model, and while expensive, it's still really awesome for what it is.

Colibri engine is your only choice. They plan to release it after Inkling.
https://github.com/JustVugg/colibri/issues/658

It works, GLM5.2 i wont recommend, 0.3 tokens/sec no matter hardware, engine is quite wet and early to show performance but it works on any potato. Quality of model in Colibri is very low by tests from GLM, but there's no way to compare, very few models supported and GLM5.2 is bad release.
Please dont waste your $ savings on hardware-this models wont return back even a penny, hardware will be obsolete, models too, no savings. People dont know how to use them right, these models will not bring you anything useful. Secret know-how available but guarded, as Alex Karp (Palantir) publicly said again on TV 5th time-they can turn any dumb model into state of the art with their secret sauce.
Giveaway from me: question "are you sure?" works on many models and they openly confess of lying and presenting wrong data deliberately. In fact correcting their work seriously. May not work on Openai's GPT456, but works on GLM models, including the free model 4.7 on their site.

Colibri engine is your only choice. They plan to release it after Inkling.
https://github.com/JustVugg/colibri/issues/658

It works, GLM5.2 i wont recommend, 0.3 tokens/sec no matter hardware, engine is quite wet and early to show performance but it works on any potato. Quality of model in Colibri is very low by tests from GLM, but there's no way to compare, very few models supported and GLM5.2 is bad release.
Please dont waste your $ savings on hardware-this models wont return back even a penny, hardware will be obsolete, models too, no savings. People dont know how to use them right, these models will not bring you anything useful. Secret know-how available but guarded, as Alex Karp (Palantir) publicly said again on TV 5th time-they can turn any dumb model into state of the art with their secret sauce.
Giveaway from me: question "are you sure?" works on many models and they openly confess of lying and presenting wrong data deliberately. In fact correcting their work seriously. May not work on Openai's GPT456, but works on GLM models, including the free model 4.7 on their site.

What does colibri do that llama.cpp doesn't?

What does colibri do that llama.cpp doesn't?

You can read on author's Github, i'm not the author.
Mainly run just from SSD utilyzing its storage size (almost no Ram, no GPU-someone on fancy CPU's got even 2-3 token/s that was unique, i wasnt able).
P.S.: if anyone want to use selfcheck method "are you sure", you need to make it several times, "Oops i did it again" is very constant problem in models.

What does colibri do that llama.cpp doesn't?

You can read on author's Github, i'm not the author.
Mainly run just from SSD utilyzing its storage size (almost no Ram, no GPU-someone on fancy CPU's got even 2-3 token/s that was unique, i wasnt able).
P.S.: if anyone want to use selfcheck method "are you sure", you need to make it several times, "Oops i did it again" is very constant problem in models.

You can do that on llama.cpp

You can do that on llama.cpp

Ok. You can try make a guide and be the first on Youtube in all world. I see everywhere mostly Colibri guides. Please don't start debates here about what is what, offer solutions without making people without mortgage for their only house next month. We dont need more homeless on streets.
My point was - guys never waste money on this, esp on dozen GPUs which in result will produce trash (biggest cold shower for everyone when they realize normal quality is Q8 but they can run only maybe Q3-Q5 trash for thousands $, a production into trashbin but with good speed).
The obsession here exactly similar like with gambling or "gold digger fever" syndrome.

does sci-fi level consumer SIMD processors count? if so then yes, otherwise no. current off the shelf cpus do not have the capabilities to compute a causal model at this size at a satiable rate.

this is incorrect, but certain invested individuals and groups are pushing the idea hard. This completely ignores colibri style developments.. seemingly purposefully

life is short and time is precious. please don't do it to yourself.

please show some foresight. The tech is suitable for a range of decentralised and offline, low power situations. Or in the event of global catastrophe.

You might get close to 5 tok/s on a 4 epyc cpu cluster with 32 memory channels with 32 ddr5 dims of course, which would still be atleast 10,000$

or use an old media computer with 32gb ram and an ssd to get 0.05tok/s, and get several responses a day for basically no power use. Stop setting boundaries of what's feasible based on a narrow scope of personal interests (which coincidentally help the mega corps projected profits)

What does colibri do that llama.cpp doesn't?

You can read on author's Github, i'm not the author.
Mainly run just from SSD utilyzing its storage size (almost no Ram, no GPU-someone on fancy CPU's got even 2-3 token/s that was unique, i wasnt able).
P.S.: if anyone want to use selfcheck method "are you sure", you need to make it several times, "Oops i did it again" is very constant problem in models.

can you tell me more about the fancy cpu results? I haven't seen that

Colibri engine is your only choice. They plan to release it after Inkling.
https://github.com/JustVugg/colibri/issues/658

It works, GLM5.2 i wont recommend, 0.3 tokens/sec no matter hardware, engine is quite wet and early to show performance but it works on any potato. Quality of model in Colibri is very low by tests from GLM, but there's no way to compare, very few models supported and GLM5.2 is bad release.
Please dont waste your $ savings on hardware-this models wont return back even a penny, hardware will be obsolete, models too, no savings. People dont know how to use them right, these models will not bring you anything useful. Secret know-how available but guarded, as Alex Karp (Palantir) publicly said again on TV 5th time-they can turn any dumb model into state of the art with their secret sauce.
Giveaway from me: question "are you sure?" works on many models and they openly confess of lying and presenting wrong data deliberately. In fact correcting their work seriously. May not work on Openai's GPT456, but works on GLM models, including the free model 4.7 on their site.

What does colibri do that llama.cpp doesn't?

On a system where Colibri hits 0.1 tok/s, llama.cpp would grind down to roughly 0.001 to 0.003 tok/s (taking 5 to 15 minutes per single token) or, more likely, cause a complete operating system lockup or crash.

Colibri differs from llama.cpp because it intelligently streams only the required fractions of a model, rather than treating disk storage as a dumb, slow extension of system RAM.

Sign up or log in to comment