Why such hype around Spark X2.5 4B?

#9
by the8leggedfriend - opened

Genuine question; I've evaluated various models in the 3B-4B range for their "intelligence" and Spark turned out not-so-well compared to to these: https://huggingface.co/collections/the8leggedfriend/awesome-models-fitting-8gb-apple-m2-silicon-gguf-mlx
I must admit I didn't do 100% rigorous benchmarks, just ran them with their recommended parameters and same quantization within coding agents:

  • gave them little code challenges in fairly non-mainstream programming languages (Common Lisp and Elixir);
  • urged them to use tools (Live REPLs) provided by the agents;
  • emphasized the reasoning process and (algorithmic) problem solving capabilities over familiarity with the programming languages;

Did I miss something? Not the right use case for Spark?
Spark's memory use / speed was pretty good though, stark contrast to Nanbeige for instance.

SparkLLM org

Thanks for the candid feedback.
Could you share a representative Lisp or Elixir task where Spark struggled, along with the quantization, inference backend, and whether thinking mode was enabled? A tool-call trace would also help us reproduce the issue and understand where we can improve.

Sign up or log in to comment