Nanbeige4.2-3B GGUF Quantizations

GGUF quantized versions of Nanbeige4.2-3B.

Original Model Card

See the original model card for details on how to use with llama.cpp and ollama (you'll need to build from their fork cause their changes didn't land in llama.cpp main yet).

Why Ollama is Bad 👎 (needs to be put down somewhere)
  1. @ngxson - https://github.com/ollama/ollama/issues/11714#issuecomment-3174632621

    "Some of llama.cpp maintainers even have to work during their vacations just to have someone else copy their work without giving any credits."

  2. @mudler - https://github.com/ollama/ollama/issues/11714#issuecomment-3175288625

    "it would have been much better if all projects that depend on @ggerganov's and the ggml team work would have upstreamed the contributions directly so anyone in the ecosystem could benefit, and avoid vendor lock-in and the duplicated efforts everywhere... consuming llama.cpp and reporting issues and upstream any change directly there. It is quite frustrating to see that the Open source scene is really getting derailed lately by this kind of bad attitude."

  3. @pwilkin - https://github.com/ollama/ollama/issues/11714#issuecomment-3175999505

    "If you build upon a technology, in the OSS world it's a good habit to actually contribute back to the technology you use if you build something new... But Ollama has, again and again, done the opposite of that - made hacky solutions of their own on top of existing llama.cpp / ggml code instead of contributing to the baseline, then taken the fixes that the ggml team has done as 'new features' or 'bugfixes' of their own platform."

  4. @Teravus - https://github.com/ollama/ollama/issues/11714#issuecomment-3176339445

    "Ollama, for sure, needs to provide something to the user that says that they're using code from llama.cpp. Usually, this is in an about box. I don't even see an about box. Therefore, doesn't look like they're complying with that... The only mention that I see of llama.cpp isn't really a 'we use this software' reference. It's just: Supported backends llama.cpp project founded by Georgi Gerganov. This doesn't seem like enough. Skirting the issue, by treating llama.cpp like a back-end."

  5. @Ggerganov - https://github.com/ggml-org/llama.cpp/pull/19324#issuecomment-3847213274

    "it's quite funny watching the ollama bros copy-pasting our bugs into their "new engine" 🤣. Let's see how long it will take them to realize."

  6. @Ggerganov - https://github.com/ollama/ollama/issues/11714#issuecomment-3172893576

    Before the model was released, the ollama devs decided to fork the ggml inference engine in order to implement gpt-oss support (#11672). In the process, they did not coordinate the changes with the upstream maintainers of ggml. As a result, the ollama implementation is not only incompatible with the vast majority of gpt-oss GGUFs that everyone else uses, but is also significantly slower and unoptimized. On the bright side, they were able to announce day-1 support for gpt-oss and get featured in the major announcements on the release day.

    Now after the model has been released, the blogs and marketing posts have circled the internet and the dust has settled, it's time for ollama to throw out their ggml fork and copy the upstream implementation (#11823). For a few days, you will struggle and wonder why none of the GGUFs work, wasting your time to figure out what is going on, without any help or even with some wrong information. But none of this matters, because soon the upstream version of ggml will be merged and ollama will once again be fast and compatible."

But hopefully here is an alternative 🤗 https://github.com/mostlygeek/llama-swap

with a happy user example: https://huggingface.co/unsloth/gpt-oss-20b-GGUF/discussions/17#68aa6eb05372ae5a8eeac9ed

Citation

@misc{nanbeige2026,
  title        = {Nanbeige4.2-3B},
  author       = {Nanbeige Team},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Nanbeige/Nanbeige4.2-3B}}
}
Downloads last month
4,775
GGUF
Model size
4B params
Architecture
nanbeige
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for owao/Nanbeige4.2-3B-GGUF

Quantized
(21)
this model