North-Mini-Code-1.0-GGUF

GGUF conversions and quantizations of CohereLabs/North-Mini-Code-1.0 for use with:

  • llama.cpp
  • LM Studio
  • Ollama
  • Jan
  • KoboldCpp
  • Text Generation WebUI
  • Open WebUI
  • Other GGUF-compatible runtimes

About the Model

North-Mini-Code-1.0 is a code-focused Mixture-of-Experts (MoE) model released by CohereLabs.

This repository provides ready-to-use GGUF conversions for local inference across a range of hardware configurations.


Available Files

Full Precision

  • North-Mini-Code-1.0-F16.gguf

Quantized Versions

  • North-Mini-Code-1.0-Q4_K_M.gguf
  • North-Mini-Code-1.0-Q5_K_M.gguf
  • North-Mini-Code-1.0-Q6_K.gguf
  • North-Mini-Code-1.0-Q8_0.gguf

Recommended Quantization

For most users:

North-Mini-Code-1.0-Q4_K_M.gguf

It offers the best balance of:

  • Quality
  • Memory usage
  • Inference speed

If you have more available RAM/VRAM, consider:

North-Mini-Code-1.0-Q5_K_M.gguf

or

North-Mini-Code-1.0-Q6_K.gguf

for slightly higher output quality.


Approximate File Sizes

F16      ~60+ GB
Q4_K_M   ~20 GB
Q5_K_M   ~23 GB
Q6_K     ~27 GB
Q8_0     ~34 GB

Actual sizes may vary slightly depending on conversion tooling versions.


Usage

llama.cpp

Prompt mode:

./llama-cli \
  -m North-Mini-Code-1.0-Q4_K_M.gguf \
  -p "Write a Python function that reverses a linked list."

Chat mode:

./llama-cli \
  -m North-Mini-Code-1.0-Q4_K_M.gguf \
  -cnv

LM Studio

  1. Download your preferred GGUF file.
  2. Open LM Studio.
  3. Import the model.
  4. Start chatting.

Ollama

Create a Modelfile:

FROM North-Mini-Code-1.0-Q4_K_M.gguf

Create the model:

ollama create north-mini-code -f Modelfile

Run it:

ollama run north-mini-code

Hardware Recommendations

Q4_K_M

Recommended minimum:

24 GB RAM

Q5_K_M

Recommended minimum:

32 GB RAM

Q6_K

Recommended minimum:

32-40 GB RAM

Q8_0

Recommended minimum:

48+ GB RAM

F16

Recommended minimum:

80+ GB RAM

Prompting Tips

This model is optimized for programming-related tasks.

Example prompts:

Implement a fast Rust HTTP server.
Explain this C++ compiler error.
Write comprehensive unit tests for the following Python code.
Convert this JavaScript function to TypeScript.
Optimize this SQL query.

Base Model

Base model:

CohereLabs/North-Mini-Code-1.0

All training, architecture, benchmarks, licensing terms, and usage restrictions belong to the original model authors.

Please refer to the original repository for official documentation and licensing information.


Conversion Details

Converted using:

llama.cpp

Generated quantizations:

F16
Q4_K_M
Q5_K_M
Q6_K
Q8_0

A tokenizer compatibility workaround was applied during conversion to support current GGUF conversion tooling.


Credits

  • Base Model: CohereLabs
  • GGUF Conversion & Quantization: NANI-Nithin
  • Tooling: llama.cpp

Repository

👉 https://huggingface.co/NANI-Nithin/north-mini-code-gguf

Downloads last month
101
GGUF
Model size
30B params
Architecture
cohere2moe
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NANI-Nithin/north-mini-code-gguf

Quantized
(37)
this model