Instructions to use Nock-AI/nock-coder-1.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nock-AI/nock-coder-1.5b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nock-AI/nock-coder-1.5b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Nock-AI/nock-coder-1.5b") model = AutoModelForCausalLM.from_pretrained("Nock-AI/nock-coder-1.5b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Nock-AI/nock-coder-1.5b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nock-AI/nock-coder-1.5b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nock-AI/nock-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nock-AI/nock-coder-1.5b
- SGLang
How to use Nock-AI/nock-coder-1.5b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nock-AI/nock-coder-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nock-AI/nock-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nock-AI/nock-coder-1.5b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nock-AI/nock-coder-1.5b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Nock-AI/nock-coder-1.5b with Docker Model Runner:
docker model run hf.co/Nock-AI/nock-coder-1.5b
nock-coder-1.5b
Nock Coder is an open coding model from NockAI. It is tuned to answer with the code first and keep any explanation to one short sentence. The 4 bit build is a 986 MB file that runs on an ordinary laptop with Ollama or llama.cpp, on Windows, macOS and Linux, with no GPU needed.
This repo holds the full model in safetensors. For the 1 GB file, see Nock-AI/nock-coder-1.5b-GGUF.
| Base model | Qwen/Qwen2.5-Coder-1.5B-Instruct (Apache 2.0) |
| Parameters | 1.54B |
| Method | LoRA fine tune, merged into the base weights |
| Context used in training | 1,024 tokens |
| License | Apache 2.0 |
Run it
With Ollama (the 1 GB build):
ollama run hf.co/Nock-AI/nock-coder-1.5b-GGUF
With transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Nock-AI/nock-coder-1.5b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Reverse a string in Python."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=300, do_sample=False)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The default system prompt in the chat template is already set to Nock Coder's. You don't need to pass one.
Results
50 short coding prompts across Python, JavaScript, TypeScript, SQL, Go, Rust, Java, C#, shell, Git, CSS and Docker. Greedy decoding, no system prompt passed.
| Base model | Nock Coder v1 | Nock Coder v2 | |
|---|---|---|---|
| Words per answer | 184 | 48 | 54 |
| Words outside the code | 133 | 17 | 16 |
| Answers that start with code | 1 of 50 | 27 of 50 | 50 of 50 |
| Answers with no code at all | 0 of 50 | 12 of 50 | 0 of 50 |
Base model and v1 were measured on the 4 bit GGUF with Ollama. v2 was measured in full precision on an earlier training run of the same recipe (same data and settings), with the same 50 prompts. This published build will be re-measured on the 4 bit GGUF with the v1 harness, and those numbers will replace these. Every raw answer is published at https://nockai.dev/evaluation/
HumanEval and Solidity compile results are not published yet, so we make no claim about code correctness.
What changed in v2
In v1, 12 of 50 answers had no code. All 12 were short one line requests, such as "Reverse a string in Python.", and two answers said they came from Alibaba Cloud, which showed the runtime was not passing our system prompt. v2:
- adds short one line requests with code answers,
- trains half the examples with the base model's default system prompt, so the behavior is in the weights and does not depend on the runtime,
- removes training rows that overlap the 50 evaluation prompts,
- updates the identity answers ("I'm Nock Coder, an open coding model from NockAI").
Training
- LoRA rank 16, alpha 32, dropout 0.05, on all seven projection matrices (q, k, v, o, gate, up, down) in all 28 blocks: 18.5M trainable parameters, 1.2% of the model.
- 2 epochs, learning rate 2e-4, cosine schedule, effective batch 16, on one free Colab T4.
- Data:
- 2,500 general coding examples from ise-uiuc/Magicoder-OSS-Instruct-75K (MIT)
- 1,500 short requests from ise-uiuc/Magicoder-Evol-Instruct-110K (Apache 2.0)
- 500 Solidity examples from AlfredPros/smart-contracts-instructions (MIT)
- 72 identity examples
- Every answer was cut to its first code block plus at most one sentence. The datasets contain text generated by other language models.
Limitations
- This is a 1.5B model. It makes mistakes, and its code can be wrong or insecure. Review it before you run it.
- Smart contract output is not audited. Have contracts reviewed by a qualified person before you deploy them.
- It is tuned for short answers, so it is a poor fit for long explanations or design discussions.
- NockAI is an independent project and is not affiliated with Robinhood, Alibaba or Nockchain.
Links
- Website and results: https://nockai.dev
- X: https://x.com/useNockAI
- Downloads last month
- 226
Model tree for Nock-AI/nock-coder-1.5b
Base model
Qwen/Qwen2.5-1.5B