Instructions to use prathamkode/particle-1.6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prathamkode/particle-1.6 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="prathamkode/particle-1.6") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("prathamkode/particle-1.6") model = AutoModelForCausalLM.from_pretrained("prathamkode/particle-1.6", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prathamkode/particle-1.6 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prathamkode/particle-1.6" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prathamkode/particle-1.6", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prathamkode/particle-1.6
- SGLang
How to use prathamkode/particle-1.6 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prathamkode/particle-1.6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prathamkode/particle-1.6", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prathamkode/particle-1.6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prathamkode/particle-1.6", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use prathamkode/particle-1.6 with Docker Model Runner:
docker model run hf.co/prathamkode/particle-1.6
Particle 1.6
Particle 1.6 is a compact (~100M) chat model trained from scratch. It uses the same architecture and pretrained base as Particle 1.0.
This release applies a second supervised fine-tune on an internal instruction dataset. That pass did not improve the model as much as expected. Everyday chat still works; factual reliability and consistency remain below what we wanted for a general-purpose assistant.
Weights are released under MIT. Training data is not included.
Quick start
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "prathamkode/particle-1.6"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
messages = [{"role": "user", "content": "hello"}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=False))
Model
| Architecture | Llama-style decoder (RoPE, SwiGLU, RMSNorm) |
| Parameters | 109.5M |
| Layers / hidden / heads | 12 / 768 / 12 |
| Context | 2048 tokens |
| Tokenizer | Custom byte-level BPE, 32k vocabulary |
| Precision | bfloat16 |
| License | MIT |
The model is trained from random initialization. It is not a fine-tune of Llama, SmolLM, or any other public checkpoint.
Training
- Pretrain โ ~2B tokens of public educational web text (same base as Particle 1.0).
- Supervised fine-tune โ an internal instruction mix intended to improve short, helpful replies.
The SFT mix is not published. It did not meet the quality bar we set for this release. Particle 1.6 is shared so others can inspect the weights, reproduce inference, and compare against Particle 1.0.
Intended use
Research, evaluation, and small demos. Suitable for studying from-scratch training at ~100M scale.
Not intended as a production assistant, a source of facts, or a coding model.
Limitations
- Small capacity: weak on reasoning, long context, and tools
- Can hallucinate or contradict itself
- English-centric
- No preference tuning or safety alignment beyond the SFT mix
- The additional SFT pass did not deliver the expected lift over 1.0
Citation
If you use these weights, please cite Particle and the public pretraining corpus used for the 1.0 base.
- Downloads last month
- 206