Instructions to use ChapAF/intent-0.1-0.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ChapAF/intent-0.1-0.5b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ChapAF/intent-0.1-0.5b:F16 # Run inference directly in the terminal: llama cli -hf ChapAF/intent-0.1-0.5b:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ChapAF/intent-0.1-0.5b:F16 # Run inference directly in the terminal: llama cli -hf ChapAF/intent-0.1-0.5b:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ChapAF/intent-0.1-0.5b:F16 # Run inference directly in the terminal: ./llama-cli -hf ChapAF/intent-0.1-0.5b:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ChapAF/intent-0.1-0.5b:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ChapAF/intent-0.1-0.5b:F16
Use Docker
docker model run hf.co/ChapAF/intent-0.1-0.5b:F16
- LM Studio
- Jan
- Ollama
How to use ChapAF/intent-0.1-0.5b with Ollama:
ollama run hf.co/ChapAF/intent-0.1-0.5b:F16
- Unsloth Desktop
- Docker Model Runner
How to use ChapAF/intent-0.1-0.5b with Docker Model Runner:
docker model run hf.co/ChapAF/intent-0.1-0.5b:F16
- Lemonade
How to use ChapAF/intent-0.1-0.5b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ChapAF/intent-0.1-0.5b:F16
Run and chat with the model
lemonade run user.intent-0.1-0.5b-F16
List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ChapAF/intent-0.1-0.5b
This repository contains a custom decoder-only model trained by the
first-order-dictionary-learning pipeline. It is an inference artifact for
the project's pretraining.model.GPT implementation, rather than a native
transformers.AutoModel checkpoint.
Files
model.safetensors: model weights only (455933952 parameters).model-f16.gguf: llama.cpp/LM Studio compatible GGUF export.config.json: the exact model and training configuration used to build the model.- tokenizer files: a Rust
LlamaTokenizerFast/fast tokenizer saved with the model. pretraining/model.pyandchat.py: the minimal custom inference code.checkpoint_metadata.json: source and conversion metadata.
Optimizer state, distributed rank state, raw training data, and W&B files are intentionally not included.
The GGUF uses llama.cpp's Arcee ReLU-squared graph. The exact training model
also applies Q/K RMSNorm, which the upstream Arcee graph does not expose, so
the GGUF is a tooling-compatible approximation. Use model.safetensors with
the included chat.py for numerically faithful inference.
Model details
- Tokenizer source:
meta-llama/Llama-2-7b-chat-hf - Vocabulary size:
32000 - Context length:
2048 - Layers / hidden size:
24 / 1152 - Thinking: off by default; explicitly request
<think>...</think>reasoning when needed.
Run locally
From the root of this repository (with PyTorch, transformers, safetensors, flash-attn/liger as available):
PYTHONPATH=. python chat.py \
--checkpoint model.safetensors \
--config config.json \
--mode chat --prompt-format sft \
--think true \
--prompt "Explain what a steering vector is." \
--max-new-tokens 256
The checkpoint uses the serialized SFT prompt format when it was produced by
the SFT run: <|system|>, <|user|>, and <|assistant|> role markers.
- Downloads last month
- 68