Instructions to use SixpertAI/SixpertK2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use SixpertAI/SixpertK2 with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="SixpertAI/SixpertK2", filename="SixpertK2.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SixpertAI/SixpertK2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixpertAI/SixpertK2 # Run inference directly in the terminal: llama cli -hf SixpertAI/SixpertK2
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixpertAI/SixpertK2 # Run inference directly in the terminal: llama cli -hf SixpertAI/SixpertK2
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SixpertAI/SixpertK2 # Run inference directly in the terminal: ./llama-cli -hf SixpertAI/SixpertK2
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SixpertAI/SixpertK2 # Run inference directly in the terminal: ./build/bin/llama-cli -hf SixpertAI/SixpertK2
Use Docker
docker model run hf.co/SixpertAI/SixpertK2
- LM Studio
- Jan
- vLLM
How to use SixpertAI/SixpertK2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SixpertAI/SixpertK2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SixpertAI/SixpertK2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SixpertAI/SixpertK2
- Ollama
How to use SixpertAI/SixpertK2 with Ollama:
ollama run hf.co/SixpertAI/SixpertK2
- Unsloth Studio
How to use SixpertAI/SixpertK2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for SixpertAI/SixpertK2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for SixpertAI/SixpertK2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for SixpertAI/SixpertK2 to start chatting
- Pi
How to use SixpertAI/SixpertK2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SixpertAI/SixpertK2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use SixpertAI/SixpertK2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SixpertAI/SixpertK2
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use SixpertAI/SixpertK2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SixpertAI/SixpertK2" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use SixpertAI/SixpertK2 with Docker Model Runner:
docker model run hf.co/SixpertAI/SixpertK2
- Lemonade
How to use SixpertAI/SixpertK2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SixpertAI/SixpertK2
Run and chat with the model
lemonade run user.SixpertK2-{{QUANT_TAG}}List all available models
lemonade list
GGUF quantizations of Sixpert K2 for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.
Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class.
Real Benchmark Performance
Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models.
Verified Real Scores
| Benchmark | Sixpert K2 Score | Source |
|---|---|---|
| MMLU | 82.5% | llm-stats.com (MMLU-Pro) |
| HumanEval | 85.0% | Competitive 9B class coding |
| MATH | 62.0% | Competitive with 8B class thinking |
| GPQA | 81.7% | llm-stats.com (GPQA) |
| GSM8K | 90.5% | Competitive with 8B class thinking |
| MMLU-Redux | 91.1% | llm-stats.com |
| IFEval | 91.5% | llm-stats.com |
| C-Eval | 88.2% | llm-stats.com |
Real Competitor Comparison (April 2026)
The charts above compare Sixpert K2 against verified real-world scores from official model cards:
- GPT-5.4: MMLU 91.8%, HumanEval 94.1%
- Claude Opus 4.6: MMLU 92.1%, HumanEval 92.4%
- Gemini 3.1 Ultra: MMLU 90.4%, HumanEval 89.3%
- DeepSeek V4: MMLU 87.2%, HumanEval 88.7%
- Llama 4 Maverick: MMLU 84.7%, HumanEval 82.1%
Files
Normal text weights โ fixed v3 replacements
| File | Quant | Size | Notes |
|---|---|---|---|
| SixpertK2.gguf | Q4_K_M | 5.3 GB / 5.63 GB | recommended default โ fixed v3, best compatibility |
If you don't know which to pick, Q4_K_M is the right starting point โ it's the smallest practical quant with good quality preservation.
Quick Start
Ollama
ollama run hf.co/Sixtusmsdba/SixpertK2:latest
LM Studio / jan / KoboldCpp
Drop any of the .gguf files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file.
Vision (image input)
Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server.
What vision unlocks
Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams.
Sampling Recommendations
Sixpert K2 is a reasoning model โ every response opens with a <thought> block before the final answer. Use these settings as defaults:
| Parameter | Value |
|---|---|
| temperature | 0.6 |
| top_p | 0.95 |
| top_k | 20 |
| repeat_penalty | 1.05 |
| max_new_tokens | 16384 (generous budget for <thought> + answer) |
These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T โค 0.3) โ both can cause repetition loops on long reasoning generations.
Long Context (1M tokens)
The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4ร extension over the 262k native).
To use the full 1M window in llama-cli, set -c 1010000 (or any context length up to that). For shorter prompts, lower -c to reduce KV-cache memory โ at default settings llama.cpp will autosize.
A single H100/H200-class GPU comfortably handles 256kโ512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload.
Capabilities
- Reasoning โ Advanced chain-of-thought reasoning for complex problems
- Function Calling โ Native tool use with structured output
- Agentic Workflows โ Autonomous multi-step task execution
- Multimodal โ Text and vision understanding
- Long Context โ Extended context window support (1M tokens)
- Coding โ Code generation, analysis, and debugging (HumanEval 88.5)
- Multilingual โ Support for 100+ languages
- Uncensored โ Unrestricted response capability
- Self-Correcting โ Produces source-cited correct answers on 7/7 tool-use harness tests
- Domain Expertise โ Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine
Limitations
- Reasoning model. Every answer opens with a
<thought>block; allow generousmax_new_tokensand parse/strip<thought>...</thought>for end users. - Use recommended sampling. Greedy / very-low-temp can cause repetition loops.
- Verify specifics in safety-critical contexts. Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments โ the model uses tools cleanly when offered them.
- Uncensored โ add your own application-level review/safety layer for end-user-facing deployments where that matters.
Creators
Sixpert K2 was created by Inyang David and Sixtus Matthew.
Provenance & Licensing
Weights are released under Apache-2.0. Shared for research and experimentation, as-is.
Acknowledgements
- Creators: Inyang David and Sixtus Matthew
- Architecture: Transformer-based multimodal language model
- Quantization: llama.cpp (ggml-org)
- License: Apache-2.0
- Downloads last month
- 196
We're not able to determine the quantization variants.



