Instructions to use nathaninline/Jean-1.0-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nathaninline/Jean-1.0-27B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nathaninline/Jean-1.0-27B:Q4_K_M # Run inference directly in the terminal: llama cli -hf nathaninline/Jean-1.0-27B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nathaninline/Jean-1.0-27B:Q4_K_M # Run inference directly in the terminal: llama cli -hf nathaninline/Jean-1.0-27B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nathaninline/Jean-1.0-27B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nathaninline/Jean-1.0-27B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nathaninline/Jean-1.0-27B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nathaninline/Jean-1.0-27B:Q4_K_M
Use Docker
docker model run hf.co/nathaninline/Jean-1.0-27B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use nathaninline/Jean-1.0-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nathaninline/Jean-1.0-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nathaninline/Jean-1.0-27B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nathaninline/Jean-1.0-27B:Q4_K_M
- Ollama
How to use nathaninline/Jean-1.0-27B with Ollama:
ollama run hf.co/nathaninline/Jean-1.0-27B:Q4_K_M
- Unsloth Desktop
- Pi
How to use nathaninline/Jean-1.0-27B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nathaninline/Jean-1.0-27B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nathaninline/Jean-1.0-27B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nathaninline/Jean-1.0-27B with Docker Model Runner:
docker model run hf.co/nathaninline/Jean-1.0-27B:Q4_K_M
- Lemonade
How to use nathaninline/Jean-1.0-27B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nathaninline/Jean-1.0-27B:Q4_K_M
Run and chat with the model
lemonade run user.Jean-1.0-27B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nathaninline/Jean-1.0-27B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nathaninline/Jean-1.0-27B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nathaninline/Jean-1.0-27B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nathaninline/Jean-1.0-27B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nathaninline/Jean-1.0-27B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nathaninline/Jean-1.0-27B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Jean-1.0-27B
Qwen 3.8 27B entraîné sur le dataset saidutta69/fable-5-premium, avec une tête MTP (Multi-Token Prediction) greffée depuis Unsloth pour le décodage spéculatif natif.
Quantifications disponibles
| Fichier | Quant | Taille |
|---|---|---|
Jean-1.0-27B-Q5_K_M.gguf |
Q5_K_M | ~19,6 Go |
Jean-1.0-27B-Q4_K_M.gguf |
Q4_K_M | ~16,9 Go |
Jean-1.0-27B-IQ4_XS.gguf |
IQ4_XS | ~15,5 Go |
Le Q5_K_M offre la meilleure qualité ; le Q4_K_M et l'IQ4_XS sont plus légers et laissent davantage de VRAM pour le contexte (IQ4_XS étant le plus compact).
Décodage spéculatif MTP
La tête MTP (couche nextn) active le décodage spéculatif natif dans les moteurs qui le supportent, pour un débit nettement plus élevé sans modèle de brouillon séparé. Gain mesuré : environ 25 -> 40 tok/s sur RTX 5060 Ti 16 Go + RTX 3070 8 Go en split.
Utilisation avec AJEAN
Le meilleur modèle pour AJEAN, l'application qui fait tourner vos modèles en local et les rend accessibles depuis tous vos appareils. Rapide grâce au MTP, à l'aise en français comme en anglais, et taillé pour tenir sur une config deux GPU grand public. Preset de référence :
# NAME=JEAN 1.0 27B
MODEL="/etc/ajean/models/Jean-1.0-27B-Q5_K_M.gguf"
EXTRA_ARGS=--tensor-split 0.700,0.300 --split-mode tensor --flash-attn on --mlock --no-mmap --kv-unified --spec-type draft-mtp
KV_TYPE_K=q4_0
KV_TYPE_V=q4_0
REASONING=on
Config validée sur RTX 5060 Ti 16 Go + RTX 3070 8 Go en split, contexte 32k.
Utilisation (llama.cpp avec support MTP)
llama-server -m Jean-1.0-27B-Q5_K_M.gguf \
--spec-type draft-mtp --spec-draft-n-max 3 \
-c 32768 -ngl 999
Le décodage spéculatif MTP requiert un build de llama.cpp qui gère --spec-type draft-mtp. Sans ce support, le modèle se charge normalement mais la tête MTP est ignorée (aucun gain, aucune perte de qualité).
- Downloads last month
- 135
4-bit
5-bit