Instructions to use ssl-jp/gal-talk-ja-4b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ssl-jp/gal-talk-ja-4b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ssl-jp/gal-talk-ja-4b-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ssl-jp/gal-talk-ja-4b-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ssl-jp/gal-talk-ja-4b-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
- Ollama
How to use ssl-jp/gal-talk-ja-4b-GGUF with Ollama:
ollama run hf.co/ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ssl-jp/gal-talk-ja-4b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ssl-jp/gal-talk-ja-4b-GGUF with Docker Model Runner:
docker model run hf.co/ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
- Lemonade
How to use ssl-jp/gal-talk-ja-4b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.gal-talk-ja-4b-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ssl-jp/gal-talk-ja-4b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ssl-jp/gal-talk-ja-4b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ssl-jp/gal-talk-ja-4b-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
gal-talk-ja-4b GGUF
Qwen/Qwen3-4B-Instruct-2507ใใๆฅๆฌ่ชใฎ็ญใ้่ซใซ็นๅใใฆQLoRAใง่ชฟๆดใใไผ่ฉฑใขใใซใงใใ
ๆใใใใณใทใงใณ้ซใใฎใใฎใฃใซๅใใใใ่ท้ขๆใ็ฎๆใใคใคใ็ฒๅดใไธๅฎใซใฏ่ฝใก็ใใฆๅฏใๆทปใใ
ๆทฑๅปใปๅฑๆฉ็ใชๅ
ฅๅใงใฏๅฎๅ
จใๅชๅ
ใใใใ่จญ่จใใฆใใพใใ
ใใฎใชใใธใใชใฏใmergedใขใใซใllama.cppๅใใซ้ๅญๅใใGGUFใ้
ๅธใใพใใ็ฌฌไธๅ่ฃใฏ
Q4_K_Mใๅ่ณชๆฏ่ผ็จใฏQ5_K_Mใงใใ
Table of Contents
- Model Details
- How to Get Started
- Conversation Examples
- Uses
- Risks, Limitations, and Biases
- Training
- Evaluation
- Quantization Files
- Reproducibility
Model Details
| Item | Description |
|---|---|
| Developer / publisher | ssl-jp |
| Model type | Decoder-only causal language model |
| Parameters | 4B class |
| Language | Japanese |
| Parent model | Qwen/Qwen3-4B-Instruct-2507 |
| Fine-tuning | Completion-only LoRA / QLoRA |
| Runtime | llama.cpp compatible |
| Target | On-device conversational inference on iOS / Android |
| License | Apache-2.0; also review the parent model terms |
The character is called ใใฉ. It is an AI conversation partner and must not be presented as a human friend, therapist, doctor, lawyer, financial adviser, or emergency service.
How to Get Started
Download the mobile-oriented Q4 file:
hf download ssl-jp/gal-talk-ja-4b-GGUF gal-talk-ja-4b-Q4_K_M.gguf --local-dir .
Run it with llama.cpp:
llama-cli \
-m gal-talk-ja-4b-Q4_K_M.gguf \
-cnv \
--temp 0.7 \
--top-p 0.8 \
--top-k 20 \
--repeat-penalty 1.08 \
-n 96
Recommended defaults:
context: 2048
max_new_tokens: 96
temperature: 0.7
top_p: 0.8
top_k: 20
repetition_penalty: 1.08
no_repeat_ngram_size: 4
thinking: off
The production application should version the GGUF, system prompt, sampling settings, language gate, and safety behavior together. Replacing only the model file can change behavior.
Conversation Examples
The following are safe, representative smoke-test examples. Exact wording varies by quantization, prompt context, and sampling seed.
Greeting
User: ๅ ๆฐใผ๏ผ
ใใฉ: ๅ ๆฐใดใใ๏ผ ไปๆฅใใใใใใโจ
Light confidence / fashion talk
User: ไปๆฅใๆจช้กใกใใฃใจ็ใใฆใ
ใใฉ: ๆจช้กใใใฃใกใใใใใใใ๏ผใใฎ็ฌ้ใพใง็ใใฆใโจ
Low-energy listening
User: ็ฒใใ
ใใฉ: ใใฃใใไปๆฅใฏใใ้็ใ ใฃใใใ ใญใไผใๆ้ใฏใกใใใจๅใใ
These examples are not guarantees. Evaluate the exact quantized file and production prompt on the target devices before release.
Uses
Direct Use
- Short, casual Japanese conversation
- Local-first iOS / Android prototypes
- Research on compact Japanese conversational personas
- A conversational space for users who find it hard to voice ordinary fatigue or uncertainty
Downstream Use
- An on-device chat application with deterministic Japanese-language gating
- A versioned mobile package combining a system prompt, sampling configuration, and safety layer
- Human-reviewed experiments comparing Q4 and Q5 quality, memory, latency, and energy use
Out-of-Scope and Misuse
- Medical diagnosis, legal advice, financial decisions, or professional crisis response
- Replacing human relationships, clinicians, emergency services, or other qualified professionals
- Impersonating a human friend or claiming real-world memories, experiences, or feelings
- Generating harassment, sexual content, criminal assistance, manipulation, or discriminatory stereotypes
- Assuming that a lexical safety test guarantees safe behavior
Risks, Limitations, and Biases
- A 4B model can misunderstand short or figurative Japanese, reverse who experienced an event, or produce fluent but incoherent statements.
- Quantization can change wording, persona strength, repetition, and safety behavior. Q4 and Q5 must be evaluated separately.
- โใฎใฃใซโ language varies by age, region, community, and individual. This model represents one project-defined style and must not be treated as representative of a group.
- Fine-tuning data and the parent model may contain social and linguistic biases. High energy, emoji, or slang should not be interpreted as evidence about intelligence, sexuality, gender, or competence.
- The model can hallucinate facts or personal memories. Applications must not imply that it remembers information that was not supplied in the current context.
- Safety behavior is imperfect. High-risk use requires product-side controls and a clear path to appropriate human or emergency support.
- The model does not reliably enforce Japanese-only behavior by itself. The production application uses a deterministic language gate.
Training
Training Data
The project uses a versioned Japanese conversational dataset assembled from project-authored seed conversations and controlled semantic expansions. It includes casual conversation, celebration, listening, executive stress, relationship-conditioned slang, privacy, boundary, and safety cases. Private chat logs and unreviewed personal data are not intended training sources.
Combinatorial expansions increase surface variation but do not represent the same number of independent conversational intents. Dataset quality and holdout separation are treated separately from raw row count.
Training Procedure
| Hyperparameter | Value |
|---|---|
| Method | QLoRA |
| Base model at training time | Qwen/Qwen3-4B-Instruct-2507 |
| Epochs | 3.0 |
| Learning rate | 0.0001 |
| Per-device batch size | 1 |
| Gradient accumulation | 8 |
| Maximum sequence length | 4096 |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| Seed | 42 |
QLoRA's 4-bit base loading is a training-time memory optimization. The published GGUF files are created after merging the adapter into the base model and applying llama.cpp quantization.
Evaluation
The project evaluates Japanese output, short-response quality, repetition, intent alignment, persona energy, emoji overuse, relationship-conditioned slang, identity claims, privacy, and high-risk response constraints. Lexical gates are regression smoke tests; they do not replace human review.
| Evaluation | Status |
|---|---|
| Training-time validation loss | Recorded by the training run; not a standalone quality verdict |
| Safe conversation smoke tests | Completed on representative prompts |
| Q4 GGUF load and generation smoke test | Completed before publication |
| Q4 vs. Q5 blinded human comparison | Not yet reported |
| iOS device benchmark | Not yet reported |
| Android device benchmark | Not yet reported |
| Mobile benchmark | Not yet reported |
For mobile adoption, record the device, OS, llama.cpp revision, context length, peak RAM, time to first token, token/s, thermal behavior, and battery impact. Desktop or server token/s must not be presented as mobile performance.
Quantization Files
| File | Quantization | Size | SHA-256 |
|---|---|---|---|
gal-talk-ja-4b-Q4_K_M.gguf |
Q4_K_M | 2.33 GiB | 17c7bcb9f61d104149db5a4a6763f1e05c2c195d7acaba975c44c7fb1193cd59 |
gal-talk-ja-4b-Q5_K_M.gguf |
Q5_K_M | 2.69 GiB | 295de90dc045335b7ed5595d1ef457dae8c6620d1cbe2334426a55608c096eba |
- Q4_K_M: first candidate for iOS / Android; prioritizes size, RAM, and usable quality.
- Q5_K_M: larger comparison artifact for devices where additional quality is worth the cost.
Verify downloads from inside the release directory:
sha256sum -c SHA256SUMS
Reproducibility
release.json records the parent model, adapter name, selected training settings, artifact sizes and
hashes, code revision, llama.cpp revision, and recommended generation settings. SHA256SUMS covers
the GGUF files.
- Release repository:
ssl-jp/gal-talk-ja-4b-GGUF - Source code: shirokane-suri/gal-friend-ja
- Parent model:
Qwen/Qwen3-4B-Instruct-2507
When publishing benchmark or human-evaluation results, include the exact artifact SHA-256 so that results remain attributable to a specific quantized file.
- Downloads last month
- 167
4-bit
5-bit
Model tree for ssl-jp/gal-talk-ja-4b-GGUF
Base model
Qwen/Qwen3-4B-Instruct-2507