Instructions to use ryujis/LLM-jp-4-33B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ryujis/LLM-jp-4-33B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use ryujis/LLM-jp-4-33B-GGUF with Ollama:
ollama run hf.co/ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
- Unsloth Studio
How to use ryujis/LLM-jp-4-33B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ryujis/LLM-jp-4-33B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ryujis/LLM-jp-4-33B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ryujis/LLM-jp-4-33B-GGUF to start chatting
- Pi
How to use ryujis/LLM-jp-4-33B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ryujis/LLM-jp-4-33B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ryujis/LLM-jp-4-33B-GGUF with Docker Model Runner:
docker model run hf.co/ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
- Lemonade
How to use ryujis/LLM-jp-4-33B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LLM-jp-4-33B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ryujis/LLM-jp-4-33B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ryujis/LLM-jp-4-33B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ryujis/LLM-jp-4-33B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
βΈ»
license: apache-2.0 base_model:
- llm-jp/llm-jp-4-33b-base library_name: llama.cpp tags:
- gguf
- quantized
- japanese
- llm-jp
- llama.cpp
- ollama language:
- ja
- en
βΈ»
LLM-jp-4-33B-Base GGUF β Q4_K_M
This repository provides an unofficial GGUF Q4_K_M quantization of:
llm-jp/llm-jp-4-33b-base
The original model was developed by the Research and Development Center for Large Language Models at the National Institute of Informatics (NII), Japan / LLM-jp.
This quantized model is not an official quantization released by LLM-jp.
Model
- Base model: llm-jp/llm-jp-4-33b-base
- Model family: LLM-jp-4
- Parameters: approximately 33.2B
- Architecture: Dense Transformer / LlamaForCausalLM
- Context length: 65,536 tokens
- Format: GGUF
- Quantization: Q4_K_M
- Original precision: BF16
- Languages: Japanese / English
- License: Apache License 2.0
File
llm-jp-4-33b-base-Q4_K_M.gguf
About this model
LLM-jp-4-33B Base is a pretrained language model developed by LLM-jp.
The Base model has undergone pre-training and mid-training, but is not a post-trained instruction-following model.
Users looking primarily for conversational or instruction-following behavior should also consider the post-trained LLM-jp-4 models released by the original developers.
Ollama
The model can be run directly from Hugging Face with recent versions of Ollama:
ollama run hf.co/ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
llama.cpp
With a recent llama.cpp installation:
llama cli -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
To start an OpenAI-compatible server:
llama serve -hf ryujis/LLM-jp-4-33B-GGUF:Q4_K_M
Download
Using the Hugging Face CLI:
hf download ryujis/LLM-jp-4-33B-GGUF
llm-jp-4-33b-base-Q4_K_M.gguf
Quantization
This repository contains a Q4_K_M GGUF conversion intended to reduce memory requirements while retaining practical model quality.
This is a community-created quantization and has not been produced or endorsed by the original LLM-jp developers.
Original model
Original repository:
llm-jp/llm-jp-4-33b-base
Please refer to the original model card for architecture details, training data information, evaluation results, risks, limitations, and citation information.
License
The original llm-jp-4-33b-base model is distributed under the Apache License 2.0.
This GGUF quantization follows the licensing terms of the original model.
Disclaimer
This is an unofficial community quantization.
The quantization process may alter model behavior or output quality compared with the original BF16 model. Users should independently evaluate outputs for their intended use.
- Downloads last month
- 103
4-bit