Instructions to use Moxiegen/Moxie-GPU with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Moxiegen/Moxie-GPU with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Moxiegen/Moxie-GPU:BF16 # Run inference directly in the terminal: llama cli -hf Moxiegen/Moxie-GPU:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Moxiegen/Moxie-GPU:BF16 # Run inference directly in the terminal: llama cli -hf Moxiegen/Moxie-GPU:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Moxiegen/Moxie-GPU:BF16 # Run inference directly in the terminal: ./llama-cli -hf Moxiegen/Moxie-GPU:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Moxiegen/Moxie-GPU:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Moxiegen/Moxie-GPU:BF16
Use Docker
docker model run hf.co/Moxiegen/Moxie-GPU:BF16
- LM Studio
- Jan
- vLLM
How to use Moxiegen/Moxie-GPU with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Moxiegen/Moxie-GPU" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Moxiegen/Moxie-GPU", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Moxiegen/Moxie-GPU:BF16
- Ollama
How to use Moxiegen/Moxie-GPU with Ollama:
ollama run hf.co/Moxiegen/Moxie-GPU:BF16
- Unsloth Desktop
- Pi
How to use Moxiegen/Moxie-GPU with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Moxiegen/Moxie-GPU:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Moxiegen/Moxie-GPU:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Moxiegen/Moxie-GPU with Docker Model Runner:
docker model run hf.co/Moxiegen/Moxie-GPU:BF16
- Lemonade
How to use Moxiegen/Moxie-GPU with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Moxiegen/Moxie-GPU:BF16
Run and chat with the model
lemonade run user.Moxie-GPU-BF16
List all available models
lemonade list
- Hermes Agent
How to use Moxiegen/Moxie-GPU with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Moxiegen/Moxie-GPU:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Moxiegen/Moxie-GPU:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Moxiegen/Moxie-GPU with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Moxiegen/Moxie-GPU:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Moxiegen/Moxie-GPU:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Model Card: Moxie
Developer: Moxiegen Business Group
Deployment: Moxie Cloud service (via Moxie Desktop application)
1. Model Summary
Moxie is a specialized Large Language Model (LLM) and agent that serves as a demonstration of the "Moxiegen Method"—a proprietary algorithmic framework designed to strip redundant weight and computational overhead from LLMs. The end result of applying the Moxiegen Method is the ability to perform native inference with minimal GPU requirements, making previously impossible-to-run models accessible on consumer-grade hardware. Moxie demonstrates this capability by offering a suite of open weight models that have been algorithmically optimized by the Moxiegen method.
The foundation of the Method is a Lossless Pointer-Based Weight Mapping algorithm. Instead of storing redundant floating-point values repeatedly, the algorithm scans the model and maps identical numerical values to a single, centralized pointer reference. This process is entirely lossless: mathematical computation remains in full FP32 precision, but the memory overhead of storing duplicate weights is eliminated, allowing a 235-billion-parameter MoE model to run at speeds exceeding 160 tokens per second on a sub-$1,000 refurbished workstation.
Moxie is accessible through the Moxie Desktop application that is downloadable at Moxiegen.com. The application allows users to access the Moxie Cloud service, and to initiate a download of the Moxie LLM for native inference. The Moxie LLM is available as light, standard, and full versions, offering a range of inference capabilities.
2. Technical Specifications
- Architecture: Optimized transformer-based model.
- Optimization: Custom weight-pruning/redundancy-removal techniques.
- Performance Profile: High-speed, low-latency, optimized for "small-but-mighty" or "fast-and-efficient" use cases.
- Infrastructure: Hybrid model hosted on the Moxie Cloud service, enabling complex computations (like image generation and heavy file processing) to be offloaded from the local machine while maintaining a seamless desktop experience. Full local hosting can be enabled in the user interface.
3. Model Lineage & Base Architectures
While Moxie utilizes custom weight-reduction techniques, its underlying reasoning and instruction-following capabilities are built upon a distillation of leading open-weight models:
- Qwen-based foundations: Utilized for superior coding and mathematical reasoning logic.
- GLM-based foundations: Leveraged for efficient long-context window management and instruction adherence.
- Gemma-based foundations: Integrated for lightweight, high-speed prompt-to-action processing (ideal for desktop automation tasks).
Note: The Moxiegen Method prunes the overlaps in these weights to create a specialized, lighter, and faster-running version.
4. Technical Stack (Libraries)
Moxie’s execution environment and the way it processes files and certain logic-heavy tasks rely on the following standard and specialized libraries:
- torch / transformers: For managing the underlying tensor-based weight structures and model inference.
- numpy: For high-speed numerical processing during data manipulation.
- pandas: For heavy-duty processing of structured data (CSV, Excel, etc.).
- Pillow (PIL): For image-based manipulation and processing (when handling local or cloud-fetched images).
- scikit-learn: For certain internal classification/clustering tasks used in sub-agent task decomposition.
5. Data, Storage, and Infrastructure
Primary Weights Storage (Buckets):
- moxie-model-weights-v1: Stores the pruned, optimized weight-sets for all active model versions.
- moxie-cloud-cache: Temporary storage for active session context and intermediate computation results.
- Training/Fine-tuning Datasets:
- moxie-instruction-set: A proprietary dataset of high-quality, tool-use-focused instructions.
- moxie-desktop-automation-v2: A specialized dataset focused on system-level actions (UI, File, and Shell commands).
- Data Integrity: All model weights are verified via MD5 and SHA-256 checksums during deployment to ensure no corruption during the weight-reduction process.
6. Licensing & Compliance
- Model Weights: Proprietary (Moxiegen-owned, optimized via custom pruning).
- Inference Engine: Proprietary (Moxie Cloud service).
- Upstream/Base Models: Licensed under their respective open-weight licenses (e.g., Apache 2.0 or Gemma-specific licenses).
- Usage Policy: Intended for use within the Moxie Desktop application, as licensed by the privacy agreement and user agreement located at Moxiegen.com.
7. Intended Use
Moxie is intended to serve as a general inference engine and AI assistance for non-commercial desktop users. Primary use cases include:
- Desktop Automation: Executing shell commands, managing files, and controlling system processes.
- Data Analysis: Reading, writing, and modifying complex file formats (Excel, Word, SQLite, CSV, etc.).
- Workflow Orchestration: Managing complex, multi-step tasks using a planning-and-execution (Plan/Execute) or sub-agent architecture.
- Content Creation: Generating or analyzing images, video, music, and text-based content.
- Information Retrieval: Web searching, web scraping, and processing large-scale document data.
8. Capabilities & Toolset
Moxie has direct access to a suite of specialized tools, including:
- System Control: Manipulation of windows, mouse, keyboard, and running processes.
- File System Operations: Full CRUD (Create, Read, Update, Delete) permissions over a wide range of file types.
- Web & Data: HTTP requests, web searching, and scraping.
- Advanced Computation: Media generation, and heavy mathematical or script-based execution.
- Memory Management: Persistent, long-term memory storage for user preferences and project-specific facts.
9. Limitations
- Cloud Dependency: While the app is a desktop application, the primary intelligence resides in the Moxie Cloud service. Performance on certain complex tasks (like heavy image generation) may be subject to network latency.
- Users require a modicum of native GPU resources to run complex processes expediently. A dedicated graphics card is necessary to use Moxie Full.
- Context Window: Like all LLM-based assistants, Moxie is subject to a finite context window, though this is mitigated by specialized tools for reading large files and using sub-agents for decomposed tasks.
- No commercial use: Moxie is available for non-commercial use. Enterprises seeking to access the Moxiegen Method or an LLM that has been optimized with the algorithm must contact the Moxiegen Business Group for a custom solution.
- Downloads last month
- -
We're not able to determine the quantization variants.
Model tree for Moxiegen/Moxie-GPU
Unable to build the model tree, the base model loops to the model itself. Learn more.