Instructions to use lgdevlop/Macaron-V1-Tall-l1-agent with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use lgdevlop/Macaron-V1-Tall-l1-agent with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="lgdevlop/Macaron-V1-Tall-l1-agent", filename="macaron-l1-agent-35b-q4_k_m.gguf", )
llm.create_chat_completion( messages = "No input example has been defined for this model task." )
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use lgdevlop/Macaron-V1-Tall-l1-agent with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M # Run inference directly in the terminal: llama cli -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M # Run inference directly in the terminal: llama cli -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Use Docker
docker model run hf.co/lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Ollama:
ollama run hf.co/lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
- Unsloth Studio
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for lgdevlop/Macaron-V1-Tall-l1-agent to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for lgdevlop/Macaron-V1-Tall-l1-agent to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for lgdevlop/Macaron-V1-Tall-l1-agent to start chatting
- Pi
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use lgdevlop/Macaron-V1-Tall-l1-agent with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Docker Model Runner:
docker model run hf.co/lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
- Lemonade
How to use lgdevlop/Macaron-V1-Tall-l1-agent with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull lgdevlop/Macaron-V1-Tall-l1-agent:Q4_K_M
Run and chat with the model
lemonade run user.Macaron-V1-Tall-l1-agent-Q4_K_M
List all available models
lemonade list
Macaron-V1-Tall (GGUF Pre-Merged Specialists)
This repository contains the GGUF-quantized, pre-merged specialist models for Macaron-V1-Tall, optimized for local deployment on consumer hardware.
๐ Acknowledgments and Credits
All foundational research, training, and architecture design belong to the original authors. This repository solely provides a community-driven quantization and deployment format.
- Original Authors: mindlab-research/Macaron-V1-Tall
- Original Repository: mindlab-research/Macaron-V1-Tall
- Base Architecture: Qwen 3.6 (35B Base + 4x 3.7B LoRA Specialists)
๐งฉ The Architectural Pivot (Why this repository exists)
The Macaron-V1-Tall architecture utilizes a Mixture-of-LoRA (MoE) pattern over a Grouped-Query Attention (GQA) base model. Currently, the standard llama.cpp ecosystem encounters tensor reshaping limitations (NotImplementedError) when attempting to dynamically apply these specific LoRA adapters at runtime.
To bypass this mathematical constraint and enable seamless local execution, this repository utilizes a Pre-Merge Strategy. Each LoRA specialist has been physically fused into the 35B base model using CPU-RAM computation (bypassing VRAM bottlenecks) prior to GGUF conversion.
The Result: Four independent, fully fused GGUF models. Instead of dynamically swapping LoRAs in VRAM, developers can leverage operating system Page Caching to swiftly alternate between these ~21GB files in RAM, enabling rapid intent-based routing without out-of-memory (OOM) errors.
๐ฆ Available Specialists (Q4_K_M Quantization)
The model have been quantized to Q4_K_M to balance perplexity and memory footprint, reducing the required storage from ~70GB (FP16) to approximately 21GB per specialist.
macaron-l1-agent-35b-q4_k_m.gguf(Agent Specialist)
Note: The experimental
mtp_num_hidden_layers(Multi-Token Prediction) metadata has been sanitized from the config to ensure strict compatibility with thellama.cpploader.
๐ Deployment & Usage (llama.cpp)
These models are heavily optimized for hybrid VRAM/RAM offloading. If you are running a consumer GPU (e.g., 16GB VRAM) backed by substantial system RAM, you must carefully balance the layers to prevent VRAM saturation.
Example of a Launch Command (llama-server)
The following command demonstrates how to load the model while explicitly offloading the heaviest MoE layers to the CPU, freeing up your VRAM for the KV cache and attention heads.
./llama-server \
-m "macaron-l1-agent-35b-q4_k_m.gguf" \
--host 127.0.0.1 \
--port 9091 \
-ngl 40 \
-ncmoe 10 \
--ctx-size 16384 \
-b 512 \
-fa on \
-ctk q4_0 \
-ctv q4_0 \
-t 16 \
--reasoning-preserve \
--no-context-shift \
-sps 0.0 \
--ctx-checkpoints 0 \
--models-max 1 \
--parallel 1 \
--no-mmap \
--metrics
- Downloads last month
- -
4-bit