Mesh LLM

Llama-3.2-1B-Instruct-Q4_K_M

Distributed GGUF inference package for Mesh LLM

Website GitHub Discord

GGUF layer package for running Llama-3.2-1B-Instruct-Q4_K_M across a local Mesh LLM cluster.

This package is derived from unsloth/Llama-3.2-1B-Instruct-GGUF and keeps the original GGUF distribution split into per-layer artifacts for distributed inference.

Highlights

Run locally Pool multiple machines OpenAI-compatible Package variant
Private inference on your hardware Split layers across peers Serve /v1/chat/completions locally Q4_K_M layer package

Model Overview

Property Value
Source model unsloth/Llama-3.2-1B-Instruct-GGUF
Model id unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
Family Llama
Parameter scale 1B
Quantization Q4_K_M
Layer count 16
Activation width not recorded
Package size 0 B
Source file Llama-3.2-1B-Instruct-Q4_K_M.gguf
Package repo meshllm/Llama-3.2-1B-Instruct-Q4_K_M-layers
License llama3.2 from unsloth/Llama-3.2-1B-Instruct-GGUF

Recommended Use

  • Local and private inference with Mesh LLM.
  • Multi-machine serving when the full GGUF is too large for one host.
  • OpenAI-compatible chat/completions workflows through Mesh LLM's local API.

For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmark notes, see the source model card: unsloth/Llama-3.2-1B-Instruct-GGUF.

Quickstart

# Run this on each machine that should contribute memory/compute.
mesh-llm serve --model "meshllm/Llama-3.2-1B-Instruct-Q4_K_M-layers" --split
# Check the mesh and discover the OpenAI-compatible model name.
curl -s http://localhost:3131/api/status
curl -s http://localhost:3131/v1/models
# Send an OpenAI-compatible chat request.
curl -s http://localhost:3131/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M",
    "messages": [{"role": "user", "content": "Write a tiny hello-world function in Rust."}],
    "max_tokens": 128
  }'

Package Variant

Property Value
Format gguf
Canonical source ref unsloth/Llama-3.2-1B-Instruct-GGUF@b69aef112e9f895e6f98d7ae0949f72ff09aa401/Llama-3.2-1B-Instruct-Q4_K_M.gguf
Source revision b69aef112e9f895e6f98d7ae0949f72ff09aa401
Source SHA-256 3f5a22426976ab26cfe84dba63c1d08391717abb1af893e10f1b2968d862dcc1
Skippy ABI not recorded
Package manifest SHA-256 109a56e5a4b47c9e8d8cf3374655ae32e6f6ef56062c9db74ec937faedb0f3e0

What Is Included

Artifact Path Contents SHA-256
Manifest model-package.json Package schema, source identity, checksums 109a56e5a4b47c9e8d8cf3374655ae32e6f6ef56062c9db74ec937faedb0f3e0

Validation

Generated by the Mesh LLM HF Jobs splitter from mesh-llm ref 293d61d8fb8be914e5244f754ceb1d67f59f177b. Each artifact is checksummed as it is written, uploaded to this repository, and removed from the job workspace before the next artifact is produced.

skippy-model-package write-package "/hf-cache/Llama-3.2-1B-Instruct-Q4_K_M.gguf" --out-dir "/tmp/meshllm-layer-job-meshllm_Llama-3.2-1B-Instruct-Q4_K_M-layers-12/package"

Links

Downloads last month
-
GGUF
Model size
60.8M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for meshllm/Llama-3.2-1B-Instruct-Q4_K_M-layers

Quantized
(5)
this model