Mesh LLM

Qwen3VL-8B-Instruct-Q4_K_M

Distributed GGUF inference package for Mesh LLM

Website GitHub Discord

GGUF layer package for running Qwen3VL-8B-Instruct-Q4_K_M across a local Mesh LLM cluster.

This package is derived from Qwen/Qwen3-VL-8B-Instruct-GGUF and keeps the original GGUF distribution split into per-layer artifacts for distributed inference.

Highlights

Run locally Pool multiple machines OpenAI-compatible Package variant
Private inference on your hardware Split layers across peers Serve /v1/chat/completions locally Q4_K_M layer package

Model Overview

Property Value
Source model Qwen/Qwen3-VL-8B-Instruct-GGUF
Model id Qwen/Qwen3-VL-8B-Instruct-GGUF:Q4_K_M
Family Qwen3
Parameter scale 8B
Quantization Q4_K_M
Layer count 36
Activation width 4096
Package size 6.7 GB
Source file Qwen3VL-8B-Instruct-Q4_K_M.gguf
Package repo namepd/Qwen3-VL-8B-Instruct-Q4_K_M-layers
License apache-2.0 from Qwen/Qwen3-VL-8B-Instruct-GGUF

Recommended Use

  • Local and private inference with Mesh LLM.
  • Multi-machine serving when the full GGUF is too large for one host.
  • OpenAI-compatible chat/completions workflows through Mesh LLM's local API.

For upstream architecture details, chat template guidance, sampling recommendations, license terms, and benchmark notes, see the source model card: Qwen/Qwen3-VL-8B-Instruct-GGUF.

Quickstart

# Run this on each machine that should contribute memory/compute.
mesh-llm serve --model "namepd/Qwen3-VL-8B-Instruct-Q4_K_M-layers" --split
# Check the mesh and discover the OpenAI-compatible model name.
curl -s http://localhost:3131/api/status
curl -s http://localhost:3131/v1/models
# Send an OpenAI-compatible chat request.
curl -s http://localhost:3131/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-VL-8B-Instruct-GGUF:Q4_K_M",
    "messages": [{"role": "user", "content": "Write a tiny hello-world function in Rust."}],
    "max_tokens": 128
  }'

Package Variant

Property Value
Format layer-package
Canonical source ref Qwen/Qwen3-VL-8B-Instruct-GGUF@f982a07559d4a2f6c8744d840bf6fccab30eea96/Qwen3VL-8B-Instruct-Q4_K_M.gguf
Source revision f982a07559d4a2f6c8744d840bf6fccab30eea96
Source SHA-256 67d1659bfe71b89d50b45a4ad1a9e5b997e5bb16ce5da66a6a6167abd569e9e2
Skippy ABI 0.1.38
Package manifest SHA-256 dd6e6a120672480455a98eab01b05ef2994db9f68f53df050f3f9fb863dd011c

What Is Included

Artifact Path Contents SHA-256
Manifest model-package.json Package schema, source identity, checksums dd6e6a120672480455a98eab01b05ef2994db9f68f53df050f3f9fb863dd011c
Metadata shared/metadata.gguf 0 tensors, 5.7 MB 9feef8d599a3ede981d6254ed23f45191439f4c3e9f8fda763d9827bcac76b27
Embeddings shared/embeddings.gguf 1 tensors, 339.5 MB df24139105af8289aa14a281dfdc6758c9fbab38b3e6d21749ae39d5e2a42c93
Output head shared/output.gguf 2 tensors, 492.5 MB a8fcd68146baac9ca49efea4d89d6eb59dbf21343988f4fa80402ee42a0af974
Transformer layers layers/layer-*.gguf 36 layer artifacts, 396 tensors, 4.1 GB see model-package.json
Projector projectors/mmproj-Qwen3VL-8B-Instruct-F16.gguf mmproj projector, 1.1 GB ca524100ebf825c9a870db1c580d03879e0da0ab2541697e2458e64891cf9d38
Projector projectors/mmproj-Qwen3VL-8B-Instruct-Q8_0.gguf mmproj projector, 717.4 MB c6ba85508d82f42590e6eb77d5340369ab6fecf107a7561d809523d8aa5f3bfd

Validation

Generated by the Mesh LLM HF Jobs splitter from mesh-llm ref main. Each artifact is checksummed as it is written, uploaded to this repository, and removed from the job workspace before the next artifact is produced.

skippy-model-package write-package "/source/Qwen3VL-8B-Instruct-Q4_K_M.gguf" --out-dir "/tmp/meshllm-layer-job-namepd_Qwen3-VL-8B-Instruct-Q4_K_M-layers-232/package"

Links

Downloads last month
462
GGUF
Model size
0.2B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for namepd/Qwen3-VL-8B-Instruct-Q4_K_M-layers

Quantized
(1)
this model