Model Card for aquaduck/Llama-3.2-3B-Instruct-GGUF

Pinned F16 GGUF of Llama-3.2-3B-Instruct (meta-llama/llama-3.2-3b-instruct), plus midpoint layer shards for staged / multi-node loading (Aquaduck Arc layer-package-v1). The shard files are not a new quantization. They are contiguous midpoint packages cut from the full F16 GGUF in this repo.

Model lineage

meta-llama/llama-3.2-3b-instruct โ””โ”€โ”€ quantized โ†’ unsloth/Llama-3.2-3B-Instruct-GGUF (F16) โ””โ”€โ”€ full GGUF + midpoint shards โ†’ aquaduck/Llama-3.2-3B-Instruct-GGUF (this repo)

Model Details

Catalog id meta-llama/llama-3.2-3b-instruct
Quantization F16
Parameters 3.2B
Native context unknown tokens
License other
Base model meta-llama/llama-3.2-3b-instruct
Ingest GGUF unsloth/Llama-3.2-3B-Instruct-GGUF

Model Description

  • Hosted by: Aquaduck (hosting and layer packaging only; base model by Meta; GGUF quant by Unsloth / llama.cpp ecosystem)
  • Shared by: Aquaduck AI
  • Model type: Causal language model (Llama-3.2-3B-Instruct), GGUF F16
  • Language(s): Multilingual (same as base)
  • License: other (inherits from meta-llama/llama-3.2-3b-instruct)
  • Finetuned from model: N/A โ€” not a fine-tune
  • Derived from: unsloth/Llama-3.2-3B-Instruct-GGUF โ† meta-llama/llama-3.2-3b-instruct

Model Sources

Files

File Role Approx. size
Llama-3.2-3B-Instruct-Q4_K_M.gguf Full-model GGUF (F16) ~2.02 GB
Llama-3.2-3B-Instruct-Q4_K_M-layers-0-14.gguf Split shard (layers 0โ€“13) ~1.17 GB
Llama-3.2-3B-Instruct-Q4_K_M-layers-14-28.gguf Split shard (layers 14โ€“27) ~1.18 GB
  • Total layers: 28
  • Valid split boundaries: 14 Filenames use exclusive end indices (layers-{start}-{endExclusive}).

Uses

Direct Use

  • Full Llama-3.2-3B-Instruct-Q4_K_M.gguf: standard single-file F16 GGUF (llama.cpp-compatible). Use this for single-node / local runs.
  • *-layers-*.gguf: Aquaduck / Arc staged loading only. These are not drop-in complete models for stock llama.cpp. Use the base modelโ€™s chat template (including thinking / instruct modes as documented on the base model card); other formats will not work correctly.

Out-of-Scope Use

  • Expecting any one shard to run as a complete model
  • Treating this repo as a new training run or re-quant
  • Uses prohibited by the other license or the base modelโ€™s model card guidance

Bias, Risks, and Limitations

Same capabilities, biases, and risks as meta-llama/llama-3.2-3b-instruct. F16 quantization can degrade quality vs. the original higher-precision releases. Layer sharding does not change weights beyond packaging.

Recommendations

Follow the base modelโ€™s docs for chat template, thinking vs instruct modes, and sampling. Prefer Llama-3.2-3B-Instruct-Q4_K_M.gguf in this repo when you do not need staged loading.

How to Get Started

These files are meant to be loaded automatically by the Aquaduck desktop app.

  1. Download the Aquaduck desktop app and sign in.
  2. Devices connected to the internet will receive a model assignment from the model catalog (meta-llama/llama-3.2-3b-instruct).
  3. Download the model from the Home view. The app will:
    • download only the assigned file from this repo (full Llama-3.2-3B-Instruct-Q4_K_M.gguf or one midpoint half)
    • keep that stage ready for serving You do not need to pick files by hand, but you may for local serving. Assignment and download are driven by model catalog metadata. The full Llama-3.2-3B-Instruct-Q4_K_M.gguf is a standard F16 GGUF. The *-layers-*.gguf files are not.

Training Details

No training. Weights come from Meta; F16 GGUF from unsloth/Llama-3.2-3B-Instruct-GGUF; this repo hosts that GGUF and (when split) packages it into midpoint layer shards.

Evaluation

No separate evals for the hosted GGUF or shards. See meta-llama/llama-3.2-3b-instruct.

Technical Specifications

  • Architecture: Llama-3.2-3B-Instruct (~3.2B params, GQA (24 Q / 8 KV heads), 28 layers, hidden dim 3072)
  • Quantization: F16
  • Packaging: pinned full F16 GGUF; optional Arc midpoint shards (*-layers-{start}-{endExclusive}.gguf)
  • Package format: layer-package-v1
  • Split: 2 stages at layer 14 (maxStages: 2)

Citation

@misc{llama323binstruct,
    title  = {Llama-3.2-3B-Instruct},
    author = {Meta},
    year   = {2026},
    url    = {https://huggingface.co/meta-llama/llama-3.2-3b-instruct}
}

Credit:

Attribution

Quantized GGUF ingested from unsloth/Llama-3.2-3B-Instruct-GGUF. Original weights: meta-llama/llama-3.2-3b-instruct. Redistributed under the base model's license.

Hosted by Aquaduck.

Model Card Contact

Aquaduck AI โ€” https://huggingface.co/aquaduck

Downloads last month
178
GGUF
Model size
2B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support