Samosa Chat Maple 2-bit SSD streaming pack

This repository contains the production artifact layout used by Samosa Chat's native Maple runtime. It is a storage-only repack of the MIT-licensed deepgrove/maple-preview-2bit-mlx checkpoint pinned at revision 361db5da5e74ff6fcdd852d478e1f266ce11013a.

The model's numerical weights are unchanged. Expert tensors are stored in fixed-size aligned records in maple-experts.bin, allowing Samosa to read only the routed experts from SSD. Non-expert tensors are stored in maple-resident.safetensors. maple-manifest.json describes and validates the packed layout.

These files are intended for the bundled samosa-maple runtime. They are not a drop-in MLX checkpoint because the original expert shards have been replaced by Samosa's streaming container.

Runtime files

  • maple-experts.bin: aligned expert records streamed on demand from SSD
  • maple-resident.safetensors: non-expert tensors retained by the runtime
  • maple-manifest.json: strict layout metadata
  • config.json and tokenizer files: model configuration and prompt encoding

Samosa verifies every artifact's byte length and SHA-256 digest before atomic installation.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for deepanwa/Samosa-Chat-Maple-2bit-SSD

Quantized
(1)
this model