Qwen2.5-14B โ€” MLBricks ElasticBit

Source: Qwen/Qwen2.5-14B
Source revision: 97e1e76335b7017d8f67c08a19d103c0504298c9
MLBricks commit: f772eb27d2d6b84764e9e90458827dac6980a88a
ElasticBit threshold: 0.01
Persistent global dtypes: preserved from the source model as loaded for conversion

Conversion result

  • Decoder layers: 48
  • Parameter-weighted effective bits: 8.115
  • Linear storage reduction: 49.25%
  • Layer artifact bytes: 11.258 GiB
  • Exact global component bytes: 2.900 GiB
  • Approximate uploaded inference payload: 14.158 GiB
  • Reconstructed end-to-end smoke test: PASS

Validation

Each newly converted layer is checked on calibration inputs, a separate validation input, and a local save/load round trip. Before complete=true, the converter reconstructs global payloads from the uploaded safetensors files, streams every uploaded ElasticBit layer through a real full-model forward, and compares final logits with the pinned FP16 source reference.

Layout

  • layers/layer_XXX.elasticbit
  • components/global_manifest.json
  • components/global_*.safetensors
  • reports/layer_XXX.json
  • reports/reconstruction_smoke.json
  • calibration/input_ids.safetensors
  • validation/input_ids.safetensors
  • elasticbit_manifest.json
  • progress.json

Layer artifacts use a minimal ElasticBit root namespace: compressed Linear paths are stored as block.*. The root class itself is not serialized. A loader should create a small nn.Module with a .block containing the pinned source-compatible decoder architecture, call ElasticBit.load(root, artifact), then use root.block in the reconstructed model while retaining root when controller access is needed.

Review the source model license before redistribution or commercial use.

Downloads last month
125
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for maxzameer/Qwen2.5-14B-ElasticBit

Base model

Qwen/Qwen2.5-14B
Finetuned
(124)
this model