granite-4.2-8b โ€” MLBricks ElasticBit

Source: ibm-granite/granite-4.2-8b
Source revision: f8de16cdcdbc6c779ca517604e050d82cc119e44
MLBricks commit: f772eb27d2d6b84764e9e90458827dac6980a88a
ElasticBit threshold: 0.01
Persistent global dtypes: preserved from the source model as loaded for conversion

Conversion result

  • Decoder layers: 40
  • Parameter-weighted effective bits: 8.175
  • Linear storage reduction: 48.87%
  • Layer artifact bytes: 6.949 GiB
  • Exact global component bytes: 1.531 GiB
  • Approximate uploaded inference payload: 8.480 GiB
  • Reconstructed end-to-end smoke test: PASS

Validation

Each newly converted layer is checked on calibration inputs, a separate validation input, and a local save/load round trip. Before complete=true, the converter reconstructs global payloads from the uploaded safetensors files, streams every uploaded ElasticBit layer through a real full-model forward, and compares final logits with the pinned FP16 source reference.

Layout

  • layers/layer_XXX.elasticbit
  • components/global_manifest.json
  • components/global_*.safetensors
  • reports/layer_XXX.json
  • reports/reconstruction_smoke.json
  • calibration/input_ids.safetensors
  • validation/input_ids.safetensors
  • elasticbit_manifest.json
  • progress.json

Layer artifacts use a minimal ElasticBit root namespace: compressed Linear paths are stored as block.*. The root class itself is not serialized. A loader should create a small nn.Module with a .block containing the pinned source-compatible decoder architecture, call ElasticBit.load(root, artifact), then use root.block in the reconstructed model while retaining root when controller access is needed.

Review the source model license before redistribution or commercial use.

Downloads last month
97
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for maxzameer/granite-4.2-8b-ElasticBit

Finetuned
(7)
this model