kleinhirn weights: GLiNER2.5-small

Weight shards of fastino/gliner2.5-small-v1 in the format the kleinhirn engine loads. kleinhirn runs small encoder classifiers in the browser on WebGPU, with a WASM-SIMD fallback on the CPU. These files are for that engine. They are not a Transformers checkpoint.

Files

Folder Content
small-upstream/f32/ full-precision weights, used by the WebGPU f32 path and the WASM path
small-upstream/f16/ half-precision weights, used by the WebGPU f16 path (needs shader-f16)

Each folder has manifest.json (tensor layout, shard sizes and sha256), tokenizer.json and the weights-*.bin shards. The engine verifies every shard against the sha256 in the manifest.

Provenance

Converted from fastino/gliner2.5-small-v1 at revision 7e6f537f10337497069276892a5ef435028252ce with convert/export_weights.py from the kleinhirn repository. The f16 shards are a cast of the fp32 weights. The tokenizer is the unchanged upstream tokenizer.

Use

Point the kleinhirn engine at a manifest URL in this repository, for example small-upstream/f16/manifest.json. The device benchmark page of the kleinhirn project loads these files to measure the engine on your device.

License and credit

Apache-2.0, the license of the original model. See LICENSE and NOTICE. GLiNER2.5-small is by Fastino. Its encoder is DeBERTa-v3-xsmall by Microsoft (MIT license). The only change here is the conversion of the tensor layout and the fp16 cast described above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for smashburger-dev/kleinhirn-weights

Finetuned
(3)
this model