GGUF

OpenELM-1.1B-Instruct โ€” GGUF (Q4_K_M)

Community Q4_K_M GGUF for Apple's OpenELM, re-hosted with measured performance for the 1bit engine on Strix Halo.

Research use only โ€” licensed under the Apple Machine Learning Research License, which restricts use to non-commercial scientific research and academic development. Commercial use, product development, or use in any commercial product/service is not permitted under this license.

Contents

  • OpenELM-1_1B-Instruct.Q4_K_M.gguf

Measured performance (Strix Halo, Vulkan)

pp512: 7356 tok/s ยท tg128: 234 tok/s

Running it

1bit serve -m OpenELM-1_1B-Instruct.Q4_K_M.gguf --device vulkan

Attribution

Downloads last month
34
GGUF
Model size
1B params
Architecture
openelm
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 1bit-MONSTER/OpenELM-1.1B-Instruct-GGUF

Quantized
(9)
this model