Instructions to use RichardErkhov/weiser_-_124M-0.4-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RichardErkhov/weiser_-_124M-0.4-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
Use Docker
docker model run hf.co/RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use RichardErkhov/weiser_-_124M-0.4-gguf with Ollama:
ollama run hf.co/RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use RichardErkhov/weiser_-_124M-0.4-gguf with Docker Model Runner:
docker model run hf.co/RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
- Lemonade
How to use RichardErkhov/weiser_-_124M-0.4-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RichardErkhov/weiser_-_124M-0.4-gguf:Q4_K_M
Run and chat with the model
lemonade run user.weiser_-_124M-0.4-gguf-Q4_K_M
List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Quantization made by Richard Erkhov.
124M-0.4 - GGUF
- Model creator: https://huggingface.co/weiser/
- Original model: https://huggingface.co/weiser/124M-0.4/
| Name | Quant method | Size |
|---|---|---|
| 124M-0.4.Q2_K.gguf | Q2_K | 0.08GB |
| 124M-0.4.IQ3_XS.gguf | IQ3_XS | 0.08GB |
| 124M-0.4.IQ3_S.gguf | IQ3_S | 0.08GB |
| 124M-0.4.Q3_K_S.gguf | Q3_K_S | 0.08GB |
| 124M-0.4.IQ3_M.gguf | IQ3_M | 0.09GB |
| 124M-0.4.Q3_K.gguf | Q3_K | 0.09GB |
| 124M-0.4.Q3_K_M.gguf | Q3_K_M | 0.09GB |
| 124M-0.4.Q3_K_L.gguf | Q3_K_L | 0.1GB |
| 124M-0.4.IQ4_XS.gguf | IQ4_XS | 0.1GB |
| 124M-0.4.Q4_0.gguf | Q4_0 | 0.1GB |
| 124M-0.4.IQ4_NL.gguf | IQ4_NL | 0.1GB |
| 124M-0.4.Q4_K_S.gguf | Q4_K_S | 0.1GB |
| 124M-0.4.Q4_K.gguf | Q4_K | 0.11GB |
| 124M-0.4.Q4_K_M.gguf | Q4_K_M | 0.11GB |
| 124M-0.4.Q4_1.gguf | Q4_1 | 0.11GB |
| 124M-0.4.Q5_0.gguf | Q5_0 | 0.11GB |
| 124M-0.4.Q5_K_S.gguf | Q5_K_S | 0.11GB |
| 124M-0.4.Q5_K.gguf | Q5_K | 0.12GB |
| 124M-0.4.Q5_K_M.gguf | Q5_K_M | 0.12GB |
| 124M-0.4.Q5_1.gguf | Q5_1 | 0.12GB |
| 124M-0.4.Q6_K.gguf | Q6_K | 0.13GB |
| 124M-0.4.Q8_0.gguf | Q8_0 | 0.17GB |
Original model description:
license: apache-2.0 datasets: - HuggingFaceFW/fineweb language: - en library_name: transformers tags: - IoT - sensor - embedded
TinyLLM
Overview
This repository hosts a small language model developed as part of the TinyLLM framework ([arxiv link]). These models are specifically designed and fine-tuned with sensor data to support embedded sensing applications. They enable locally hosted language models on low-computing-power devices, such as single-board computers. The models, based on the GPT-2 architecture, are trained using Nvidia's H100 GPUs. This repo provides base models that can be further fine-tuned for specific downstream tasks related to embedded sensing.
Model Information
- Parameters: 124M (Hidden Size = 768)
- Architecture: Decoder-only transformer
- Training Data: Up to 10B tokens from the SHL and Fineweb datasets, combined in a 4:6 ratio
- Input and Output Modality: Text
- Context Length: 1024
Acknowledgements
We want to acknowledge the open-source frameworks llm.c and llama.cpp and the sensor dataset provided by SHL, which were instrumental in training and testing these models.
Usage
The model can be used in two primary ways:
With Hugging Face’s Transformers Library
from transformers import pipeline import torch path = "tinyllm/124M-0.4" prompt = "The sea is blue but it's his red sea" generator = pipeline("text-generation", model=path,max_new_tokens = 30, repetition_penalty=1.3, model_kwargs={"torch_dtype": torch.bfloat16}, device_map="auto") print(generator(prompt)[0]['generated_text'])With llama.cpp Generate a GGUF model file using this tool and use the generated GGUF file for inferencing.
python3 convert_hf_to_gguf.py models/mymodel/
Disclaimer
This model is intended solely for research purposes.
- Downloads last month
- 57
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit