teknium/OpenHermes-2.5
Viewer • Updated • 1M • 21k • 913
How to use Meridian-MRM/Zedev-134m with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="Meridian-MRM/Zedev-134m") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Meridian-MRM/Zedev-134m")
model = AutoModelForCausalLM.from_pretrained("Meridian-MRM/Zedev-134m", device_map="auto")How to use Meridian-MRM/Zedev-134m with llama.cpp:
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Meridian-MRM/Zedev-134m:F16 # Run inference directly in the terminal: llama cli -hf Meridian-MRM/Zedev-134m:F16
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Meridian-MRM/Zedev-134m:F16 # Run inference directly in the terminal: llama cli -hf Meridian-MRM/Zedev-134m:F16
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Meridian-MRM/Zedev-134m:F16 # Run inference directly in the terminal: ./llama-cli -hf Meridian-MRM/Zedev-134m:F16
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Meridian-MRM/Zedev-134m:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Meridian-MRM/Zedev-134m:F16
docker model run hf.co/Meridian-MRM/Zedev-134m:F16
How to use Meridian-MRM/Zedev-134m with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Meridian-MRM/Zedev-134m"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Meridian-MRM/Zedev-134m",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'docker model run hf.co/Meridian-MRM/Zedev-134m:F16
How to use Meridian-MRM/Zedev-134m with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "Meridian-MRM/Zedev-134m" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Meridian-MRM/Zedev-134m",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "Meridian-MRM/Zedev-134m" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Meridian-MRM/Zedev-134m",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'How to use Meridian-MRM/Zedev-134m with Ollama:
ollama run hf.co/Meridian-MRM/Zedev-134m:F16
How to use Meridian-MRM/Zedev-134m with Docker Model Runner:
docker model run hf.co/Meridian-MRM/Zedev-134m:F16
How to use Meridian-MRM/Zedev-134m with Lemonade:
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Meridian-MRM/Zedev-134m:F16
lemonade run user.Zedev-134m-F16
lemonade list
Base causal language model with 134.47M parameters.
| Field | Value |
|---|---|
| Parameters | 134.47M |
| Layers | 26 |
| Hidden size | 640 |
| Attention | GQA (10 Query / 2 KV heads, head_dim 64) |
| Intermediate size | 1536 SwiGLU |
| Norm | RMSNorm (eps 1e-5) + per-head QK-Norm |
| RoPE theta | 10000.0 |
| Context length | 1024 |
| Vocab size | 50304 (50261 used, padded to /128) |
| Tokenizer | GPT-2 BPE |
| Embeddings | Tied (lm_head == embed_tokens) |
| HF Layout | Qwen3ForCausalLM |
| Parameter | Value |
|---|---|
| Training Time | 17 hours, 41 minutes |
| Trained on Context Length | 1024 |
| Tokens seen | 17.652B~ |
| Source | Tokens | License |
|---|---|---|
| Wikipedia (20231101.en) | 4.585B | CC BY-SA 4.0 |
| Gutenberg | 4.500B | Public Domain |
| DCLM-baseline s1 | 4.384B | CC-BY-4.0 |
| DCLM-baseline s2 | 4.261B | CC-BY-4.0 |
| OpenHermes | 0.385B | Apache 2.0 |
| Total | 18.115B |
Evaluated with lm-evaluation-harness (0-shot, float16, batch size 16).
| Benchmark | Metric | Cagliari | Zedev | Improvement |
|---|---|---|---|---|
| ARC-Challenge | acc_norm | 22.78% | 25.51% | 🟢 +2.73% |
| acc | 19.03% | 23.12% | 🟢 +4.09% | |
| ARC-Easy | acc | 38.30% | 45.62% | 🟢 +7.32% |
| acc_norm | 35.40% | 41.67% | 🟢 +6.27% | |
| BoolQ | acc | 41.44% | 45.84% | 🟢 +4.40% |
| HellaSwag | acc_norm | 27.41% | 31.17% | 🟢 +3.76% |
| acc | 27.04% | 28.72% | 🟢 +1.68% | |
| OpenBookQA | acc_norm | 28.20% | 29.00% | 🟢 +0.80% |
| acc | 14.80% | 17.40% | 🟢 +2.60% | |
| PIQA | acc | 58.27% | 61.43% | 🟢 +3.16% |
| acc_norm | 55.98% | 60.94% | 🟢 +4.96% | |
| WinoGrande | acc | 50.36% | 49.64% | 🔴 -0.72% |
| BLiMP | acc | 80.22% | 80.32% | 🟢 +0.10% |
| WikiText-2 | word_perplexity (↓) | 36.95 | 29.67 | 🟢 -7.28 |
| byte_perplexity (↓) | 1.96 | 1.89 | 🟢 -0.07 | |
| bits_per_byte (↓) | 0.97 | 0.9146 | 🟢 -0.0554 |
| Token | ID |
|---|---|
<|endoftext|> |
50256 (EOS/BOS/PAD) |
<|im_start|> |
50257 |
<|im_end|> |
50258 |
<think> |
50259 |
</think> |
50260 |
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Meridian-MRM/Zedev-134m")
model = AutoModelForCausalLM.from_pretrained(
"Meridian-MRM/Zedev-134m", torch_dtype=torch.float16
).cuda().eval()
prompt = "The old lighthouse keeper walked to the edge of the cliff and"
ids = tok(prompt, return_tensors="pt").input_ids.cuda()
with torch.no_grad():
out = model.generate(
ids, max_new_tokens=128, do_sample=True,
temperature=0.7, top_k=40, top_p=0.95,
repetition_penalty=1.15, pad_token_id=50256,
)
print(tok.decode(out[0], skip_special_tokens=True))
./llama-cli -m zedev-134m-f16.gguf \
-p "The old lighthouse keeper walked to the edge of the cliff and" \
-n 128 --temp 0.7 --top-k 40 --top-p 0.95 --repeat-penalty 1.15
Apache 2.0.
@misc{zedev134m,
title = {Zedev-134M},
author = {Meridian-MRM},
howpublished = {\url{https://huggingface.co/Meridian-MRM/zedev-134m}}
}