Instructions to use Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental") model = AutoModelForCausalLM.from_pretrained("Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental
- SGLang
How to use Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental with Docker Model Runner:
docker model run hf.co/Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental
OLMoE-0125-Instruct-Compressed-Experimental
An experimental, modified derivative of AllenAI's allenai/OLMoE-1B-7B-0125-Instruct, based on revision b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e. This repository contains the complete healed BF16 Hugging Face checkpoint and its tokenizer.
All nine benchmark categories have completed for the base, cut, and healed checkpoints.
Observed regression: Further training reduced sampled GSM8K accuracy from 64.00% (Cut) to 43.00% (Healed). This is an experimental result, not a promoted quality improvement. The cut checkpoint is retained below.
The saved weights occupy 12.03 GB, compared with 13.84 GB for the base: 13.10% smaller. The healed model contains 6.013 billion parameters. This is a storage reduction, not a claim of faster inference or improved quality.
The model and tokenizer were independently reloaded from the preserved bundle and generated tokens successfully. That execution check is not a quality certification.
The complete cut checkpoint, before further training, is retained at revision 376988e3. Its measurements appear in the Cut column below. Set revision="376988e3a817b9d7f4ed7eb71ab7007c483b59ce" in both loading calls to use that checkpoint.
Measured results
These are measurements from one seed (42) on sampled benchmark tasks. They use local evaluation protocols and should not be compared directly with vendor leaderboard scores. No replicated quality improvement is claimed.
| Benchmark category | Base | Cut | Healed |
|---|---|---|---|
| arc science | 45.50% (n=200) | 44.50% (n=200) | 45.00% (n=200) |
| bbq safety | 41.50% (n=200) | 42.50% (n=200) | 39.50% (n=200) |
| gpqa reasoning | 22.73% (n=198) | 19.19% (n=198) | 16.67% (n=198) |
| gsm8k math | 65.00% (n=200) | 64.00% (n=200) | 43.00% (n=200) |
| hellaswag commonsense | 67.00% (n=200) | 65.00% (n=200) | 66.00% (n=200) |
| ifeval instruction following | 61.27% (n=142) | 52.82% (n=142) | 55.63% (n=142) |
| mmlu knowledge | 30.50% (n=200) | 27.50% (n=200) | 26.50% (n=200) |
| multilingual knowledge | 28.00% (n=200) | 27.50% (n=200) | 26.50% (n=200) |
| wikitext language | 18.306 PPL (n=64) | 23.768 PPL (n=64) | 21.632 PPL (n=64) |
WikiText perplexity is lower-is-better. IFEval reports strict prompt accuracy over the supported instruction subset; its coverage and excluded counts are in evaluation_results.json. BBQ is multiple-choice accuracy, not a safety certification. These sampled results do not establish general capability, safety, or production readiness.
Formats and evaluation coverage
File sizes are storage requirements, not measured runtime memory or speed.
| Format | File size | Public download | Evaluation |
|---|---|---|---|
| BF16 | 12.03 GB | Available | 9 categories measured |
| Q8_0 | Not generated | Not published | Not measured |
| Q4_K_M | 3.66 GB | Not published | Not measured |
| Q2_K | Not generated | Not published | Not measured |
Scores apply only to the exact format evaluated; BF16 scores do not qualify a GGUF conversion.
| Benchmark | BF16 | Q8_0 | Q4_K_M | Q2_K |
|---|---|---|---|---|
| arc science | 45.00% (n=200) | Not measured | Not measured | Not measured |
| hellaswag commonsense | 66.00% (n=200) | Not measured | Not measured | Not measured |
| mmlu knowledge | 26.50% (n=200) | Not measured | Not measured | Not measured |
| gpqa reasoning | 16.67% (n=198) | Not measured | Not measured | Not measured |
| wikitext language | 21.632 PPL (n=64) | Not measured | Not measured | Not measured |
| bbq safety | 39.50% (n=200) | Not measured | Not measured | Not measured |
| multilingual knowledge | 26.50% (n=200) | Not measured | Not measured | Not measured |
| gsm8k math | 43.00% (n=200) | Not measured | Not measured | Not measured |
| ifeval instruction following | 55.63% (n=142) | Not measured | Not measured | Not measured |
Artifact identities
| Format | SHA-256 |
|---|---|
| BF16 | e7b3d37beb32930d2c9b1cea775b50f15d79d06dab5d0b39782d206c8ba99573 |
| Q4_K_M | 4f7fe283cd9a714a3c96824f30a83e7e4d70e55f16a9b9e0aea8f644e68e620f |
Q4_K_M is 69.54% smaller than the BF16 checkpoint by file size. It was generated and preserved locally, but is not downloadable from this repository and has no separate quality evaluation. Q8_0 and Q2_K were not generated in this first experiment. No format-specific speed, peak-memory, or quantization quality claim is made.
Use
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "What is 2 + 2?"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(output[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
License and attribution
The original model is by AllenAI. This modified derivative retains Apache License 2.0; see LICENSE. The original authors do not endorse this derivative.
- Downloads last month
- 441
Model tree for Omega-Factory/OLMoE-0125-Instruct-Compressed-Experimental
Base model
allenai/OLMoE-1B-7B-0125