Instructions to use muhamedkamil/berna-prototype-162m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use muhamedkamil/berna-prototype-162m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="muhamedkamil/berna-prototype-162m")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("muhamedkamil/berna-prototype-162m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use muhamedkamil/berna-prototype-162m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "muhamedkamil/berna-prototype-162m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muhamedkamil/berna-prototype-162m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/muhamedkamil/berna-prototype-162m
- SGLang
How to use muhamedkamil/berna-prototype-162m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "muhamedkamil/berna-prototype-162m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muhamedkamil/berna-prototype-162m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "muhamedkamil/berna-prototype-162m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muhamedkamil/berna-prototype-162m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use muhamedkamil/berna-prototype-162m with Docker Model Runner:
docker model run hf.co/muhamedkamil/berna-prototype-162m
Berna Prototype 162M
Decoder-only Transformer trained from scratch on 98M tokens.
Model Details
| Property | Value |
|---|---|
| Parameters | 162.42M |
| Architecture | Decoder-only |
| Layers | 12 |
| Hidden size | 768 |
| Intermediate size | 3072 |
| Attention heads | 12 |
| Vocab size | 32000 |
| Context length | 2048 |
| Position encoding | RoPE |
| Normalization | RMSNorm |
| Activation | SwiGLU |
| Precision | BF16 |
Training Details
| Property | Value |
|---|---|
| Hardware | RTX 5090 |
| Time | 40 minutes |
| Tokens | 98,295,808 |
| Throughput | 40,815 tok/s |
| Optimizer | AdamW |
| Scheduler | Cosine warmup |
| Peak LR | 6e-4 |
| Batch size | 16 x 512 |
| Steps | 12000 |
| Final loss | 5.58 |
| Perplexity | 265 |
Architecture Features
- RMSNorm pre-normalization
- RoPE position embeddings
- SwiGLU activation
- Separate LM head
- Causal attention mask
- No bias terms
Usage
pip install torch transformers
from transformers import AutoConfig, AutoModelForCausalLM
from transformers import PreTrainedTokenizerFast
from modeling_berna import BernaConfig, BernaForCausalLM
AutoConfig.register("berna", BernaConfig)
AutoModelForCausalLM.register(BernaConfig, BernaForCausalLM)
model = AutoModelForCausalLM.from_pretrained(".", local_files_only=True)
tokenizer = PreTrainedTokenizerFast.from_pretrained(".", local_files_only=True)
model = model.to("cuda")
Generate
prompt = "Hello"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(output[0]))
Status
Prototype. Trained on 98M tokens only. Text quality is limited at this stage.
| Quality | Tokens | Progress |
|---|---|---|
| Basic words | 300M | 33% |
| Phrases | 1B | 10% |
| Sentences | 3B | 3% |
| Coherent | 40B | 0.25% |
Roadmap
- Environment setup
- BPE tokenizer (32k)
- Llama-style architecture
- Atomic checkpointing
- Kill -9 recovery
- HF-compatible export
- Train on 98M tokens
- Scale to 1.5B
License
This model is released under the Berna Custom License. See LICENSE for full terms.
Key restrictions:
- Name "Berna" must be preserved (no renaming).
- Modification of weights is exclusive to the creator.
- Commercial use requires separate written license.
For commercial licensing: contact @muhamedkamil on HuggingFace.
- Downloads last month
- 362