Instructions to use vtava/SmolLM2-135M-AMCeNN-Top2-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/SmolLM2-135M-AMCeNN-Top2-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/SmolLM2-135M-AMCeNN-Top2-v2")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/SmolLM2-135M-AMCeNN-Top2-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/SmolLM2-135M-AMCeNN-Top2-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/SmolLM2-135M-AMCeNN-Top2-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-AMCeNN-Top2-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vtava/SmolLM2-135M-AMCeNN-Top2-v2
- SGLang
How to use vtava/SmolLM2-135M-AMCeNN-Top2-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/SmolLM2-135M-AMCeNN-Top2-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-AMCeNN-Top2-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/SmolLM2-135M-AMCeNN-Top2-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-AMCeNN-Top2-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vtava/SmolLM2-135M-AMCeNN-Top2-v2 with Docker Model Runner:
docker model run hf.co/vtava/SmolLM2-135M-AMCeNN-Top2-v2
SmolLM2-135M-AMCeNN-Top2-v2
Research artifact from TinyCeNN-LM. Architecture: smollm2-amcenn-top2-v2.
Architecture
- Architecture/run type:
smollm2-amcenn-top2-v2 - Base model:
HuggingFaceTB/SmolLM2-135M - Dataset:
HuggingFaceFW/fineweb-edu (sample-10BT) - Source code: https://github.com/vtavakkoli/TinyCeNN-LM
Latest saved results
| Metric | Value |
|---|---|
status |
trained |
stop_reason |
token_budget |
context_length |
128 |
feature_dim |
128 |
num_shards |
8 |
top_k |
2 |
trainable |
26,926,110 |
trainable_percent |
19.9602 |
last_training_ce |
5.42189 |
last_distillation_kl |
4.96675 |
mean_route_mix |
-0.0024306 |
mean_router_entropy |
2.07669 |
elapsed_minutes |
46.2677 |
peak_vram_gib |
1.63402 |
evaluation_performed |
0 |
The Hugging Face repository keeps timestamped run artifacts under runs/. This preserves training reports, configs and run metadata independently of the temporary Colab filesystem.
Saved experiment files
.hf_live_redacted/smollm2_amcenn_v2_config.json.hf_live_redacted/smollm2_amcenn_v2_training_report.json.hf_live_redacted/tokenizer_config.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/.hf_live_redacted/smollm2_amcenn_v2_config.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/.hf_live_redacted/smollm2_amcenn_v2_training_report.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/.hf_live_redacted/tokenizer_config.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/smollm2_amcenn_v2_config.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/smollm2_amcenn_v2_training_report.json.hf_run_archive/smollm2-amcenn-top2-v2-20260913T111214Z/tokenizer_config.jsonsmollm2_amcenn_v2_config.jsonsmollm2_amcenn_v2_training_report.jsontokenizer_config.json
Reproducibility
Run the matching notebook from the TinyCeNN-LM repository. Colab notebooks use a Hugging Face write token from the HF_TOKEN Colab Secret; tokens should never be pasted into notebook source.
Limitations
This is a research checkpoint. Metrics saved here are the metrics produced by the corresponding training notebook/script; unless explicitly marked as held-out evaluation, they should not be treated as publication-grade benchmark results. Generation quality can differ substantially from the base model.
Citation
If you use this experimental checkpoint, cite the TinyCeNN-LM repository and the upstream base model.
Model tree for vtava/SmolLM2-135M-AMCeNN-Top2-v2
Base model
HuggingFaceTB/SmolLM2-135M