Instructions to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/Tiny-LLM-PDelta3-GDN2-InputRoute")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute
- SGLang
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Tiny-LLM-PDelta3-GDN2-InputRoute", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vtava/Tiny-LLM-PDelta3-GDN2-InputRoute with Docker Model Runner:
docker model run hf.co/vtava/Tiny-LLM-PDelta3-GDN2-InputRoute
Tiny-LLM-PDelta3-GDN2-InputRoute
Research artifact from TinyCeNN-LM. Architecture: TinyCeNN-LM experiment.
Architecture
- Architecture/run type:
TinyCeNN-LM experiment - Base model:
arnir0/Tiny-LLM - Dataset:
HuggingFaceFW/fineweb-edu - Source code: https://github.com/vtavakkoli/TinyCeNN-LM
Latest saved results
| Metric | Value |
|---|---|
feature_dim |
96 |
The Hugging Face repository keeps timestamped run artifacts under runs/. This preserves training reports, configs and run metadata independently of the temporary Colab filesystem.
Saved experiment files
pdelta3_config.jsontiny_llm_pdelta3_report.json
Reproducibility
Run the matching notebook from the TinyCeNN-LM repository. Colab notebooks use a Hugging Face write token from the HF_TOKEN Colab Secret; tokens should never be pasted into notebook source.
Limitations
This is a research checkpoint. Metrics saved here are the metrics produced by the corresponding training notebook/script; unless explicitly marked as held-out evaluation, they should not be treated as publication-grade benchmark results. Generation quality can differ substantially from the base model.
Citation
If you use this experimental checkpoint, cite the TinyCeNN-LM repository and the upstream base model.
Model tree for vtava/Tiny-LLM-PDelta3-GDN2-InputRoute
Base model
arnir0/Tiny-LLM