Instructions to use SadSinglePringle/Obscuris-V1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SadSinglePringle/Obscuris-V1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SadSinglePringle/Obscuris-V1", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SadSinglePringle/Obscuris-V1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SadSinglePringle/Obscuris-V1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SadSinglePringle/Obscuris-V1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SadSinglePringle/Obscuris-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/SadSinglePringle/Obscuris-V1
- SGLang
How to use SadSinglePringle/Obscuris-V1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SadSinglePringle/Obscuris-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SadSinglePringle/Obscuris-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SadSinglePringle/Obscuris-V1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SadSinglePringle/Obscuris-V1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use SadSinglePringle/Obscuris-V1 with Docker Model Runner:
docker model run hf.co/SadSinglePringle/Obscuris-V1
ObscurisV1
ObscurisV1 is the first model from Umbral, a 100.8M-parameter dense causal Transformer created by Umbral. It uses 16 Transformer layers, width 512, eight attention heads, RoPE, GELU feed-forward blocks, tied input/output embeddings, and a 32,768-token byte-BPE tokenizer. The maximum trained context length is 512 tokens.
This model exists as an initial baseline model for comparison with future architectures, but in the proccess highlighted training problems we'll fix for future iterations.
This package contains the post-trained identity-calibration checkpoint. It has been trained to identify itself as: "I am ObscurisV1, created by Umbral." However due to training shortcomings the model identity is fairly brittle.
Loading
This is a custom architecture, so load it with trust_remote_code=True:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "YOUR_USERNAME/ObscurisV1"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).cuda().eval()
prompt = "System:\nYou are ObscurisV1, created by Umbral.\nUser:\nWhat is your name?\nAssistant:\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=64, do_sample=True, temperature=0.55, top_p=0.85)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Limitations
This is a small experimental model. It can produce fluent but incorrect information, especially in long factual explanations. It should not be used for high-stakes decisions.
The reference implementation is optimized for compatibility and correctness, not maximum throughput. It currently does not implement a KV cache.