ObscurisV1

ObscurisV1 is the first model from Umbral, a 100.8M-parameter dense causal Transformer created by Umbral. It uses 16 Transformer layers, width 512, eight attention heads, RoPE, GELU feed-forward blocks, tied input/output embeddings, and a 32,768-token byte-BPE tokenizer. The maximum trained context length is 512 tokens.

This model exists as an initial baseline model for comparison with future architectures, but in the proccess highlighted training problems we'll fix for future iterations.

This package contains the post-trained identity-calibration checkpoint. It has been trained to identify itself as: "I am ObscurisV1, created by Umbral." However due to training shortcomings the model identity is fairly brittle.

Loading

This is a custom architecture, so load it with trust_remote_code=True:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "YOUR_USERNAME/ObscurisV1"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).cuda().eval()

prompt = "System:\nYou are ObscurisV1, created by Umbral.\nUser:\nWhat is your name?\nAssistant:\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=64, do_sample=True, temperature=0.55, top_p=0.85)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Limitations

This is a small experimental model. It can produce fluent but incorrect information, especially in long factual explanations. It should not be used for high-stakes decisions.

The reference implementation is optimized for compatibility and correctness, not maximum throughput. It currently does not implement a KV cache.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support