Instructions to use qikp/silia-v2-hf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use qikp/silia-v2-hf with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="qikp/silia-v2-hf", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("qikp/silia-v2-hf", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use qikp/silia-v2-hf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "qikp/silia-v2-hf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qikp/silia-v2-hf", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/qikp/silia-v2-hf
- SGLang
How to use qikp/silia-v2-hf with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "qikp/silia-v2-hf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qikp/silia-v2-hf", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "qikp/silia-v2-hf" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qikp/silia-v2-hf", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use qikp/silia-v2-hf with Docker Model Runner:
docker model run hf.co/qikp/silia-v2-hf
Silia v2 — native Hugging Face port
A drop-in Hugging Face port of Srijan-Srivastava/Silia-v2,
a 524,672-parameter "Silu in Attention" (Silia) language model. The weights, the
tokenizer, and the forward pass are bit-exact against the reference
implementation.
Original work: "Tiny Scale Is All I Can Spare To Play With Transformer." by Srijan Srivastava (MIT), from the upstream repository. This is a port, not a new model: no weights were retrained or altered. Downstream metrics reported for the original checkpoint (HellaSwag, PIQA, LAMBADA) are in the upstream model card; they are the author's numbers, not re-measured here.
Loads through the standard Auto* classes — modular_siliav2.py is picked up via
auto_map in config.json plus trust_remote_code=True.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("silia-v2-hf", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("silia-v2-hf", trust_remote_code=True)
out = model.generate(
**tok("the", return_tensors="pt"),
max_new_tokens=32,
do_sample=False, # `generation_config.json` samples by default
pad_token_id=510,
)
print(tok.decode(out[0, 3:]))
# ' parts of the United States, the United States, and the Unit'
<|actor|> is the model's BOS — the reference's context sink, id 511. The
tokenizer prepends it for you, so pass the bare prompt. Do not also write
<|actor|> yourself, or you get two sinks. See
documented differences.
No pip install of this directory is required, and the repo is self-contained:
there is no requirements.txt, and the tokenizer does not need o512.bin at
runtime.
Files
| File | Purpose |
|---|---|
modular_siliav2.py |
All SiliaV2* classes. Self-contained, absolute imports only. |
config.json |
architectures: ["SiliaV2ForCausalLM"], auto_map, model_type: "siliav2". |
model.safetensors |
31 tensors, fp32. The tied lm_head is omitted. |
tokenizer.json |
WordPiece model + GPT-4-pattern pre-tokenizer, 512 tokens. |
tokenizer_config.json |
TokenizersBackend; <|end-text|>/<|actor|> as added tokens, <|actor|> as bos_token. |
generation_config.json |
do_sample: true, temperature: 0.8, top_k: 50, bos_token_id: 511, eos_token_id: 510. |
Architecture
Silia fuses attention and the SwiGLU feed-forward into one Hydra Latent Attention unit. A block is:
u, v = a1(x).chunk(2, dim=-1) # a1: n_embd -> 2 * d_ff
y = u * F.silu(v)
return x + a2(y) # a2: 2 * d_ff -> n_embd
There is no pre-layer-norm and a single residual. a1/a2 each split their
n_head heads into two halves with opposite roles:
- XLA (even heads) — ordinary causal attention over a KV cache, with a NeoX-style half-split RoPE on Q and K only, and a "GLA-style" output projection onto the normal value direction.
- AFT (odd heads) — Apple's Attention Free Transformer
(arXiv:2105.14103):
y_t = (Σ_{i≤t} exp(k_i) v_i) / (Σ_{i≤t} exp(k_i) + 1e-6), gated bysigmoid(q).
The AFT recurrence is a pure running sum, so decoding is O(1) per step. Both halves
are cached, using a hybrid cache layer that holds a full-attention KV state and two
recurrent states (Σexp(k) and Σexp(k)·v).
| Property | Value |
|---|---|
| Parameters | 524,672 (unique; lm_head is tied to wte) |
| Blocks | 3 |
n_embd / head_dim |
64 / 64 |
n_head |
2 (1 XLA + 1 AFT) |
d_ff (a.k.a. d_model) |
256 |
| Context | 1024 |
| Vocab | 512 |
| Checkpoint tensors | 32 → 31 saved after tying |
Fidelity
Verified against the original implementation — the reference graph in
Silia-v2/model.py and tokenizer in Silia-v2/encoder.py — with a 25-test /
2,069-subtest parity suite (not shipped here):
- Logits are bit-exact.
max |reference − port| = 0.000e+00withattn_implementation="sdpa", at sequence lengths 1, 7, 64 and 256, on both the cached and uncached paths, for a model built by hand frommodel.bin— and again at lengths 1, 3, 16 and 128 for the shipped artifacts loaded throughfrom_pretrained. With"eager"the difference is ~2.7e-4 in practice, which is sdpa-vs-eager fp32 kernel noise amplified byexp(k); the test allowsrtol=1e-2, atol=1e-3. - The tokenizer is id-for-id exact on the original project's prose files, on the specials, on >100-character single chunks, and on a 2,000-string mixed-script fuzz set. The port was additionally checked against an 11-file corpus (75,538 tokens) and every one of the 510 vocabulary tokens in 4 contexts.
- Greedy decoding matches the reference token for token — asserted over 12 steps in the suite, and spot-checked to 96 steps out of band.
The tokenizer is not BPE
The reference Encoder._encode_chunk has no merge table. It walks left to right,
repeatedly extending the current token while the concatenation is in the vocabulary —
greedy longest-match, i.e. maximal munch. (BPE is used only to train the
vocabulary, in the reference's own tokenizer-training code; encoding never consults
it.)
So "the" is ['th', 'e'] here, whereas a rank-ordered BPE model would give
['t', 'he'] — a different token sequence, and therefore out-of-distribution for a
model trained on the reference segmentation. HF WordPiece with an empty
continuing_subword_prefix is maximal munch, so the port is exact without any
custom tokenizer class. max_input_chars_per_word is raised to 10⁶ because a single
GPT-4 pre-tokenizer chunk can exceed the 100-character default. unk_token is
omitted: all 256 single bytes are in the vocabulary, so longest match never fails.
The pre-tokenizer is the shared GPT-4 split pattern from the reference. Note that
Encoder.load() sets self.pattern but never recompiles self.compiled_pattern, so
the pattern stored in o512.bin is inert; it equals GPT4_SPLIT_PATTERN anyway.
Documented differences from the reference
These are deliberate, and none of them change the logits for a given input.
- No per-step
<|actor|>sink prepend. The reference'sSilia.generateprepends the actor sink token to the front of the whole context on every step (so a growing sequence reads[<|actor|>, a, b, c, ...]). That is equivalent to keeping the token at position 0 with a KV cache, so the port does not reproduce the per-step prepend. Instead<|actor|>is registered as the BOS (bos_token/bos_token_id: 511), and the tokenizer prepends it once at encode time — the same convention Gemma, Llama and friends use, and the same single sink at position 0 that the reference produces. Pass the bare prompt. Adding<|actor|>yourself yields two sinks, exactly as'<bos>the'does under Gemma.add_special_tokens=Falsesuppresses it if you need the bare sequence. num_hidden_layerscounts attention units, not blocks. Each block has two attention units, so the config carriesnum_hidden_layers == 2 * num_blocks(6, for 3 blocks) to give the cache one layer per unit. The true block count isconfig.num_blocks, which is authoritative on reload — that is what keeps a saved config idempotent instead of doubling the layer count on every load.- The AFT half ignores
attention_mask.attention_free_transformerhas no padding mask, because the recurrence is a cumulative sum rather than a weight-and-sum and masking it would change the state carried into the next position. The XLA half usescreate_causal_masknormally. Batch only equal-length sequences, or pad on the left and slice. - Special tokens are always recognized. The reference's
Encoder.encodetakes anallowed_specialpolicy — the default is"none_raise", which asserts if a special token appears in the text, while"none"treats it as ordinary text. The HF tokenizer always honors added special tokens. To compare against the reference, passallowed_special="all". - The loss is shifted cross-entropy. The reference computes
cross_entropy(logits, targets)with no shift, i.e. it scores position t against the label for t rather than t+1. That is not the usual autoregressive objective, so this port keeps HF's standard behaviour. The logits already match exactly, so nothing else is affected. - Vestigial GPT-2 config fields.
SiliaV2ConfigsubclassesGPT2Config, sosummary_*,add_cross_attention,reorder_and_upcast_attn,scale_attn_weights,layer_norm_epsilonand friends appear inconfig.jsonand are ignored. GPT-2 is used only as plumbing — the config,GenerationMixin, and v5's tied-weights convention. No GPT-2 compute is retained; the substantive math comes fromLlamaRotaryEmbedding,apply_rotary_pos_embandeager_attention_forward. - A warm-up context bug in the reference's
inference.pyis not reproduced. Line 33 draws a random prompt asrandom.randint(0, len(enc.vocab) + len(enc.special_tokens))— i.e.randint(0, 512), which can emit id 512, one past the end of the 512-token vocabulary. This port always supplies a real prompt.
Where the paper and the code disagree
The published description of Silia v2 does not match the released code. This port follows the code, since the weights were trained against it.
| Paper | Code (and this port) |
|---|---|
| §2.2: RoPE applied to Q, K, and V | Q and K only, head_dim half-split |
§2.2: RMSNorm(C) after the concat |
No norm after the concat |
§2.3: SiLU(u) ⊙ v |
u * silu(v) — operands reversed |
| §5.1: "byte-pair encoding" | Maximal munch (no merge table) |
| §5.1 sample output | Contains a literal <|end-text|> |
Two further details are easy to get wrong and are worth stating explicitly:
- RMSNorm is gain-less, and its epsilon is
torch.finfo(dtype).eps(1.19e-7 in fp32) — not 1e-5 or 1e-6.config.rms_norm_epsis thereforenullby default and resolved at runtime from the activation dtype. - The RoPE
sinterm is sign-mirrored relative to HF'srotate_half. The reference rotates[x₁·cos + x₂·sin, −x₁·sin + x₂·cos]whererotate_halfgives[q₁·cos − q₂·sin, q₂·cos + q₁·sin], so this port passes-sin. The discrepancy is invisible at position 0 (wheresin = 0), which is exactly why it survived the first round of tests.
Implementation note
SiliaV2PreTrainedModel deliberately has no _init_weights override.
from_pretrained builds the model on the meta device, and inv_freq is a
non-persistent buffer, so it is absent from the checkpoint and must be recomputed
during weight initialization. Overriding _init_weights replaces — rather than
extends — the inherited implementation that performs that recomputation, which
leaves inv_freq as uninitialized memory. The resulting model still has the right
parameter count, the right weights, and the right tied head, but scrambled every
forward pass. A dedicated parity test guards this.
- Downloads last month
- 603
Model tree for qikp/silia-v2-hf
Base model
Srijan-Srivastava/Silia-v2