Silia v2 — native Hugging Face port

A drop-in Hugging Face port of Srijan-Srivastava/Silia-v2, a 524,672-parameter "Silu in Attention" (Silia) language model. The weights, the tokenizer, and the forward pass are bit-exact against the reference implementation.

Original work: "Tiny Scale Is All I Can Spare To Play With Transformer." by Srijan Srivastava (MIT), from the upstream repository. This is a port, not a new model: no weights were retrained or altered. Downstream metrics reported for the original checkpoint (HellaSwag, PIQA, LAMBADA) are in the upstream model card; they are the author's numbers, not re-measured here.

Loads through the standard Auto* classes — modular_siliav2.py is picked up via auto_map in config.json plus trust_remote_code=True.

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("silia-v2-hf", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("silia-v2-hf", trust_remote_code=True)

out = model.generate(
    **tok("the", return_tensors="pt"),
    max_new_tokens=32,
    do_sample=False,          # `generation_config.json` samples by default
    pad_token_id=510,
)
print(tok.decode(out[0, 3:]))
# ' parts of the United States, the United States, and the Unit'

<|actor|> is the model's BOS — the reference's context sink, id 511. The tokenizer prepends it for you, so pass the bare prompt. Do not also write <|actor|> yourself, or you get two sinks. See documented differences.

No pip install of this directory is required, and the repo is self-contained: there is no requirements.txt, and the tokenizer does not need o512.bin at runtime.

Files

File Purpose
modular_siliav2.py All SiliaV2* classes. Self-contained, absolute imports only.
config.json architectures: ["SiliaV2ForCausalLM"], auto_map, model_type: "siliav2".
model.safetensors 31 tensors, fp32. The tied lm_head is omitted.
tokenizer.json WordPiece model + GPT-4-pattern pre-tokenizer, 512 tokens.
tokenizer_config.json TokenizersBackend; <|end-text|>/<|actor|> as added tokens, <|actor|> as bos_token.
generation_config.json do_sample: true, temperature: 0.8, top_k: 50, bos_token_id: 511, eos_token_id: 510.

Architecture

Silia fuses attention and the SwiGLU feed-forward into one Hydra Latent Attention unit. A block is:

u, v = a1(x).chunk(2, dim=-1)   # a1: n_embd -> 2 * d_ff
y = u * F.silu(v)
return x + a2(y)                # a2: 2 * d_ff -> n_embd

There is no pre-layer-norm and a single residual. a1/a2 each split their n_head heads into two halves with opposite roles:

  • XLA (even heads) — ordinary causal attention over a KV cache, with a NeoX-style half-split RoPE on Q and K only, and a "GLA-style" output projection onto the normal value direction.
  • AFT (odd heads) — Apple's Attention Free Transformer (arXiv:2105.14103): y_t = (Σ_{i≤t} exp(k_i) v_i) / (Σ_{i≤t} exp(k_i) + 1e-6), gated by sigmoid(q).

The AFT recurrence is a pure running sum, so decoding is O(1) per step. Both halves are cached, using a hybrid cache layer that holds a full-attention KV state and two recurrent states (Σexp(k) and Σexp(k)·v).

Property Value
Parameters 524,672 (unique; lm_head is tied to wte)
Blocks 3
n_embd / head_dim 64 / 64
n_head 2 (1 XLA + 1 AFT)
d_ff (a.k.a. d_model) 256
Context 1024
Vocab 512
Checkpoint tensors 32 → 31 saved after tying

Fidelity

Verified against the original implementation — the reference graph in Silia-v2/model.py and tokenizer in Silia-v2/encoder.py — with a 25-test / 2,069-subtest parity suite (not shipped here):

  • Logits are bit-exact. max |reference − port| = 0.000e+00 with attn_implementation="sdpa", at sequence lengths 1, 7, 64 and 256, on both the cached and uncached paths, for a model built by hand from model.bin — and again at lengths 1, 3, 16 and 128 for the shipped artifacts loaded through from_pretrained. With "eager" the difference is ~2.7e-4 in practice, which is sdpa-vs-eager fp32 kernel noise amplified by exp(k); the test allows rtol=1e-2, atol=1e-3.
  • The tokenizer is id-for-id exact on the original project's prose files, on the specials, on >100-character single chunks, and on a 2,000-string mixed-script fuzz set. The port was additionally checked against an 11-file corpus (75,538 tokens) and every one of the 510 vocabulary tokens in 4 contexts.
  • Greedy decoding matches the reference token for token — asserted over 12 steps in the suite, and spot-checked to 96 steps out of band.

The tokenizer is not BPE

The reference Encoder._encode_chunk has no merge table. It walks left to right, repeatedly extending the current token while the concatenation is in the vocabulary — greedy longest-match, i.e. maximal munch. (BPE is used only to train the vocabulary, in the reference's own tokenizer-training code; encoding never consults it.)

So "the" is ['th', 'e'] here, whereas a rank-ordered BPE model would give ['t', 'he'] — a different token sequence, and therefore out-of-distribution for a model trained on the reference segmentation. HF WordPiece with an empty continuing_subword_prefix is maximal munch, so the port is exact without any custom tokenizer class. max_input_chars_per_word is raised to 10⁶ because a single GPT-4 pre-tokenizer chunk can exceed the 100-character default. unk_token is omitted: all 256 single bytes are in the vocabulary, so longest match never fails.

The pre-tokenizer is the shared GPT-4 split pattern from the reference. Note that Encoder.load() sets self.pattern but never recompiles self.compiled_pattern, so the pattern stored in o512.bin is inert; it equals GPT4_SPLIT_PATTERN anyway.

Documented differences from the reference

These are deliberate, and none of them change the logits for a given input.

  1. No per-step <|actor|> sink prepend. The reference's Silia.generate prepends the actor sink token to the front of the whole context on every step (so a growing sequence reads [<|actor|>, a, b, c, ...]). That is equivalent to keeping the token at position 0 with a KV cache, so the port does not reproduce the per-step prepend. Instead <|actor|> is registered as the BOS (bos_token/bos_token_id: 511), and the tokenizer prepends it once at encode time — the same convention Gemma, Llama and friends use, and the same single sink at position 0 that the reference produces. Pass the bare prompt. Adding <|actor|> yourself yields two sinks, exactly as '<bos>the' does under Gemma. add_special_tokens=False suppresses it if you need the bare sequence.
  2. num_hidden_layers counts attention units, not blocks. Each block has two attention units, so the config carries num_hidden_layers == 2 * num_blocks (6, for 3 blocks) to give the cache one layer per unit. The true block count is config.num_blocks, which is authoritative on reload — that is what keeps a saved config idempotent instead of doubling the layer count on every load.
  3. The AFT half ignores attention_mask. attention_free_transformer has no padding mask, because the recurrence is a cumulative sum rather than a weight-and-sum and masking it would change the state carried into the next position. The XLA half uses create_causal_mask normally. Batch only equal-length sequences, or pad on the left and slice.
  4. Special tokens are always recognized. The reference's Encoder.encode takes an allowed_special policy — the default is "none_raise", which asserts if a special token appears in the text, while "none" treats it as ordinary text. The HF tokenizer always honors added special tokens. To compare against the reference, pass allowed_special="all".
  5. The loss is shifted cross-entropy. The reference computes cross_entropy(logits, targets) with no shift, i.e. it scores position t against the label for t rather than t+1. That is not the usual autoregressive objective, so this port keeps HF's standard behaviour. The logits already match exactly, so nothing else is affected.
  6. Vestigial GPT-2 config fields. SiliaV2Config subclasses GPT2Config, so summary_*, add_cross_attention, reorder_and_upcast_attn, scale_attn_weights, layer_norm_epsilon and friends appear in config.json and are ignored. GPT-2 is used only as plumbing — the config, GenerationMixin, and v5's tied-weights convention. No GPT-2 compute is retained; the substantive math comes from LlamaRotaryEmbedding, apply_rotary_pos_emb and eager_attention_forward.
  7. A warm-up context bug in the reference's inference.py is not reproduced. Line 33 draws a random prompt as random.randint(0, len(enc.vocab) + len(enc.special_tokens)) — i.e. randint(0, 512), which can emit id 512, one past the end of the 512-token vocabulary. This port always supplies a real prompt.

Where the paper and the code disagree

The published description of Silia v2 does not match the released code. This port follows the code, since the weights were trained against it.

Paper Code (and this port)
§2.2: RoPE applied to Q, K, and V Q and K only, head_dim half-split
§2.2: RMSNorm(C) after the concat No norm after the concat
§2.3: SiLU(u) ⊙ v u * silu(v) — operands reversed
§5.1: "byte-pair encoding" Maximal munch (no merge table)
§5.1 sample output Contains a literal <|end-text|>

Two further details are easy to get wrong and are worth stating explicitly:

  • RMSNorm is gain-less, and its epsilon is torch.finfo(dtype).eps (1.19e-7 in fp32) — not 1e-5 or 1e-6. config.rms_norm_eps is therefore null by default and resolved at runtime from the activation dtype.
  • The RoPE sin term is sign-mirrored relative to HF's rotate_half. The reference rotates [x₁·cos + x₂·sin, −x₁·sin + x₂·cos] where rotate_half gives [q₁·cos − q₂·sin, q₂·cos + q₁·sin], so this port passes -sin. The discrepancy is invisible at position 0 (where sin = 0), which is exactly why it survived the first round of tests.

Implementation note

SiliaV2PreTrainedModel deliberately has no _init_weights override. from_pretrained builds the model on the meta device, and inv_freq is a non-persistent buffer, so it is absent from the checkpoint and must be recomputed during weight initialization. Overriding _init_weights replaces — rather than extends — the inherited implementation that performs that recomputation, which leaves inv_freq as uninitialized memory. The resulting model still has the right parameter count, the right weights, and the right tied head, but scrambled every forward pass. A dedicated parity test guards this.

Downloads last month
603
Safetensors
Model size
525k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for qikp/silia-v2-hf

Finetuned
(1)
this model

Dataset used to train qikp/silia-v2-hf

Space using qikp/silia-v2-hf 1

Paper for qikp/silia-v2-hf