ÇABA 50M — Kumru

An experimental ÇABA v0-core causal language model with about 49.96M trainable parameters.

This base model continues from Efe2898/caba-kumru-50m. It is not instruction-tuned or a chat assistant.

Architecture

  • No self-attention layers. Each block uses a causal depthwise convolution and three recurrent associative matrix banks.
  • Fast bank writes each token; medium and slow banks receive fixed-window, delayed promotions.
  • The v0 promotion score uses a stop-gradient normalized association residual as unresolvedness. It is not a learned future-utility estimate.
  • Medium/slow content selection and admission are separate. Soft admission does not claim FLOP skipping.
  • The prototype uses an explicit delta-rule update. It is not claimed to be identical to the official Gated DeltaNet-2 kernel.
  • Training config: hidden 512, 8 blocks, 8 heads, 1280 SwiGLU width, 8/64-token promotion clocks.

Data and tokenizer

  • Tokenizer: vngrs-ai/Kumru-2B-Base at revision 55711ea224e4bf5d4e11a4baf79130ae73785ece (vocabulary 50,176; EOS/packed separator ID 3).
  • Dataset: Efe2898/tokenized at revision 6cb993aa63064cd89b511bb164e8c6e5512902ed, raw little-endian uint16 token shards.
  • Manifest target for Turkish is currently 0. This run included unmanifested shards: true; selected sources/weights and shard paths are recorded in training/training_config.json.
  • The dataset repository does not declare a redistribution license in its card. This model repo is public; confirm data rights before sharing it publicly.

Use with Transformers

The repo includes the custom configuration and model implementation. Loading custom Hub code requires trust_remote_code=True; review the code before enabling it.

from transformers import AutoTokenizer, AutoModelForCausalLM

repo_id = "Efe2898/caba-v2"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)
prompt = "Türkiye'de bilim ve teknoloji"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.9)
print(tokenizer.decode(output[0], skip_special_tokens=True))

The custom .generate() supports greedy decoding and top-k/top-p sampling. This first model is a research checkpoint, not a quality or speed guarantee. Please retain the provenance and access conditions of the training data.

Training record

Training and evaluation logs, the exact configuration, and the Colab training entry point are in training/.

Downloads last month
178
Safetensors
Model size
50M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support