Wasl-1 — وَصَل
A compact experimental Transformer for Egyptian cash-on-delivery order workflows.
Wasl-1 explores modeling order events, delivery-address descriptions, and customer communication signals for e-commerce fulfillment support. It uses a custom 10.16-million-parameter causal language model trained from scratch on template-generated Egyptian Arabic text.
The intended application is to turn order events into a structured state containing a delivery-intent score, return-risk tier, and fulfillment route. The supplied training notebook does not establish reliable structured-state generation or real-world return-risk prediction. Its 256-character truncation can remove the entire target state before training.
- Organization: Racore
- Model repository: racore/Wasl-1
- Notebook name:
WASAL_1_Colab_Notebook.ipynb - Implementation: custom PyTorch
SiyaqFamilyTransformer - Status: research prototype
This card documents the supplied notebook and its saved outputs. The current remote repository contents and checkpoint were not independently verified while preparing this card. File names below follow the notebook's export function.
Model details
| Property | Value |
|---|---|
| Parameters | 10,163,520, as reported in the notebook |
| Architecture | Decoder-only causal Transformer |
| Layers | 8 |
| Hidden dimension | 320 |
| Attention heads | 8 |
| Head dimension | 40 |
| Feed-forward dimension | 1,088 |
| Activation | GELU |
| Normalization | RMSNorm |
| Position encoding | RoPE, theta = 500,000 |
| Embedding vocabulary capacity | 4,096 |
| Fitted tokenizer entries | 94 in the saved run |
| Tokenization | Character-level lookup |
| Weight tying | Input embeddings and output projection share weights |
| Training sequence cap | 256 characters; 255 next-token input positions |
| Configured maximum positions | 262,144; not a validated context length |
FastByteTokenizer is the notebook's class name, but its implementation iterates over Unicode characters, not UTF-8 bytes. Strings such as <bos>, <eos>, [EVENT], and [STATE] are encoded character by character rather than recognized as special tokens. Unknown characters map to <unk>.
The configuration includes dropout=0.05, but the model does not apply a dropout layer or pass a nonzero attention dropout probability. Attention is dense and causal. No KV cache, streaming-memory mechanism, or long-context benchmark is implemented. Dataset streaming is separate from model context support.
Intended input and output
Training examples are serialized as:
<bos>[EVENT] {event_1}
[EVENT] {event_2}
[EVENT] {event_3}
[EVENT] {event_4}
[STATE]
{state_json}<eos>
The notebook inserts a space before the newline between events. To reproduce that formatting:
event_text = " \n".join(f"[EVENT] {event}" for event in events)
prompt = f"<bos>{event_text}\n[STATE]\n"
Intended Arabic state fields:
| Field | Meaning |
|---|---|
اسم_العميل |
Customer or merchant name |
المنطقة_والحي |
Region and district |
العنوان_المفصل |
Address details |
قيمة_الشحنة |
Order value |
استجابة_العميل |
Communication signal |
مؤشر_جدية_الاستلام |
Synthetic intent score |
مستوى_مخاطرة_المرتجع |
Synthetic risk tier |
مسار_التنفيذ_المعتمد |
Route label selected by the generator |
القرار_التشغيلي_الملزم |
Templated operational action |
These are dataset targets, not demonstrated model outputs. A generated percentage is not a calibrated probability of successful delivery.
Usage
The export function saves a PyTorch state dictionary, a configuration, and a custom vocabulary. It does not export a Transformers AutoModel integration, auto_map, tokenizer implementation, or .generate() method. Do not assume that pipeline() or AutoModelForCausalLM.from_pretrained() can load this export.
Install:
pip install torch huggingface_hub
1. Define the checkpoint-compatible architecture
Run this block once in Python or Colab. It reproduces the notebook's model definition without running training or publishing.
import torch
import torch.nn as nn
import torch.nn.functional as F
from dataclasses import dataclass
@dataclass
class SiyaqFamilyConfig:
vocab_size: int = 4096
d_model: int = 320
n_layers: int = 8
n_heads: int = 8
d_ff: int = 1088
max_position_embeddings: int = 262144 # Configuration value; long-context quality is unverified
rope_theta: float = 500000.0
dropout: float = 0.05
learning_rate: float = 3e-4
weight_decay: float = 0.01
class RMSNorm(nn.Module):
def __init__(self, dim: int, eps: float = 1e-6):
super().__init__()
self.eps = eps
self.weight = nn.Parameter(torch.ones(dim))
def forward(self, x):
return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps) * self.weight
class RotaryEmbedding(nn.Module):
def __init__(self, dim: int, theta: float = 500000.0):
super().__init__()
inv_freq = 1.0 / (theta ** (torch.arange(0, dim, 2).float() / dim))
self.register_buffer("inv_freq", inv_freq, persistent=False)
def forward(self, x, seq_len: int):
t = torch.arange(seq_len, device=x.device, dtype=self.inv_freq.dtype)
freqs = torch.outer(t, self.inv_freq)
emb = torch.cat((freqs, freqs), dim=-1)
return emb.cos()[None, None, :, :], emb.sin()[None, None, :, :]
def apply_rotary_pos_emb(q, k, cos, sin):
def rotate_half(x):
return torch.cat((-x[..., x.shape[-1] // 2:], x[..., :x.shape[-1] // 2]), dim=-1)
cos = cos[:, :, :q.shape[2], :]
sin = sin[:, :, :q.shape[2], :]
return (q * cos) + (rotate_half(q) * sin), (k * cos) + (rotate_half(k) * sin)
class CausalSelfAttention(nn.Module):
def __init__(self, config: SiyaqFamilyConfig):
super().__init__()
self.n_heads = config.n_heads
self.head_dim = config.d_model // config.n_heads
self.q_proj = nn.Linear(config.d_model, config.d_model, bias=False)
self.k_proj = nn.Linear(config.d_model, config.d_model, bias=False)
self.v_proj = nn.Linear(config.d_model, config.d_model, bias=False)
self.out_proj = nn.Linear(config.d_model, config.d_model, bias=False)
def forward(self, x, cos, sin):
B, S, C = x.shape
q = self.q_proj(x).view(B, S, self.n_heads, self.head_dim).transpose(1, 2)
k = self.k_proj(x).view(B, S, self.n_heads, self.head_dim).transpose(1, 2)
v = self.v_proj(x).view(B, S, self.n_heads, self.head_dim).transpose(1, 2)
q, k = apply_rotary_pos_emb(q, k, cos, sin)
out = F.scaled_dot_product_attention(q, k, v, is_causal=True)
return self.out_proj(out.transpose(1, 2).contiguous().view(B, S, C))
class TransformerBlock(nn.Module):
def __init__(self, config: SiyaqFamilyConfig):
super().__init__()
self.norm1 = RMSNorm(config.d_model)
self.attn = CausalSelfAttention(config)
self.norm2 = RMSNorm(config.d_model)
self.mlp = nn.Sequential(
nn.Linear(config.d_model, config.d_ff, bias=False),
nn.GELU(),
nn.Linear(config.d_ff, config.d_model, bias=False)
)
def forward(self, x, cos, sin):
x = x + self.attn(self.norm1(x), cos, sin)
return x + self.mlp(self.norm2(x))
class SiyaqFamilyTransformer(nn.Module):
def __init__(self, config: SiyaqFamilyConfig):
super().__init__()
self.config = config
self.embed = nn.Embedding(config.vocab_size, config.d_model)
self.rope = RotaryEmbedding(config.d_model // config.n_heads, config.rope_theta)
self.layers = nn.ModuleList([TransformerBlock(config) for _ in range(config.n_layers)])
self.norm = RMSNorm(config.d_model)
self.lm_head = nn.Linear(config.d_model, config.vocab_size, bias=False)
self.lm_head.weight = self.embed.weight
def forward(self, input_ids):
B, S = input_ids.shape
cos, sin = self.rope(input_ids, S)
x = self.embed(input_ids)
for layer in self.layers:
x = layer(x, cos, sin)
return self.lm_head(self.norm(x))
2. Load the exported checkpoint and run a completion
Run this after the architecture block in the same session. Public repositories need no token; authenticate separately if access is restricted. These download calls require the three exported files to exist in the repository.
import json
from dataclasses import fields
from huggingface_hub import hf_hub_download
repo_id = "racore/Wasl-1"
# For reproducibility, replace "main" with a verified repository commit SHA.
revision = "main"
def download(name):
return hf_hub_download(repo_id=repo_id, filename=name, revision=revision)
with open(download("config.json"), encoding="utf-8") as f:
raw_config = json.load(f)
allowed = {f.name for f in fields(SiyaqFamilyConfig)}
config = SiyaqFamilyConfig(**{k: v for k, v in raw_config.items() if k in allowed})
with open(download("vocab.json"), encoding="utf-8") as f:
c2i = json.load(f)
i2c = {int(v): k for k, v in c2i.items()}
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = SiyaqFamilyTransformer(config)
state_dict = torch.load(download("pytorch_model.bin"), map_location="cpu", weights_only=True)
model.load_state_dict(state_dict, strict=True)
model.to(device).eval()
def encode(text):
# Match the notebook exactly: character lookup, not byte encoding.
return [c2i.get(c, c2i["<unk>"]) for c in text]
def decode(ids):
return "".join(i2c.get(int(i), "") for i in ids if int(i) not in (0, 1, 2, 3))
@torch.inference_mode()
def complete(prompt, max_new_tokens=64):
ids = encode(prompt)
if not ids:
raise ValueError("Prompt must not be empty")
if len(ids) + max_new_tokens > 256:
raise ValueError("This smoke test stays within the training length of 256 characters")
x = torch.tensor([ids], dtype=torch.long, device=device)
generated = []
for _ in range(max_new_tokens):
token = int(model(x)[0, -1].argmax().item())
if token == c2i["<eos>"]:
break
generated.append(token)
x = torch.cat([x, torch.tensor([[token]], device=device)], dim=1)
if decode(generated).endswith("<eos>"):
break
return decode(generated)
# Text-completion smoke test; this is not a validated delivery-risk prediction.
prompt = "<bos>[EVENT] طلب شراء جديد"
print(complete(prompt))
The example demonstrates checkpoint loading and greedy text completion. It does not supply fabricated expected results. The full four-event prompt typically exceeds the training length; increasing the guard is not evidence of reliable long-context inference. This simple loop recomputes the prefix on each step and is not a latency-optimized serving implementation.
Training data
The notebook generates synthetic examples from predefined pools. Saved output reports 97 city/district entries, 104 merchant/customer names, 35 address descriptions, and 30 communication signals.
The executed training code requests 40,000 training examples and 10,000 validation examples through generate_sample() with its default Arabic setting. A separate stream_dataset() supports alternating Arabic and English and defaults to one million generated records, but it is not used by the shown training loop. Therefore this card does not claim one-million-example training or validated English support. Random combinations are not guaranteed unique.
Labels are heuristic: the generator multiplies a signal's base score by an address weight, converts to an integer, and clips to 5–99. Scores at least 75 receive the low-risk tier; scores 45–74 receive the medium-risk tier; lower scores receive the critical tier. Route labels come separately from the signal template and may not agree with the score-derived action.
The fourth input event explicitly includes the target score. Consequently, full-example evaluation could reward copying the score rather than inferring it. Training and validation use the same template pools, and no duplicate removal or disjoint-template split is implemented.
Training configuration
| Setting | Value |
|---|---|
| Objective | Next-character cross-entropy over the serialized sequence |
| Optimizer | AdamW |
| Initial learning rate | 0.0003 |
| Weight decay | 0.01 |
| Batch size | 8 |
| Configured epochs | 3 |
| Schedule | Cosine annealing, minimum LR 0.00001 |
| Gradient clipping | 1.0 |
| Padding loss | Ignored at token ID 0 |
| Seed | 42 |
| Checkpoint selection | Lowest mean validation batch loss |
The objective is not masked to the state response. The tokenizer is fitted on only the first 500 training examples, excluding the explicit serialization wrapper.
Recorded evaluation
The attached notebook contains these saved validation results; they were not rerun for this card:
| Epoch | Validation next-character cross-entropy |
|---|---|
| 1 | 0.1032 |
| 2 | 0.0980 |
Although the loop is configured for three epochs, the supplied output does not show an epoch-three result. Validation loss is averaged over batches of truncated examples. It is not delivery accuracy, JSON validity, route accuracy, or a return-rate reduction measurement.
Critical training limitation: examples are truncated to 256 characters from the beginning. An inspection of 100 newly generated examples placed the start of the JSON target after approximately 388–433 characters. All those inspected examples lost the entire target state at the configured truncation length. Low validation loss can therefore reflect predictable event prefixes rather than learning to generate the intended decisions.
No measured evidence is provided for 96.5% accuracy, 12–15% returns, sub-5-ms inference, per-order cost, or reliable 250k-token processing. The notebook's comparison table contains manually written claims, not benchmark output; they are intentionally excluded as performance results here.
Limitations and next evaluation steps
- Preserve the complete input and target during training, and measure target-token coverage before retraining.
- Remove the precomputed target score from the input when evaluating inference from raw customer signals.
- Hold out templates and entities, remove duplicates across splits, and evaluate on consented, de-identified real order outcomes.
- Report JSON validity, field accuracy, route macro-F1, score error/calibration, and measured latency on named hardware.
- Audit errors across regions and address styles: handcrafted address weights can encode geographic or socioeconomic bias.
- Keep shipment cancellation, deposit requests, and other customer-impacting actions under human review until validated. The model itself does not contact customers, collect deposits, or dispatch parcels.
Exported files
| File | Purpose |
|---|---|
pytorch_model.bin |
Exported state dictionary |
config.json |
Custom architecture configuration |
vocab.json |
Character-to-ID dictionary |
README.md |
Model card |
The notebook exporter selects checkpoints/best_model.pt if present; otherwise it exports the current in-memory model. These are the expected export files, not an independently verified remote inventory.
License
The supplied notebook declares Apache-2.0. This card retains that declaration. Include the corresponding license text in the repository for distribution.
بالعربي
وَصَل تجربة بحثية لموديل صغير يخدم إجراءات طلبات الدفع عند الاستلام في مصر. الهدف هو تحويل أحداث الطلب ورسائل العميل إلى حالة منظمة تساعد فريق التشغيل. النسخة الموثقة هنا اتدربت على نصوص صناعية عربية، وفيها مشكلة قصّ العينات قبل الوصول للإجابة المستهدفة. لذلك نتائج الـloss الحالية لا تثبت دقة تقييم جدية العميل، ولا انخفاض المرتجعات، ولا دعم سياق 250 ألف رمز. مثال الاستخدام أعلاه لتحميل الأوزان وتجربة إكمال النص، ويحتاج تقييمًا وإعادة تدريب صحيحة قبل استخدامه لاتخاذ قرارات تشغيلية.
- Downloads last month
- -