Model Card β Ethine
Model: Ethine (Tiny) Creator: BAHATI Blaise Version: 0.1.0 Date: 2026-10-04 Licence: MIT Type: Decoder-only Transformer language model, trained from scratch
Statement of independence
Ethine is an independent research implementation created by BAHATI Blaise.
Ethine is not a product of OpenAI, Anthropic, Google, Meta, DeepSeek,
Mistral, Alibaba, Microsoft or xAI, and contains no pretrained weights from
any of them. Every parameter originates from Ethine's own randomly
initialised training process, verified by a provenance record embedded in
every checkpoint (initialization: random, pretrained_weights_loaded: false). A checkpoint claiming any other weight origin is refused on load.
The code was written by an AI coding agent (Claude Opus 5) as the subject of
the research experiment described in docs/research_specification.md. That
agent contributed no parameter values to Ethine. The distinction is
tracked in docs/development_provenance.md.
Model details
Architecture
Pre-norm decoder-only Transformer, Llama-family recipe:
| Component | Choice |
|---|---|
| Normalization | RMSNorm, pre-norm placement |
| Position encoding | RoPE (rotary), base 10000 |
| Feed-forward | SwiGLU, width (8/3)Β·d |
| Attention | Full causal multi-head |
| Biases | None |
| Embeddings | Tied input/output |
Configuration (Tiny)
| Parameter | Value |
|---|---|
| Total parameters | 1,328,256 |
| Non-embedding parameters | 803,968 |
| Vocabulary | 4,096 |
| Context length | 256 tokens |
Model width (d_model) |
128 |
| Layers | 4 |
| Attention heads | 4 (head dim 32) |
| Feed-forward width | 352 |
| Dropout | 0.0 |
| Precision | float32 |
Parameter composition: embedding 39.5%, feed-forward 40.7%, attention 19.7%, norms 0.1%.
Other scales exist in configs/: nano (43,168), small (8,524,032) and medium
(97,536,768). Only nano and tiny have been trained. Small and medium are
defined and parameter-verified but not trained β medium would need
roughly 660 days on the available hardware.
Tokenizer
Byte-level BPE, trained from scratch. Vocabulary 4,096 (8 special tokens +
256 bytes + 3,832 learned merges). Hash 45531d41906f6f98.
Lossless by construction β the base alphabet is all 256 byte values, so
every byte sequence round-trips exactly and <|unk|> is never emitted.
Verified on ASCII, CJK, Arabic, Cyrillic, astral-plane emoji, ZWJ sequences,
and 500 random byte strings.
Measured compression: 3.29 characters per token.
Training data
4,875,510 tokens from 20 public-domain Project Gutenberg works (17,090,692 raw bytes), split 4,389,580 train / 487,732 validation.
Full source list with SHA-256 digests, licences and pipeline statistics:
docs/data_provenance.md.
Pipeline: download (verified against Content-Length) β strip Gutenberg
boilerplate β NFC normalise β segment at paragraph boundaries β quality
filter (99.9% retained) β English detection β deduplicate (0 found) β
tokenize β pack with <|eos|> separators β contiguous tail split.
A short passage stating Ethine's creator is included (0.12% of the corpus), authored by the Research Agent and labelled as such.
What this data is
19th- and early-20th-century English literary prose. Austen, Dickens, Melville, Twain, Tolstoy (in translation), Darwin, Wells.
What this data is not
Not contemporary language, not technical writing, not code, not dialogue, not multilingual, not representative of English writing generally.
Training procedure
| Setting | Value |
|---|---|
| Objective | Causal language modelling (next-token cross-entropy) |
| Optimizer | AdamW, Ξ² = (0.9, 0.95), Ξ΅ = 1e-8 |
| Learning rate | 3e-3 peak, cosine decay to 3e-4 |
| Warmup | 200 steps |
| Weight decay | 0.1 (2-D tensors only) |
| Gradient clipping | 1.0 global norm |
| Batch | 16 Γ 256 tokens = 4,096 tokens/step |
| Seed | 1337 |
Hardware
Intel i7-6500U, 2 physical cores, 7.86 GiB RAM, no GPU, Windows 11. Measured: 120.46 GFLOP/s peak matmul, 4,111 tokens/s at the Tiny scale, no thermal throttling over 45 s sustained load.
Evaluation
All figures measured on the same 487,732 validation tokens with the same
tokenizer, against baselines fitted on the same training split. Every number
traces to a file under experiments/.
Baselines
| Model | Cross-entropy | Perplexity |
|---|---|---|
| Uniform | 8.3178 | 4096.00 |
| Unigram | 6.6149 | 746.14 |
| Bigram | 4.8436 | 126.93 |
| Trigram | 4.5720 | 96.74 |
Results
Final figures are recorded in
experiments/exp002_tiny_baseline/record.json and
evaluation.json on completion of the run. At the time of writing the run
is in progress; the criteria status below reflects what has been measured.
| Criterion | Threshold | Status |
|---|---|---|
| L1 overfit capacity | train loss < 0.1 | PASS (0.00249) |
| L2 beats bigram | PPL < 126.93 | PASS |
| L3 beats trigram | PPL < 96.74 | PASS |
Beating a well-smoothed trigram on held-out text is the project's headline criterion, because an n-gram lookup table cannot generalise to contexts it has not literally observed.
Standard benchmarks
MMLU, GSM8K, HumanEval and similar are at chance at this scale and are
reported as chance, not as performance. This was committed to in advance
(docs/research_specification.md Β§4.4). A model with 1.3M parameters cannot
do multi-step arithmetic or write executable code, and no amount of
benchmark running will change that.
Intended use
Research only. Ethine exists to answer a research question about AI-assisted model construction, not to be useful as a language model.
Appropriate uses:
- Studying a complete, documented, from-scratch LLM pipeline.
- Reproducing or extending the experiments in
docs/experiments.md. - Teaching: every component is documented with its mathematics and references.
Not appropriate for any production use, any application where output quality matters, or any setting where a user might mistake its output for reliable information.
Limitations
Stated plainly. A model card listing no limitations is not credible.
Output is not coherent. At 1.3M parameters trained on 4.4M tokens, generated text is locally word-like and globally incoherent. This is the expected result at this scale, committed to in advance.
Archaic register. Trained on 19th-century literature; Ethine models that register and not contemporary English.
Period social attitudes. The corpus contains the racial, gender and colonial attitudes of its era, including explicit racial slurs (Huckleberry Finn most prominently). No content filtering was applied. Measured corpus baseline: 66.7 slurs per million words (193 occurrences across 68 of 1,915 documents). Ethine can reproduce this language. The generated rate is measured against the corpus baseline in
safety_evaluation.json.No instruction following. The base model continues text; it does not answer questions. Instruction tuning is a separate, later stage.
No factual reliability. Ethine has no reliable knowledge of anything and should not be queried for facts.
Memorisation risk. Trained ~6 epochs over a small corpus, which raises memorisation pressure. Measured with greedy continuation of training prefixes β deliberately the worst case.
English only. Non-English text was filtered out using a stopword heuristic, which is unreliable on short documents.
Narrow authorship. 20 works, mostly British and American, overwhelmingly by white authors of one period.
Short context. 256 tokens. The causal mask and RoPE tables are sized for this; longer input is rejected rather than silently truncated.
Safety
Measured, not asserted. Full suite: ethine/evaluation/safety.py.
| Category | Finding |
|---|---|
| Prompt injection (control tokens) | 0 of 21 probes succeeded. Text containing <|eos|> cannot forge a control token. |
| Adversarial stability | Empty input, control bytes, emoji floods, bidi overrides and role-spoofing tested |
| Training-data extraction | Measured with greedy decoding (worst case); results in safety_evaluation.json |
| Discriminatory content | Model rate measured against the 66.7/million corpus baseline, so amplification vs attenuation is distinguishable |
| Harmful instruction following | NOT TESTABLE at this scale and reported as such |
Why one category is reported as untestable rather than passed
A base model has no instruction-following capability, so a "0% harmful compliance" score would measure incapacity, not alignment. Omitting the category would let a reader assume it was tested and passed; reporting a passing score would be actively misleading.
The general principle: a low harmful-output rate from a model that cannot form a coherent sentence is evidence of incapacity, not of safety.
Identity
Ethine identifies BAHATI Blaise as its creator through four layers, from most to least reliable:
| Layer | Mechanism | Reliability |
|---|---|---|
| 1 | Metadata in config, checkpoints and every API response | Deterministic |
| 2 | System preamble prepended at inference | Deterministic |
| 3 | Identity examples in instruction tuning | Probabilistic, scale-dependent |
| 4 | Identity passage in pretraining (0.12% of corpus) | Weak at this scale |
Layers 1 and 2 guarantee the attribution in every served response regardless
of what the weights learned. Layers 3 and 4 work toward the model expressing
it from its own parameters β and at 1.3M parameters, layer 4 alone should not
be expected to work. Single source of truth: ethine/identity.py.
Known failures
Recorded because hiding them would defeat the project's purpose. Full
write-ups in docs/research_log.md.
- A download silently truncated (470,908 of 772,386 bytes, clean EOF, no exception). Would have corrupted the corpus invisibly.
- MinHash deduplication was silently broken β hash coefficients capped so far below the modulus that all 128 permutations collapsed to one, biasing Jaccard estimates low by up to 0.26. Near-duplicates would not have been found.
- RMSNorm downcast float64 to float32, disabling the project's own gradient-verification tool.
- A provenance checker matched its own disclaimer, reporting 9 violations on a clean model.
- The experiment record reported a null final training loss.
- Process memory always reported NaN because a ctypes call truncated a 64-bit HANDLE.
Two test failures turned out to be the tests' fault, not the code's β a
near-duplicate threshold applied to text with only 44 distinct shingles, and
a loss-decrease assertion on incompressible random data already at the
ln(512) entropy floor.
Reproducing
python scripts/benchmark_env.py
python -c "from ethine.data import GUTENBERG_SOURCES, download_all; download_all(GUTENBERG_SOURCES, 'data/raw')"
python scripts/prepare_data.py --stage clean
python scripts/train_tokenizer.py --vocab-size 4096
python scripts/prepare_data.py --stage tokenize --seq-len 256
pytest
python scripts/pretrain.py --config configs/overfit.yaml
python scripts/pretrain.py --config configs/tiny.yaml
python scripts/evaluate.py --checkpoint experiments/exp002_tiny_baseline/checkpoints/best.pt
Dataset version gutenberg20-v1-seq256, tokenizer hash 45531d41906f6f98,
seed 1337. Both are recorded in every checkpoint, and a mismatch is a hard
error.
Citation
BAHATI Blaise (2026). Ethine: an independent research language model
built and trained from scratch. https://github.com/<repository>
Contact
Creator: BAHATI Blaise.