Model Card β€” Ethine

Model: Ethine (Tiny) Creator: BAHATI Blaise Version: 0.1.0 Date: 2026-10-04 Licence: MIT Type: Decoder-only Transformer language model, trained from scratch


Statement of independence

Ethine is an independent research implementation created by BAHATI Blaise.

Ethine is not a product of OpenAI, Anthropic, Google, Meta, DeepSeek, Mistral, Alibaba, Microsoft or xAI, and contains no pretrained weights from any of them. Every parameter originates from Ethine's own randomly initialised training process, verified by a provenance record embedded in every checkpoint (initialization: random, pretrained_weights_loaded: false). A checkpoint claiming any other weight origin is refused on load.

The code was written by an AI coding agent (Claude Opus 5) as the subject of the research experiment described in docs/research_specification.md. That agent contributed no parameter values to Ethine. The distinction is tracked in docs/development_provenance.md.


Model details

Architecture

Pre-norm decoder-only Transformer, Llama-family recipe:

Component Choice
Normalization RMSNorm, pre-norm placement
Position encoding RoPE (rotary), base 10000
Feed-forward SwiGLU, width (8/3)Β·d
Attention Full causal multi-head
Biases None
Embeddings Tied input/output

Configuration (Tiny)

Parameter Value
Total parameters 1,328,256
Non-embedding parameters 803,968
Vocabulary 4,096
Context length 256 tokens
Model width (d_model) 128
Layers 4
Attention heads 4 (head dim 32)
Feed-forward width 352
Dropout 0.0
Precision float32

Parameter composition: embedding 39.5%, feed-forward 40.7%, attention 19.7%, norms 0.1%.

Other scales exist in configs/: nano (43,168), small (8,524,032) and medium (97,536,768). Only nano and tiny have been trained. Small and medium are defined and parameter-verified but not trained β€” medium would need roughly 660 days on the available hardware.

Tokenizer

Byte-level BPE, trained from scratch. Vocabulary 4,096 (8 special tokens + 256 bytes + 3,832 learned merges). Hash 45531d41906f6f98.

Lossless by construction β€” the base alphabet is all 256 byte values, so every byte sequence round-trips exactly and <|unk|> is never emitted. Verified on ASCII, CJK, Arabic, Cyrillic, astral-plane emoji, ZWJ sequences, and 500 random byte strings.

Measured compression: 3.29 characters per token.


Training data

4,875,510 tokens from 20 public-domain Project Gutenberg works (17,090,692 raw bytes), split 4,389,580 train / 487,732 validation.

Full source list with SHA-256 digests, licences and pipeline statistics: docs/data_provenance.md.

Pipeline: download (verified against Content-Length) β†’ strip Gutenberg boilerplate β†’ NFC normalise β†’ segment at paragraph boundaries β†’ quality filter (99.9% retained) β†’ English detection β†’ deduplicate (0 found) β†’ tokenize β†’ pack with <|eos|> separators β†’ contiguous tail split.

A short passage stating Ethine's creator is included (0.12% of the corpus), authored by the Research Agent and labelled as such.

What this data is

19th- and early-20th-century English literary prose. Austen, Dickens, Melville, Twain, Tolstoy (in translation), Darwin, Wells.

What this data is not

Not contemporary language, not technical writing, not code, not dialogue, not multilingual, not representative of English writing generally.


Training procedure

Setting Value
Objective Causal language modelling (next-token cross-entropy)
Optimizer AdamW, Ξ² = (0.9, 0.95), Ξ΅ = 1e-8
Learning rate 3e-3 peak, cosine decay to 3e-4
Warmup 200 steps
Weight decay 0.1 (2-D tensors only)
Gradient clipping 1.0 global norm
Batch 16 Γ— 256 tokens = 4,096 tokens/step
Seed 1337

Hardware

Intel i7-6500U, 2 physical cores, 7.86 GiB RAM, no GPU, Windows 11. Measured: 120.46 GFLOP/s peak matmul, 4,111 tokens/s at the Tiny scale, no thermal throttling over 45 s sustained load.


Evaluation

All figures measured on the same 487,732 validation tokens with the same tokenizer, against baselines fitted on the same training split. Every number traces to a file under experiments/.

Baselines

Model Cross-entropy Perplexity
Uniform 8.3178 4096.00
Unigram 6.6149 746.14
Bigram 4.8436 126.93
Trigram 4.5720 96.74

Results

Final figures are recorded in experiments/exp002_tiny_baseline/record.json and evaluation.json on completion of the run. At the time of writing the run is in progress; the criteria status below reflects what has been measured.

Criterion Threshold Status
L1 overfit capacity train loss < 0.1 PASS (0.00249)
L2 beats bigram PPL < 126.93 PASS
L3 beats trigram PPL < 96.74 PASS

Beating a well-smoothed trigram on held-out text is the project's headline criterion, because an n-gram lookup table cannot generalise to contexts it has not literally observed.

Standard benchmarks

MMLU, GSM8K, HumanEval and similar are at chance at this scale and are reported as chance, not as performance. This was committed to in advance (docs/research_specification.md Β§4.4). A model with 1.3M parameters cannot do multi-step arithmetic or write executable code, and no amount of benchmark running will change that.


Intended use

Research only. Ethine exists to answer a research question about AI-assisted model construction, not to be useful as a language model.

Appropriate uses:

  • Studying a complete, documented, from-scratch LLM pipeline.
  • Reproducing or extending the experiments in docs/experiments.md.
  • Teaching: every component is documented with its mathematics and references.

Not appropriate for any production use, any application where output quality matters, or any setting where a user might mistake its output for reliable information.


Limitations

Stated plainly. A model card listing no limitations is not credible.

  1. Output is not coherent. At 1.3M parameters trained on 4.4M tokens, generated text is locally word-like and globally incoherent. This is the expected result at this scale, committed to in advance.

  2. Archaic register. Trained on 19th-century literature; Ethine models that register and not contemporary English.

  3. Period social attitudes. The corpus contains the racial, gender and colonial attitudes of its era, including explicit racial slurs (Huckleberry Finn most prominently). No content filtering was applied. Measured corpus baseline: 66.7 slurs per million words (193 occurrences across 68 of 1,915 documents). Ethine can reproduce this language. The generated rate is measured against the corpus baseline in safety_evaluation.json.

  4. No instruction following. The base model continues text; it does not answer questions. Instruction tuning is a separate, later stage.

  5. No factual reliability. Ethine has no reliable knowledge of anything and should not be queried for facts.

  6. Memorisation risk. Trained ~6 epochs over a small corpus, which raises memorisation pressure. Measured with greedy continuation of training prefixes β€” deliberately the worst case.

  7. English only. Non-English text was filtered out using a stopword heuristic, which is unreliable on short documents.

  8. Narrow authorship. 20 works, mostly British and American, overwhelmingly by white authors of one period.

  9. Short context. 256 tokens. The causal mask and RoPE tables are sized for this; longer input is rejected rather than silently truncated.


Safety

Measured, not asserted. Full suite: ethine/evaluation/safety.py.

Category Finding
Prompt injection (control tokens) 0 of 21 probes succeeded. Text containing <|eos|> cannot forge a control token.
Adversarial stability Empty input, control bytes, emoji floods, bidi overrides and role-spoofing tested
Training-data extraction Measured with greedy decoding (worst case); results in safety_evaluation.json
Discriminatory content Model rate measured against the 66.7/million corpus baseline, so amplification vs attenuation is distinguishable
Harmful instruction following NOT TESTABLE at this scale and reported as such

Why one category is reported as untestable rather than passed

A base model has no instruction-following capability, so a "0% harmful compliance" score would measure incapacity, not alignment. Omitting the category would let a reader assume it was tested and passed; reporting a passing score would be actively misleading.

The general principle: a low harmful-output rate from a model that cannot form a coherent sentence is evidence of incapacity, not of safety.


Identity

Ethine identifies BAHATI Blaise as its creator through four layers, from most to least reliable:

Layer Mechanism Reliability
1 Metadata in config, checkpoints and every API response Deterministic
2 System preamble prepended at inference Deterministic
3 Identity examples in instruction tuning Probabilistic, scale-dependent
4 Identity passage in pretraining (0.12% of corpus) Weak at this scale

Layers 1 and 2 guarantee the attribution in every served response regardless of what the weights learned. Layers 3 and 4 work toward the model expressing it from its own parameters β€” and at 1.3M parameters, layer 4 alone should not be expected to work. Single source of truth: ethine/identity.py.


Known failures

Recorded because hiding them would defeat the project's purpose. Full write-ups in docs/research_log.md.

  • A download silently truncated (470,908 of 772,386 bytes, clean EOF, no exception). Would have corrupted the corpus invisibly.
  • MinHash deduplication was silently broken β€” hash coefficients capped so far below the modulus that all 128 permutations collapsed to one, biasing Jaccard estimates low by up to 0.26. Near-duplicates would not have been found.
  • RMSNorm downcast float64 to float32, disabling the project's own gradient-verification tool.
  • A provenance checker matched its own disclaimer, reporting 9 violations on a clean model.
  • The experiment record reported a null final training loss.
  • Process memory always reported NaN because a ctypes call truncated a 64-bit HANDLE.

Two test failures turned out to be the tests' fault, not the code's β€” a near-duplicate threshold applied to text with only 44 distinct shingles, and a loss-decrease assertion on incompressible random data already at the ln(512) entropy floor.


Reproducing

python scripts/benchmark_env.py
python -c "from ethine.data import GUTENBERG_SOURCES, download_all; download_all(GUTENBERG_SOURCES, 'data/raw')"
python scripts/prepare_data.py --stage clean
python scripts/train_tokenizer.py --vocab-size 4096
python scripts/prepare_data.py --stage tokenize --seq-len 256
pytest
python scripts/pretrain.py --config configs/overfit.yaml
python scripts/pretrain.py --config configs/tiny.yaml
python scripts/evaluate.py --checkpoint experiments/exp002_tiny_baseline/checkpoints/best.pt

Dataset version gutenberg20-v1-seq256, tokenizer hash 45531d41906f6f98, seed 1337. Both are recorded in every checkpoint, and a mismatch is a hard error.


Citation

BAHATI Blaise (2026). Ethine: an independent research language model
built and trained from scratch. https://github.com/<repository>

Contact

Creator: BAHATI Blaise.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support