Instructions to use omurberaisik/NoTokenLM-Gen-4.6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use omurberaisik/NoTokenLM-Gen-4.6 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="omurberaisik/NoTokenLM-Gen-4.6", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("omurberaisik/NoTokenLM-Gen-4.6", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use omurberaisik/NoTokenLM-Gen-4.6 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "omurberaisik/NoTokenLM-Gen-4.6" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "omurberaisik/NoTokenLM-Gen-4.6", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/omurberaisik/NoTokenLM-Gen-4.6
- SGLang
How to use omurberaisik/NoTokenLM-Gen-4.6 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "omurberaisik/NoTokenLM-Gen-4.6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "omurberaisik/NoTokenLM-Gen-4.6", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "omurberaisik/NoTokenLM-Gen-4.6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "omurberaisik/NoTokenLM-Gen-4.6", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use omurberaisik/NoTokenLM-Gen-4.6 with Docker Model Runner:
docker model run hf.co/omurberaisik/NoTokenLM-Gen-4.6
- NoTokenLM-Gen-4.6
- Usage
- What is this checkpoint, actually?
- How well does it actually write? (1,000-prompt evaluation)
- Pronoun agreement, measured automatically
- Real-world benchmark results
- Gen-4.6 vs Gen-3.6 at a glance
- What it knows, and what it doesn't
- Language behavior
- What it's actually good at
- What it's not good at, and why
- How to actually run this thing
- Architecture details
- Training data
- Method notes
- Not run
- License
- Usage
NoTokenLM-Gen-4.6
A 36.8-million-parameter, byte-level, tokenizer-free language model trained on an English text mix that includes about 20% math. No subword vocabulary, no BPE — just raw UTF-8 bytes in, raw UTF-8 bytes out.
This is part of the NoTokenLM family: a series of small models built around one guiding question — how much can a genuinely small model do, if the architecture and training are done carefully, without leaning on scale to cover for weak design?
If you're looking for a model that answers questions, does math, or holds a conversation — this isn't that, and this card will tell you exactly why not. If you're curious what a ~37M-parameter byte-level transformer does at an early checkpoint of a broad English + math mix — keep reading.
Read this first. At this checkpoint Gen-4.6 is the larger model but not the better one. It is clearly weaker than the 13.9M-parameter Gen-3.6 at short story continuation (53.7% vs 78.0% fully coherent) and statistically indistinguishable from it on BLiMP and LAMBADA. The comparison section below has the numbers and the most likely reasons.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO_ID = "omurberaisik/NoTokenLM-Gen-4.6"
model = AutoModelForCausalLM.from_pretrained(REPO_ID, trust_remote_code=True).eval()
tokenizer = AutoTokenizer.from_pretrained(REPO_ID, trust_remote_code=True)
text = model.generate_bytes("Once upon a time, there was a little girl named Mia.",
max_new_tokens=200, temperature=0.5, top_k=40)
print(text)
trust_remote_code=True is required — this model ships its own architecture code (byte-level I/O, RoPE, RMSNorm, SwiGLU, plus a short input convolution, QK-norm, value residual and per-head output gates) alongside the weights. There is no subword tokenizer; AutoTokenizer here is a thin byte<->id mapping provided purely so the model loads the standard Hugging Face way.
Recommended sampling settings are in How to actually run this thing below.
What is this checkpoint, actually?
Gen-4.6 is a from-scratch pretraining run of a 36.8M-parameter byte-level transformer (15 layers, d_model=448). This card describes step 9,000, and every number in it comes from that checkpoint's EMA (exponential moving average) weights.
Training volume: 9,000 steps × (batch 8 × gradient accumulation 8 × 1,024 bytes) ≈ 0.59 GB of text. The best validation loss recorded during the run was 0.9233 nats per byte (≈1.33 bits per byte) on a held-out split of the training mix. That split comes from the rolling-chunk mix of this run, so it is not comparable to Gen-3.6's 0.8915 (different data). Pretraining only — no instruction tuning.
Optimizer. A hybrid: Muon (lr 0.02, momentum 0.95, Nesterov, 5 Newton–Schulz steps) on the attention and feed-forward weight matrices, and AdamW on everything else (embedding lr 0.2 with no weight decay; norms, gates, convolution and scalars lr 3e-4 with weight decay 0.1; betas 0.9 / 0.95). 500 warmup steps, plateau-based learning-rate decay (halve the rate if the smoothed loss does not improve by 0.1 over 150 steps, floored at 5% of peak), gradient clipping at 1.0, EMA decay 0.999. Data was streamed in rolling 500 MB chunks.
This is the first NoTokenLM release trained with Muon, and it also changes the data mix (adds math, drops French), so the comparison with earlier cards is not a controlled experiment.
How well does it actually write? (1,000-prompt evaluation)
We generated 1,000 completions — 1,000 distinct prompts, one generation each, temperature 0.5, top-k 40, 35 new bytes — and every one of the 1,000 outputs was read and sorted into one of three categories. The prompt set, prompt order, ids and generation settings are identical to those used for Gen-3.6 and Gen-4.5. No filtering, no cherry-picking.
Grading criteria (strict: "grammatically fine but says nothing coherent" does not count as a win):
| Category | Definition | Gen-4.6 | Share | Gen-3.6 |
|---|---|---|---|---|
| Fully coherent | Correct grammar and the sentence makes sense — characters, objects and pronouns tracked correctly. | 537 | 53.7% | 78.0% |
| Grammar OK, meaning breaks down | Well-formed sentences that drift into a pronoun/gender mismatch, a non-sequitur detail, or an object inconsistent with the scene. | 413 | 41.3% | 21.5% |
| Grammar breaks down | The sentence structure itself collapses — repetition, a garbled or malformed clause. | 50 | 5.0% | 0.5% |
Result: 53.7% fully coherent, 95.0% grammatically correct overall (Gen-3.6: 78.0% and 99.5%).
Both categories of failure grew: more outputs have intact grammar but broken meaning, and ten times as many have broken grammar. The document-end byte was sampled and re-drawn in 6 of 1,000 outputs.
The label for every output is in gen46_9000_1000_test_outputs.json under the category field.
Real examples, unedited
FULLY COHERENT
"Just before sunset, the young fox was cleaning up at the market after a long day."
-> " The young fox was so happy that he"
"One cold winter day, Sadie planted a small seed in the meadow."
-> " She was so excited to explore the "
GRAMMAR OK, MEANING BREAKS DOWN
"One sunny morning, Silas was looking for a bag of apples in the backyard."
-> " She wanted to bake her apples to s" <- pronoun doesn't match
"Without any warning, Owen spent the whole morning in the toy store."
-> " During the warning, Owen tried to " <- picks up a prompt word, turns it into nonsense
GRAMMAR BREAKS DOWN
"Grace was trying to fix a candle, while the kind turtle noticed a bright kite was missing."
-> " The kind kind of the kind was so c"
"That very night, Owen was cleaning up on the farm road after a long day."
-> " The owen was detected and the comp" <- a name becomes a common noun
"June was walking in the meadow and started to sing."
-> " The June was all sent to the meado" <- same
Failure modes that are new compared with Gen-3.6
- Register drift inside 35 bytes. Story prompts sometimes continue as dates or encyclopedic fragments ("In the summer of 2009, ...", "On 15 July 1840, ...", "commissioned by The St..."). Gen-3.6 stayed in story mode far more reliably.
- Characters or objects appearing from nowhere (a new name, a "young girl named Alex", a new object that the scene never introduced).
- Names reinterpreted as common nouns ("The owen", "The June"), usually a sign of grammar collapse.
Pronoun agreement, measured automatically
A generation can be fluent and still give a named character the wrong pronoun. We took every prompt in the 1,000-prompt set that has a single named character with a clearly gendered name and checked the first he/she pronoun in the completion against that name. Wren and Nova are excluded as ambiguous, as are prompts with two named characters. The same script run on Gen-3.6's published outputs gives 77.7%, close to that card's 77.2%, so the two columns are comparable.
| Gen-4.6 | Gen-3.6 (same script) | |
|---|---|---|
| All names | 71.8% (287 / 400) | 77.7% (391 / 503) |
| Male names | 73.5% (147 / 200) | 92.0% (253 / 275) |
| Female names | 70.0% (140 / 200) | 60.5% (138 / 228) |
Two things changed. First, Gen-4.6 produces a he/she at all in only 400 of 1,000 completions (Gen-3.6: 503), so the sample is smaller. Second, the pattern by gender flipped: female names improved overall, male names got much worse.
The split by name is still sharp:
| Female names | Pronoun matches |
|---|---|
| Hazel, Nora, Willow, Poppy, June, Zoe | 43.4% (23 / 53) |
| Anna, Ellie, Lily, Ivy, Luna, Mia, Sadie, Daisy, Rosie, Grace, Aria, Ruby | 79.6% (117 / 147) |
Rarer female names are better than in Gen-3.6 (43% vs 8.5%) but still well below the common ones. And unlike Gen-3.6, the male side is now weak for rarer names too (Silas 3/13, Milo 2/8, Ezra 2/5). Common names work; rare names are close to a coin flip.
Real-world benchmark results
The coherence grading above is our own methodology. To see how the model does on external benchmarks, we ran it on two established suites. Neither involves code or math. Sample sizes are noted; the sampling error is not small — read differences of a few points as noise.
BLiMP (Benchmark of Linguistic Minimal Pairs)
BLiMP scores whether a model assigns higher likelihood to a grammatical sentence than to a minimally-different ungrammatical one, across 67 phenomena. We ran all 67 paradigms with 30 randomly sampled pairs each (2,010 pairs, random.Random(0)). Scoring is the sum of log-probabilities of the sentence bytes, conditioned on a leading document-end byte (0x00).
Sample caveat. The pair sample is not identical to the one in the Gen-3.6 / Gen-4.5 cards (that seed was not available). The baselines are therefore reported on this sample too, and they land within 1 point of their published full-benchmark values, so the sample is representative.
| Model | Params | BLiMP (full benchmark / earlier cards) | BLiMP (this 2,010-pair sample) |
|---|---|---|---|
| 5-gram | — | 61.2% | 60.2% |
| NoTokenLM-Gen-4.5 | 20M | 65.4% (that card's sample) | — |
| NoTokenLM-Gen-4.6 | 36.8M | — | 66.9% (95% CI 64.8–68.9) |
| NoTokenLM-Gen-3.6 | 13.9M | 68.9% (that card's sample) | — |
| Transformer-XL | — | 69.6% | 70.0% |
| LSTM | — | 69.8% | 69.8% |
| GPT-2 (BLiMP-paper outputs) | — | 83.0% | 83.3% |
Gen-4.6 sits between the 5-gram and the LSTM / Transformer-XL baselines. Given the sampling error (about ±2 points), it is statistically indistinguishable from Gen-3.6 and Gen-4.5; nearly tripling the parameters did not show up here. On identical pairs, paradigm by paradigm:
| Comparison | Gen-4.6 ahead | behind | tied |
|---|---|---|---|
| 5-gram | 37 | 25 | 5 |
| LSTM | 20 | 35 | 12 |
| Transformer-XL | 25 | 35 | 7 |
| GPT-2 | 5 | 57 | 5 |
Strengths and weaknesses follow the same shape as in earlier cards. 11 of 67 paradigms score 90% or higher: existential_there_quantifiers_1 and principle_A_case_1 (both 100%); determiner_noun_agreement_1, irregular_past_participle_adjectives, principle_A_domain_1, wh_vs_that_no_gap_long_distance (all 96.7%); determiner_noun_agreement_with_adjective_1 and passive_1 (93.3%); passive_2, regular_plural_subject_verb_agreement_1 and sentential_negation_npi_licensor_present (90%).
The lowest are wh_vs_that_with_gap_long_distance (3.3%), sentential_negation_npi_scope (16.7%), left_branch_island_echo_question (23.3%), wh_vs_that_with_gap (26.7%), tough_vs_raising_1 (33.3%), and coordinate_structure_constraint_complex_left_branch and existential_there_quantifiers_2 (both 36.7%).
Several of these come in pairs whose scores mirror each other (wh_vs_that_no_gap_long_distance 96.7% vs ..._with_gap_... 3.3%; npi_licensor_present 90% vs npi_scope 16.7%; the two existential_there_quantifiers paradigms). That is the signature of a lexical shortcut rather than a structural rule: the model prefers a particular word or pattern in both versions of the pair.
By phenomenon (n = pairs): irregular forms 90.0 (60), determiner–noun agreement 82.5 (240), anaphor agreement 76.7 (60), subject–verb agreement 76.1 (180), quantifiers 75.0 (120), binding 71.9 (210), argument structure 68.5 (270), control/raising 64.0 (150), filler–gap 62.9 (210), ellipsis 61.7 (60), NPI licensing 49.5 (210), island effects 47.5 (240).
The per-paradigm table, including same-sample baseline scores and win/loss flags, is in blimp_all_models_comparison_en.csv. Each of the 2,010 pairs with log-probabilities is in gen46_blimp_2010_pairs.json. Unlike the Gen-3.6 card, this is not a full 67,000-pair run: with only 30 pairs per paradigm, single-paradigm scores carry a large error (roughly ±8 points), so read the per-paradigm numbers as indications, not measurements.
LAMBADA (long-context final-word prediction)
LAMBADA gives a paragraph and asks for the final word. We scored 500 passages sampled from the LAMBADA test split (random.Random(0)) with greedy decoding: the first word of the continuation, punctuation stripped, lowercase, exact match. The dataset is the lowercased, space-tokenized version; context is truncated to the last 1,023 bytes. The earlier cards' seed was not available, so these 500 passages are not guaranteed to be the same ones.
| Model | Params | LAMBADA accuracy |
|---|---|---|
| NoTokenLM-Gen-4.6 | 36.8M | 10.0% (50 / 500, 95% Wilson CI 7.7–12.9) |
| NoTokenLM-Gen-3.6 (its own 500) | 13.9M | 11.4% (57 / 500, CI 8.9–14.5) |
| NoTokenLM-Gen-4.5 (same 500 as Gen-3.6) | 20M | 8.0% (40 / 500) |
| GPT-2-small | 124M | ~46% |
| GPT-2-large | 774M | ~59% |
| GPT-3 (zero-shot) | 175B | 76.2% |
Gen-4.6 is within sampling error of Gen-3.6 and Gen-4.5. In absolute terms this is still very far below even GPT-2-small. Per-passage predictions are in gen46_lambada_500_results.json. HellaSwag was not evaluated.
Bits per byte on held-out text
A model-agnostic measure of how well it predicts raw text (lower is better; not comparable across tokenizers or other corpora):
| Text | Bits per byte |
|---|---|
| 200 fiction passages (LAMBADA validation, first 1,000 bytes each; 68,536 bytes) | 1.910 |
| 2,010 grammatical BLiMP sentences (leading document-end byte as context) | 1.985 |
| First 201 of those sentences | 2.017 |
Gen-3.6's card reports 1.891 and 1.967 on its own passage and sentence selections. The selections differ, so these are only roughly comparable, but both point the same way: Gen-4.6 is not better at predicting fiction. A clean 87-byte sentence scored 1.35 bpb as a spot check.
Gen-4.6 vs Gen-3.6 at a glance
| Gen-4.6 (step 9k) | Gen-3.6 (step 13k) | |
|---|---|---|
| Parameters | 36.8M | 13.9M |
| Bytes of training text seen | ~0.59 GB | ~0.85 GB |
| Fully coherent (1,000 prompts) | 53.7% | 78.0% |
| Grammatically correct | 95.0% | 99.5% |
| Pronoun agreement (all names) | 71.8% | 77.7% |
| BLiMP (2,010-pair sample) | 66.9% | 68.9% (a different sample) |
| LAMBADA (500 passages) | 10.0% | 11.4% |
| Bits per byte, LAMBADA validation | 1.910 | 1.891 |
| Languages | English | English + French |
| Math in the data | ~20% | none |
Why is the bigger model not better? We did not run the experiments needed to say, so these are hypotheses, not findings:
- Fewer bytes seen. 0.59 GB vs 0.85 GB, with a model nearly three times the size. Bigger models need more data to pay off, and this checkpoint is early.
- Less story data. About 12% of the target mix is story text (TinyStories plus Cosmopedia stories) against about 15% for Gen-3.6, while the web share is larger. The register drift into dates and encyclopedic text fits this.
- A fifth of the mix is math (OpenWebMath, MathInstruct), which does nothing for story coherence, BLiMP, or LAMBADA.
- Different optimizer and a changed protocol. Gen-4.6 uses Muon, the labels come from a different annotator, and some samples differ. The comparisons are close, not controlled.
What it knows, and what it doesn't
The earlier cards' exact probe prompts were not available, so these probes use new prompt sets of the same kind and size (temperature 0.5, top-k 40, 4 seeds each; a hit means the expected answer appears in the first ~25 bytes). They are not directly comparable with the Gen-3.6 numbers.
| Probe | Result |
|---|---|
| Capital cities ("The capital of France is", 8 prompts × 4 seeds) | 0 / 32 — every prompt continues "the same as the ..." |
| Arithmetic (10 prompts × 4 seeds; the first number must equal the answer) | 5 / 40 — hits on 12 / 4 = (2), 3 + 4 =, 1 + 1 =, 10 - 3 =; the model then continues with unrelated equations or "The answer is C" |
| Mixed facts, opposites, sequences, code, context reuse (25 prompts × 4 seeds) | 21 / 100 — mostly idioms ("Once upon a time", "Thank you very much", "A, B, C, D"); it fails opposites, colors, weekdays/months and code |
Next-item probability, measured directly (no sampling; exact prompts from the Gen-3.6 card, probability of the target string starting with a space):
| Prompt | Target | Gen-4.6 | Gen-3.6 |
|---|---|---|---|
1, 2, 3, 4, |
5 | 46.4% | 35.6% |
10, 20, 30, 40, |
50 | 26.3% | 17.6% |
2, 4, 6, 8, |
10 | 12.2% | 20.7% |
5, 10, 15, 20, |
25 | 3.6% | 3.5% |
Monday, Tuesday, Wednesday, |
Thursday | 1.1% | 0.0% |
Counting patterns are picked up; the rule behind them is not (compare 2, 4, 6, 8 with 5, 10, 15, 20). Math-style text shows up in the output (equations, "The answer is A") because of the OpenWebMath and MathInstruct share, but correct arithmetic does not: the few hits are a leading-digit coincidence followed by more equations. Facts, rules and reasoning are out of reach.
Language behavior
Gen-4.6 is trained on English text only (plus math). The data pipeline has no language-ID filter, but non-English behavior was not evaluated for this checkpoint, and none is targeted — do not expect usable French or any other language. (Gen-3.6 included French; Gen-4.6 does not.)
What it's actually good at
- Well-formed English at the sentence level. 95.0% of the 1,000 outputs are grammatically correct, and local agreement is strong (11 of 67 BLiMP paradigms at 90%+, determiner–noun agreement 82.5%, irregular forms 90%).
- Common story openings. On concrete character-and-action prompts, many continuations are usable for a clause or two ("The young fox was so happy that he…").
- Counting patterns. The most common sequences get substantial probability (
1, 2, 3, 4,→ 5 at 46%). - Matching Gen-3.6 on BLiMP-style grammar and LAMBADA with a different data mix — useful as a baseline for later runs, not as a win.
What it's not good at, and why
- Story coherence. Only about half of short continuations are fully coherent, and one in twenty has broken grammar.
- Pronoun and gender consistency. 71.8% overall; rare names (male and female) are close to random.
- Register drift. Story prompts can jump to dates and encyclopedic text within 35 bytes.
- No usable world knowledge. 0 / 32 on capital cities. About 0.59 GB of training text at 36.8M parameters is far too little to store facts. It writes plausible-sounding text that is wrong.
- No usable arithmetic. It imitates the form of equations and "The answer is …" without computing.
- Long-range structure. LAMBADA at 10.0%, NPI licensing at 49.5% and island effects at 47.5% on BLiMP, and two paradigms far below chance (3.3% and 16.7%), point at the same gap.
- Rare malformed output. Names turned into common nouns ("The owen"), repeated words ("The kind kind of the kind"), and prompt words turned into nonsense.
How to actually run this thing
Recommended: temperature 0.5, top-k 40. The 1,000-prompt evaluation and the probes above were produced at these settings (BLiMP is likelihood-based and LAMBADA is greedy). Greedy decoding on open-ended prompts tends to fall into repetition loops; much higher temperatures increase drift and malformed output. Keep generations short: coherence is already shaky at 35 bytes, and the context window is 1,024 bytes.
Architecture details
| Parameters | ~36.8M |
| Layers | 15 |
| d_model | 448 |
| Attention heads | 14 (head dim 32) |
| Feedforward dim | 1,216 (SwiGLU) |
| Vocabulary | 256 (raw bytes, no tokenizer) |
| Context length | 1,024 bytes |
| Position encoding | RoPE (base 10000) |
| Normalization | RMSNorm |
| Attention extras | QK-norm, value residual, per-head sigmoid output gate |
| Input stage | Causal depthwise convolution (kernel 4) over the byte embeddings |
| Output layer | Weight-tied to the input embedding |
| Optimizer | Muon (attention / feed-forward matrices) + AdamW (everything else), EMA 0.999 |
| Training | Byte-level next-token prediction, from scratch (pretraining only — no instruction tuning) |
Training data
Target byte shares (the loader normalizes the weights and balances the stream toward them; the shares actually realized in the 0.59 GB seen were not measured):
| Category | Source | Share |
|---|---|---|
| General English (79.8%) | FineWeb-Edu (sample-100BT) |
44.9% |
Wikipedia (20231101.en) |
11.2% | |
| Cosmopedia — stories | 9.0% | |
| Cosmopedia — Khan Academy-style text | 7.9% | |
| TinyStories | 3.4% | |
Simple English Wikipedia (20231101.simple) |
3.4% | |
| Math (20.2%) | OpenWebMath | 14.6% |
MathInstruct (Q: … / A: … format) |
5.6% |
Documents are kept if they are 200–20,000 bytes, at least half alphabetic characters, and under a non-ASCII byte cap (8% for general sources, 15% for math), with a hash-based filter for repeated documents and a filter that drops documents made mostly of repeated lines. There is no language-ID filter, no code corpus, and no instruction-following data.
Method notes
- The PyTorch environment could not be installed in the evaluation sandbox. The checkpoint was loaded without PyTorch and the architecture was re-implemented in NumPy (fp32, CPU, EMA weights from
last.pt); the code is ingen46_eval_code.zip. Internal checks pass (KV-cache and full-pass logits agree to about 3e-5, generated text is fluent, bits per byte is in the expected range), but it was not cross-checked against the PyTorch forward pass. Training-time evaluations used bf16 autocast, this one used fp32. Individual samples will not match a PyTorch run byte for byte. - Sampling. Top-k mask on logits / temperature, then
rng.choicewithdefault_rng(1000 + prompt_id). The random-number consumption pattern differs from the earlier cards' code, so individual samples differ from theirs even at identical seeds.
Not run
- The exact BLiMP pair sample and 201-sentence bits-per-byte set from the earlier cards (their seeds / ids were not available), and the earlier LAMBADA passage ids.
- A full 67,000-pair BLiMP run (only the 2,010-pair sample).
- HellaSwag (not evaluated for any NoTokenLM card).
- Non-English behavior (not applicable; Gen-4.6 is English-only).
- Any experiment isolating why Gen-4.6 trails Gen-3.6 on story coherence.
Part of the NoTokenLM family — small models, built and evaluated honestly.
License
This project is licensed under the Apache License 2.0. Please refer to the LICENSE file for the full text of the license.
- Downloads last month
- 208
