NeuralAI Nae1 banner

🧠 NeuralAI β€” Nae1

Trained from scratch. Every weight owned.
A 100M-parameter decoder pretrained end-to-end on 768M FineWeb tokens β€” no distillation, no fine-tune of someone else's base.

Nae1 100.1M params PPL 148.84 20000 steps Apache 2.0


πŸš€ Quick facts

Property Value
Architecture Llama-style decoder β€” RoPE, GQA, SwiGLU
Parameters 100,111,872 (100.1M)
Layers / hidden / heads 12 layers Β· 768 hidden Β· 12 heads Β· 4 KV heads (GQA)
Vocabulary 439 tokens (ByteLevel BPE, sliced from a 32,000-target config)
Context 2,048 trained Β· evaluated on n_ctx=512
Training data 768M tokens β€” FineWeb CC-MAIN-2013-20, 400k documents, 8 shards
Training schedule 20,000 steps — five cosine cycles (0→4k→8k→12k→16k→20k) · Kaggle T4
Final training loss 24.15 @ step 20,000 (22.96 @ step 16,000 β†’ 24.15 @ step 20,000; final-step loss is batch-noisy β€” held-out PPL decides)
Held-out perplexity 148.84 (Q4_K_M, 992 docs / 1,027,766 tokens, teacher-forced)
Formats in this repo f32 GGUF (305 MB) Β· Q4_K_M GGUF (45 MB) Β· config + tokenizer
License Apache 2.0

πŸ“ˆ Training

Nae1 training loss step 4000 to 8000

Nae1 is pretrained from random initialization in five cosine cycles: 0 β†’ 4,000 Β· 4,000 β†’ 8,000 Β· 8,000 β†’ 12,000 Β· 12,000 β†’ 16,000 Β· 16,000 β†’ 20,000, each resumed from the previous checkpoint. The chart below shows cycle two β€” loss falls from ~25.9 to a best of 22.88 (step 7,050), finishing at 24.90 at step 8,000 as the learning rate anneals to 1.2e-11; cycles three and four carry training loss to 22.96 at step 16,000, and cycle five finishes at 24.15 at step 20,000 (batch-noisy final step; held-out PPL keeps improving: 151.45 β†’ 148.84).

  • Data pipeline: FineWeb parquet β†’ chunked into 513-token windows served as zero-copy views (peak-RAM-safe at 768M tokens)
  • Checkpoints: every 250 steps; this release ships the final step-20000 checkpoint

πŸ“Š Evaluation

Held-out perplexity comparison across Nae1 checkpoints

Teacher-forced perplexity on a held-out FineWeb slice (992 documents / 1,027,766 tokens, n_ctx=512), scored token-by-token with exact row alignment β€” not the echo=True logprob path, which mispairs positions and inflates NLL.

Checkpoint (Q4_K_M) Held-out PPL
step-2000 192.97
step-4000 (resumed run) 188.89
step-4000 (from scratch) 179.17
step-8000 165.92
step-12000 154.25
step-16000 151.45
step-20000 β€” this release 148.84

Raw report: eval_step20000_q4_n512.json β€” mean NLL 5.0029 nats/token (prior releases kept alongside: eval_step16000_q4_n512.json, eval_step8000_q4_n512.json). Fixed probe ("The quick brown fox", 10 tokens): 169.63 β€” the 10-token probe is noisy; held-out decides.

⚠️ Perplexity is the trust signal here, not raw capability. At 100M params over 768M tokens, expect exploratory, often garbled text β€” this is a research-lineage model, not a chat assistant.

Logits Parity Notice

Logits parity with the Nae1 native forward pass is not achievable by design: llama.cpp uses RMSNorm (no bias) while the native model uses LayerNorm (with bias), so conversion skips the bias tensors. Tensor integrity is fully verified (109/109 tensors match the source safetensors, max diff 0.00e+00) β€” small logit drift is a runtime constraint, not a conversion bug.


πŸ› οΈ Usage

llama.cpp (recommended)

llama-server -m nae1-llama-Q4_K_M.gguf -c 512 --host 127.0.0.1 --port 8080

llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="nae1-llama-Q4_K_M.gguf",
    n_ctx=512,
    verbose=False,
)

out = llm("The future of AI", max_tokens=32, temperature=0.7)
print(out["choices"][0]["text"])

Files

File What it is
nae1-llama-Q4_K_M.gguf 4-bit quant, 4.80 BPW, 45 MB β€” the recommended artifact
nae1-llama-f32.gguf Full-precision GGUF, 305 MB β€” for requantization / research
config.json Architecture config as trained
tokenizer/ Native 439-token vocab + merges + manifest
eval_step20000_q4_n512.json The exact evaluation report cited above
eval_step16000_q4_n512.json Prior release (step-16000), kept for lineage
eval_step8000_q4_n512.json Prior release (step-8000), kept for lineage

🧰 What Is NeuralAI?

NeuralAI is a local-first, private generative AI engine built by De'Andrew Preston Harris. The mission: your AI, on your hardware, under your control. Nae1 is the project's first end-to-end pretrained model β€” trained from random init on public data, converted with verified tensor integrity, and evaluated with a reproducible protocol.


⚠️ Limitations

  • Scale: 100M parameters pretrained on 768M tokens is a research checkpoint β€” long-form reasoning, coding, and factual recall are limited.
  • Vocabulary: a 439-token BPE trained on this corpus; out-of-domain text will tokenize inefficiently.
  • No chat template: base completion only β€” it was never instruction-tuned (SFT run in progress).
  • No internet access: pair with a tool layer if you need live data.

πŸ‘€ Who Created NeuralAI?

NeuralAI was built from resilience, fatherhood, and the belief that personal computing deserves personal intelligence. Every release is handcrafted, iterated, and documented in the open.


πŸ“– Citation

@software{neuralai_nae1_2026,
  author       = {Harris, De'Andrew Preston},
  title        = {NeuralAI β€” Nae1},
  year         = {2026},
  url          = {https://huggingface.co/Subject-Emu-5259/NeuralAI-Nae1},
  version      = {step-20000},
  description  = {A 100M-parameter language model pretrained from scratch on 768M FineWeb tokens}
}

Built with discipline by De'Andrew Preston Harris. Maintained in the open. Updated whenever the model, dataset, or project state changes.

Downloads last month
-
GGUF
Model size
76.2M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support