Instructions to use Subject-Emu-5259/NeuralAI-Nae1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Subject-Emu-5259/NeuralAI-Nae1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Use Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Subject-Emu-5259/NeuralAI-Nae1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Subject-Emu-5259/NeuralAI-Nae1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Subject-Emu-5259/NeuralAI-Nae1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Ollama
How to use Subject-Emu-5259/NeuralAI-Nae1 with Ollama:
ollama run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Subject-Emu-5259/NeuralAI-Nae1 with Docker Model Runner:
docker model run hf.co/Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
- Lemonade
How to use Subject-Emu-5259/NeuralAI-Nae1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Subject-Emu-5259/NeuralAI-Nae1:Q4_K_M
Run and chat with the model
lemonade run user.NeuralAI-Nae1-Q4_K_M
List all available models
lemonade list
- Atomic Chat
π§ NeuralAI β Nae1
Trained from scratch. Every weight owned.
A 100M-parameter decoder pretrained end-to-end on 768M FineWeb tokens β no distillation, no fine-tune of someone else's base.
π Quick facts
| Property | Value |
|---|---|
| Architecture | Llama-style decoder β RoPE, GQA, SwiGLU |
| Parameters | 100,111,872 (100.1M) |
| Layers / hidden / heads | 12 layers Β· 768 hidden Β· 12 heads Β· 4 KV heads (GQA) |
| Vocabulary | 439 tokens (ByteLevel BPE, sliced from a 32,000-target config) |
| Context | 2,048 trained Β· evaluated on n_ctx=512 |
| Training data | 768M tokens β FineWeb CC-MAIN-2013-20, 400k documents, 8 shards |
| Training schedule | 20,000 steps β five cosine cycles (0β4kβ8kβ12kβ16kβ20k) Β· Kaggle T4 |
| Final training loss | 24.15 @ step 20,000 (22.96 @ step 16,000 β 24.15 @ step 20,000; final-step loss is batch-noisy β held-out PPL decides) |
| Held-out perplexity | 148.84 (Q4_K_M, 992 docs / 1,027,766 tokens, teacher-forced) |
| Formats in this repo | f32 GGUF (305 MB) Β· Q4_K_M GGUF (45 MB) Β· config + tokenizer |
| License | Apache 2.0 |
π Training
Nae1 is pretrained from random initialization in five cosine cycles: 0 β 4,000 Β· 4,000 β 8,000 Β· 8,000 β 12,000 Β· 12,000 β 16,000 Β· 16,000 β 20,000, each resumed from the previous checkpoint. The chart below shows cycle two β loss falls from ~25.9 to a best of 22.88 (step 7,050), finishing at 24.90 at step 8,000 as the learning rate anneals to 1.2e-11; cycles three and four carry training loss to 22.96 at step 16,000, and cycle five finishes at 24.15 at step 20,000 (batch-noisy final step; held-out PPL keeps improving: 151.45 β 148.84).
- Data pipeline: FineWeb parquet β chunked into 513-token windows served as zero-copy views (peak-RAM-safe at 768M tokens)
- Checkpoints: every 250 steps; this release ships the final step-20000 checkpoint
π Evaluation
Teacher-forced perplexity on a held-out FineWeb slice (992 documents / 1,027,766 tokens, n_ctx=512), scored token-by-token with exact row alignment β not the echo=True logprob path, which mispairs positions and inflates NLL.
| Checkpoint (Q4_K_M) | Held-out PPL |
|---|---|
| step-2000 | 192.97 |
| step-4000 (resumed run) | 188.89 |
| step-4000 (from scratch) | 179.17 |
| step-8000 | 165.92 |
| step-12000 | 154.25 |
| step-16000 | 151.45 |
| step-20000 β this release | 148.84 |
Raw report: eval_step20000_q4_n512.json β mean NLL 5.0029 nats/token (prior releases kept alongside: eval_step16000_q4_n512.json, eval_step8000_q4_n512.json). Fixed probe ("The quick brown fox", 10 tokens): 169.63 β the 10-token probe is noisy; held-out decides.
β οΈ Perplexity is the trust signal here, not raw capability. At 100M params over 768M tokens, expect exploratory, often garbled text β this is a research-lineage model, not a chat assistant.
Logits Parity Notice
Logits parity with the Nae1 native forward pass is not achievable by design: llama.cpp uses RMSNorm (no bias) while the native model uses LayerNorm (with bias), so conversion skips the bias tensors. Tensor integrity is fully verified (109/109 tensors match the source safetensors, max diff 0.00e+00) β small logit drift is a runtime constraint, not a conversion bug.
π οΈ Usage
llama.cpp (recommended)
llama-server -m nae1-llama-Q4_K_M.gguf -c 512 --host 127.0.0.1 --port 8080
llama-cpp-python
from llama_cpp import Llama
llm = Llama(
model_path="nae1-llama-Q4_K_M.gguf",
n_ctx=512,
verbose=False,
)
out = llm("The future of AI", max_tokens=32, temperature=0.7)
print(out["choices"][0]["text"])
Files
| File | What it is |
|---|---|
nae1-llama-Q4_K_M.gguf |
4-bit quant, 4.80 BPW, 45 MB β the recommended artifact |
nae1-llama-f32.gguf |
Full-precision GGUF, 305 MB β for requantization / research |
config.json |
Architecture config as trained |
tokenizer/ |
Native 439-token vocab + merges + manifest |
eval_step20000_q4_n512.json |
The exact evaluation report cited above |
eval_step16000_q4_n512.json |
Prior release (step-16000), kept for lineage |
eval_step8000_q4_n512.json |
Prior release (step-8000), kept for lineage |
π§° What Is NeuralAI?
NeuralAI is a local-first, private generative AI engine built by De'Andrew Preston Harris. The mission: your AI, on your hardware, under your control. Nae1 is the project's first end-to-end pretrained model β trained from random init on public data, converted with verified tensor integrity, and evaluated with a reproducible protocol.
β οΈ Limitations
- Scale: 100M parameters pretrained on 768M tokens is a research checkpoint β long-form reasoning, coding, and factual recall are limited.
- Vocabulary: a 439-token BPE trained on this corpus; out-of-domain text will tokenize inefficiently.
- No chat template: base completion only β it was never instruction-tuned (SFT run in progress).
- No internet access: pair with a tool layer if you need live data.
π€ Who Created NeuralAI?
- Founder & Lead Architect: De'Andrew Preston Harris (D. Harris / Dre)
- Hugging Face: @Subject-Emu-5259
- GitHub: @Subject-Emu-5259
- LinkedIn: linkedin.com/in/deandrewharris94
- Location: Memphis, Tennessee / West Memphis, Arkansas
- Education: AI Software Engineering at Maestro College
NeuralAI was built from resilience, fatherhood, and the belief that personal computing deserves personal intelligence. Every release is handcrafted, iterated, and documented in the open.
π Citation
@software{neuralai_nae1_2026,
author = {Harris, De'Andrew Preston},
title = {NeuralAI β Nae1},
year = {2026},
url = {https://huggingface.co/Subject-Emu-5259/NeuralAI-Nae1},
version = {step-20000},
description = {A 100M-parameter language model pretrained from scratch on 768M FineWeb tokens}
}
Built with discipline by De'Andrew Preston Harris. Maintained in the open. Updated whenever the model, dataset, or project state changes.
- Downloads last month
- -
4-bit
32-bit