🧠 NeuralAI · Mamba K3

NeuralAI's 2.8B Mamba SSM — the third and largest model in the NeuralAI line (K1 130M → K2 793M → K3 2.8B), fine-tuned from state-spaces/mamba-2.8b for assistant-style chat with worked reasoning, code, math, writing, and safety behaviors.

Model Details

  • Architecture: Mamba SSM (state-spaces/mamba-2.8b)
  • Parameters: 2.8B
  • Status: SFT queued — weights will land here when training completes
  • Fine-tune: LoRA r=16, α=32 on Mamba in_proj, bf16, cosine LR 2e-4
  • Dataset: NeuralAI curated set — 3,068 examples (computed-answer math, reasoning puzzles, code, safety, writing, factual, multi-turn)
  • Chat template: NeuralAI [INST] (same as Mamba K1)
  • Creator: De'Andrew Preston Harris, NeuralAI

Status

  • This page is created ahead of the SFT run (per the model-line order in github.com/Subject-Emu-5259/NeuralAI).
  • Training is paused pending GPU allocation on the Colab runtime; the dataset and training script are ready to launch the moment a GPU is available.
  • When the SFT run completes, the merged weights + tokenizer + config are pushed here automatically.

Training Command (when GPU is available)

BASE_MODEL=state-spaces/mamba-2.8b HF_REPO=Subject-Emu-5259/NeuralAI-Mamba-K3 \
  python training/finetune_run.py

Prompt Format

Same as Mamba K1 — the NeuralAI [INST] chat template:

<|endoftext|>[INST] {user message} [/INST]

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Subject-Emu-5259/NeuralAI-Mamba-K3

Finetuned
(2)
this model