Franken-MoE-Large-Base

Franken-MoE-Large-Base is a high-performance native Qwen 3 MoE (Mixture-of-Experts) upcycled from Qwen/Qwen3-0.6B.

It combines a dense foundation pre-trained on 18T tokens with an expanded SwiGLU Mixture-of-Experts architecture using official Qwen3MoeForCausalLM specifications.

Notice

This model has not been re-trained. Expect broken outputs, looping, and other general weirdness. A healed and fine-tuned version soon though!

Architecture Highlights

  • Base Foundation: Qwen/Qwen3-0.6B (28 Layers, Hidden Dim $D=1024$, Intermediate Dim $H=3072$, Vocab Size 151,936)
  • Total Parameters: 4.56B
  • Active Parameters per Token: 860M (Compute cost and generation speed of a <1B model)
  • MoE Topology: 16 routed experts, Top-2 active per token (num_experts_per_tok=2), normalized top-k probabilities (norm_topk_prob=True)
  • Attention: Grouped-Query Attention (GQA 16:8) with QK-Norm (q_norm, k_norm)
  • Context Length: 40,960 tokens (RoPE $ heta = 1,000,000$)
  • Official Format: Uses fused 3D tensor layout (gate_up_proj and down_proj) without legacy shared experts.

Direct Compatibility

  • Unsloth: 100% plug-and-play in Google Colab / Unsloth notebooks (FastLanguageModel.from_pretrained("DrRiceIO7/Franken-MoE-Large-Base"))
  • llama.cpp / GGUF: Convert directly using convert_hf_to_gguf.py
  • Hugging Face Transformers: Loads directly with AutoModelForCausalLM.from_pretrained("DrRiceIO7/Franken-MoE-Large-Base")
  • vLLM / SGLang / Ollama: Native out-of-the-box support

Quick Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "DrRiceIO7/Franken-MoE-Large-Base"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "<|im_start|>user\nExplain how photosynthesis works in plants.<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.7, top_p=0.85, do_sample=True)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DrRiceIO7/Franken-MoE-Large-Base

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1262)
this model