A newer version of this model is available: Neura-Tech-AI/Nexa-AI-4x4B-Instruct

Nexa-AI-4x4B-Instruct

A collaborative open-source large language model featuring an advanced Mixture of Experts (MoE) architecture, developed by Neura Tech AI and Lumina AI.

Overview

Nexa-AI-4x4B-Instruct is an instruction-tuned, state-of-the-art Mixture of Experts (MoE) model built using the framework and foundations of the Qwen3 architecture family. Featuring a dedicated setup of 4 routing experts each scaled at 4B parameters, it provides exceptional computation balancing and advanced contextual intelligence.

This project is jointly developed by:

  • Neura Tech AI
  • Lumina AI

Nexa AI focuses on delivering a capable multilingual AI assistant with strong performance in:

  • High-efficiency compute routing (MoE)
  • Complex logical reasoning & thinking loops
  • Advanced coding assistance & software engineering tasks
  • Mathematics & multi-step academic problem solving
  • Tool calling, automation, & autonomous AI agents
  • Deep multilingual understanding (including English, Hindi, Chinese, and more)

Model Architecture & Details

Unlike standard dense models, this variant leverages a modern sparse MoE setup that maps optimized parameter routing to specialized neural layers dynamically during runtime.

  • Model Name: Nexa-AI-4x4B-Instruct
  • Base Architecture: Transformer Decoder (Sparse Mixture of Experts)
  • Parameters: ~16B Total Parameters
  • Total Experts: 4 Specialists
  • Expert Size: ~4 Billion parameters per expert
  • Active Parameters: ~4B parameters routed per token (Dynamic Routing)
  • Context Length: 262,144 Tokens
  • License: Apache-2.0

Developers

Project: Nexa AI
Developed by:

  • Neura Tech AI
  • Lumina AI

Base Model Acknowledgment

We sincerely thank the Qwen Team for releasing the foundational Qwen3 model family under the Apache 2.0 License, which served as the structural baseline for this unified architectural development.

Features

  • Sparse MoE Architecture: Dynamic token routing through 4 separate specialized expert spaces for optimized intelligence per parameter.
  • Advanced Thinking & Alignment: Highly tuned instruction-following wrapper optimized for structured responses.
  • High-Quality Code Generation: Native compatibility with complex data tracking, script generations, and error handling.
  • Vast Multi-turn Context: Massive 262k context space allows keeping track of long documentation or terminal outputs seamlessly.
  • Native Tool Integration: Solid structure built for external API calls, JSON function routing, and systematic agentic pipelines.

Performance

Benchmark GPT-4.1-nano-2025-04-14 Qwen3-30B-A3B Non-Thinking Qwen3-4B Non-Thinking Nexa-AI-4x4B-Instruct
Knowledge
MMLU-Pro 62.8 69.1 58.0 69.6
MMLU-Redux 80.2 84.1 77.3 84.2
GPQA 50.3 54.8 41.7 62.0
SuperGPQA 32.2 42.2 32.0 42.8
Reasoning
AIME25 22.7 21.6 19.1 47.4
HMMT25 9.7 12.0 12.1 31.0
ZebraLogic 14.8 33.2 35.2 80.2
LiveBench 20241125 41.5 59.4 48.4 63.0
Coding
LiveCodeBench v6 (25.02-25.05) 31.5 29.0 26.4 35.1
MultiPL-E 76.3 74.6 66.6 76.2
Aider-Polyglot 9.8 24.4 13.8 12.9
Alignment
IFEval 74.5 83.5 81.2 83.8
Arena-Hard v2* 15.9 24.8 9.5 43.4
Creative Writing v3 72.7 68.1 53.6 83.5
WritingBench 66.9 72.2 68.5 83.4
Agent
BFCL-v3 53.0 58.6 57.6 61.9
TAU1-Retail 23.5 38.3 24.3 48.7
TAU1-Airline 14.0 18.0 16.0 32.0
TAU2-Retail - 31.6 28.1 40.4
TAU2-Airline - 18.0 12.0 24.0
TAU2-Telecom - 18.4 17.5 13.2
Multilingualism
MultiIF 60.7 70.8 61.3 79.1
MMLU-ProX 56.2 65.1 49.6 61.6
INCLUDE 58.6 67.8 53.8 60.1
PolyMATH 15.6 23.3 16.6 31.1

*: For reproducibility, we report the win rates evaluated by GPT-4.1.


© 2026 Neura Tech AI & Lumina AI. All rights reserved.

Downloads last month
151
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lumina-Ai-Official/Nexa-AI-4x4B-Instruct

Finetuned
(3)
this model