Instructions to use SkMasud58/Nox-Alpha with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SkMasud58/Nox-Alpha with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SkMasud58/Nox-Alpha") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("SkMasud58/Nox-Alpha") model = AutoModelForCausalLM.from_pretrained("SkMasud58/Nox-Alpha", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SkMasud58/Nox-Alpha with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SkMasud58/Nox-Alpha" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SkMasud58/Nox-Alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SkMasud58/Nox-Alpha
- SGLang
How to use SkMasud58/Nox-Alpha with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SkMasud58/Nox-Alpha" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SkMasud58/Nox-Alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SkMasud58/Nox-Alpha" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SkMasud58/Nox-Alpha", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use SkMasud58/Nox-Alpha with Docker Model Runner:
docker model run hf.co/SkMasud58/Nox-Alpha
β‘ Nox Alpha: Pushing the Limits of Causal Reasoning and Sovereign Multilingual Synthesis
Developed by Falcon Intelligence | Founder & Chief AI Architect: SK Masud Rahman
Nox Alpha is the foundational high-speed reasoning model of the sovereign NOX intelligence ecosystem, custom-architected by SK Masud Rahman. Designed for sub-second causal inference (TTFT < 320ms), stateful zero-leakage scratchpad reasoning, and native multilingual synthesis across English, Bengali, Hindi, and Urdu.
π Highlights
- Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent encumbranceβideal for academic research, proprietary workflows, and commercial deployment.
- 2-Tier Model Hierarchy: Nox Alpha acts as the high-speed everyday causal reasoning and code synthesis engine, while Alpha Gen 1 provides apex formal mathematical proofs and deep scientific research.
- Dynamic Reasoning Effort: Seamlessly integrates with the Falcon Adaptive Router, allowing controllable reasoning budgets that dynamically scale compute proportional to task complexity.
- Production Zero-Leakage Scratchpad: An air-tight state machine suppresses rough internal thought traces (
<think>), delivering clean, formatted output directly to client streams. - Native Indic Multilingual Sovereignty: Eliminates historical Devanagari transliteration bias, generating authentic native script for Bengali, Hindi, Urdu, and English.
- Sub-Second Latency: Delivers sub-320ms Time-to-First-Token on the sovereign Mumbai cluster accelerated with NVIDIA RTX 4090 GPUs.
π Live Interactive Testing (Guest Mode Active)
Evaluate Nox Alpha and Alpha Gen 1 directly in the official interactive playground with zero registration:
π β‘ Click Here to Launch Nox Sovereign Playground (Guest Mode Active)
π Evaluation Results
All base models are evaluated under the Open LLM Leaderboard v2 protocol and standard benchmark suites. Scores within 0.3 of each other are considered equivalent.
Comprehensive Benchmark Evaluation Matrix
| Benchmark Task | Evaluation Metric | Alpha Gen 1 | Nox Alpha | DeepSeek-R1 (671B) | Claude 3.5 Sonnet | GPT-4o |
|---|---|---|---|---|---|---|
| MATH-500 | Strict LaTeX Pass@1 | 96.8% | 94.2% | 97.3% | 93.8% | 94.6% |
| AIME 2024 | Olympiad Pass@1 | 84.2% | 78.4% | 79.8% | 68.4% | 53.3% |
| GPQA Diamond | Doctoral STEM CoT | 72.8% | 68.2% | 71.5% | 65.0% | 66.8% |
| HumanEval | Zero-Shot Python | 92.4% | 89.6% | 90.2% | 93.7% | 90.2% |
| MMLU-Pro | 5-shot CoT Reasoning | 91.2% | 88.4% | 90.8% | 89.2% | 88.6% |
| LiveCodeBench (v5) | Hard Contest Problems | 86.4% | 82.5% | 85.0% | 80.2% | 76.5% |
| Indic-MMLU (Bengali) | 5-shot STEM CoT | 87.4% | 85.6% | 74.2% | 76.8% | 78.1% |
| Indic-MMLU (Hindi) | 5-shot STEM CoT | 89.1% | 87.2% | 76.5% | 79.1% | 80.4% |
| Inference Latency | Time-to-First-Token | 480ms | < 320ms | 980ms | 750ms | 820ms |
Analysis: Alpha Gen 1 excels at formal mathematics (AIME: 84.2%, GPQA: 72.8%), while Claude 3.5 Sonnet edges slightly on HumanEval Python syntax (93.7%), DeepSeek-R1 full model leads slightly on MATH-500 (97.3%), and Nox Alpha establishes the fastest streaming response (< 320ms TTFT).
π§ System Architecture
Architectural Components
- Falcon Adaptive Router: A neural complexity classifier (scoring 0β100) routes fast casual dialogue and standard queries to Nox Alpha and escalates rigorous mathematical theorems and research queries to Alpha Gen 1.
- Stateful Scratchpad Leakage Guard: Suppresses preliminary chain-of-thought tokens, eliminating internal reasoning leaks while preserving rigorous logical derivation.
- Multilingual Tokenizer Vocab: Engineered for Indic linguistic parity, ensuring accurate tokenization and generation across Bengali, Hindi, Urdu, and English.
π» Inference with Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "SkMasud58/Nox-Alpha"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "Explain quantum entanglement in Bengali and write the Bell state equation."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024, temperature=0.2)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
π Citation
@article{rahman2026noxalpha,
title={Nox Alpha: Pushing the Limits of Causal Reasoning and Sovereign Multilingual Synthesis},
author={Rahman, SK Masud},
journal={Falcon Intelligence Technical Report},
volume={1},
year={2026},
url={https://huggingface.co/SkMasud58/Nox-Alpha}
}
Developed with β€οΈ by Falcon Intelligence β’ Sovereign AI for Humanity.
- Downloads last month
- 350