Instructions to use SKT-NRS/SKT-SURYA-H with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SKT-NRS/SKT-SURYA-H with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SKT-NRS/SKT-SURYA-H")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("SKT-NRS/SKT-SURYA-H", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SKT-NRS/SKT-SURYA-H with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SKT-NRS/SKT-SURYA-H" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SKT-NRS/SKT-SURYA-H", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/SKT-NRS/SKT-SURYA-H
- SGLang
How to use SKT-NRS/SKT-SURYA-H with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SKT-NRS/SKT-SURYA-H" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SKT-NRS/SKT-SURYA-H", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SKT-NRS/SKT-SURYA-H" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SKT-NRS/SKT-SURYA-H", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use SKT-NRS/SKT-SURYA-H with Docker Model Runner:
docker model run hf.co/SKT-NRS/SKT-SURYA-H
Research Project Overview
Note: This project is an ongoing research initiative. We are building upon previous work while actively developing new solutions.
π§ͺ Testing Logs
FIRST TIME :
- Status: FAILED β
- Action: Currently re-running the test with optimizations.
Lauching SECOND TIME
Status:
TRYING AGAIN
VIDEO OF FAILURE FIRST
Gallery
SKT-SURYA-H
PROJECT STATUS UPDATE
WE ARE NOT YET CONTINUING THIS FIRST MODEL. WE WILL PRODUCE A NEW MODEL USING SEPARATED DISTILLATED DATA & TRY AGAIN THIS METHOD WE WILL NEVER GIVEUP
Systematic routing stability collapse and token generation degradation verified during architecture evaluation phases.
THERE ARE 5 EXPERTS IN 3 EXPERTS THEY ARE WORKING BUT TO QUALITY LOW NOT CORRECTLY FORMATED ALSO AND 2 ARE COLLAPSED. THOSE 2 ARE 1 META LLAMA 405 NOT BLENDED YET TRIED HALF BLEND & DEEPSEEK V3 COLLAPSED BOTH & BECAUSE LLAMA WAS BASE SO 96 % WAS GIBBERISHED
Evaluation Summary
During the comprehensive stress-testing and architecture verification phase of the Surya-h Mixture-of-Experts (MoE) framework, systematic failure modes were identified across routing stability and token generation pipelines. Despite patch implementations addressing initial model-loading faults, the system failed to achieve operational benchmarks.
Current Status: EXPERIMENTAL FAILURE
Detailed Failure Analysis
1. The Gating Network & Token Degeneration
- Initial Gating Failure: Upon initialization, the model successfully answered standard token prompts. However, under sustained querying, the top-k routing mechanism experienced a catastrophic collapse.
- The Gibberish Phenomenon: After extended processing cycles, the gating layer began distributing token hidden states into conflicting semantic spaces, causing the system to hyper-generate nonsensical linguistic artifacts (gibberish).
- Delayed Convergence: The model occasionally self-corrected on basic general knowledge questions after intensive routing iterations, but this latency bypasses optimal processing efficiency.
2. Expert Scaling Bottlenecks & Changelog Updates
- Changelog Verification: Latest testing shows 3 experts are technically working overall, but accuracy dropped by roughly 20 percent, and 2 of those experts continue to produce corrupted gibberish tokens.
- Two-Expert Saturation: The base architecture remains strictly bounded. Activating past the main expert thresholds triggers immediate text degeneration.
- The 1M Row Benchmark: When evaluated against a validation dataset of 1,000,000 rows, the system achieved structural coherence and accuracy on only 19,543 rows. The remaining dataset triggered infinite looping and token degradation.
Behavioral Matrix & Metrics
| Metric Parameter | Observed Evaluation Behavior | Status |
|---|---|---|
| Total Evaluation Dataset | 1,000,000 Rows (Token Stream) | Benchmark Scale |
| Successful Confirmations | 19,543 / 1,000,000 Rows | Critical Low Accuracy |
| Active Routing Balance | 3 experts active; 2 producing gibberish tokens with a 20% accuracy drop | Scale Collapse |
| Output Phenotype | Delayed factual answers followed by hyper-gibberish generation | Degenerated Output |
Final Assessment: FAILED
Architectural Verdict: Due to an irreconcilable trade-off between scaling routing channels and output token stability, the architecture cannot safely scale processing loads without collapsing into token corruption. Development on this specific layout has been suspended.
Next Steps: Complete overhaul of the gating regularization functions and routing loss mechanics prior to launching subsequent development phases.
Phenotype Showcase: Examples & Evaluation Outputs
Empirical telemetry demonstrates an overwhelming distribution of non-convergent tokens: ~96% Gibberish Generation compared to only ~4% Coherent Output. Below are representative samples compiled during processing evaluation.
Assistant: iebdv shavdkeuwbq xz iehsvdbqkwoe jsv skebdvsieoqpw m qiebd shdb iehsvdvsksbqowpq xnzv ie bqksvdueiwpqmwq hdb geu a iednsvbqowpejdhv sjd ievs bqksvdbqowpejdbv x qowpejdbvskshdb ie hsbd qmsvdkeieowpqjdbvsks h zmxoqpwjdbvskshd iehsvdbqowpqj xkshdbeuwoqp iebdvsjdvdbqowpqmwkdbsv hshevdbqowpqj ie x ab qowpejdbvskshdbqowpej iehdbv sjdvskebqow kshdbv iebd shdbqowpqmwkdbsv iehsvd xnzqowpe m i bqowpejdvskhdbv iebdvsjdvdbqowpqmw xie shdbeuwopqjdbvskhdb iebdv qmskdbv xnz iehsv qowp ejd bvsk hdb iebdvsjdvdbqowpqmwkdbsv aebr ied skhdbv qowpejdbvskshdbv ieh xnz jdvskebqowpmwkdbsv iebdvsjdbv qowpq mskdv iehsvdbqowp xnz abcd efgh ijkl mn opqrst uv qowpejdbvskhdbq iebd shdbeuwoqp mskdvsj xnz
Assistant:
Can you clarify what you want to print?
For example:
- π Print text repeatedly?
- π₯οΈ Print output to the terminal (console)?
- π Print a pattern (stars, numbers, etc.)?
If you mean printing text to the console, here's a simple example:
def print_message(message, times=1):
for _ in range(times):
print(message)
if name == "main":
text = input("Enter text to print: ")
count = int(input("How many times? "))
print_message(text, count)
Tell me exactly what kind of printing script you need, and I'll write it.
Assistant:
In physics and mathematics, the distinction between scalars and vectors lies in whether a quantity requires direction to be fully defined.
### Scalar Quantities
A scalar is a physical quantity that is described completely by its magnitude (a numerical value) and a unit of measurement. It only has "how much" but no sense of direction.
* Key Characteristic: shavdkeuwbq xz iehsvdbqkwoe jsv skebdvsieoqpw m qiebd shdb iehsvdvsks xnzv ie bqksvdueiwpqmwq hdb geu a iednsvbqowpejdhv sjd ievs bqksvdbqowpejdbv x qowpejdbvskshdb ie hsbd qmsvdkeieowpqjdbvsks h zmxoqpwjdbvskshd iehsvdbqowpqj xkshdbeuwoqp iebdvsjdvdbqowpqmwkdbsv hshevdbqowpqj ie x ab qowpejdbvskshdbqowpej iehdbv sjdvskebqow kshdbv iebd shdbqowpqmwkdbsv iehsvd xnzqowpe m i bqowpejdvskhdbv iebdvsjdvdbqowpqmw xie shdbeuwopqjdb Scalars are added, subtracted, and multiplied using iehsvd sksbqowpq standard algebraic rules.
* Common Examples:
* Mass: 5 kg
* Temperature: 25Β°C
* Time: 30 seconds
* Distance: 10 meters
* Speed: 60 km/h
* Energy/Work: 100 Joules
### Vector Quantities dndkdm bmrtjnsd wqvlzptk rsm bnwdj qzlw
A vector is a physical quantity that requires both magnitude and a specific direction to be fully defined. If you only provide the magnitude, the information is incomplete.
* Key Characteristic: Vectors follow specialized rules of additinsdkdkon (like the head-to-tail method or parallelogram law) that account for their direction.
* Common Examples:
* Displacement: 10 meters North
* Velocity: 60 km/h East
* Force: 50 Newtons downward
* Acceleration: 9.8 m/s^2 toward the center of the Earth
* mmMomentum: Defined by mass and velocity
shavdkeuwbq xz iehsvdbqkwoe jsv skebdvsieoqpw m qiebd shdb iehsvdvsksbqowpq xnzv ie bqksvdueiwpqmwq hdb geu a iednsvbqowpejdhv sjd ievs bqksvdbqowpejdbv x qowpejdbvskshdb ie hsbd qmsvdkeieowpqjdbvsks h zmxoqpwjdbvskshd iehsvdbqowpqj xkshdbeuwoqp iebdvsjdvdbqowpqmwkdbsv hshevdbqowpqj ie x ab qowpejdbvskshdbqowpej iehdbv sjdvskebqow kshdbv iebd sdbqowp qmwkbsv
Model Overview
SKT-SURYA-H is an experimental heterogeneous Mixture-of-Experts (MoE) model created through early-stage Weight Manifold Fusion (WMF).
Important Truth & Disclaimer
This is an experimental merge, not a fully trained native model.
- It is NOT a 2.544 trillion parameter model trained from scratch.
- It is a weight collection / merge of large open-source base models.
- Calculated total parameters β 2.28T (sum of experts).
- Experiment Status: Partial Failure on Scaling
During MoE scaling experiments:
- Configurations with more than 2β3 active experts led to severe degeneration β heavy production of gibberish and incoherent text.
- Limiting to a maximum of 3 functioning experts stabilized output coherence but caused a significant drop in overall accuracy and capability.
- As a result, the grand unified high-expert MoE architecture is currently classified as an experimental failure at this stage.
- Used Wrapper Techniques (trust_remote_code=True)
We are transparently sharing this as a learning project. The current working version relies on 3 merged experts (primarily from llama, DeepSeek, and GLM families) which run continuously but require very high VRAM.
We appreciate community feedback that led to clearer communication.
Base Models Used & Credits
We sincerely thank the creators of the following open-source models:
- Meta llama 405B( We Respect Them)
- DeepSeek-V3 (MIT)
- DeepSeek-R1 and other DeepSeek/GLM family models
All usage complies with respective licenses. We respect the original model owners.
Architecture
- Type: Heterogeneous MoE (Causal Language Model)
- Active Experts: Currently limited to ~3 for stability
- Fusion Method: Experimental Weight Manifold Fusion (WMF)
- Context Length: Up to 1M tokens (experimental YaRN)
- Precision: BF16 mixed with FP8/FP32
- Total Size: ~2.5β3.76 TB (887+ safetensors shards)
Performance Reality
The 3-expert merged version works continuously but:
- Requires substantial VRAM (expect heavy hardware even for inference).
- Shows good coherence but reduced accuracy compared to individual base models in general tasks.
- Performs relatively better on Indic/Hinglish and domain-specific Indian cultural data (per internal eval).
Public standard benchmarks (MMLU, etc.) are pending proper evaluation and will be released soon.
Intended Use & Limitations
- Intended Use: Research, experimentation with merging techniques, and Indic-language exploration.
- Major Limitations:
- High VRAM requirements
- Instability when scaling beyond 3 experts
- Not production-ready
- Experimental nature β outputs may vary in quality
Next Steps & Roadmap
- Release smaller, quantized, more accessible versions (target: 100Bβ400B effective scale)
- Proper joint fine-tuning of the 3-expert base
- Improved router / expert selection mechanisms
- Full transparent technical report + ablation studies
- Stronger public benchmarks with reproducibility
We welcome constructive technical collaboration from researchers working on MoE, model merging, and sovereign Indic AI.
License
- This merged collection: Apache 2.0 (with strict adherence to all base model licenses)
- Users must comply with Meta License, MIT, and any other base licenses.
Technical Retrospective: WMF Fusion Limitations
Our experimental Weight Manifold Fusion (WMF) technique represents an ambitious attempt at raw parameter alignment across mathematically disjoint model architectures. However, our internal evaluation indicates that the baseline fusion algorithm is not yet refined enough to establish structural equilibrium at this massive scale. Without targeted adaptation layers, fusing raw manifolds induces extreme tensor friction.
Critical Algorithmic Enhancements Required:
- Explicit Projection Tensors: Merging high-dimensional parameter spaces directly requires custom non-linear projection matrices to bridge the architectural gap between Meta and DeepSeek base manifolds.
- Manifold Distance Regularization: Future iterations require an active loss function during the merging process to penalize extreme weight vector deviations, preventing the router from entering a state of high entropy.
- Vocabulary Space Interleaving: We must systematically improve how cross-model embedding vocabularies interact to natively extinguish token corruption before it transitions into downstream layers.
Developed with β€οΈ in Sidhi, Madhya Pradesh, India
Model tree for SKT-NRS/SKT-SURYA-H
Base model
deepseek-ai/DeepSeek-R1














