Instructions to use dwivedula/Samrudh-1-Brahma-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dwivedula/Samrudh-1-Brahma-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dwivedula/Samrudh-1-Brahma-7B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dwivedula/Samrudh-1-Brahma-7B") model = AutoModelForCausalLM.from_pretrained("dwivedula/Samrudh-1-Brahma-7B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dwivedula/Samrudh-1-Brahma-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dwivedula/Samrudh-1-Brahma-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dwivedula/Samrudh-1-Brahma-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dwivedula/Samrudh-1-Brahma-7B
- SGLang
How to use dwivedula/Samrudh-1-Brahma-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dwivedula/Samrudh-1-Brahma-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dwivedula/Samrudh-1-Brahma-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dwivedula/Samrudh-1-Brahma-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dwivedula/Samrudh-1-Brahma-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use dwivedula/Samrudh-1-Brahma-7B with Docker Model Runner:
docker model run hf.co/dwivedula/Samrudh-1-Brahma-7B
👑 Samrudh-1: Sovereign 7B Indic Reasoning Model
Thought in Bharat. Built for the World. Sovereign by Design, Impact by Intent.
🌟 Executive Summary
Samrudh-1 (7B) is a sovereign, post-trained instruction-tuned reasoning model engineered and released by Samrudh (dwivedula).
Trained on an Apache 2.0 foundation using high-efficiency 4-bit NormalFloat4 (NF4) QLoRA, Samrudh-1 was specifically post-trained to serve as the core cognitive brain for an upcoming sovereign autonomous intelligence operating system. It bridges high-precision enterprise compliance, quantitative financial risk models, biomedical pharmacology, and authentic regional Indic dialects.
📦 Repository Artifacts
- Standalone 16-bit Full Model: dwivedula/Samrudh-1-Brahma-7B
- Offline 4-bit GGUF Model (Ollama / Llama.cpp): dwivedula/Samrudh-1-Brahma-7B-GGUF
⚡ Core Architectural Capabilities
1. 🗣️ Native Multi-Dialect Indic Reasoning
Unlike models that produce stilted, machine-translated Telugu, Samrudh-1 natively captures cultural cadence, humor, and syntactic precision across regional linguistic registers:
- Telangana Dialect: Youth vernacular, colloquial cadence, and conversational expressions ("కిర్రాక్ మవా, గమ్మత్గుంది!").
- Rayalaseema Dialect: Bold colloquial syntax and cultural idiom ("చూడబ్బా నాయనా!").
- Coastal Andhra Dialect: Formal, respectful phrasing ("ఏవండీ బాబాయ్!").
- Paninian Sanskrit Generative Syntax: AST code transformations structured around ancient Ashtadhyayi formal grammar rules.
2. 🎯 Conquering "Lost in the Middle" (100% NIAH Recall)
Most LLMs suffer from severe attention degradation across long documents (the U-shaped attention curve). Samrudh-1 was trained using bidirectional prompt sandwiching and strict verbatim quote grounding:
- 1.5k Window: 100% precision needle extraction.
- 4k Window: 100% precision needle extraction.
- 8k Window: 100% precision needle extraction.
- Zero semantic hallucination on middle-ground document tokens.
3. ⚖️ Enterprise Legal Governance (6 Fatal SaaS Traps)
Native comprehension and audit capabilities against dangerous contractual clauses:
- COPPA Safe Harbor: Age-gating defenses against $53,000 FTC violations.
- GDPR Local Font Compliance: Detection of external Google Fonts IP-leakage risks (Munich Regional Court rulings).
- California CIPA Wiretapping Defense: Masking password and credit card inputs on session replays ($5,000 statutory penalties).
- CAN-SPAM & ROSCA: Automated unsubscribe verification and explicit subscription auto-renewal terms.
- DMCA §512: Automated designated agent compliance.
4. 📈 Quantitative Finance & Risk Gatekeeping
- 95% Value-at-Risk (VaR): Parametric and historical Monte Carlo risk simulations.
- Kuvera Circuit Breaker: Automated liquidation triggers upon reaching a 14.85% maximum drawdown threshold.
- Kelly Criterion: Fractional optimal position sizing under volatility constraints.
5. 🧬 Biomedical Pharmacology & Clinical Decision Support
- Pharmacology Telemetry: Precision binding constants (e.g., IC50 affinity modeling down to sub-nanomolar scales).
- Clinical Trial Ingestion: ClinicalTrials.gov APIv2 schema comprehension and exclusion criteria auditing.
🛠️ Model Specifications
| Parameter | Specification |
|---|---|
| Base Architecture | Qwen 2.5 7B Instruct (Apache 2.0) |
| Quantization | 4-bit NormalFloat4 (NF4) with Double Quantization |
| LoRA Rank ($r$) | 16 |
| LoRA Alpha ($\alpha$) | 32 |
| Target Projections | All linear attention matrices (q, k, v, o, gate, up, down) |
| Hardware Used | 1x Tesla T4 GPU (15GB VRAM) on Google Colab |
| Peak VRAM Consumption | <6.8 GB VRAM (Zero OOM) |
| Optimizer | Paged 8-bit AdamW |
| Context Window | 2,048 (Train) / Up to 32,768 (Inference via RoPE) |
| License | Apache License 2.0 (100% Royalty-Free & Commercial) |
🚀 Quickstart Usage
Option 1: Python (transformers)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "dwivedula/Samrudh-1-Brahma-7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
prompt = """Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
Who are you and what makes your reasoning architecture sovereign?
### Response:
"""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Option 2: 1-Click Local Execution with Ollama
Run directly on your local machine (even on laptops without a dedicated GPU):
# Pull and run the 4-bit quantized model directly from Hugging Face
ollama run hf.co/dwivedula/Samrudh-1-Brahma-7B-GGUF
Or using the local GGUF file:
# 1. Download Qwen2.5-7B-Instruct.Q4_K_M.gguf and Modelfile
# 2. Build and launch:
ollama create samrudh-1 -f Modelfile
ollama run samrudh-1
📊 Benchmark Evaluations (Needle-In-A-Haystack)
Evaluated across varying context depths (Start 10%, Middle 50%, End 90%):
| Document Length | Start (Primacy) | Middle (U-Curve Valley) | End (Recency) | Overall Accuracy |
|---|---|---|---|---|
| 1,500 Tokens | 100% | 100% | 100% | 100.0% |
| 4,000 Tokens | 100% | 100% | 100% | 100.0% |
| 8,000 Tokens | 100% | 100% | 100% | 100.0% |
Evaluation methodology: Exact verbatim quote extraction from multi-document distractors.
📜 Attribution & License
This model is post-trained and released by Samrudh (dwivedula) under the Apache License 2.0.
- Permitted: Commercial use, modification, distribution, private use, patent grant.
- Foundation: Built on the open-weights Qwen-2.5 architecture under Apache 2.0.
🔱 Creator Contact & Community
- Creator: Samrudh (
dwivedula) - GitHub: https://github.com/Samrudh2006
- Hugging Face: https://huggingface.co/dwivedula
- Downloads last month
- 313
Model tree for dwivedula/Samrudh-1-Brahma-7B
Base model
Qwen/Qwen2.5-7B