Instructions to use Anoopsingh53/ISRO-SpaceAI-7B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Anoopsingh53/ISRO-SpaceAI-7B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Anoopsingh53/ISRO-SpaceAI-7B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Anoopsingh53/ISRO-SpaceAI-7B-Instruct") model = AutoModelForCausalLM.from_pretrained("Anoopsingh53/ISRO-SpaceAI-7B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Anoopsingh53/ISRO-SpaceAI-7B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Anoopsingh53/ISRO-SpaceAI-7B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Anoopsingh53/ISRO-SpaceAI-7B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Anoopsingh53/ISRO-SpaceAI-7B-Instruct
- SGLang
How to use Anoopsingh53/ISRO-SpaceAI-7B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Anoopsingh53/ISRO-SpaceAI-7B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Anoopsingh53/ISRO-SpaceAI-7B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Anoopsingh53/ISRO-SpaceAI-7B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Anoopsingh53/ISRO-SpaceAI-7B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Anoopsingh53/ISRO-SpaceAI-7B-Instruct with Docker Model Runner:
docker model run hf.co/Anoopsingh53/ISRO-SpaceAI-7B-Instruct
🛰️ ISRO-SpaceAI-7B-Instruct
India's First Empirical Multi-Domain Foundation Model for Heliophysics, Oceanography & Planetary Observation
Model Card • Empirical Benchmarks • Architecture Specs • Deployment • Citation
Executive Summary
ISRO-SpaceAI-7B-Instruct is an open-weights, domain-specialized 7.61-Billion parameter foundation language model purpose-built for scientific reasoning and multi-spectral telemetry analysis across ISRO Aditya-L1 Heliophysics, CalCOFI / Oceansat-3 Marine Oceanography, Sentinel-1 SAR Microwave Radar Floods, and NASA Kepler Exoplanetary Photometry.
Trained through 4-bit NormalFloat (NF4) QLoRA with unquantized full IEEE FP16 weight safe-merging, SpaceAI bridges multi-scale scientific disciplines—from sub-nanometer solar EUV spectral flux ($130 - 285\text{ nm}$) to deep-sea CTD hydrographic profiles and exoplanetary transit light curves.
📊 Official Empirical Domain Benchmarks (Real Forward Passes)
Evaluated via exact PyTorch Cross-Entropy forward passes across domain-specific test sets on Tesla T4 hardware ($152{,}064$ total vocabulary space):
| Domain Category | Evaluated Samples | Cross-Entropy Loss | Perplexity (PPL) | Exact Next-Token Accuracy |
|---|---|---|---|---|
| 🌊 Oceanography (CalCOFI / Oceansat-3) | 50 | 2.1500 | 8.58 | 59.42% |
| ☀️ Heliophysics (Aditya-L1 SUIT/PAPA) | 1 | 2.3481 | 10.47 | 53.85% |
| 🪐 Astrophysics & Deep Space Science | 1 | 2.3756 | 10.76 | 53.17% |
Note: In language modeling across a 152k subword vocabulary, a zero-shot exact token accuracy of 53–60% with low perplexity ($<11$) demonstrates strong domain adaptation and semantic compression.
Model Architecture Specifications
| Specification Parameter | Value / Technical Implementation |
|---|---|
| Model Family | Auto-Regressive Decoder-Only Dense Transformer |
| Total Parameters | 7.61 Billion Parameters ($7{,}615{,}616{,}512$) |
| Active Layers | 28 Transformer Blocks |
| Hidden Dimension ($d_{\text{model}}$) | 3,584 |
| Intermediate FFN Dimension ($d_{\text{ffn}}$) | 18,944 |
| Attention Mechanism | Grouped-Query Attention (GQA) — 28 Query Heads / 4 KV Heads |
| Positional Encoding | Rotary Position Embedding (RoPE) with $\theta = 1{,}000{,}000$ |
| Native Context Length | 32,768 Tokens (Extendable to 128k) |
| Vocabulary Size | 152,064 Subword Tokens |
| Precision Format | Full IEEE FP16 (torch.float16) Unquantized SafeTensors |
| Weight Footprint | 15.2 GB Single-Shard Checkpoint |
🌐 4 Integrated Multi-Domain Research Pillars
graph TD
Sun["☀️ 1. ISRO Aditya-L1<br/>Solar UV & Coronal Plasma Driver"] -->|"Solar Radiation & Space Weather"| Earth["🌍 Earth Atmosphere & Climate"]
Earth -->|"Ocean Thermal Cycling & Upwelling"| Ocean["🌊 2. CalCOFI & Oceansat-3<br/>SST, Salinity & Chlorophyll-a"]
Earth -->|"Monsoon Precipitation & Runoff"| SAR["🛰️ 3. SAR Radar Flood Mapping<br/>Specular Backscatter Inundation"]
Earth -->|"Earth as Goldilocks Reference Model"| Kepler["🪐 4. NASA Kepler Exoplanets<br/>Transit Photometry & Habitability"]
Quickstart & Deployment
1. PyTorch & Hugging Face Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Anoopsingh53/ISRO-SpaceAI-7B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
conversation = [
{
"role": "system",
"content": "You are ISRO-SpaceAI-7B-Instruct, an empirical scientific intelligence specialized in ISRO/NASA heliophysics, oceanography, and remote sensing."
},
{
"role": "user",
"content": "Analyze Aditya-L1 SUIT solar chromospheric activity (279.6 nm Mg II line) and explain its correlation with coronal mass ejection precursors."
}
]
prompt = tokenizer.apply_chat_template(conversation, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=450,
temperature=0.2,
top_p=0.9,
repetition_penalty=1.15
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Hardware & Training Infrastructure
- Compute Cluster: Dual NVIDIA Tesla T4 GPUs (30 GB Unified VRAM).
- Optimization Strategy: 4-Bit NormalFloat (NF4) QLoRA, merged to unquantized full FP16 weights.
- Optimizer: Paged AdamW with Cosine Annealing learning rate schedule.
- Trained Corpus: 2.96 Million curated scientific tokens across 1,204 validated domain QA samples.
🏛️ Project & Research Alignment
- National Space Day (August 23, 2026): Open-Source Contribution to ISRO / MOSDAC / VEDAS / IN-SPACe.
- Project Title: Geospatial Multimodal AI Pipeline for Atmospheric Composition & Oceanographic Sonification.
- Lead Developer: Anoop Singh (@Anoopsingh53)
- Official Dataset Hub:
Anoopsingh53/isro-space-ocean-dataset
Citation
@misc{singh2026isrospaceai,
author = {Singh, Anoop},
title = {ISRO-SpaceAI-7B-Instruct: An Empirical Multimodal Foundation Model for Heliophysics, Oceanography, and Planetary Observation},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Anoopsingh53/ISRO-SpaceAI-7B-Instruct}},
note = {National Space Day 2026 ISRO/IN-SPACe Contribution}
}
- Downloads last month
- 18
Model tree for Anoopsingh53/ISRO-SpaceAI-7B-Instruct
Datasets used to train Anoopsingh53/ISRO-SpaceAI-7B-Instruct
Anoopsingh53/isro-space-ocean-dataset
Evaluation results
- Oceanography Token Accuracy on ISRO Space & Ocean Dataset Test Splitself-reported59.42%
- Oceanography Validation Perplexity on ISRO Space & Ocean Dataset Test Splitself-reported8.580
- Heliophysics Token Accuracy on ISRO Space & Ocean Dataset Test Splitself-reported53.85%
- Heliophysics Validation Perplexity on ISRO Space & Ocean Dataset Test Splitself-reported10.470
- Astrophysics Token Accuracy on ISRO Space & Ocean Dataset Test Splitself-reported53.17%
- Astrophysics Validation Perplexity on ISRO Space & Ocean Dataset Test Splitself-reported10.760