Samudra

Model Introduction

Samudra is a global ocean emulator developed by the M2LInES team.

Paper: Samudra: An AI Global Ocean Emulator for Climate

https://doi.org/10.1029/2024GL114318

Model Description

Samudra predicts global ocean states on an approximately one-degree grid with a five-day time step. It is designed to emulate the evolution of the OM4 ocean circulation model with a deep neural network.

Use Cases

Scenario Description
Global ocean simulation Train the model on OM4 data following the 77-state-channel and 4-forcing-channel Samudra protocol.
Local quick validation Use synthetic NPZ data to check training, inference, and ocean-field visualization.
ModelScope / OneCode execution Download the standalone model package, install dependencies, and run the scripts directly.
Multi-GPU training Launch PyTorch DistributedDataParallel with torchrun.

Usage Guide

1. OneCode Usage

Experience intelligent one-click AI4S programming through the OneCode online environment:

Click to Experience Intelligent One-Click AI4S Programming

2. Manual Installation and Usage

Hardware Requirements

  • A GPU or DCU is recommended.
  • CPU can be used for import and small-scale connectivity verification; full training and inference will be slow.
  • DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended.

Download the Model Package

hf download OneScience-Group/Samudra --local-dir ./Samudra
cd Samudra

Install the Runtime Environment

DCU Environment

# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai

GPU Environment

# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai

Training Data Introduction

The official training data is generated by NOAA/GFDL OM4. The M2LInES project provides the dataset and documentation:

https://huggingface.co/datasets/M2LInES/Samudra-OM4

The dataset must be converted to the native NPZ layout expected by this repository. When real OM4 data is unavailable, generate a synthetic fixture for pipeline validation:

python scripts/fake_data.py

The synthetic fixture contains 77 prognostic state channels and 4 boundary-forcing channels and is not suitable for scientific evaluation.

Training

Single GPU:

python scripts/train.py

Multi-GPU:

torchrun --nproc_per_node=8 scripts/train.py

The default checkpoint is saved to data/checkpoints/model_bak.pth.

Training Weights

This repository provides a weight/ directory for Samudra checkpoints. The weight files will be uploaded soon and are expected to be available in the near future.

Inference

Inference performs an autoregressive rollout from data/test.npz and reads data/checkpoints/model_bak.pth by default:

python scripts/inference.py

Predictions are written to result/output/prediction.npz.

Evaluation and Visualization

python scripts/result.py

The default outputs are result/forecast_maps.png and result/temperature_profile.png.

Official OneScience Resources

Citation and License

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train OneScience-Group/Samudra