YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

SDG โ€” Distributed Sharded Pipeline

Generate synthetic data with vLLM + uncertainty-quantification validation, distributed across multiple GPUs via SLURM sharding with a shared MongoDB inference cache.

Prerequisites

pip install -r sdg/requirements.txt

Install MongoDB (download binary):

curl -O https://fastdl.mongodb.org/linux/mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
tar xzf mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
cp mongodb-linux-x86_64-ubuntu2204-8.0.4/bin/mongod /u/nlp/anaconda/main/anaconda3/envs/tonyreasoningtraces/bin/
mongod --version

Quick Start (Single GPU, No Sharding)

python -m sdg.generate --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml --limit 100

This requires mongo_uri set in the config YAML (see below).

Distributed Sharded Run

1. Start MongoDB

Launch a MongoDB server on a cluster node (no GPU needed) with max TTL of 21 days:

nlprun -a tonyreasoningtraces --time 21-0 -m john3 --exclude john17 -c 4 -g 0 --memory 64g \
  -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
  --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
 
nlprun -a tonyreasoningtraces --time 21-0 -q jag -p high -m jagupard39 --exclude john17 -c 4 -g 0 --memory 64g \
  -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
  --job-name sdg-serve "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"

Or locally:

python -m sdg.launch_mongodb --port 27017 --dbpath /tmp/sdg_mongo

It will print mongo-uri: mongodb://<hostname>:27017 to stdout (and mongo.log).

2. Set mongo_uri and num_shards in your config

Grab the hostname from mongo.log, then edit your YAML config (e.g. sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml):

mongo_uri: "mongodb://<hostname>:27017"
num_shards: 32

3. Launch shards

conda activate tonyreasoningtraces && cd /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data

Dry run first to verify commands:

python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch --dry-run

Then launch for real:

python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch

This reads num_shards from the config and submits that many nlprun jobs (1 GPU each). Each job processes ceil(total_seeds / num_shards) examples and writes to output/{experiment_name}/shards/shard_NNN/. You can override with --num-shards N on the CLI.

4. Check status

python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check

Prints a table like:

Shard  Status       Seeds   Passed   Failed
000    completed     2930     2100      830
001    running          -        -        -
002    failed           -        -        -
...
Overall: 20/32 completed, 8/32 running, 4/32 failed

When all shards are completed, check automatically:

  • Combines all shard output.jsonl files into output/{experiment_name}/output.jsonl
  • Aggregates all stats.json into a combined output/{experiment_name}/stats.json
  • Uploads to HuggingFace at teetone/{experiment_name}

5. Rerun failed shards

python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun

This re-launches only failed shards (skips completed, running, and not_started).

Output Directory Structure

output/{experiment_name}/
  launch_meta.json           # Shard metadata
  config.yaml                # Config copy
  output.jsonl               # Combined output (after all shards done)
  stats.json                 # Aggregated stats (after all shards done)
  logs/
    shard_000.log
    shard_001.log
    ...
  shards/
    shard_000/
      output.jsonl
      stats.json
    shard_001/
      output.jsonl
      stats.json
    ...

Example: OpenThoughts4 with 32 shards

# 1. Start MongoDB on john (no GPU, 16 GB RAM)
nlprun -a tonyreasoningtraces -q john --exclude john17 -c 4 -g 0 --memory 64g \
  -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
  --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"

# 2. Grab the hostname from the job output, then set mongo_uri in your config
#    mongo_uri: "mongodb://<hostname>:27017"

# 3. Launch
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch

# 4. Monitor
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check

# 5. Rerun any failures
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun

CLI Reference

Command Description
python -m sdg.generate --config CONFIG Run pipeline on a single GPU
python -m sdg.generate --config CONFIG --shard-id I --num-shards N Run a single shard
python -m sdg.launch_mongodb Start MongoDB server
python -m sdg.launcher --config CONFIG launch Launch all shards via nlprun
python -m sdg.launcher --config CONFIG launch --dry-run Print commands without executing
python -m sdg.launcher --config CONFIG check Check shard status, combine if done
python -m sdg.launcher --config CONFIG rerun Rerun failed shards only
python -m sdg.launcher --config CONFIG rerun --dry-run Show which failed shards would be relaunched

Launch Jobs

Qwen3 4B

Launch OpenThoughts3 Math53K

# n=1, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy3_round1.yaml launch

# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy1_round1.yaml launch

# n=8, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy1_round1.yaml launch

# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy3_round1.yaml launch

# n=1, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy5_round1.yaml launch

# n=4, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n4_valredundancy5_round1.yaml launch

# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_round1.yaml launch

# n=8, no_filter (no UQ validation)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_no_filter_round1.yaml launch

# n=8, val_redundancy=5, all_valid (keep all samples that pass UQ)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_allvalid_round1.yaml launch

Launch OpenThoughts4 Science26K

# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n4_valredundancy3_round1.yaml launch

# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n1_valredundancy1_round1.yaml launch

# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy3_round1.yaml launch

# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch

# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_no_filter_round1.yaml launch

Launch OpenThoughts4 Code9K

# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy3_round1.yaml launch

# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n1_valredundancy1_round1.yaml launch

# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy3_round1.yaml launch

# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch

# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_no_filter_round1.yaml launch

Qwen3 8B

Launch OpenThoughts4 Science26K (Qwen3-8B)

# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch

# n=4, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n4_valredundancy5_round1.yaml launch

Launch OpenThoughts4 Code9K (Qwen3-8B)

# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch

Launch OpenThoughts4 Math219K (Qwen3-8B)

# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_math219K_instill_n8_valredundancy5_round1.yaml launch

Other

Check logs

scp -r tonyhlee@scdt.stanford.edu:/nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data/output/qwen3_0.6b_openthoughts3_math53K_instill_n8_valredundancy5_round1  /Users/tonyhlee/Dev/virtual-world-data/
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support