YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
SDG โ Distributed Sharded Pipeline
Generate synthetic data with vLLM + uncertainty-quantification validation, distributed across multiple GPUs via SLURM sharding with a shared MongoDB inference cache.
Prerequisites
pip install -r sdg/requirements.txt
Install MongoDB (download binary):
curl -O https://fastdl.mongodb.org/linux/mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
tar xzf mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
cp mongodb-linux-x86_64-ubuntu2204-8.0.4/bin/mongod /u/nlp/anaconda/main/anaconda3/envs/tonyreasoningtraces/bin/
mongod --version
Quick Start (Single GPU, No Sharding)
python -m sdg.generate --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml --limit 100
This requires mongo_uri set in the config YAML (see below).
Distributed Sharded Run
1. Start MongoDB
Launch a MongoDB server on a cluster node (no GPU needed) with max TTL of 21 days:
nlprun -a tonyreasoningtraces --time 21-0 -m john3 --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
nlprun -a tonyreasoningtraces --time 21-0 -q jag -p high -m jagupard39 --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-serve "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
Or locally:
python -m sdg.launch_mongodb --port 27017 --dbpath /tmp/sdg_mongo
It will print mongo-uri: mongodb://<hostname>:27017 to stdout (and mongo.log).
2. Set mongo_uri and num_shards in your config
Grab the hostname from mongo.log, then edit your YAML config (e.g. sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml):
mongo_uri: "mongodb://<hostname>:27017"
num_shards: 32
3. Launch shards
conda activate tonyreasoningtraces && cd /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data
Dry run first to verify commands:
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch --dry-run
Then launch for real:
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch
This reads num_shards from the config and submits that many nlprun jobs (1 GPU each). Each job processes ceil(total_seeds / num_shards) examples and writes to output/{experiment_name}/shards/shard_NNN/. You can override with --num-shards N on the CLI.
4. Check status
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check
Prints a table like:
Shard Status Seeds Passed Failed
000 completed 2930 2100 830
001 running - - -
002 failed - - -
...
Overall: 20/32 completed, 8/32 running, 4/32 failed
When all shards are completed, check automatically:
- Combines all shard
output.jsonlfiles intooutput/{experiment_name}/output.jsonl - Aggregates all
stats.jsoninto a combinedoutput/{experiment_name}/stats.json - Uploads to HuggingFace at
teetone/{experiment_name}
5. Rerun failed shards
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun
This re-launches only failed shards (skips completed, running, and not_started).
Output Directory Structure
output/{experiment_name}/
launch_meta.json # Shard metadata
config.yaml # Config copy
output.jsonl # Combined output (after all shards done)
stats.json # Aggregated stats (after all shards done)
logs/
shard_000.log
shard_001.log
...
shards/
shard_000/
output.jsonl
stats.json
shard_001/
output.jsonl
stats.json
...
Example: OpenThoughts4 with 32 shards
# 1. Start MongoDB on john (no GPU, 16 GB RAM)
nlprun -a tonyreasoningtraces -q john --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
# 2. Grab the hostname from the job output, then set mongo_uri in your config
# mongo_uri: "mongodb://<hostname>:27017"
# 3. Launch
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch
# 4. Monitor
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check
# 5. Rerun any failures
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun
CLI Reference
| Command | Description |
|---|---|
python -m sdg.generate --config CONFIG |
Run pipeline on a single GPU |
python -m sdg.generate --config CONFIG --shard-id I --num-shards N |
Run a single shard |
python -m sdg.launch_mongodb |
Start MongoDB server |
python -m sdg.launcher --config CONFIG launch |
Launch all shards via nlprun |
python -m sdg.launcher --config CONFIG launch --dry-run |
Print commands without executing |
python -m sdg.launcher --config CONFIG check |
Check shard status, combine if done |
python -m sdg.launcher --config CONFIG rerun |
Rerun failed shards only |
python -m sdg.launcher --config CONFIG rerun --dry-run |
Show which failed shards would be relaunched |
Launch Jobs
Qwen3 4B
Launch OpenThoughts3 Math53K
# n=1, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy3_round1.yaml launch
# n=1, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy5_round1.yaml launch
# n=4, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n4_valredundancy5_round1.yaml launch
# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_no_filter_round1.yaml launch
# n=8, val_redundancy=5, all_valid (keep all samples that pass UQ)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_allvalid_round1.yaml launch
Launch OpenThoughts4 Science26K
# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n4_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy3_round1.yaml launch
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_no_filter_round1.yaml launch
Launch OpenThoughts4 Code9K
# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy3_round1.yaml launch
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_no_filter_round1.yaml launch
Qwen3 8B
Launch OpenThoughts4 Science26K (Qwen3-8B)
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch
# n=4, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n4_valredundancy5_round1.yaml launch
Launch OpenThoughts4 Code9K (Qwen3-8B)
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch
Launch OpenThoughts4 Math219K (Qwen3-8B)
# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_math219K_instill_n8_valredundancy5_round1.yaml launch
Other
Check logs
scp -r tonyhlee@scdt.stanford.edu:/nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data/output/qwen3_0.6b_openthoughts3_math53K_instill_n8_valredundancy5_round1 /Users/tonyhlee/Dev/virtual-world-data/