YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Diversity Selection
This folder contains the helper scripts for selecting high-quality and diverse Tulu samples with:
final_score_continuousas the quality score.BAAI/bge-large-en-v1.5as the semantic embedding model.- Deita's score-first diversity filter as the final selector.
Files
prepare_deita_diversity_inputs.py: builds a top-N candidate JSON from the original parquet and score CSV.embed_with_bge.py: creates a Deita-compatible embedding pickle with BGE.run_deita_diversity_selection.sh: runs BGE embedding and Deita filtering.top_50k_by_final_score.json: default quality-ranked candidate pool.top_50k_by_final_score_for_embed.json: Deita-style conversation copy kept for compatibility/reference.
Usage
From the project root:
conda activate tokenclean
bash diversity_selection/run_deita_diversity_selection.sh
Common options:
GPU=0 THRESHOLD=0.85 DATA_SIZE=10000 BGE_BATCH_SIZE=128 \
bash diversity_selection/run_deita_diversity_selection.sh
Regenerate the top-50k candidate pool:
python diversity_selection/prepare_deita_diversity_inputs.py
Force embedding regeneration:
FORCE_REEMBED=1 bash diversity_selection/run_deita_diversity_selection.sh
The default output is:
diversity_selection/top_10k_by_final_score_diverse.json
The default BGE embedding cache is:
diversity_selection/top_50k_by_final_score_bge_embeddings.pkl
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support