YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Diversity Selection

This folder contains the helper scripts for selecting high-quality and diverse Tulu samples with:

  1. final_score_continuous as the quality score.
  2. BAAI/bge-large-en-v1.5 as the semantic embedding model.
  3. Deita's score-first diversity filter as the final selector.

Files

  • prepare_deita_diversity_inputs.py: builds a top-N candidate JSON from the original parquet and score CSV.
  • embed_with_bge.py: creates a Deita-compatible embedding pickle with BGE.
  • run_deita_diversity_selection.sh: runs BGE embedding and Deita filtering.
  • top_50k_by_final_score.json: default quality-ranked candidate pool.
  • top_50k_by_final_score_for_embed.json: Deita-style conversation copy kept for compatibility/reference.

Usage

From the project root:

conda activate tokenclean
bash diversity_selection/run_deita_diversity_selection.sh

Common options:

GPU=0 THRESHOLD=0.85 DATA_SIZE=10000 BGE_BATCH_SIZE=128 \
  bash diversity_selection/run_deita_diversity_selection.sh

Regenerate the top-50k candidate pool:

python diversity_selection/prepare_deita_diversity_inputs.py

Force embedding regeneration:

FORCE_REEMBED=1 bash diversity_selection/run_deita_diversity_selection.sh

The default output is:

diversity_selection/top_10k_by_final_score_diverse.json

The default BGE embedding cache is:

diversity_selection/top_50k_by_final_score_bge_embeddings.pkl
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support