YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

CFG3f Dataset Collection

This repository contains three versions of the CFG3f synthetic dataset, generated using Context-Free Grammar rules for language modeling research.

πŸ“Š Dataset Overview

Version Description Train Size Val Size Total Size Avg Tokens/Sample
Balanced Equal probability for all grammar rules 800K samples 80K samples 437MB ~261 tokens
Moderate Moderate imbalance (10:1 weight ratio) 800K samples 80K samples 192MB ~114 tokens
Extreme Extreme imbalance (100:1 weight ratio) 800K samples 80K samples 150MB ~89 tokens

πŸ”§ Dataset Structure

cfg3f_datasets/
β”œβ”€β”€ balanced/           # Balanced generation (all rules equal probability)
β”‚   β”œβ”€β”€ train/         # 800K training samples
β”‚   └── val/           # 80K validation samples
β”œβ”€β”€ moderate/          # Moderate imbalance (10:1 ratio)
β”‚   β”œβ”€β”€ train/         # 800K training samples  
β”‚   └── val/           # 80K validation samples
└── extreme/           # Extreme imbalance (100:1 ratio)
    β”œβ”€β”€ train/         # 800K training samples
    └── val/           # 80K validation samples

🎯 Grammar Rules (CFG3f)

The dataset is generated using a hierarchical context-free grammar with 4 levels:

Level 4 (Highest Abstraction)

22 -> 20 21 | 20 19 21 | 21 19 19 | 20 20
21 -> 18 17 | 17 16 | 16 17 18 | 16 18
20 -> 16 17 | 17 16 18
19 -> 18 16 18 | 17 18 | 18 18

Level 3

18 -> 14 15 13 | 15 13 13 | 13 15
17 -> 14 15 | 15 14 | 15 14 13
16 -> 13 15 13 | 15 15 | 14 13 | 14 14

Level 2

15 -> 10 11 11 | 11 11 10 | 10 10 | 12 12 11
14 -> 12 10 12 | 10 12 12 | 12 11 | 10 12
13 -> 11 12 | 10 12 11 | 12 11 12

Level 1

12 -> 7 9 7 | 9 8 | 8 8 9
11 -> 8 8 | 9 7 | 9 7 7
10 -> 8 9 9 | 9 7 9 | 7 9 9

Level 0 (Terminals)

9 -> '3' '3' | '1' '1' | '1' '2' '1'
8 -> '3' '3' '1' | '3' '1' '1' | '1' '2'
7 -> '3' '2' | '3' '2' '2' | '3' '1' '2' | '2' '2' '1'

βš–οΈ Weight Configurations

Balanced Version

  • All grammar rules have equal probability
  • Explores all possible generation paths uniformly
  • Results in longer, more diverse sentences

Moderate Version (10:1 ratio)

MODERATE_WEIGHTS = {
    '9': [1, 10, 1],   # Prefer "1 1" over others
    '8': [1, 1, 10],   # Prefer "1 2" over others
    '7': [10, 1, 1, 1], # Prefer "3 2" over others
    # ... similar patterns for all rules
}

Extreme Version (100:1 ratio)

EXTREME_WEIGHTS = {
    '9': [1, 100, 1],   # Heavily prefer "1 1"
    '8': [1, 1, 100],   # Heavily prefer "1 2"
    '7': [100, 1, 1, 1], # Heavily prefer "3 2"
    # ... similar patterns for all rules
}

πŸ”€ Data Format

  • File format: Binary (.bin) files compatible with NanoGPT
  • Token encoding:
    • '1' β†’ 0
    • '2' β†’ 1
    • '3' β†’ 2
    • '[BOS]' β†’ 3
    • '[EOS]' β†’ 4
  • Header: 256 int32 values followed by uint16 tokens
  • Vocabulary size: 7 tokens (including special tokens)

πŸ“ˆ Key Differences

Aspect Balanced Moderate Extreme
Rule Selection Uniform 10:1 bias 100:1 bias
Min Branch Prob 25-33% 8.33% 0.97%
Sentence Complexity High Medium Low
Rare Patterns Common Present Very Rare
Research Focus Baseline Moderate imbalance Long-tail distribution

πŸš€ Usage Example

import numpy as np

def read_bin_file(filename):
    with open(filename, 'rb') as f:
        # Read header
        header = np.frombuffer(f.read(256 * 4), dtype=np.int32)
        magic, version, num_tokens = header[0], header[1], header[2]
        
        # Read tokens
        tokens = np.frombuffer(f.read(), dtype=np.uint16)
        return tokens

# Load training data
train_tokens = read_bin_file("balanced/train/cfg3f_train_000000.bin")
print(f"Loaded {len(train_tokens)} tokens")

πŸŽ“ Research Applications

  • Long-tail Learning: Study how models handle imbalanced distributions
  • Grammar Induction: Test models' ability to learn hierarchical structure
  • Generalization: Evaluate performance on rare vs common patterns
  • Synthetic Benchmarks: Controlled experiments with known ground truth

πŸ“ Citation

If you use this dataset in your research, please cite:

@dataset{cfg3f_dataset_2024,
  title={CFG3f: Context-Free Grammar Dataset with Controlled Imbalance},
  author={Shuche Wang},
  year={2024},
  url={https://huggingface.co/datasets/Shuche/cfg3f_dataset}
}

πŸ”— Related Work

This dataset is part of research on:

  • Sharpness-aware optimization
  • Imbalanced learning in language models
  • Synthetic data generation for NLP

πŸ“„ License

MIT License - Feel free to use for research and educational purposes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support