YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
CFG3f Dataset Collection
This repository contains three versions of the CFG3f synthetic dataset, generated using Context-Free Grammar rules for language modeling research.
π Dataset Overview
| Version | Description | Train Size | Val Size | Total Size | Avg Tokens/Sample |
|---|---|---|---|---|---|
| Balanced | Equal probability for all grammar rules | 800K samples | 80K samples | 437MB | ~261 tokens |
| Moderate | Moderate imbalance (10:1 weight ratio) | 800K samples | 80K samples | 192MB | ~114 tokens |
| Extreme | Extreme imbalance (100:1 weight ratio) | 800K samples | 80K samples | 150MB | ~89 tokens |
π§ Dataset Structure
cfg3f_datasets/
βββ balanced/ # Balanced generation (all rules equal probability)
β βββ train/ # 800K training samples
β βββ val/ # 80K validation samples
βββ moderate/ # Moderate imbalance (10:1 ratio)
β βββ train/ # 800K training samples
β βββ val/ # 80K validation samples
βββ extreme/ # Extreme imbalance (100:1 ratio)
βββ train/ # 800K training samples
βββ val/ # 80K validation samples
π― Grammar Rules (CFG3f)
The dataset is generated using a hierarchical context-free grammar with 4 levels:
Level 4 (Highest Abstraction)
22 -> 20 21 | 20 19 21 | 21 19 19 | 20 20
21 -> 18 17 | 17 16 | 16 17 18 | 16 18
20 -> 16 17 | 17 16 18
19 -> 18 16 18 | 17 18 | 18 18
Level 3
18 -> 14 15 13 | 15 13 13 | 13 15
17 -> 14 15 | 15 14 | 15 14 13
16 -> 13 15 13 | 15 15 | 14 13 | 14 14
Level 2
15 -> 10 11 11 | 11 11 10 | 10 10 | 12 12 11
14 -> 12 10 12 | 10 12 12 | 12 11 | 10 12
13 -> 11 12 | 10 12 11 | 12 11 12
Level 1
12 -> 7 9 7 | 9 8 | 8 8 9
11 -> 8 8 | 9 7 | 9 7 7
10 -> 8 9 9 | 9 7 9 | 7 9 9
Level 0 (Terminals)
9 -> '3' '3' | '1' '1' | '1' '2' '1'
8 -> '3' '3' '1' | '3' '1' '1' | '1' '2'
7 -> '3' '2' | '3' '2' '2' | '3' '1' '2' | '2' '2' '1'
βοΈ Weight Configurations
Balanced Version
- All grammar rules have equal probability
- Explores all possible generation paths uniformly
- Results in longer, more diverse sentences
Moderate Version (10:1 ratio)
MODERATE_WEIGHTS = {
'9': [1, 10, 1], # Prefer "1 1" over others
'8': [1, 1, 10], # Prefer "1 2" over others
'7': [10, 1, 1, 1], # Prefer "3 2" over others
# ... similar patterns for all rules
}
Extreme Version (100:1 ratio)
EXTREME_WEIGHTS = {
'9': [1, 100, 1], # Heavily prefer "1 1"
'8': [1, 1, 100], # Heavily prefer "1 2"
'7': [100, 1, 1, 1], # Heavily prefer "3 2"
# ... similar patterns for all rules
}
π Data Format
- File format: Binary (.bin) files compatible with NanoGPT
- Token encoding:
'1'β 0'2'β 1'3'β 2'[BOS]'β 3'[EOS]'β 4
- Header: 256 int32 values followed by uint16 tokens
- Vocabulary size: 7 tokens (including special tokens)
π Key Differences
| Aspect | Balanced | Moderate | Extreme |
|---|---|---|---|
| Rule Selection | Uniform | 10:1 bias | 100:1 bias |
| Min Branch Prob | 25-33% | 8.33% | 0.97% |
| Sentence Complexity | High | Medium | Low |
| Rare Patterns | Common | Present | Very Rare |
| Research Focus | Baseline | Moderate imbalance | Long-tail distribution |
π Usage Example
import numpy as np
def read_bin_file(filename):
with open(filename, 'rb') as f:
# Read header
header = np.frombuffer(f.read(256 * 4), dtype=np.int32)
magic, version, num_tokens = header[0], header[1], header[2]
# Read tokens
tokens = np.frombuffer(f.read(), dtype=np.uint16)
return tokens
# Load training data
train_tokens = read_bin_file("balanced/train/cfg3f_train_000000.bin")
print(f"Loaded {len(train_tokens)} tokens")
π Research Applications
- Long-tail Learning: Study how models handle imbalanced distributions
- Grammar Induction: Test models' ability to learn hierarchical structure
- Generalization: Evaluate performance on rare vs common patterns
- Synthetic Benchmarks: Controlled experiments with known ground truth
π Citation
If you use this dataset in your research, please cite:
@dataset{cfg3f_dataset_2024,
title={CFG3f: Context-Free Grammar Dataset with Controlled Imbalance},
author={Shuche Wang},
year={2024},
url={https://huggingface.co/datasets/Shuche/cfg3f_dataset}
}
π Related Work
This dataset is part of research on:
- Sharpness-aware optimization
- Imbalanced learning in language models
- Synthetic data generation for NLP
π License
MIT License - Feel free to use for research and educational purposes.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support