Hugging Face
Models
Datasets
Spaces
Community
Docs
Enterprise
Pricing
Log In
Sign Up
nabeelshan
/
rlhf-gpt2-pipeline
like
0
Text Generation
Transformers
Safetensors
Dahoas/synthetic-instruct-gptj-pairwise
English
gpt2
rlhf
reinforcement-learning
ppo
reward-model
instruction-tuning
Eval Results
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Deploy
Use this model
main
rlhf-gpt2-pipeline
/
ppo_aligned_final
/
special_tokens_map.json
Nabeel Shan
Add tokenizer files
b461de7
2 months ago
raw
Copy download link
history
blame
contribute
delete
Safe
131 Bytes
{
"bos_token"
:
"<|endoftext|>"
,
"eos_token"
:
"<|endoftext|>"
,
"pad_token"
:
"<|endoftext|>"
,
"unk_token"
:
"<|endoftext|>"
}