4.35 MB
3 files
Updated about 19 hours ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 2.46 kB xet | 19463de8 | |
| README.md | 1.32 kB xet | 0fc57417 | |
| dataset.jsonl | 4.34 MB xet | 6bca9cf2 |
A small synthetic dataset of instruct prompts from various fandoms, answered by anime / game characters instructed to respond in a creative writing style.
Data was created by a mix of Claude 3.7, Deepseek-v3 & Gemini Flash.
Dataset has been cleaned for major slop. There's a decent amount of one word (testaments, fractures etc.) slop in here, but generally this one is pretty clean and has a positive impact on writing style for models it's applied to.
Creation Process:
- Character is randomly selected from a CSV of scraped Danbooru characters
- Fandom page for character is scraped if exists and evaluated for quality
- A second fandom is selected at random, uses the Special:Random url to get a random page (i.e. https://naruto.fandom.com/wiki/Special:Random)
- This page is scraped, summarized and not evaluated for quality
- An LLM is tasked to come up with a prompt based on the summarized page, the prompt will ask it to make an incorrect statement question or correct one
- LLM is now tasked to answer the question as the character selected in step 1 and provided the same summary, randomly additional instructions are inserted
- System prompt is heavily reduced or stripped entirely and conversation is saved
- Python script is run to identify the level of slop and remove the most egregious examples
- Total size
- 4.35 MB
- Files
- 3
- Last updated
- Sep 21
- Pre-warmed CDN
- US EU US EU