Buckets:
331 kB
3 files
Updated 3 months ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 2.46 kB xet | 19463de8 | |
| README.md | 872 Bytes xet | 791a6858 | |
| java-dataset.jsonl | 327 kB xet | 5d7968c9 |
Java Coding Dataset
This dataset contains high-quality Java code samples designed for fine-tuning coding-focused language models. It includes a diverse set of examples such as utility functions, class definitions, interface implementations, and exception handling.
Dataset Details
- Number of samples: 520 (and growing)
- Purpose: Fine-tuning LLMs to generate accurate and idiomatic Java code
- Content: Functions, classes, interfaces, exception handling
- License: MIT License
Usage
You can use this dataset with Hugging Face libraries or any fine-tuning pipeline compatible with JSONL data.
Example to load with datasets library:
from datasets import load_dataset
dataset = load_dataset("Hoglet-33/java-coding-dataset")
print(dataset["train"][0])
- Total size
- 331 kB
- Files
- 3
- Last updated
- Jun 20
- Pre-warmed CDN
- US EU US EU