You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3 0.6B — germanweb

12B-token pretraining run from the Qwen3-0.6B base model.

Base model

Qwen/Qwen3-0.6B. These checkpoints are base models, not instruction-tuned or chat-tuned models.

Training data

  • training data: GermanWeb — 12B-token GermanWeb subset, tokenized with the Qwen3 tokenizer.

Training configuration

| Context length | 4,096 tokens | | Global batch | 512 sequences | | Micro batch | 8 sequences per rank | | Steps / target | 5,722 steps / approximately 12B tokens | | Optimizer | Distributed Adam, weight decay 0.1 | | Learning rate | 3e-4 peak; 3e-5 minimum; cosine or WSD schedule | | Precision | BF16 with FP8 current scaling | | GPUs | 8, data parallelism 8 | | Seed | 42 |

The repository contains checkpoint revisions named step-XXXXXXX; main is the final checkpoint. Optimizer states are not part of these HF exports. The exact source paths, revision mapping, and publication code are maintained in the private KletterMix_Ablations repository.

Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RuHae/KletterMix-Ablations-Qwen3-0.6B-GermanWeb

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1243)
this model