Text Generation
Transformers
Safetensors
English
checkpoint
checkpoints

description

#1
by Banaxi-Tech - opened
AppleMind AI org

Best name ever!

AppleMind AI org

Best name ever!

Thanks! But expect training to take a long time... Training will take about a day on CPU

AppleMind AI org

Also I'm optimizing it so before, it has 5 seconds a step. And after, i get 1 second a step!

AppleMind AI org

Best name ever!

Thanks! But expect training to take a long time... Training will take about a day on CPU

how many params?

AppleMind AI org

Best name ever!

Thanks! But expect training to take a long time... Training will take about a day on CPU

how many params?

1.64M (was planned to be 1M)
and for the training tokens
~20M

AppleMind AI org

you need to get a GPU!

AppleMind AI org

used 3060 is good

AppleMind AI org

you need to get a GPU!

i just use cpu or colab pro, gpus are so expensive these days

AppleMind AI org

you ahve colab pro?

AppleMind AI org

you ahve colab pro?

yes

AppleMind AI org

i can do these runtimes:
CPU
A100 GPU
L4 GPU
T4 GPU
v6e-1 TPU
v5e-1 TPU

(h100 and g4 were grayed out)

AppleMind AI org

A100 is best there

AppleMind AI org

I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay

AppleMind AI org

I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay

5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)

AppleMind AI org

I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay

5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)

Why chinchilla optimal?

I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay

5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)

Why chinchilla optimal?

i want appleminds parameter sizes and training token budgets to follow the classic chinchilla compute optimal scaling rule. overtraining a model that tiny is unusual for me, so i prefer keeping the token-to-parameter ratio around 20:1.

AppleMind AI org

I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay

5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)

Why chinchilla optimal?

i want appleminds parameter sizes and training token budgets to follow the classic chinchilla compute optimal scaling rule. overtraining a model that tiny is unusual for me, so i prefer keeping the token-to-parameter ratio around 20:1.

ok thats fine but i already did some tests and, Chinchilla models at that size a worse than random.

AppleMind AI org

You can still do i just wanted to tell you

AppleMind AI org

ok i will test 2 models: one with 20:1 token to param ratio and one with 6K:1 token to param ratio if the 6K one is better then i might use it more but i will still use 20:1 for cpu because storage limits

6K:1 will be much better. Id reccomend 22k:1 it was the best for our models. Heres our benchmarks for 0.9M model:

Benchmark Chinchilla (18M) 10% (20B) Δ
ARC-Easy 26.64 26.98 +0.34
PIQA 49.78 53.54 +3.76
ARC-Challenge 26.54 22.27 −4.27
HellaSwag 24.88 29.01 +4.13
Average 31.96 32.95 +0.99
AppleMind AI org

ok the 6K one was surprisingly better

AppleMind AI org

i will use 300:1 because training will take too long with 6K:1

AppleMind AI org

some logs:
W0812 17:46:31.906000 9262 torch/_inductor/utils.py:1731] [1/0_1] Not enough SMs to use max_autotune_gemm mode
[transformers] loss_type=None was set in the config but it is unrecognized. Using the default loss: ForCausalLMLoss.
FineWeb-Edu: 1,000,131 / 100,000,000 tokens
FineWeb-Edu: 2,001,441 / 100,000,000 tokens
FineWeb-Edu: 3,001,677 / 100,000,000 tokens
FineWeb-Edu: 4,003,530 / 100,000,000 tokens
FineWeb-Edu: 5,003,864 / 100,000,000 tokens
FineWeb-Edu: 6,014,497 / 100,000,000 tokens
FineWeb-Edu: 7,014,690 / 100,000,000 tokens
FineWeb-Edu: 8,014,892 / 100,000,000 tokens
FineWeb-Edu: 9,015,759 / 100,000,000 tokens
FineWeb-Edu: 10,016,701 / 100,000,000 tokens
FineWeb-Edu: 11,042,613 / 100,000,000 tokens
FineWeb-Edu: 12,043,456 / 100,000,000 tokens
FineWeb-Edu: 13,044,297 / 100,000,000 tokens

Step 100/2,289 | Loss 10.8201 | LR 2.632e-05 | 147,790 tok/s | 13,107,200 tokens | 4.37% | 0.02h
VRAM: 1.57 GB allocated / 4.69 GB reserved

Saving checkpoint-100...
Writing model shards: 100% 1/1 [00:00<00:00, 65.09it/s]

UPLOADING CHECKPOINT 100

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-100
FineWeb-Edu: 14,051,468 / 100,000,000 tokens
FineWeb-Edu: 15,057,858 / 100,000,000 tokens
FineWeb-Edu: 16,058,350 / 100,000,000 tokens
FineWeb-Edu: 17,058,910 / 100,000,000 tokens
FineWeb-Edu: 18,059,120 / 100,000,000 tokens
FineWeb-Edu: 19,059,214 / 100,000,000 tokens
FineWeb-Edu: 20,060,034 / 100,000,000 tokens
FineWeb-Edu: 21,061,034 / 100,000,000 tokens
FineWeb-Edu: 22,061,243 / 100,000,000 tokens
FineWeb-Edu: 23,061,857 / 100,000,000 tokens
FineWeb-Edu: 24,068,376 / 100,000,000 tokens
FineWeb-Edu: 25,068,451 / 100,000,000 tokens
FineWeb-Edu: 26,069,133 / 100,000,000 tokens

Step 200/2,289 | Loss 10.7533 | LR 2.990e-05 | 204,624 tok/s | 26,214,400 tokens | 8.74% | 0.04h
VRAM: 1.57 GB allocated / 4.69 GB reserved

Saving checkpoint-200...
Writing model shards: 100% 1/1 [00:00<00:00, 36.37it/s]

UPLOADING CHECKPOINT 200

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-200
FineWeb-Edu: 27,075,071 / 100,000,000 tokens
FineWeb-Edu: 28,080,714 / 100,000,000 tokens
FineWeb-Edu: 29,080,972 / 100,000,000 tokens
FineWeb-Edu: 30,085,136 / 100,000,000 tokens
FineWeb-Edu: 31,085,689 / 100,000,000 tokens
FineWeb-Edu: 32,086,415 / 100,000,000 tokens
FineWeb-Edu: 33,095,884 / 100,000,000 tokens
FineWeb-Edu: 34,096,837 / 100,000,000 tokens
FineWeb-Edu: 35,097,143 / 100,000,000 tokens
FineWeb-Edu: 36,097,343 / 100,000,000 tokens
FineWeb-Edu: 37,104,064 / 100,000,000 tokens
FineWeb-Edu: 38,104,207 / 100,000,000 tokens
FineWeb-Edu: 39,105,073 / 100,000,000 tokens

Step 300/2,289 | Loss 10.6609 | LR 2.952e-05 | 205,417 tok/s | 39,321,600 tokens | 13.11% | 0.06h
VRAM: 1.57 GB allocated / 4.69 GB reserved

Saving checkpoint-300...
Writing model shards: 100% 1/1 [00:00<00:00, 37.12it/s]

UPLOADING CHECKPOINT 300

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-300
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-100
FineWeb-Edu: 40,107,907 / 100,000,000 tokens
FineWeb-Edu: 41,108,319 / 100,000,000 tokens
FineWeb-Edu: 42,108,687 / 100,000,000 tokens

AppleMind AI org

fyi this is colab L4

AppleMind AI org

more logs agaaa

AppleMind 1.0 Mini
Colab L4 High-RAM
300M TOKEN PRETRAINING

Device: cuda
GPU: NVIDIA L4
VRAM: 22.03 GB
CUDA: 12.8

Total training tokens: 300,000,000
Target ratio: ~300:1

======================================================================
DATASET MIXTURE

FineWeb-Edu sample-10BT : 100,000,000
FineWeb-HQ : 100,000,000
SmolLM-Corpus : 100,000,000

TOTAL : 300,000,000

TinyStories: REMOVED
WikiText: REMOVED
Local datasets: REMOVED
Dataset cycling: DISABLED

======================================================================
LOADING GPT-2 TOKENIZER

Vocabulary: 50,260
Special tokens added: 3

======================================================================
CREATING APPLEMIND

Parameters: 1,020,480
Trainable: 1,020,480

Target tokens: 300,000,000
Target token/parameter ratio: 293.98:1

======================================================================
SETTING UP HUGGING FACE

Logged in as:
GGUFGuy

Checkpoint repository ready:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints

Final model repository ready:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini
Precision: BF16

======================================================================
TRAINING CONFIGURATION

Parameters: 1,020,480
Training tokens: 300,000,000
Tokens/parameter: 293.98:1

Micro batch size: 64
Gradient accumulation: 8
Effective batch: 512
Context length: 256
Tokens/micro batch: 16,384
Tokens/optimizer step: 131,072
Optimizer steps: 2,289
Warmup steps: 114

FineWeb-Edu: 100M
FineWeb-HQ: 100M
SmolLM: 100M

Checkpoint repository:
AppleMind-AI/AppleMind-1.0-Mini-Checkpoints

Final repository:
AppleMind-AI/AppleMind-1.0-Mini

======================================================================
STARTING 300M-TOKEN PRETRAINING

Opening FineWeb-Edu sample-10BT...
Resolving data files: 100% 2410/2410 [00:00<00:00, 23080.23it/s]

STREAMING FineWeb-Edu

Target tokens: 100,000,000
Progress state: 0

FineWeb-Edu: 1,000,131 / 100,000,000 tokens
FineWeb-Edu: 2,001,441 / 100,000,000 tokens
FineWeb-Edu: 3,001,677 / 100,000,000 tokens
FineWeb-Edu: 4,003,530 / 100,000,000 tokens
FineWeb-Edu: 5,003,864 / 100,000,000 tokens
FineWeb-Edu: 6,014,497 / 100,000,000 tokens
FineWeb-Edu: 7,014,690 / 100,000,000 tokens
FineWeb-Edu: 8,014,892 / 100,000,000 tokens
FineWeb-Edu: 9,015,759 / 100,000,000 tokens
FineWeb-Edu: 10,016,701 / 100,000,000 tokens
FineWeb-Edu: 11,042,613 / 100,000,000 tokens
FineWeb-Edu: 12,043,456 / 100,000,000 tokens
FineWeb-Edu: 13,044,297 / 100,000,000 tokens

Step 100/2,289 | Loss 10.8201 | LR 2.632e-05 | 100,666 tok/s | 13,107,200 tokens | 4.37% | 0.04h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-100...
Writing model shards: 100% 1/1 [00:00<00:00, 11.08it/s]

UPLOADING CHECKPOINT 100

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-100
FineWeb-Edu: 14,051,468 / 100,000,000 tokens
FineWeb-Edu: 15,057,858 / 100,000,000 tokens
FineWeb-Edu: 16,058,350 / 100,000,000 tokens
FineWeb-Edu: 17,058,910 / 100,000,000 tokens
FineWeb-Edu: 18,059,120 / 100,000,000 tokens
FineWeb-Edu: 19,059,214 / 100,000,000 tokens
FineWeb-Edu: 20,060,034 / 100,000,000 tokens
FineWeb-Edu: 21,061,034 / 100,000,000 tokens
FineWeb-Edu: 22,061,243 / 100,000,000 tokens
FineWeb-Edu: 23,061,857 / 100,000,000 tokens
FineWeb-Edu: 24,068,376 / 100,000,000 tokens
FineWeb-Edu: 25,068,451 / 100,000,000 tokens
FineWeb-Edu: 26,069,133 / 100,000,000 tokens

Step 200/2,289 | Loss 10.7533 | LR 2.990e-05 | 100,525 tok/s | 26,214,400 tokens | 8.74% | 0.07h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-200...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]

UPLOADING CHECKPOINT 200

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-200
FineWeb-Edu: 27,075,071 / 100,000,000 tokens
FineWeb-Edu: 28,080,714 / 100,000,000 tokens
FineWeb-Edu: 29,080,972 / 100,000,000 tokens
FineWeb-Edu: 30,085,136 / 100,000,000 tokens
FineWeb-Edu: 31,085,689 / 100,000,000 tokens
FineWeb-Edu: 32,086,415 / 100,000,000 tokens
FineWeb-Edu: 33,095,884 / 100,000,000 tokens
FineWeb-Edu: 34,096,837 / 100,000,000 tokens
FineWeb-Edu: 35,097,143 / 100,000,000 tokens
FineWeb-Edu: 36,097,343 / 100,000,000 tokens
FineWeb-Edu: 37,104,064 / 100,000,000 tokens
FineWeb-Edu: 38,104,207 / 100,000,000 tokens
FineWeb-Edu: 39,105,073 / 100,000,000 tokens

Step 300/2,289 | Loss 10.6609 | LR 2.952e-05 | 100,422 tok/s | 39,321,600 tokens | 13.11% | 0.11h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-300...
Writing model shards: 100% 1/1 [00:00<00:00, 10.96it/s]

UPLOADING CHECKPOINT 300

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-300
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-100
FineWeb-Edu: 40,107,907 / 100,000,000 tokens
FineWeb-Edu: 41,108,319 / 100,000,000 tokens
FineWeb-Edu: 42,108,687 / 100,000,000 tokens
FineWeb-Edu: 43,108,744 / 100,000,000 tokens
FineWeb-Edu: 44,109,584 / 100,000,000 tokens
FineWeb-Edu: 45,109,641 / 100,000,000 tokens
FineWeb-Edu: 46,110,333 / 100,000,000 tokens
FineWeb-Edu: 47,110,722 / 100,000,000 tokens
FineWeb-Edu: 48,111,125 / 100,000,000 tokens
FineWeb-Edu: 49,111,530 / 100,000,000 tokens
FineWeb-Edu: 50,113,218 / 100,000,000 tokens
FineWeb-Edu: 51,113,865 / 100,000,000 tokens
FineWeb-Edu: 52,114,128 / 100,000,000 tokens

Step 400/2,289 | Loss 10.5818 | LR 2.887e-05 | 100,639 tok/s | 52,428,800 tokens | 17.48% | 0.14h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-400...
Writing model shards: 100% 1/1 [00:00<00:00, 10.91it/s]

UPLOADING CHECKPOINT 400

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-400
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-200
FineWeb-Edu: 53,114,351 / 100,000,000 tokens
FineWeb-Edu: 54,115,268 / 100,000,000 tokens
FineWeb-Edu: 55,115,867 / 100,000,000 tokens
FineWeb-Edu: 56,118,747 / 100,000,000 tokens
FineWeb-Edu: 57,118,792 / 100,000,000 tokens
FineWeb-Edu: 58,129,696 / 100,000,000 tokens
FineWeb-Edu: 59,129,772 / 100,000,000 tokens
FineWeb-Edu: 60,130,463 / 100,000,000 tokens
FineWeb-Edu: 61,130,846 / 100,000,000 tokens
FineWeb-Edu: 62,131,974 / 100,000,000 tokens
FineWeb-Edu: 63,145,958 / 100,000,000 tokens
FineWeb-Edu: 64,146,051 / 100,000,000 tokens
FineWeb-Edu: 65,146,257 / 100,000,000 tokens

Step 500/2,289 | Loss 10.5081 | LR 2.797e-05 | 100,116 tok/s | 65,536,000 tokens | 21.85% | 0.18h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-500...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]

UPLOADING CHECKPOINT 500

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-500
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-300
FineWeb-Edu: 66,147,176 / 100,000,000 tokens
FineWeb-Edu: 67,153,903 / 100,000,000 tokens
FineWeb-Edu: 68,163,652 / 100,000,000 tokens
FineWeb-Edu: 69,164,783 / 100,000,000 tokens
FineWeb-Edu: 70,166,355 / 100,000,000 tokens
FineWeb-Edu: 71,166,545 / 100,000,000 tokens
FineWeb-Edu: 72,167,733 / 100,000,000 tokens
FineWeb-Edu: 73,168,881 / 100,000,000 tokens
FineWeb-Edu: 74,175,780 / 100,000,000 tokens
FineWeb-Edu: 75,176,435 / 100,000,000 tokens
FineWeb-Edu: 76,177,413 / 100,000,000 tokens
FineWeb-Edu: 77,177,431 / 100,000,000 tokens
FineWeb-Edu: 78,179,729 / 100,000,000 tokens

Step 600/2,289 | Loss 10.4357 | LR 2.682e-05 | 100,691 tok/s | 78,643,200 tokens | 26.21% | 0.22h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-600...
Writing model shards: 100% 1/1 [00:00<00:00, 11.04it/s]

UPLOADING CHECKPOINT 600

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-600
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-400
FineWeb-Edu: 79,187,841 / 100,000,000 tokens
FineWeb-Edu: 80,187,924 / 100,000,000 tokens
FineWeb-Edu: 81,188,888 / 100,000,000 tokens
FineWeb-Edu: 82,189,255 / 100,000,000 tokens
FineWeb-Edu: 83,189,443 / 100,000,000 tokens
FineWeb-Edu: 84,192,547 / 100,000,000 tokens
FineWeb-Edu: 85,192,626 / 100,000,000 tokens
FineWeb-Edu: 86,193,916 / 100,000,000 tokens
FineWeb-Edu: 87,194,129 / 100,000,000 tokens
FineWeb-Edu: 88,196,386 / 100,000,000 tokens
FineWeb-Edu: 89,203,193 / 100,000,000 tokens
FineWeb-Edu: 90,203,499 / 100,000,000 tokens
FineWeb-Edu: 91,204,016 / 100,000,000 tokens

Step 700/2,289 | Loss 10.3676 | LR 2.546e-05 | 100,440 tok/s | 91,750,400 tokens | 30.58% | 0.25h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-700...
Writing model shards: 100% 1/1 [00:00<00:00, 10.87it/s]

UPLOADING CHECKPOINT 700

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-700
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-500
FineWeb-Edu: 92,205,445 / 100,000,000 tokens
FineWeb-Edu: 93,206,406 / 100,000,000 tokens
FineWeb-Edu: 94,215,721 / 100,000,000 tokens
FineWeb-Edu: 95,216,609 / 100,000,000 tokens
FineWeb-Edu: 96,218,844 / 100,000,000 tokens
FineWeb-Edu: 97,219,420 / 100,000,000 tokens
FineWeb-Edu: 98,219,961 / 100,000,000 tokens
FineWeb-Edu: 99,220,378 / 100,000,000 tokens

FineWeb-Edu complete: 100,000,000 tokens
Opening FineWeb-HQ...
Resolving data files: 100% 9246/9246 [00:00<00:00, 24460.23it/s]

STREAMING FineWeb-HQ

Target tokens: 100,000,000
Progress state: 0

FineWeb-HQ: 1,000,237 / 100,000,000 tokens
FineWeb-HQ: 2,000,353 / 100,000,000 tokens
FineWeb-HQ: 3,000,648 / 100,000,000 tokens
FineWeb-HQ: 4,001,578 / 100,000,000 tokens

Step 800/2,289 | Loss 10.3038 | LR 2.391e-05 | 95,154 tok/s | 104,857,600 tokens | 34.95% | 0.29h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-800...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]

UPLOADING CHECKPOINT 800

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-800
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-600
FineWeb-HQ: 5,003,165 / 100,000,000 tokens
FineWeb-HQ: 6,003,192 / 100,000,000 tokens
FineWeb-HQ: 7,003,245 / 100,000,000 tokens
FineWeb-HQ: 8,005,095 / 100,000,000 tokens
FineWeb-HQ: 9,005,500 / 100,000,000 tokens
FineWeb-HQ: 10,005,995 / 100,000,000 tokens
FineWeb-HQ: 11,006,275 / 100,000,000 tokens
FineWeb-HQ: 12,006,703 / 100,000,000 tokens
FineWeb-HQ: 13,008,348 / 100,000,000 tokens
FineWeb-HQ: 14,039,910 / 100,000,000 tokens
FineWeb-HQ: 15,040,042 / 100,000,000 tokens
FineWeb-HQ: 16,040,179 / 100,000,000 tokens
FineWeb-HQ: 17,040,530 / 100,000,000 tokens

Step 900/2,289 | Loss 10.2507 | LR 2.221e-05 | 99,896 tok/s | 117,964,800 tokens | 39.32% | 0.33h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-900...
Writing model shards: 100% 1/1 [00:00<00:00, 10.81it/s]

UPLOADING CHECKPOINT 900

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-900
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-700
FineWeb-HQ: 18,045,133 / 100,000,000 tokens
FineWeb-HQ: 19,053,155 / 100,000,000 tokens
FineWeb-HQ: 20,053,155 / 100,000,000 tokens
FineWeb-HQ: 21,053,631 / 100,000,000 tokens
FineWeb-HQ: 22,061,664 / 100,000,000 tokens
FineWeb-HQ: 23,061,940 / 100,000,000 tokens
FineWeb-HQ: 24,062,978 / 100,000,000 tokens
FineWeb-HQ: 25,063,229 / 100,000,000 tokens
FineWeb-HQ: 26,073,081 / 100,000,000 tokens
FineWeb-HQ: 27,073,220 / 100,000,000 tokens
FineWeb-HQ: 28,073,681 / 100,000,000 tokens
FineWeb-HQ: 29,073,966 / 100,000,000 tokens
FineWeb-HQ: 30,074,272 / 100,000,000 tokens

Step 1,000/2,289 | Loss 10.1952 | LR 2.039e-05 | 99,484 tok/s | 131,072,000 tokens | 43.69% | 0.36h
VRAM: 1.57 GB allocated / 15.40 GB reserved

Saving checkpoint-1000...
Writing model shards: 100% 1/1 [00:00<00:00, 10.97it/s]

UPLOADING CHECKPOINT 1000

Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-1000
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-800
FineWeb-HQ: 31,075,196 / 100,000,000 tokens
FineWeb-HQ: 32,075,679 / 100,000,000 tokens

Sign up or log in to comment