Instructions to use AppleMind-AI/AppleMind-1.0-Mini-Checkpoints with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AppleMind-AI/AppleMind-1.0-Mini-Checkpoints with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AppleMind-AI/AppleMind-1.0-Mini-Checkpoints")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AppleMind-AI/AppleMind-1.0-Mini-Checkpoints", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AppleMind-AI/AppleMind-1.0-Mini-Checkpoints with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints
- SGLang
How to use AppleMind-AI/AppleMind-1.0-Mini-Checkpoints with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AppleMind-AI/AppleMind-1.0-Mini-Checkpoints", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AppleMind-AI/AppleMind-1.0-Mini-Checkpoints with Docker Model Runner:
docker model run hf.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints
description
Best name ever!
Best name ever!
Thanks! But expect training to take a long time... Training will take about a day on CPU
Also I'm optimizing it so before, it has 5 seconds a step. And after, i get 1 second a step!
Best name ever!
Thanks! But expect training to take a long time... Training will take about a day on CPU
how many params?
Best name ever!
Thanks! But expect training to take a long time... Training will take about a day on CPU
how many params?
1.64M (was planned to be 1M)
and for the training tokens
~20M
you need to get a GPU!
used 3060 is good
you need to get a GPU!
i just use cpu or colab pro, gpus are so expensive these days
you ahve colab pro?
you ahve colab pro?
yes
i can do these runtimes:
CPU
A100 GPU
L4 GPU
T4 GPU
v6e-1 TPU
v5e-1 TPU
(h100 and g4 were grayed out)
A100 is best there
I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay
I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay
5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)
I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay
5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)
Why chinchilla optimal?
I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay
5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)Why chinchilla optimal?
i want appleminds parameter sizes and training token budgets to follow the classic chinchilla compute optimal scaling rule. overtraining a model that tiny is unusual for me, so i prefer keeping the token-to-parameter ratio around 20:1.
I would recommend you, for your first good model do this: 5M parameters trained on 30B tokens of epfml/FineWeb-HQ trained on A100 with a 4e-3 lr and adamw. and WSD decay
5M model will be medium for applemind 1.0
and i use a 20:1 token to param ratio (chinchilla optimal)Why chinchilla optimal?
i want appleminds parameter sizes and training token budgets to follow the classic chinchilla compute optimal scaling rule. overtraining a model that tiny is unusual for me, so i prefer keeping the token-to-parameter ratio around 20:1.
ok thats fine but i already did some tests and, Chinchilla models at that size a worse than random.
You can still do i just wanted to tell you
ok i will test 2 models: one with 20:1 token to param ratio and one with 6K:1 token to param ratio if the 6K one is better then i might use it more but i will still use 20:1 for cpu because storage limits
6K:1 will be much better. Id reccomend 22k:1 it was the best for our models. Heres our benchmarks for 0.9M model:
| Benchmark | Chinchilla (18M) | 10% (20B) | Δ |
|---|---|---|---|
| ARC-Easy | 26.64 | 26.98 | +0.34 |
| PIQA | 49.78 | 53.54 | +3.76 |
| ARC-Challenge | 26.54 | 22.27 | −4.27 |
| HellaSwag | 24.88 | 29.01 | +4.13 |
| Average | 31.96 | 32.95 | +0.99 |
ok the 6K one was surprisingly better
i will use 300:1 because training will take too long with 6K:1
some logs:
W0812 17:46:31.906000 9262 torch/_inductor/utils.py:1731] [1/0_1] Not enough SMs to use max_autotune_gemm mode
[transformers] loss_type=None was set in the config but it is unrecognized. Using the default loss: ForCausalLMLoss.
FineWeb-Edu: 1,000,131 / 100,000,000 tokens
FineWeb-Edu: 2,001,441 / 100,000,000 tokens
FineWeb-Edu: 3,001,677 / 100,000,000 tokens
FineWeb-Edu: 4,003,530 / 100,000,000 tokens
FineWeb-Edu: 5,003,864 / 100,000,000 tokens
FineWeb-Edu: 6,014,497 / 100,000,000 tokens
FineWeb-Edu: 7,014,690 / 100,000,000 tokens
FineWeb-Edu: 8,014,892 / 100,000,000 tokens
FineWeb-Edu: 9,015,759 / 100,000,000 tokens
FineWeb-Edu: 10,016,701 / 100,000,000 tokens
FineWeb-Edu: 11,042,613 / 100,000,000 tokens
FineWeb-Edu: 12,043,456 / 100,000,000 tokens
FineWeb-Edu: 13,044,297 / 100,000,000 tokens
Step 100/2,289 | Loss 10.8201 | LR 2.632e-05 | 147,790 tok/s | 13,107,200 tokens | 4.37% | 0.02h
VRAM: 1.57 GB allocated / 4.69 GB reserved
Saving checkpoint-100...
Writing model shards: 100% 1/1 [00:00<00:00, 65.09it/s]
UPLOADING CHECKPOINT 100
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-100
FineWeb-Edu: 14,051,468 / 100,000,000 tokens
FineWeb-Edu: 15,057,858 / 100,000,000 tokens
FineWeb-Edu: 16,058,350 / 100,000,000 tokens
FineWeb-Edu: 17,058,910 / 100,000,000 tokens
FineWeb-Edu: 18,059,120 / 100,000,000 tokens
FineWeb-Edu: 19,059,214 / 100,000,000 tokens
FineWeb-Edu: 20,060,034 / 100,000,000 tokens
FineWeb-Edu: 21,061,034 / 100,000,000 tokens
FineWeb-Edu: 22,061,243 / 100,000,000 tokens
FineWeb-Edu: 23,061,857 / 100,000,000 tokens
FineWeb-Edu: 24,068,376 / 100,000,000 tokens
FineWeb-Edu: 25,068,451 / 100,000,000 tokens
FineWeb-Edu: 26,069,133 / 100,000,000 tokens
Step 200/2,289 | Loss 10.7533 | LR 2.990e-05 | 204,624 tok/s | 26,214,400 tokens | 8.74% | 0.04h
VRAM: 1.57 GB allocated / 4.69 GB reserved
Saving checkpoint-200...
Writing model shards: 100% 1/1 [00:00<00:00, 36.37it/s]
UPLOADING CHECKPOINT 200
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-200
FineWeb-Edu: 27,075,071 / 100,000,000 tokens
FineWeb-Edu: 28,080,714 / 100,000,000 tokens
FineWeb-Edu: 29,080,972 / 100,000,000 tokens
FineWeb-Edu: 30,085,136 / 100,000,000 tokens
FineWeb-Edu: 31,085,689 / 100,000,000 tokens
FineWeb-Edu: 32,086,415 / 100,000,000 tokens
FineWeb-Edu: 33,095,884 / 100,000,000 tokens
FineWeb-Edu: 34,096,837 / 100,000,000 tokens
FineWeb-Edu: 35,097,143 / 100,000,000 tokens
FineWeb-Edu: 36,097,343 / 100,000,000 tokens
FineWeb-Edu: 37,104,064 / 100,000,000 tokens
FineWeb-Edu: 38,104,207 / 100,000,000 tokens
FineWeb-Edu: 39,105,073 / 100,000,000 tokens
Step 300/2,289 | Loss 10.6609 | LR 2.952e-05 | 205,417 tok/s | 39,321,600 tokens | 13.11% | 0.06h
VRAM: 1.57 GB allocated / 4.69 GB reserved
Saving checkpoint-300...
Writing model shards: 100% 1/1 [00:00<00:00, 37.12it/s]
UPLOADING CHECKPOINT 300
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-300
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-100
FineWeb-Edu: 40,107,907 / 100,000,000 tokens
FineWeb-Edu: 41,108,319 / 100,000,000 tokens
FineWeb-Edu: 42,108,687 / 100,000,000 tokens
fyi this is colab L4
more logs agaaa
AppleMind 1.0 Mini
Colab L4 High-RAM
300M TOKEN PRETRAINING
Device: cuda
GPU: NVIDIA L4
VRAM: 22.03 GB
CUDA: 12.8
Total training tokens: 300,000,000
Target ratio: ~300:1
======================================================================
DATASET MIXTURE
FineWeb-Edu sample-10BT : 100,000,000
FineWeb-HQ : 100,000,000
SmolLM-Corpus : 100,000,000
TOTAL : 300,000,000
TinyStories: REMOVED
WikiText: REMOVED
Local datasets: REMOVED
Dataset cycling: DISABLED
======================================================================
LOADING GPT-2 TOKENIZER
Vocabulary: 50,260
Special tokens added: 3
======================================================================
CREATING APPLEMIND
Parameters: 1,020,480
Trainable: 1,020,480
Target tokens: 300,000,000
Target token/parameter ratio: 293.98:1
======================================================================
SETTING UP HUGGING FACE
Logged in as:
GGUFGuy
Checkpoint repository ready:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints
Final model repository ready:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini
Precision: BF16
======================================================================
TRAINING CONFIGURATION
Parameters: 1,020,480
Training tokens: 300,000,000
Tokens/parameter: 293.98:1
Micro batch size: 64
Gradient accumulation: 8
Effective batch: 512
Context length: 256
Tokens/micro batch: 16,384
Tokens/optimizer step: 131,072
Optimizer steps: 2,289
Warmup steps: 114
FineWeb-Edu: 100M
FineWeb-HQ: 100M
SmolLM: 100M
Checkpoint repository:
AppleMind-AI/AppleMind-1.0-Mini-Checkpoints
Final repository:
AppleMind-AI/AppleMind-1.0-Mini
======================================================================
STARTING 300M-TOKEN PRETRAINING
Opening FineWeb-Edu sample-10BT...
Resolving data files: 100% 2410/2410 [00:00<00:00, 23080.23it/s]
STREAMING FineWeb-Edu
Target tokens: 100,000,000
Progress state: 0
FineWeb-Edu: 1,000,131 / 100,000,000 tokens
FineWeb-Edu: 2,001,441 / 100,000,000 tokens
FineWeb-Edu: 3,001,677 / 100,000,000 tokens
FineWeb-Edu: 4,003,530 / 100,000,000 tokens
FineWeb-Edu: 5,003,864 / 100,000,000 tokens
FineWeb-Edu: 6,014,497 / 100,000,000 tokens
FineWeb-Edu: 7,014,690 / 100,000,000 tokens
FineWeb-Edu: 8,014,892 / 100,000,000 tokens
FineWeb-Edu: 9,015,759 / 100,000,000 tokens
FineWeb-Edu: 10,016,701 / 100,000,000 tokens
FineWeb-Edu: 11,042,613 / 100,000,000 tokens
FineWeb-Edu: 12,043,456 / 100,000,000 tokens
FineWeb-Edu: 13,044,297 / 100,000,000 tokens
Step 100/2,289 | Loss 10.8201 | LR 2.632e-05 | 100,666 tok/s | 13,107,200 tokens | 4.37% | 0.04h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-100...
Writing model shards: 100% 1/1 [00:00<00:00, 11.08it/s]
UPLOADING CHECKPOINT 100
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-100
FineWeb-Edu: 14,051,468 / 100,000,000 tokens
FineWeb-Edu: 15,057,858 / 100,000,000 tokens
FineWeb-Edu: 16,058,350 / 100,000,000 tokens
FineWeb-Edu: 17,058,910 / 100,000,000 tokens
FineWeb-Edu: 18,059,120 / 100,000,000 tokens
FineWeb-Edu: 19,059,214 / 100,000,000 tokens
FineWeb-Edu: 20,060,034 / 100,000,000 tokens
FineWeb-Edu: 21,061,034 / 100,000,000 tokens
FineWeb-Edu: 22,061,243 / 100,000,000 tokens
FineWeb-Edu: 23,061,857 / 100,000,000 tokens
FineWeb-Edu: 24,068,376 / 100,000,000 tokens
FineWeb-Edu: 25,068,451 / 100,000,000 tokens
FineWeb-Edu: 26,069,133 / 100,000,000 tokens
Step 200/2,289 | Loss 10.7533 | LR 2.990e-05 | 100,525 tok/s | 26,214,400 tokens | 8.74% | 0.07h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-200...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]
UPLOADING CHECKPOINT 200
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-200
FineWeb-Edu: 27,075,071 / 100,000,000 tokens
FineWeb-Edu: 28,080,714 / 100,000,000 tokens
FineWeb-Edu: 29,080,972 / 100,000,000 tokens
FineWeb-Edu: 30,085,136 / 100,000,000 tokens
FineWeb-Edu: 31,085,689 / 100,000,000 tokens
FineWeb-Edu: 32,086,415 / 100,000,000 tokens
FineWeb-Edu: 33,095,884 / 100,000,000 tokens
FineWeb-Edu: 34,096,837 / 100,000,000 tokens
FineWeb-Edu: 35,097,143 / 100,000,000 tokens
FineWeb-Edu: 36,097,343 / 100,000,000 tokens
FineWeb-Edu: 37,104,064 / 100,000,000 tokens
FineWeb-Edu: 38,104,207 / 100,000,000 tokens
FineWeb-Edu: 39,105,073 / 100,000,000 tokens
Step 300/2,289 | Loss 10.6609 | LR 2.952e-05 | 100,422 tok/s | 39,321,600 tokens | 13.11% | 0.11h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-300...
Writing model shards: 100% 1/1 [00:00<00:00, 10.96it/s]
UPLOADING CHECKPOINT 300
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-300
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-100
FineWeb-Edu: 40,107,907 / 100,000,000 tokens
FineWeb-Edu: 41,108,319 / 100,000,000 tokens
FineWeb-Edu: 42,108,687 / 100,000,000 tokens
FineWeb-Edu: 43,108,744 / 100,000,000 tokens
FineWeb-Edu: 44,109,584 / 100,000,000 tokens
FineWeb-Edu: 45,109,641 / 100,000,000 tokens
FineWeb-Edu: 46,110,333 / 100,000,000 tokens
FineWeb-Edu: 47,110,722 / 100,000,000 tokens
FineWeb-Edu: 48,111,125 / 100,000,000 tokens
FineWeb-Edu: 49,111,530 / 100,000,000 tokens
FineWeb-Edu: 50,113,218 / 100,000,000 tokens
FineWeb-Edu: 51,113,865 / 100,000,000 tokens
FineWeb-Edu: 52,114,128 / 100,000,000 tokens
Step 400/2,289 | Loss 10.5818 | LR 2.887e-05 | 100,639 tok/s | 52,428,800 tokens | 17.48% | 0.14h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-400...
Writing model shards: 100% 1/1 [00:00<00:00, 10.91it/s]
UPLOADING CHECKPOINT 400
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-400
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-200
FineWeb-Edu: 53,114,351 / 100,000,000 tokens
FineWeb-Edu: 54,115,268 / 100,000,000 tokens
FineWeb-Edu: 55,115,867 / 100,000,000 tokens
FineWeb-Edu: 56,118,747 / 100,000,000 tokens
FineWeb-Edu: 57,118,792 / 100,000,000 tokens
FineWeb-Edu: 58,129,696 / 100,000,000 tokens
FineWeb-Edu: 59,129,772 / 100,000,000 tokens
FineWeb-Edu: 60,130,463 / 100,000,000 tokens
FineWeb-Edu: 61,130,846 / 100,000,000 tokens
FineWeb-Edu: 62,131,974 / 100,000,000 tokens
FineWeb-Edu: 63,145,958 / 100,000,000 tokens
FineWeb-Edu: 64,146,051 / 100,000,000 tokens
FineWeb-Edu: 65,146,257 / 100,000,000 tokens
Step 500/2,289 | Loss 10.5081 | LR 2.797e-05 | 100,116 tok/s | 65,536,000 tokens | 21.85% | 0.18h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-500...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]
UPLOADING CHECKPOINT 500
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-500
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-300
FineWeb-Edu: 66,147,176 / 100,000,000 tokens
FineWeb-Edu: 67,153,903 / 100,000,000 tokens
FineWeb-Edu: 68,163,652 / 100,000,000 tokens
FineWeb-Edu: 69,164,783 / 100,000,000 tokens
FineWeb-Edu: 70,166,355 / 100,000,000 tokens
FineWeb-Edu: 71,166,545 / 100,000,000 tokens
FineWeb-Edu: 72,167,733 / 100,000,000 tokens
FineWeb-Edu: 73,168,881 / 100,000,000 tokens
FineWeb-Edu: 74,175,780 / 100,000,000 tokens
FineWeb-Edu: 75,176,435 / 100,000,000 tokens
FineWeb-Edu: 76,177,413 / 100,000,000 tokens
FineWeb-Edu: 77,177,431 / 100,000,000 tokens
FineWeb-Edu: 78,179,729 / 100,000,000 tokens
Step 600/2,289 | Loss 10.4357 | LR 2.682e-05 | 100,691 tok/s | 78,643,200 tokens | 26.21% | 0.22h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-600...
Writing model shards: 100% 1/1 [00:00<00:00, 11.04it/s]
UPLOADING CHECKPOINT 600
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-600
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-400
FineWeb-Edu: 79,187,841 / 100,000,000 tokens
FineWeb-Edu: 80,187,924 / 100,000,000 tokens
FineWeb-Edu: 81,188,888 / 100,000,000 tokens
FineWeb-Edu: 82,189,255 / 100,000,000 tokens
FineWeb-Edu: 83,189,443 / 100,000,000 tokens
FineWeb-Edu: 84,192,547 / 100,000,000 tokens
FineWeb-Edu: 85,192,626 / 100,000,000 tokens
FineWeb-Edu: 86,193,916 / 100,000,000 tokens
FineWeb-Edu: 87,194,129 / 100,000,000 tokens
FineWeb-Edu: 88,196,386 / 100,000,000 tokens
FineWeb-Edu: 89,203,193 / 100,000,000 tokens
FineWeb-Edu: 90,203,499 / 100,000,000 tokens
FineWeb-Edu: 91,204,016 / 100,000,000 tokens
Step 700/2,289 | Loss 10.3676 | LR 2.546e-05 | 100,440 tok/s | 91,750,400 tokens | 30.58% | 0.25h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-700...
Writing model shards: 100% 1/1 [00:00<00:00, 10.87it/s]
UPLOADING CHECKPOINT 700
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-700
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-500
FineWeb-Edu: 92,205,445 / 100,000,000 tokens
FineWeb-Edu: 93,206,406 / 100,000,000 tokens
FineWeb-Edu: 94,215,721 / 100,000,000 tokens
FineWeb-Edu: 95,216,609 / 100,000,000 tokens
FineWeb-Edu: 96,218,844 / 100,000,000 tokens
FineWeb-Edu: 97,219,420 / 100,000,000 tokens
FineWeb-Edu: 98,219,961 / 100,000,000 tokens
FineWeb-Edu: 99,220,378 / 100,000,000 tokens
FineWeb-Edu complete: 100,000,000 tokens
Opening FineWeb-HQ...
Resolving data files: 100% 9246/9246 [00:00<00:00, 24460.23it/s]
STREAMING FineWeb-HQ
Target tokens: 100,000,000
Progress state: 0
FineWeb-HQ: 1,000,237 / 100,000,000 tokens
FineWeb-HQ: 2,000,353 / 100,000,000 tokens
FineWeb-HQ: 3,000,648 / 100,000,000 tokens
FineWeb-HQ: 4,001,578 / 100,000,000 tokens
Step 800/2,289 | Loss 10.3038 | LR 2.391e-05 | 95,154 tok/s | 104,857,600 tokens | 34.95% | 0.29h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-800...
Writing model shards: 100% 1/1 [00:00<00:00, 11.02it/s]
UPLOADING CHECKPOINT 800
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-800
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-600
FineWeb-HQ: 5,003,165 / 100,000,000 tokens
FineWeb-HQ: 6,003,192 / 100,000,000 tokens
FineWeb-HQ: 7,003,245 / 100,000,000 tokens
FineWeb-HQ: 8,005,095 / 100,000,000 tokens
FineWeb-HQ: 9,005,500 / 100,000,000 tokens
FineWeb-HQ: 10,005,995 / 100,000,000 tokens
FineWeb-HQ: 11,006,275 / 100,000,000 tokens
FineWeb-HQ: 12,006,703 / 100,000,000 tokens
FineWeb-HQ: 13,008,348 / 100,000,000 tokens
FineWeb-HQ: 14,039,910 / 100,000,000 tokens
FineWeb-HQ: 15,040,042 / 100,000,000 tokens
FineWeb-HQ: 16,040,179 / 100,000,000 tokens
FineWeb-HQ: 17,040,530 / 100,000,000 tokens
Step 900/2,289 | Loss 10.2507 | LR 2.221e-05 | 99,896 tok/s | 117,964,800 tokens | 39.32% | 0.33h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-900...
Writing model shards: 100% 1/1 [00:00<00:00, 10.81it/s]
UPLOADING CHECKPOINT 900
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-900
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-700
FineWeb-HQ: 18,045,133 / 100,000,000 tokens
FineWeb-HQ: 19,053,155 / 100,000,000 tokens
FineWeb-HQ: 20,053,155 / 100,000,000 tokens
FineWeb-HQ: 21,053,631 / 100,000,000 tokens
FineWeb-HQ: 22,061,664 / 100,000,000 tokens
FineWeb-HQ: 23,061,940 / 100,000,000 tokens
FineWeb-HQ: 24,062,978 / 100,000,000 tokens
FineWeb-HQ: 25,063,229 / 100,000,000 tokens
FineWeb-HQ: 26,073,081 / 100,000,000 tokens
FineWeb-HQ: 27,073,220 / 100,000,000 tokens
FineWeb-HQ: 28,073,681 / 100,000,000 tokens
FineWeb-HQ: 29,073,966 / 100,000,000 tokens
FineWeb-HQ: 30,074,272 / 100,000,000 tokens
Step 1,000/2,289 | Loss 10.1952 | LR 2.039e-05 | 99,484 tok/s | 131,072,000 tokens | 43.69% | 0.36h
VRAM: 1.57 GB allocated / 15.40 GB reserved
Saving checkpoint-1000...
Writing model shards: 100% 1/1 [00:00<00:00, 10.97it/s]
UPLOADING CHECKPOINT 1000
Checkpoint uploaded to:
https://huggingface.co/AppleMind-AI/AppleMind-1.0-Mini-Checkpoints/tree/main/checkpoint-1000
Removing old local checkpoint: /content/applemind-1.0-mini-checkpoints/checkpoint-800
FineWeb-HQ: 31,075,196 / 100,000,000 tokens
FineWeb-HQ: 32,075,679 / 100,000,000 tokens