LLM D6 (~0.8B)

LLM D6 is the 6th model in the series, built on a custom LLaMA-3 architecture. It features 0.8B parameters and was pre-trained from scratch on 300 Billion tokens of high-quality, diverse data spanning English, Russian, programming code, and mathematical reasoning.

This repository (firdavsus/LLM_multimodal) contains both the Pre-trained base weights and the Fine-tuned instruct/chat weights, along with training logs and evaluation benchmarks.


Benchmarks

benchmark not available

Fine Tuning Grapth

benchmark not available

πŸ’‘ Key Highlights

  • Bilingual & Domain-Focused: Pre-trained on a balanced mix of English, Russian, Code, and Mathematics.
  • Custom Tokenizer: Trained from scratch on the dataset with a 52,000 vocabulary size, yielding higher token compression efficiency for Russian and technical syntax.
  • High Token Efficiency: Achieves competitive performance at 300B tokens compared to models trained on 9T–18T tokens.
  • Dual Weights Included: Structured subdirectories for both pre-train and fine-tune stages.

πŸ“ Model Architecture & Specifications

Parameter Value
Architecture LLaMA-3 style (LlamaForCausalLM)
Total Parameters ~800M (0.8B)
Hidden Dimension ($d_{\text{model}}$) 1536
Number of Layers 24
Head Dimension 128
Attention Heads 12
Vocabulary Size 52,000 (Custom Trained)
Context Pre-training Tokens 300 Billion Tokens

πŸ“ Repository Structure

firdavsus/LLM_multimodal/
β”œβ”€β”€ pre-train/                 # Pre-trained base weights & tokenizer
β”œβ”€β”€ fine-tune/                 # Supervised Fine-Tuned (SFT) weights & configs
β”œβ”€β”€ benchmarks.png             # Evaluation comparison curves & charts
β”œβ”€β”€ training_curves_finetune.png # SFT Loss, LR, and Gradient stability plots
└── README.md                  # Model Card

πŸ“Š Evaluation & Benchmarks

The model was evaluated using lm-evaluation-harness across zero-shot and few-shot tasks, comparing Custom 0.8B (300B tokens) against Qwen 2.5 0.5B (18.0T tokens) and Llama 3.2 1B (9.0T tokens).

Final Benchmark Performance Summary

Benchmark Task Metric Custom 0.8B (300B Tokens) Qwen 2.5 0.5B (18,000B Tokens) Llama 3.2 1B (9,000B Tokens)
PIQA Physical Commonsense 69.3% 69.8% 74.5%
HellaSwag General Reasoning 47.8% 52.2% 64.1%
XWinograd RU Russian Language Understanding 58.7% 56.8% 66.0%
ARC-Challenge Grade-School Science 30.6% 32.2% 37.2%
MMLU Multi-task Accuracy 25.5% 47.4% 37.8%

Note: Despite being trained on significantly fewer tokens (0.3T vs 18T/9T), LLM D6 demonstrates strong bilingual capability, outperforming Qwen 2.5 0.5B on Russian Winograd evaluation (58.7% vs 56.8%) and maintaining competitive common-sense physical reasoning (69.3% on PIQA).


πŸ“‰ Fine-Tuning Dynamics

Fine-tuning was conducted over 3,600 steps using a cosine learning rate decay schedule with a peak learning rate of $2.0 \times 10^{-5}$ and initial linear warmup.

  • Loss Convergence: Both training and evaluation loss decayed smoothly down to ~1.16, showing consistent learning without over-fitting.
  • Gradient Stability: Gradient norm remained tightly bounded between 0.5 and 1.5 throughout the execution.

πŸš€ How to Use

Loading the Fine-Tuned Model

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "firdavsus/LLM_multimodal"

# Load Custom Tokenizer and Model Weights from fine-tune folder
tokenizer = AutoTokenizer.from_pretrained(model_id, subfolder="fine-tune")
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    subfolder="fine-tune",
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "Solve the following problem step by step: What is the derivative of f(x) = x^3 + 2x?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs, 
    max_new_tokens=256, 
    temperature=0.7, 
    top_p=0.9
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Loading the Pre-Trained Base Weights

tokenizer = AutoTokenizer.from_pretrained("firdavsus/LLM_multimodal", subfolder="pre-train")
model = AutoModelForCausalLM.from_pretrained("firdavsus/LLM_multimodal", subfolder="pre-train")

⚠️ Limitations & Intended Use

  • Primary Intended Use: General text generation, bilingual English/Russian translation, code assistance, and basic mathematical reasoning.
  • Limitations: Due to its lightweight 0.8B parameter size, complex multi-step reasoning (e.g., MMLU domain knowledge) may show limitations compared to much larger models. Users should implement system prompts or validation guardrails for mission-critical deployments.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support