Yejin Korean 8B v3 Base

A Korean-English bilingual foundation language model with 8 billion parameters, pre-trained from scratch on a large-scale curated bilingual corpus.

Model Details

  • Model Name: Yejin-Korean-8B-v3-Base
  • Parameters: ~7.52B (Tied Word Embeddings)
  • Architecture: Transformer Decoder with QK-Norm, GQA, and SwiGLU
  • Hidden Size: 4096
  • Layers: 32
  • Attention Heads: 32 (Query) / 8 (Key-Value) โ€” Grouped Query Attention (4:1)
  • Head Dimension: 128
  • Intermediate Size: 14,336
  • Context Length: 4096 tokens
  • RoPE Base (Theta): 500,000.0
  • Vocabulary Size: 64,000 (Optimized Korean-English BPE)
  • Precision: BF16 (bfloat16)

Training

Pre-trained from scratch using PyTorch FSDP on an 8ร— NVIDIA H200 GPU cluster. The model was trained on a balanced, curated corpus consisting of extensive Korean web, encyclopedia, academic, news, and conversational datasets, alongside high-quality English educational and technical corpora.

Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "cooler8/yejin-korean-8b-v3-base"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "๋Œ€ํ•œ๋ฏผ๊ตญ์˜ ์ˆ˜๋„๋Š”"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=64,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
    repetition_penalty=1.1
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Notice

This is a pre-trained base model without instruction alignment. For interactive chat and assistant capabilities, please refer to the upcoming SFT / DPO aligned releases.

License

Apache 2.0

Downloads last month
432
Safetensors
Model size
7B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support