Saxo's picture
Update README.md
3a9fdb2 verified
metadata
library_name: transformers
license: apache-2.0
base_model: meta-llama/Meta-Llama-3-8B-Instruct
datasets:
  - Saxo/total_ko_train_set_1_without_wiki_with_orca
language:
  - ko
  - en
  - ja
  - zh
pipeline_tag: text-generation

Model Card for Model ID

AI ์™€ ๋น…๋ฐ์ดํ„ฐ ๋ถ„์„ ์ „๋ฌธ ๊ธฐ์—…์ธ Linkbricks์˜ ๋ฐ์ดํ„ฐ์‚ฌ์ด์–ธํ‹ฐ์ŠคํŠธ์ธ ์ง€์œค์„ฑ(Saxo) ์ด์‚ฌ๊ฐ€ meta-llama/Meta-Llama-3-8B๋ฅผ ๋ฒ ์ด์Šค๋ชจ๋ธ๋กœ GCP์ƒ์˜ H100-80G 8๊ฐœ๋ฅผ ํ†ตํ•ด SFT-DPO ํ›ˆ๋ จํ•œ ํ•œ๊ธ€ ๊ธฐ๋ฐ˜ LLAMA3-8b 8๊ฐœ์˜ MoE(Mixture of Expert)๋ชจ๋ธ. ํ† ํฌ๋‚˜์ด์ €๋Š” ๋ผ๋งˆ3๋ž‘ ๋™์ผํ•˜๋ฉฐ ํ•œ๊ธ€ VOCA ํ™•์žฅ์€ ํ•˜์ง€ ์•Š์€ ๋ฒ„์ „ ์ž…๋‹ˆ๋‹ค. ์ผ๋ฐ˜์งˆ์˜์‘๋‹ต(์ฑ„ํŒ…)-์˜๋ฃŒ-๊ตฐ์‚ฌ-ํ•œ์ค‘์ผ๋ฒˆ์—ญ-์ฝ”๋”ฉ ๊ฐ ํŠนํ™” LLM์„ ํ†ตํ•ฉ

Dr. Yunsung Ji (Saxo), a data scientist at Linkbricks, a company specializing in AI and big data analytics, trained the meta-llama/Meta-Llama-3-8B base model on 8 H100-60Gs on GCP for 4 hours of instructional training (8000 Tokens). Accelerate, Deepspeed Zero-3 libraries were used.

www.linkbricks.com, www.linkbricks.vc