localmodels
/

OpenAssistant-SFT-7-LLaMA-30B-ggml

Model card Files Files and versions

OpenAssistant-SFT-7-LLaMA-30B-ggml / README.md

ukim4

Duplicate from localmodels/LLM

5f8b044 over 1 year ago

preview code

raw

history blame

4.14 kB

	---
	duplicated_from: localmodels/LLM
	---
	# OpenAssistant SFT 7 LLaMA 30B ggml

	From: https://huggingface.co/OpenAssistant/oasst-sft-7-llama-30b-xor

	---

	### Original llama.cpp quant methods: `q4_0, q4_1, q5_0, q5_1, q8_0`

	Quantized using an older version of llama.cpp and compatible with llama.cpp from May 19, commit 2d5db48.

	### k-quant methods: `q2_K, q3_K_S, q3_K_M, q3_K_L, q4_K_S, q4_K_M, q5_K_S, q6_K`

	Quantization methods compatible with latest llama.cpp from June 6, commit 2d43387.

	---

	## Provided files
	\| Name \| Quant method \| Bits \| Size \| Max RAM required \| Use case \|
	\| ---- \| ---- \| ---- \| ---- \| ---- \| ----- \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q2_K.bin \| q2_K \| 2 \| 13.60 GB \| 16.10 GB \| New k-quant method. Uses GGML_TYPE_Q4_K for the attention.vw and feed_forward.w2 tensors, GGML_TYPE_Q2_K for the other tensors. \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q3_K_L.bin \| q3_K_L \| 3 \| 17.20 GB \| 19.70 GB \| New k-quant method. Uses GGML_TYPE_Q5_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else GGML_TYPE_Q3_K \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q3_K_M.bin \| q3_K_M \| 3 \| 15.64 GB \| 18.14 GB \| New k-quant method. Uses GGML_TYPE_Q4_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else GGML_TYPE_Q3_K \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q3_K_S.bin \| q3_K_S \| 3 \| 13.98 GB \| 16.48 GB \| New k-quant method. Uses GGML_TYPE_Q3_K for all tensors \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q4_0.bin \| q4_0 \| 4 \| 18.30 GB \| 20.80 GB \| Original llama.cpp quant method, 4-bit. \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q4_K_M.bin \| q4_K_M \| 4 \| 19.57 GB \| 22.07 GB \| New k-quant method. Uses GGML_TYPE_Q6_K for half of the attention.wv and feed_forward.w2 tensors, else GGML_TYPE_Q4_K \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q4_K_S.bin \| q4_K_S \| 4 \| 18.30 GB \| 20.80 GB \| New k-quant method. Uses GGML_TYPE_Q4_K for all tensors \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q5_0.bin \| q5_0 \| 5 \| 22.37 GB \| 24.87 GB \| Original llama.cpp quant method, 5-bit. Higher accuracy, higher resource usage and slower inference. \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q5_1.bin \| q5_1 \| 5 \| 24.40 GB \| 26.90 GB \| Original llama.cpp quant method, 5-bit. Even higher accuracy, resource usage and slower inference. \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q5_K_M.bin \| q5_K_M \| 5 \| 23.02 GB \| 25.52 GB \| New k-quant method. Uses GGML_TYPE_Q6_K for half of the attention.wv and feed_forward.w2 tensors, else GGML_TYPE_Q5_K \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q5_K_S.bin \| q5_K_S \| 5 \| 22.37 GB \| 24.87 GB \| New k-quant method. Uses GGML_TYPE_Q5_K for all tensors \|
	\| OpenAssistant-SFT-7-Llama-30B.ggmlv3.q6_K.bin \| q6_K \| 6 \| 26.69 GB \| 29.19 GB \| New k-quant method. Uses GGML_TYPE_Q8_K - 6-bit quantization - for all tensors \|

	---

	# OpenAssistant LLaMA 30B SFT 7

	### Configuration

	```
	llama-30b-sft-7:
	dtype: fp16
	log_dir: "llama_log_30b"
	learning_rate: 1e-5
	model_name: /home/ubuntu/Open-Assistant/model/model_training/.saved/llama-30b-super-pretrain/checkpoint-3500
	#model_name: OpenAssistant/llama-30b-super-pretrain
	output_dir: llama_model_30b
	deepspeed_config: configs/zero3_config_sft.json
	weight_decay: 0.0
	residual_dropout: 0.0
	max_length: 2048
	use_flash_attention: true
	warmup_steps: 20
	gradient_checkpointing: true
	gradient_accumulation_steps: 12
	per_device_train_batch_size: 2
	per_device_eval_batch_size: 3
	eval_steps: 101
	save_steps: 485
	num_train_epochs: 4
	save_total_limit: 3
	use_custom_sampler: true
	sort_by_length: false
	#save_strategy: steps
	save_strategy: epoch
	datasets:
	- oasst_export:
	lang: "bg,ca,cs,da,de,en,es,fr,hr,hu,it,nl,pl,pt,ro,ru,sl,sr,sv,uk"
	input_file_path: 2023-04-12_oasst_release_ready_synth.jsonl.gz
	val_split: 0.05
	- vicuna:
	val_split: 0.05
	max_val_set: 800
	fraction: 1.0
	- dolly15k:
	val_split: 0.05
	max_val_set: 300
	- grade_school_math_instructions:
	val_split: 0.05
	- code_alpaca:
	val_split: 0.05
	max_val_set: 250
	```

	- OASST dataset paper: https://arxiv.org/abs/2304.07327