SagaLM Logo

SagaLM-slm2

SagaLM is a QLoRA/SFT fine-tuned Large Language Model derived from Qwen/Qwen2.5-7B-Instruct.

Important

The underlying architecture remains compatible with the Qwen2/Qwen2.5 Transformers implementation. SagaLM is the model and assistant identity; the internal model_type is intentionally not renamed to ensure compatibility with the Transformers ecosystem.

Training

  • Base model: Qwen/Qwen2.5-7B-Instruct
  • Context length: 2,048 tokens
  • Fine-tuning method: QLoRA / Supervised Fine-Tuning (SFT)
  • LoRA rank: 16
  • LoRA alpha: 32
  • LoRA dropout: 0.05
  • Effective batch size: 4
  • Training steps: 1,200
  • Completion-only loss (prompt tokens masked to -100)
  • Long-answer token filtering with short identity-grounding examples
  • Training mixture includes instruction following, conversations, reasoning, mathematics, and coding datasets

Identity

SagaLM is trained to identify itself as SagaLM rather than Qwen, ChatGPT, Claude, or another named assistant.

Generation Recommendation

  • max_new_tokens: 1024
  • temperature: 0.7
  • top_p: 0.9
  • top_k: 50
  • repetition_penalty: 1.05
Downloads last month
651
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for venkateshchsagalm/SagaLM-slm2

Base model

Qwen/Qwen2.5-7B
Finetuned
(3041)
this model

Space using venkateshchsagalm/SagaLM-slm2 1