mitra-qwen35-it

The general-purpose instruction/chat model of the Dharmamitra Qwen3.5 family: a multi-turn SFT of mitra-qwen35-base-stage2 for Buddhist-studies conversation — closed-book Buddhism Q&A plus translation and translation-refinement assistance for classical languages (Sanskrit, Tibetan, Buddhist Chinese, Pāli).

This is the recommended entry point if you want to talk to the mitra family; use the stage-2 base for raw translation pipelines and the embedder for retrieval.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("buddhist-nlp/mitra-qwen35-it")
model = AutoModelForCausalLM.from_pretrained(
    "buddhist-nlp/mitra-qwen35-it", dtype=torch.bfloat16, device_map="cuda"
)
messages = [{"role": "user", "content":
    "What is the difference between śamatha and vipaśyanā?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
                                 return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Tibetan may be given in Wylie transliteration; Sanskrit and Pāli in IAST.

Training details

  • Base: buddhist-nlp/mitra-qwen35-base-stage2 (9B; ~30B tokens Buddhist CPT
    • translation-capability SFT)
  • Multi-turn SFT (TRL, completion-only loss over full chat histories), LR 1e-5, on a combined corpus of closed-book Buddhism Q&A (knowledge distilled into weights) and translation/refinement tasks across zh/sa/bo/pi

Citation

If you use this model, please cite the Dharmamitra project (https://dharmamitra.org). A technical report is in preparation.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for buddhist-nlp/mitra-qwen35-it

Finetuned
(1)
this model

Collection including buddhist-nlp/mitra-qwen35-it