mitra-qwen35-base-stage2

The 9B foundation model of the Dharmamitra stack: Qwen3.5-9B after large-scale domain adaptation to classical Buddhist literature. Stage 2 adds translation-specific capabilities on top of the stage-1 pretraining: bidirectional parallel translation training (Sanskrit↔English, Sanskrit↔German, and related classical pairs) plus full production translation/research turnarounds at 16k context. This is the model that buddhist-nlp/mitra-qwen35-embedder was finetuned from, released for research use as a starting point for Buddhist-NLP downstream tasks (translation, retrieval finetunes, QA, philology tooling).

Training

Two stages on top of Qwen3.5-9B:

  1. Continued pretraining (~30B tokens, 8k context) on classical Buddhist corpora: Sanskrit, Tibetan (Wylie transliteration), Buddhist Chinese, and Pāli source texts with related secondary literature.
  2. Stage-2 finetuning (16k context, 1 epoch) on a mixture of bidirectional parallel translation data (Sanskrit↔English, Sanskrit↔German, and related pairs), full production translation/research turnarounds, monolingual replay, and general instructions.

The result is an instruction-following model with strong classical-language competence in reading, translating, and discussing Buddhist source texts.

Conventions

  • Tibetan input/output is Wylie transliteration, not Tibetan script.
  • Sanskrit and Pāli are IAST romanization.
  • A chat template is bundled (tokenizer.apply_chat_template).

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("buddhist-nlp/mitra-qwen35-base-stage2")
model = AutoModelForCausalLM.from_pretrained(
    "buddhist-nlp/mitra-qwen35-base-stage2", dtype=torch.bfloat16, device_map="cuda"
)

messages = [{"role": "user", "content":
    "Translate into English: evaṃ mayā śrutam ekasmin samaye"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
                                 return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Relation to other releases

Citation

If you use this model, please cite the Dharmamitra project (https://dharmamitra.org). A technical report is in preparation.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for buddhist-nlp/mitra-qwen35-base-stage2

Finetunes
1 model

Collection including buddhist-nlp/mitra-qwen35-base-stage2