YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Czech Question Answering Model
Overview
A question answering model fine-tuned on Czech language data using mBART-large. The model takes questions in Czech and generates responses in Czech.
Model Details
- Base Model: facebook/nllb-200-distilled-1.3B
- Training Method: LoRA fine-tuning
- BLEU Score: 19.90
- Input Language: Czech
- Output Language: Czech
- Sequence Length: 64 tokens
Training Configuration
LoRA Config:
- r=16
- lora_alpha=32
- target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]
- lora_dropout=0.1
Training Args:
- learning_rate=2e-4
- per_device_train_batch_size=8
- gradient_accumulation_steps=2
- num_train_epochs=3
- weight_decay=0.01
- warmup_steps=100
Usage
from transformers import M2M100ForConditionalGeneration, AutoTokenizer
# Load model and tokenizer
model_name = "koushikkanch/large-model"
model = M2M100ForConditionalGeneration.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Example usage
question = "Je tu obsazeno?"
inputs = tokenizer(question, return_tensors="pt")
outputs = model.generate(**inputs)
answer = tokenizer.decode(outputs[0], skip_special_tokens=True)
Hardware Requirements
- GPU with minimum 12GB VRAM recommended
- Tested on NVIDIA T4 GPU
Limitations
- Current version shows some issues with response coherency
- Model occasionally generates non-Czech responses
- Further optimization needed for better quality responses
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support