YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Czech Question Answering Model

Overview

A question answering model fine-tuned on Czech language data using mBART-large. The model takes questions in Czech and generates responses in Czech.

Model Details

  • Base Model: facebook/nllb-200-distilled-1.3B
  • Training Method: LoRA fine-tuning
  • BLEU Score: 19.90
  • Input Language: Czech
  • Output Language: Czech
  • Sequence Length: 64 tokens

Training Configuration

LoRA Config:
- r=16
- lora_alpha=32
- target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]
- lora_dropout=0.1

Training Args:
- learning_rate=2e-4
- per_device_train_batch_size=8
- gradient_accumulation_steps=2
- num_train_epochs=3
- weight_decay=0.01
- warmup_steps=100

Usage

from transformers import M2M100ForConditionalGeneration, AutoTokenizer

# Load model and tokenizer
model_name = "koushikkanch/large-model"
model = M2M100ForConditionalGeneration.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Example usage
question = "Je tu obsazeno?"
inputs = tokenizer(question, return_tensors="pt")
outputs = model.generate(**inputs)
answer = tokenizer.decode(outputs[0], skip_special_tokens=True)

Hardware Requirements

  • GPU with minimum 12GB VRAM recommended
  • Tested on NVIDIA T4 GPU

Limitations

  • Current version shows some issues with response coherency
  • Model occasionally generates non-Czech responses
  • Further optimization needed for better quality responses

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support