Instructions to use 1017qud/korean-document-qa-enhanced with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 1017qud/korean-document-qa-enhanced with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("question-answering", model="1017qud/korean-document-qa-enhanced")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForQuestionAnswering tokenizer = AutoTokenizer.from_pretrained("1017qud/korean-document-qa-enhanced") model = AutoModelForQuestionAnswering.from_pretrained("1017qud/korean-document-qa-enhanced", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Korean Document QA Enhanced
기존 model_10000을 출발점으로 KorQuAD v1.0의 새 10,000문항을 learning rate 2e-5, 1 epoch로 추가 fine-tuning한 한국어 추출형 질의응답 모델입니다.
평가 결과
KorQuAD dev에서 seed 42로 고정한 같은 500문항을 사용했습니다.
| 모델 | EM | 문자 F1 | 정확 경계 |
|---|---|---|---|
| 기존 model_10000 | 0.808 | 0.9122 | 0.782 |
| 이 모델 | 0.822 | 0.9256 | 0.800 |
confidence는 모델 선정 기준에서 제외했습니다. F1, EM, 정확 경계와 소규모 무응답 경고 결과를 함께 사용했습니다.
한계
- 답이 없는 질문이 포함된 데이터셋으로 학습하지 않았습니다.
- 직접 만든 무응답 질문 10개에서 경고 정확도 80%였지만 소규모 보조 결과입니다.
- 실제 PDF에서 confidence가 높아도 짧은 오답을 출력한 사례가 있었습니다.
- 영어 문맥에 한국어로 질문할 때 답을 놓치거나 정답 후보에도 경고가 나타날 수 있습니다.
- Downloads last month
- 9