ko-hand-ocr
한글·영문이 섞인 손글씨 한 줄을 읽는다. 손글씨 데이터셋 없이, 합성 그림만으로 처음부터 학습한 41M 모델. Apache-2.0 — 가중치도 같다.
GitHub · PyPI · 만든 과정 (PDF) · 한국어 · English
맨 위 표는 공개된 다른 OCR 여섯(ko-trocr · PaddleOCR PP-OCRv5 · EasyOCR · Tesseract 5 ·
PaddleOCR-VL-1.6 · Qwen3-VL-2B)과 같은 시험지·같은 PC·같은 순간에 견준 것이다 — 누구를 어떻게
돌렸는지는 아래. 그 아래는 우리 두 모델의 항목별 성적이고,
tools/verify.py 로 글씨체마다 약 960줄씩 잰 값이다. 단일 모델은 41M 모델 하나로 읽는다.
앙상블은 모델 3개가 같이 읽고 서로 닮은 답을 고른다 — 더 정확하고 일곱 배 느리다.
받기와 쓰기
이 저장소의 가중치는 GitHub v0.4.1 릴리스의 zip 과 같은 파일이다 (sha256 같음).
| 자리 | 무엇 | 크기 |
|---|---|---|
/ (맨 위) |
단일 모델 — 41M, 디코더 8층. 대부분 이것으로 충분하다 | 156MB |
ensemble/ |
앙상블 — 모델 3개(36M + 41M + 41M)가 같이 읽고 서로 닮은 답을 고른다. 바닥을 올리고 7배 느리다 | 450MB |
무게만으로는 안 읽힌다. 자모 어휘와 줄 자르기가 kohandocr
꾸러미에 있으니 먼저 설치한다. transformers 의 pipeline 으로는 못 연다.
pip install ko-hand-ocr
from huggingface_hub import snapshot_download
from kohandocr.reader import Reader
# 단일 모델 (156MB)
folder = snapshot_download("localdeel/ko-hand-ocr", ignore_patterns=["ensemble/*", "assets/*"])
reader = Reader(folder, device="cpu")
# 사진 한 장을 통째로. 줄 자르기까지 해 준다.
reader.read_photo(open("scan.jpg", "rb").read())
# -> ['보안점검표', '부서 : 포테토뭉부서', '이름 : 감자밭']
# 이미 줄을 잘라 두었다면 그림(PIL)을 묶음으로 넘긴다.
reader.read([cell_a, cell_b, cell_c])
# 앙상블 (450MB) — ensemble/also/ 가 있으면 Reader 가 알아서 모델 3개로 읽는다
folder = snapshot_download("localdeel/ko-hand-ocr", allow_patterns=["ensemble/*"])
reader = Reader(folder + "/ensemble", device="cpu")
후보 목록 안에서만 고르게 가두는 both(), 읽은 값과 믿음값을 같이 주는 read_trust()(믿음값은
모델끼리 견주어 재므로 앙상블에서만 나온다) 같은 나머지 쓰는 법은 GitHub README 에 있다.
LM Studio · Ollama · vLLM 과 같이 쓰기
모델로 올리지는 못한다. 셋은 글을 이어 쓰는 LLM 을 돌리는 엔진이다. LM Studio·Ollama 는 llama.cpp 의 GGUF 만, vLLM 은 자기가 아는 구조만 돌린다. 이 모델은 그림 인코더에 교차 주의 디코더를 단 TrOCR 꼴이고 자모 어휘와 줄 자르기가 따로 있어서, GGUF 로 바꿀 수도 그 목록에 넣을 수도 없다. 대신 그 엔진들과 같은 말(API)을 하는 서버와 MCP 도구를 넣었다.
pip install "ko-hand-ocr[mcp]" # 서버만 쓸 거면 [mcp] 는 빼도 된다
서버 — OpenAI·Ollama 와 같은 말
ko-hand-ocr-serve # 허깅페이스에서 단일 모델(156MB)을 받아 127.0.0.1:8765 에 띄운다
ko-hand-ocr-serve --ensemble # 앙상블(450MB)
| 부르는 쪽 | 이렇게 |
|---|---|
openai 클라이언트, vLLM 을 부르던 코드 |
base_url="http://127.0.0.1:8765/v1" 로 바꾼다. 그림은 image_url 에 base64 로 |
ollama CLI |
OLLAMA_HOST=127.0.0.1:8765 를 주고 ollama run ko-hand-ocr "C:\scan.jpg" · ollama list |
| curl | curl --data-binary @scan.jpg http://127.0.0.1:8765/read |
import base64
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8765/v1", api_key="none")
picture = "data:image/jpeg;base64," + base64.b64encode(open("scan.jpg", "rb").read()).decode()
reply = client.chat.completions.create(model="ko-hand-ocr", messages=[
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": picture}}]}])
print(reply.choices[0].message.content) # 줄마다 한 줄
- 글로 물은 것은 쓰지 않는다. 그림을 줄로 잘라 읽고 줄마다 한 줄씩 돌려줄 뿐이다. 요약·정리는 못 한다 — 그건 채팅 모델 몫이다(아래 MCP).
- 그림 주소(
http://…)는 받으러 가지 않는다. base64 로 실어 보낸다. 아무 데도 나가지 않는다. - 기본은 이 PC(
127.0.0.1)에서만 열린다. 비밀번호가 없으니--host 0.0.0.0은 믿는 망에서만.
LM Studio — MCP 도구로
LM Studio(0.3.17 부터)의 채팅 모델은 바깥 도구를 부를 수 있다(MCP). ~/.lmstudio/mcp.json 에:
{"mcpServers": {"ko-hand-ocr": {"command": "ko-hand-ocr-mcp"}}}
ko-hand-ocr-mcp 가 PATH 에 없으면 전체 경로를 적는다(가상환경의 Scripts\ko-hand-ocr-mcp.exe).
그 뒤 도구를 부를 줄 아는 채팅 모델에게 "C:\scan.jpg 읽고 표로 정리해 줘" 라고 하면, 모델이
read_handwriting 으로 글자를 받아 정리한다. 글자를 알아보는 것은 이 모델, 정리는 채팅
모델이다. 채팅 창에 붙인 그림은 도구에게 오지 않으므로 파일 경로로 준다. 처음 부를 때 모델을
받는다(156MB, 이 PC 에서 20~30초). 그다음은 사진 한 장에 1초 안팎이다.
확인한 것과 안 한 것. ollama CLI 0.34(list·show·run), 공식 openai 클라이언트(한
번에·흘려 받기), MCP SDK 2.3 손님으로 불러 보았다. LM Studio 채팅 창과 Open WebUI 화면 안에서
끝까지 해 본 것은 아직 아니다.
vLLM 자체에 넣으려면 플러그인으로 모델 구조를 새로 써야 한다. 41M 모델이라 GPU 서버로 얻을 것이 작아서 하지 않았다 — 위 서버가 vLLM 과 같은 OpenAI 주소를 낸다.
한눈에
| 읽는 것 | 사진 속 손글씨 한 줄 (한글 + 영문) |
| 모델 | ViT-Small 인코더(ImageNet, Apache-2.0) + 8층 TrOCR 디코더(처음부터 학습) |
| 어휘 | 자모 229개 — 음절 11,172자를 외우지 않는다. 못 본 글자도 쓴다 |
| 파라미터 | 41M — ko-trocr(213.7M)의 1/5 |
| 가중치 | 156MB 단일 모델 · 450MB 앙상블(모델 3개) (따로 받는다 — GitHub Releases · 허깅페이스) |
| 꾸러미 | 119KB — 코드만 |
| 붙이기 | OpenAI·Ollama 와 같은 말을 하는 서버(ko-hand-ocr-serve) · LM Studio 에 붙는 MCP 도구(ko-hand-ocr-mcp) |
| 속도 | 줄당 0.18초 · 사진 한 장 1.04초 — GPU 없이 CPU 만으로 |
| 메모리 | 1.07GB 단일 모델 · 1.51GB 앙상블 |
| 성적 | 학습에 안 쓴 글씨체 여섯 벌 평균 95.6% (앙상블 95.6%) · 가장 낮은 벌 90.4% (앙상블 90.7%) · 손글씨 사진 11장 평균 94.9% (앙상블 96.6%) · 17개 가운데 95% 넘김 11개 (앙상블 12개) |
| 남과 견주면 | 같은 시험지에서 공개 OCR 여섯 가운데 가장 잘 읽는 것보다 글씨체 6벌 +17.7%p(ko-trocr), 사진 +15.7%p(ko-trocr) — 단일 모델 기준 |
| 학습 자료 | 합성 — OFL 글꼴로 그때그때 그린다. 디스크에 남지 않는다 |
| 라이선스 | Apache-2.0, 가중치도 같다 — 자료에서 물려받는 약관이 없다 |
| 필요한 것 | Python 3.10+ · PyTorch 2.5+ |
되는 것 — 그리고 안 되는 것
된다
- 사진 한 장을 통째로. 줄을 찾아 자르고 각각 읽는다 (
read_photo) - 이미 잘린 줄 그림을 묶음으로 (
read) - 가둬서 읽기 — 후보 목록 안에서만 고르게 하고, 자유롭게 읽은 것과 맞는지 알려 준다 (
both) - 앙상블로 읽기 — 모델을 여럿 두고 서로 닮은 답을 고른다
- 아무 데도 보내지 않는다. CPU 에서 완전히 혼자 돈다
- LM Studio·Ollama·vLLM 쪽 도구와 붙는다 — 같은 말을 하는 서버와 MCP 도구(위 「LM Studio · Ollama · vLLM 과 같이 쓰기」)
안 된다
- 표 안의 칸을 라벨과 값으로 갈라 주지 않는다.
read_photo는 줄까지만 자른다 - 세로쓰기, 칸 병합은 다루지 않는다
- 읽을 것이 없는 칸(전체가 낙서인 것)도 무언가 읽어 낸다
왜 만들었나
우리가 찾아본 가운데 한국어 손글씨를 노리고 공개된 모델은 ddobokki/ko-trocr 하나였는데,
학습 자료가 AI Hub 라 이용 목적·재배포에 제약이 붙는다. 사내 도구에 넣거나 공개 꾸러미로
배포하려면 걸린다. 두루 쓰는 OCR(PaddleOCR·EasyOCR·Tesseract)과 큰 VLM 은 손글씨 한
줄에서 크게 떨어진다 — 같은 시험지에서 그 가운데 가장 잘 읽는 것이 글씨체 6벌
77.6%, 사진 78.5% 다(아래). 그래서 이 일을 하는 모델을
처음부터 다시 만들었다 — 부품마다 출처를 따져 볼 수 있게.
다른 OCR 과 같은 자로 견주면
상대는 여섯이다. 한국어 OCR 로 사람들이 실제로 꺼내 쓰는 것을 갈래마다 세웠다.
| 상대 | 갈래 | 크기 | 어떻게 돌렸나 |
|---|---|---|---|
ddobokki/ko-trocr |
한국어 TrOCR (AI Hub 로 배움) | 214M | transformers, float32, 빔 5, 길이 상한 16 -> 64 로 풀어 줌 |
| PaddleOCR PP-OCRv5 | 줄 인식기 (korean_PP-OCRv5_mobile_rec) |
3.3M | paddleocr 3.7.0 / paddle 3.4.0, TextRecognition |
| EasyOCR | 줄 인식기 (korean_g2) |
4.0M | easyocr 1.7.2, Reader(["ko", "en"]).recognize() |
| Tesseract 5 | 줄 인식기 (LSTM, 파라미터 수는 공개하지 않는다) | 파일 6MB | tesseract 5.5.3, --oem 1 --psm 7 -l kor+eng |
| PaddleOCR-VL-1.6 | OCR 전용 VLM | 906M | transformers 5.16.1, bfloat16, 물음 「OCR:」(모델 카드) |
| Qwen3-VL-2B-Instruct | 두루 쓰는 VLM | 2.1B | transformers 5.16.1, bfloat16, 한글 물음 — 셋을 대 보고 가장 잘 읽은 것 |
모두 모델 카드의 사용법 그대로다. 엔진마다 따로 깐 환경에서 돌리고
(tools/rival_worker.py), 시간은 그 엔진이 잰 값만 센다(그림을 넘기는 값은 뺀다).
견줌이 성립하려면 모델만 다르고 나머지는 같아야 한다. 사진은 우리 줄
자르기(kohandocr.page)로 자른 같은 칸을 모두에게 넣었고, 글씨체 시험지는
synth.render 가 그린 날것을 넣었다 — 우리 전처리(64×640 letterbox)를
씌워서 주면 상대만 두 번 찌그러진다. 자·GPU·잰 순간이 모두 같다. 무엇을 어떻게
쟀는지는 tools/vs.py 머리말에 다 적혀 있고, 아래 숫자는 전부 그것이 남긴
runs/VS.json 에서 나왔다.
실제 손글씨 사진 · 34줄 209글자
| 자모 닮음 | 글자 오류율(CER) | 줄 통째로 일치 | 가장 나쁜 장 | |
|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 96.52% | 8.61% | 67.6% | 90.28% |
| ko-hand-ocr 단일 모델 (41M) | 94.24% | 13.40% | 61.8% | 83.03% |
| ddobokki/ko-trocr (214M) | 78.50% | 39.23% | 26.5% | 58.56% |
| PaddleOCR PP-OCRv5 (3.3M) | 71.75% | 39.71% | 8.8% | 34.81% |
| Qwen3-VL-2B (2.1B) | 69.86% | 38.28% | 8.8% | 46.86% |
| PaddleOCR-VL-1.6 (906M) | 61.71% | 44.02% | 2.9% | 18.52% |
| EasyOCR (4.0M) | 66.88% | 55.02% | 2.9% | 46.19% |
| Tesseract 5 (6MB) | 13.21% | 98.56% | 0.0% | 0.00% |
안 배운 글씨체 6벌 · 670줄 4708글자
| 자모 닮음 | 글자 오류율(CER) | 줄 통째로 일치 | 가장 나쁜 벌 | |
|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 95.65% | 8.28% | 73.0% | 89.83% |
| ko-hand-ocr 단일 모델 (41M) | 95.36% | 8.45% | 72.2% | 89.02% |
| ddobokki/ko-trocr (214M) | 77.62% | 41.67% | 29.4% | 65.24% |
| PaddleOCR PP-OCRv5 (3.3M) | 72.28% | 40.40% | 28.4% | 52.43% |
| Qwen3-VL-2B (2.1B) | 66.02% | 79.89% | 19.9% | 50.02% |
| PaddleOCR-VL-1.6 (906M) | 62.97% | 62.34% | 18.1% | 43.05% |
| EasyOCR (4.0M) | 56.35% | 64.21% | 9.7% | 36.95% |
| Tesseract 5 (6MB) | 17.42% | 98.13% | 0.9% | 7.29% |
두루 읽나 — 24벌 · 648줄 3768글자
| 자모 닮음 | 글자 오류율(CER) | 줄 통째로 일치 | 가장 나쁜 벌 | |
|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 97.71% | 3.42% | 85.5% | 92.46% |
| ko-hand-ocr 단일 모델 (41M) | 97.20% | 3.98% | 84.7% | 89.14% |
| ddobokki/ko-trocr (214M) | 79.94% | 32.38% | 42.0% | 46.52% |
| PaddleOCR PP-OCRv5 (3.3M) | 75.27% | 30.02% | 39.5% | 18.66% |
| Qwen3-VL-2B (2.1B) | 67.81% | 62.98% | 26.1% | 24.46% |
| PaddleOCR-VL-1.6 (906M) | 66.69% | 45.94% | 24.8% | 27.33% |
| EasyOCR (4.0M) | 64.38% | 48.91% | 19.1% | 19.28% |
| Tesseract 5 (6MB) | 19.98% | 95.12% | 2.2% | 3.07% |
글꼴 480벌 전부 · 벌당 12줄 4320줄 22560글자
| 고르기 6벌 6벌 |
두루 24벌 24벌 |
학습에 쓴 벌 450벌 |
전부 480벌 |
95% 넘긴 벌 | 80% 밑인 벌 | |
|---|---|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 99.15% | 99.00% | 98.22% | 98.27% | 436벌 | 5벌 |
| ko-hand-ocr 단일 모델 (41M) | 98.49% | 98.58% | 98.06% | 98.09% | 434벌 | 4벌 |
| ddobokki/ko-trocr (214M) | 85.29% | 83.29% | 82.48% | 82.56% | 46벌 | 158벌 |
| PaddleOCR PP-OCRv5 (3.3M) | 80.07% | 78.48% | 78.02% | 78.07% | 86벌 | 198벌 |
| Qwen3-VL-2B (2.1B) | 65.17% | 67.35% | 69.32% | 69.17% | 50벌 | 289벌 |
| PaddleOCR-VL-1.6 (906M) | — | — | — | — | — | — |
| EasyOCR (4.0M) | 67.73% | 68.17% | 67.15% | 67.21% | 9벌 | 319벌 |
| Tesseract 5 (6MB) | 21.56% | 19.28% | 20.50% | 20.45% | 0벌 | 480벌 |
- 상대 여섯 가운데 가장 잘 읽는 것도 큰 차이로 뒤진다. 안 배운 글씨체 6벌에서 상대 으뜸은 ko-trocr 77.62%, 우리 단일 모델(41M)은 95.36% 다. 두루 24벌에서도 ko-trocr 79.94% 대 97.20%.
- 사진 열한 장은 아무도 본 적이 없다. 거기가 가장 깨끗한 자리이고, 우리 앙상블 96.52% · 단일 모델 94.24% 대 상대 으뜸 ko-trocr 78.50% 다. 줄을 통째로 맞힌 비율은 67.6% 대 26.5%(상대 가운데 가장 높은 값).
- 큰 VLM 이라고 손글씨 한 줄을 잘 읽지는 않는다. Qwen3-VL-2B(2.1B)는 6벌 66.02%, PaddleOCR-VL-1.6(906M)은 62.97% 다. 우리보다 스무 배에서 쉰 배 크지만 이 일에 맞춰 배운 모델이 아니다. 문서 전체를 읽는 일이라면 이야기가 다르다(재지 않았다).
- Tesseract 는 손글씨를 거의 못 읽는다(6벌 17.42%, 사진 13.21%). 인쇄 글자용 엔진이라 예상한 대로다. 기준선으로 둔다.
- 글꼴을 하나도 안 빼고 480벌을 읽혀도 우리 98.27%(앙상블) · 98.09%(단일 모델) 대 상대 으뜸 ko-trocr 82.56%. 우리가 본 적 없는 24벌(99.00%)이 학습에 쓴 450벌(98.22%)보다 높다 — 외워서 올라간 숫자가 아니라는 뜻이다. PaddleOCR-VL 은 가장 싼 묶음이 줄당 1초 안팎이라 480벌은 재지 않았다(표에서 빈칸).
- 그래도 무너지는 벌이 있다. 480벌 가운데 우리 단일 모델의 가장 낮은 벌이 47.31%, 앙상블은 63.21% 다. 여섯 벌만 보면 이런 벌이 안 보인다. 벌당 12줄뿐이라 한 벌의 값은 크게 흔들리므로 무리 합계와 분포로 본다.
- 단일 모델은 사진 바닥이 0.3.0 보다 낮다. 가장 나쁜 장이 hand-06
83.03% 다 — 세 줄짜리 장의 한 줄(
전자금융TF서약)이다. 앙상블에서는 다른 모델이 받아 주어 90.28% 그대로다. - 0.4.1 단일 모델은 글씨체 쪽만 올랐다. 0.4.0(v83)과 견주면 여섯 벌 95.20 ->
95.36%, 스물네 벌 96.87 -> 97.20%, 480벌 98.00 -> 98.09% 로 세 자에서 다 올랐다.
이 표의 사진 쪽은 94.95 -> 94.24% 로 내려갔다 — 34줄 가운데 통째로 맞힌 줄이
22 -> 21줄, 한 줄이 갈린 폭이다(
tools/verify.py로는 95.0 대 94.9%).
어디서 갈리나
갈래별 글자 오류율 — 낮을수록 좋다
| 라틴 섞임 66줄 |
서식 라벨 60줄 |
숫자 섞임 84줄 |
한글 낱말·이름 304줄 |
한글 문장 156줄 |
|
|---|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 14.1% | 8.3% | 9.6% | 4.7% | 8.1% |
| ko-hand-ocr 단일 모델 (41M) | 14.5% | 7.6% | 8.7% | 5.5% | 8.6% |
| ddobokki/ko-trocr (214M) | 59.5% | 30.3% | 49.4% | 25.0% | 46.1% |
| PaddleOCR PP-OCRv5 (3.3M) | 53.7% | 32.1% | 43.7% | 33.5% | 40.9% |
| Qwen3-VL-2B (2.1B) | 58.7% | 96.5% | 49.1% | 104.5% | 82.1% |
| PaddleOCR-VL-1.6 (906M) | 75.7% | 65.9% | 45.6% | 81.7% | 51.4% |
| EasyOCR (4.0M) | 68.5% | 55.6% | 64.8% | 64.7% | 64.1% |
| Tesseract 5 (6MB) | 90.8% | 103.3% | 82.3% | 114.5% | 96.0% |
줄 길이별 글자 오류율 — 낮을수록 좋다
| 1~5자 286줄 |
6~10자 270줄 |
11~20자 108줄 |
21자 이상 6줄 |
|
|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 6.1% | 8.4% | 8.3% | 17.6% |
| ko-hand-ocr 단일 모델 (41M) | 6.5% | 8.9% | 8.1% | 15.7% |
| ddobokki/ko-trocr (214M) | 30.7% | 35.0% | 50.7% | 92.1% |
| PaddleOCR PP-OCRv5 (3.3M) | 42.9% | 33.8% | 40.1% | 92.1% |
| Qwen3-VL-2B (2.1B) | 101.2% | 84.0% | 59.0% | 85.6% |
| PaddleOCR-VL-1.6 (906M) | 96.3% | 50.8% | 55.5% | 58.8% |
| EasyOCR (4.0M) | 70.4% | 58.5% | 63.0% | 97.2% |
| Tesseract 5 (6MB) | 114.6% | 98.9% | 86.0% | 98.6% |
- 라틴 약어에서 갈린다. 글자 오류율이 우리 14.5%(단일 모델) 대 상대 가운데
가장 나은 PaddleOCR PP-OCRv5 53.7% 다. 서식에 섞이는 영문은 대개
약어(
TF,OCR,Codex)인데, 한글과 영문을 한 줄에서 같이 읽는 데서 갈린다. - 길수록 갈린다. 21자 이상에서 우리 15.7% 대 상대 가운데 가장 나은 PaddleOCR-VL-1.6 58.8%. ko-trocr 는 인코더가 384×384 정사각형이라 긴 줄을 통째로 눌러 넣는다. 우리는 64×640 이다. 다만 21자 이상은 6줄뿐이라 이 칸 하나로 결론 내면 안 된다 — 방향만 본다.
무엇을 치렀나
빠르기와 덩치
| 파라미터 | 받는 크기 | 줄/초 (GPU) | 줄/초 (CPU) | 30줄 쪽 (GPU) | 30줄 쪽 (CPU) | 꼭대기 VRAM | |
|---|---|---|---|---|---|---|---|
| ko-hand-ocr 앙상블 (118M) | 118M | 450MB | 6.9 | 0.66 | 4.76초 | 45.8초 | 1772MB |
| ko-hand-ocr 단일 모델 (41M) | 41M | 156MB | 34.2 | 3.08 | 1.31초 | 10.2초 | 1052MB |
| ddobokki/ko-trocr (214M) | 214M | 408MB | 5.3 | 0.52 | 6.09초 | 58.1초 | 1798MB |
| PaddleOCR PP-OCRv5 (3.3M) | 3.3M | 13MB | 336.5 | 14.68 | 0.52초 | 2.5초 | 111MB |
| Qwen3-VL-2B (2.1B) | 2.1B | 4058MB | 6.5 | 0.47 | 5.02초 | 64.1초 | 5724MB |
| PaddleOCR-VL-1.6 (906M) | 906M | 1828MB | 0.9 | 0.04 | 32.80초 | 680.7초 | 1937MB |
| EasyOCR (4.0M) | 4.0M | 15MB | 32.5 | 14.88 | 1.36초 | 2.5초 | 70MB |
| Tesseract 5 (6MB) | — | 6MB | — | 22.25 | — | 1.8초 | — |
- 가장 빠른 것은 우리가 아니다. CPU 에서 Tesseract 5 22.2줄/초, EasyOCR 14.9줄/초, PaddleOCR PP-OCRv5 14.7줄/초 로 우리 단일 모델(3.1줄/초)보다 빠르다. 대신 6벌에서 Tesseract 5 17.4%, EasyOCR 56.4%, PaddleOCR PP-OCRv5 72.3% 를 읽는다. 맞히는 값과 빠르기를 같이 봐야 한다 — 서식에서 틀린 답은 빈칸보다 나쁘다.
- 같은 일을 노린 ko-trocr(214M)와 견주면 CPU 에서 5.9배 빠르다(서른 줄 쪽 10.2초 대 58.1초). GPU 없는 사무용 PC 에 넣을 수 있느냐가 여기서 갈린다.
- CPU 값은 이 노트북의 그때 상태를 크게 탄다. 같은 날 같은 칸·같은 방식으로 단일
모델이 아침 4.6줄/초, 이 판 3.1줄/초, 저녁 2.6줄/초였고 ko-trocr 와의 배수도
6.3·5.9·4.4배로 움직였다. 그래서 아래 「속도와 사양」(
bench.py로 따로 잰 것)의 줄당 0.18초와 숫자가 다르다. 줄/초 자체보다 같은 판 안의 순서를 보고, 배수는 4~6배로 읽는다 — ko-trocr 보다 빠르다는 것은 세 번 모두 같았다. - 큰 VLM 은 GPU 가 있어야 쓸 수 있다. CPU 에서 Qwen3-VL-2B 는 0.47줄/초, PaddleOCR-VL-1.6 은 0.04줄/초라 서른 줄 쪽이 64초, 681초다. GPU 에서도 6.5줄/초, 0.9줄/초로 우리 단일 모델(34.2줄/초)보다 느리다. CPU 에서 줄당 2초·23초라 1줄 묶음만 한 번 쟀다.
- 묶음 크기는 짐작하지 않고 잰다. 모델마다 가장 싼 묶음으로 읽혔다. 이 카드에 안 들어가는 묶음은 재지 않는다 — ko-trocr 의 32줄 묶음이 그렇고, 자가 적어 놓은 까닭은 「앞 묶음 4642MB(가중치 848MB)에서 활성값을 두 배 하면 8437MB 로 카드 8151MB 를 넘는다」이다. 안 잰 값을 잰 것처럼 적지 않는다.
같지 않은 것 — 어느 쪽에 유리한지까지 적는다
| 무엇이 | 어떻게 | 누구에게 유리 |
|---|---|---|
| 시험 글꼴 | 여섯·스물네 벌은 우리 학습에서만 뺐다. 상대가 무엇을 보고 배웠는지 우리는 모른다 | 상대 |
| 줄 자르기 | 상대에게 우리 것을 빌려주었다(모두 한 줄 그림을 받는다) | 상대 |
| 길이 상한 | ko-trocr 의 기본은 max_length: 16(글자 단위). VLM 둘과 같이 64 토큰으로 풀어 주었다. Qwen3-VL 이 상한에 닿은 줄은 모두 되풀이로 넘친 줄(11 11 11 …)이었고 긴 정답이 잘린 줄은 없었다(글씨체 670줄을 다시 읽혀 보니 상한에 닿은 24줄이 모두 그랬다) |
ko-trocr |
| 물음 | Qwen3-VL 은 정해 둔 OCR 물음이 없어 셋을 대 보고 가장 잘 읽는 것을 주었다 | Qwen3-VL |
| 정밀도 | VLM 둘은 나온 정밀도(bfloat16) 그대로, 나머지는 float32. ko-trocr 는 float16 빠르기도 따로 쟀다 | — |
| VLM 의 CPU | 줄당 몇 초에서 수십 초라 1줄 묶음을 한 번만 쟀다. 8줄로 묶으면 더 빠를 수도 있다 | ko-hand-ocr |
| 480벌 | PaddleOCR-VL 은 줄당 1초 안팎이라 480벌 전부는 안 쟀다(빈칸) | — |
| 만든 목적 | 줄 인식기 셋은 주로 인쇄·장면 글자, VLM 둘은 문서 전체, ko-trocr 는 AI Hub 의 손글씨·공공행정문서. 손글씨 한 줄 하나만 노리고 만든 것은 우리 쪽이다 | ko-hand-ocr |
| 모델 수 | 우리 앙상블은 모델 셋(36M 하나 + 41M 둘, 118M). 크기로 견주려면 「단일 모델」 줄을 본다 | ko-hand-ocr |
| 학습에 쓴 글꼴 | 480벌 가운데 450벌은 우리가 본 글씨체다. 그래서 무리를 갈라 적는다 | ko-hand-ocr |
상대들이 나쁜 모델이라는 뜻이 아니다. 인쇄 서류·장면 글자·문서 전체를 읽으라고 만든 모델에 손글씨 한 줄을 들이민 것이고, 그 일에서는 이쪽이 낫다는 것뿐이다. 반대 방향(인쇄 서류 전체, 표, 장면 글자)은 재지 않았고, 재면 우리가 질 가능성이 크다.
잰 자리: NVIDIA GeForce RTX 5060 Laptop GPU · torch 2.14.0+cu130 · 2026-10-07 17:32. 우리 두 모델과 ko-trocr 의 정확도는 그날 아침에 잰 값과 글자 하나까지 같다 — 같은 모델을 같은 자로 다시 잰 것이다. 빠르기는 그때그때 기계 상태를 타므로 같은 순간에 잰 것끼리만 견준다.
얼마나 맞히나
두 가지로 잰다. 둘 다 봐야 한다. 아래는 전부
python tools/verify.py 로 낸 값이다(글씨체 벌마다 약 960줄, 사진 11장).
| 안 배운 글씨체 (평균) | 가장 낮은 글씨체 | 손글씨 사진 11장 | 가장 나쁜 장 | 95% 넘긴 항목 | |
|---|---|---|---|---|---|
| 단일 모델 (0.4.1) | 95.6% | 90.4% | 94.9% | 84.2% | 11 / 17 |
| 앙상블 (모델 3개) | 95.6% | 90.7% | 96.6% | 90.3% | 12 / 17 |
| 0.4.0 단일 모델 (v83) | 95.5% | 90.1% | 95.0% | 83.0% | 10 / 17 |
| 0.3.0 단일 모델 (v62) | 94.6% | 88.5% | 95.4% | 90.3% | |
| 0.3.0 앙상블 | 95.3% | 90.3% | 95.8% | 90.3% | 10 / 17 |
글씨체는 벌마다 약 960줄로 잰다(0.4.0 을 낼 때는 380줄이었다). 기울임 줄은 지난 릴리스다. 0.4.0 단일 모델은 같은 960줄 자로 다시 쟀고, 0.3.0 줄은 그때의 380줄 자로 잰 값이라 차이에는 자가 촘촘해진 몫도 들어 있다.
0.4.1 에서 바뀐 것은 단일 모델 하나다. 앙상블(모델 3개)이 다음 글자마다 내는 확률을 단일 모델이 따라 배우게 하고(증류), 그렇게 배운 모델을 0.4.0 의 v83 과 반씩 섞었다. 크기·구조가 그대로라 빠르기도 같다. 17개 가운데 95% 를 넘긴 것이 10 -> 11개(hand-07 94.4 -> 95.8%), 글씨체 여섯 벌이 모두 조금씩 올랐고(가장 낮은 벌 90.1 -> 90.4%), 가장 나쁜 장 hand-06 이 83.0 -> 84.2% 다. 대신 hand-08 이 100 -> 96.7% 로 내려갔다. 크게 나아진 것은 아니다 — 사진 평균은 95.0 대 94.9% 로 같다. 같은 방법을 세기만 바꿔 다시 해 보거나(몫 0.25·1.0), 드롭아웃·무게 지수 평균·SAM 으로 해 본 것은 모두 v83 을 못 넘었다.
0.4.0 은 0.3.0 에서 무엇이 나아졌나. 단일 모델(v83)은 글씨체 평균 94.6 -> 95.5%, 가장 낮은 벌 88.5 -> 90.1% 로 올랐다. 대신 사진 바닥이 내려갔다(hand-06 90.3 -> 83.0%, 세 줄짜리 장의 한 줄). 앙상블은 바닥을 그대로 지키고 사진 평균이 95.8 -> 96.6% 로 올랐다. 17개 항목(글씨체 6 + 사진 11) 가운데 95% 를 넘긴 것이 앙상블 10 -> 12개다.
자모 단위 닮음이고, 사진 쪽은 사진 한 장에 한 표다.
사진 숫자를 그대로 기대하면 안 된다. 만든 사람의 사진 11장 35칸으로 잰 값이라 폭이 ±5%p 다. 남의 손씨에 더 가까운 값은 '안 배운 글씨체' 쪽이다 — 학습에 한 번도 안 쓴 글꼴 여섯 벌로만 시험한 것이라 폭이 절반이다(±2~3%p).
앙상블의 값은 평균이 아니라 바닥에 있다. 글씨체 평균은 이제 같은데 (95.6 대 95.6) 가장 낮은 글씨체가 90.4% 에서 90.7% 로 올라가고, 단일 모델이 84.2% 로 흘린 장(hand-06)이 90.3% 가 된다. 한 모델이 아슬아슬하게 읽는 칸을 다른 모델이 받아 주기 때문이다.
대신 모든 장이 좋아지지는 않는다. 단일 모델이 95.2% 로 넘긴 장이 92.9% 로 내려가기도 한다(hand-10) — 두 모델이 같이 틀리면 맞게 읽은 한 모델이 밀린다. 그리고 일곱 배 느리다. 한 장이라도 사람 손으로 넘어가면 곤란한 자리에는 앙상블을, 많은 양을 빨리 훑어야 하면 단일 모델을 쓴다.
글씨체마다 (벌마다 약 960줄)
합계 하나만 보면 무너지는 글씨체가 가려진다. 실제로 9%p 가까이 벌어진다.
| 글씨체 | 단일 모델 (0.4.1) | 앙상블 | 0.4.0 단일 모델 (v83) |
|---|---|---|---|
| 동해독도 (EastSeaDokdo) | 90.4% | 90.7% | 90.1% |
| 기랑해랑 (KirangHaerang) | 94.7% | 94.7% | 94.7% |
| 하이멜로디 (HiMelody) | 95.6% | 95.4% | 95.4% |
| 나눔펜스크립트 (NanumPenScript) | 95.9% | 95.9% | 95.7% |
| 감자꽃 (GamjaFlower) | 98.0% | 98.0% | 97.9% |
| 싱글데이 (SingleDay) | 99.0% | 98.9% | 98.9% |
넘어야 하는 것은 평균이 아니라 가장 낮은 벌이다. 여섯 벌이 모두 90% 를 넘긴 뒤(2026-09-18) 목표를 95% 로 올렸다. 지금은 넷이 넘고 동해독도와 기랑해랑이 남았다. 이 두 벌은 획이 뭉개지고 생략되는 손씨다.
여섯 벌 다 학습에 한 번도 안 쓴 글꼴이고, 모두 SIL Open Font License 다.
python tools/holdout.py <모델폴더> --per-font --sample sample.png
더 보기
- 학습 방법, 시험지 그림, 속도와 사양(
tools/bench.py), 돌아가는 곳: GitHub README - 부품마다의 출처와 약관: PROVENANCE.md
- 견줌을 다시 재는 도구:
tools/vs.py,tools/rival_worker.py
라이선스
Apache-2.0. 가중치도 같다. 학습에 쓴 글꼴의 라이선스는 가중치에 옮아붙지 않는다 — 가중치는 글자 모양의 저작물이 아니라 그림에서 배운 값이다. 근거는 PROVENANCE.md 에 부품별로 적어 두었다.
English
Reads one line of handwritten Korean + English. A 41M model trained from scratch on synthetic data — no handwriting dataset, no inherited terms. Apache-2.0, weights included. Full English README: README.en.md.
Download and use
The weights here are the same files as the zips on the GitHub v0.4.1 release (same sha256).
| Path | What | Size |
|---|---|---|
/ (root) |
Single model — 41M, 8-layer decoder. Enough for most uses | 156MB |
ensemble/ |
Ensemble — three models (36M + 41M + 41M) read together and the answer they agree on wins. Raises the floor, 7x slower | 450MB |
The weights alone do not read anything: the jamo vocabulary and the line cutter live in the
kohandocr package, so install it first. The
transformers pipeline cannot open this model.
pip install ko-hand-ocr
from huggingface_hub import snapshot_download
from kohandocr.reader import Reader
folder = snapshot_download("localdeel/ko-hand-ocr", ignore_patterns=["ensemble/*", "assets/*"])
reader = Reader(folder, device="cpu")
reader.read_photo(open("scan.jpg", "rb").read()) # a whole photo: cuts lines, reads each
# ensemble
folder = snapshot_download("localdeel/ko-hand-ocr", allow_patterns=["ensemble/*"])
reader = Reader(folder + "/ensemble", device="cpu")
Use it with LM Studio, Ollama and vLLM
It cannot be loaded into them as a model. All three are engines for LLMs that continue text: LM Studio and Ollama run llama.cpp GGUF files, and vLLM runs the architectures it knows. This model is a TrOCR-style image encoder with a cross-attention decoder, plus its own jamo vocabulary and line cutter, so it neither converts to GGUF nor fits vLLM's list. Instead the package ships a server that speaks the same APIs and an MCP tool.
pip install "ko-hand-ocr[mcp]" # drop [mcp] if you only need the server
Server — the OpenAI and Ollama APIs
ko-hand-ocr-serve # fetches the single model (156MB) from Hugging Face, serves 127.0.0.1:8765
ko-hand-ocr-serve --ensemble # the ensemble (450MB)
| Caller | How |
|---|---|
openai client, code written for vLLM |
set base_url="http://127.0.0.1:8765/v1"; send the image as base64 in image_url |
ollama CLI |
set OLLAMA_HOST=127.0.0.1:8765, then ollama run ko-hand-ocr "C:\scan.jpg" · ollama list |
| curl | curl --data-binary @scan.jpg http://127.0.0.1:8765/read |
import base64
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8765/v1", api_key="none")
picture = "data:image/jpeg;base64," + base64.b64encode(open("scan.jpg", "rb").read()).decode()
reply = client.chat.completions.create(model="ko-hand-ocr", messages=[
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": picture}}]}])
print(reply.choices[0].message.content) # one line per handwritten line
- Text in the prompt is ignored. The image is cut into lines and each line comes back as one line of text. It cannot summarise or tidy up — that is the chat model's job (MCP, below).
- Image URLs (
http://…) are not fetched. Send base64. Nothing leaves the machine. - By default it listens on this PC only (
127.0.0.1). There is no password, so use--host 0.0.0.0only on a network you trust.
LM Studio — as an MCP tool
Since 0.3.17, chat models in LM Studio can call outside tools (MCP). In ~/.lmstudio/mcp.json:
{"mcpServers": {"ko-hand-ocr": {"command": "ko-hand-ocr-mcp"}}}
If ko-hand-ocr-mcp is not on PATH, give the full path (Scripts\ko-hand-ocr-mcp.exe in the
virtual environment). Then ask a chat model that can call tools to "read C:\scan.jpg and put it in
a table": it calls read_handwriting for the text and does the tidying itself. This model
recognises the characters; the chat model organises them. An image pasted into the chat window
goes to the chat model, not the tool, so pass a file path. The first call downloads the model
(156MB, 20–30 s here); after that a photo takes about a second.
What was checked and what was not. The ollama CLI 0.34 (list, show, run), the official
openai client (whole and streamed replies) and an MCP SDK 2.3 client were all run against it. A full
run inside the LM Studio chat window or Open WebUI has not been done yet.
Putting it inside vLLM itself would mean writing the architecture as a plugin. For a 41M model a GPU server buys little, so that was not done — the server above exposes the same OpenAI address vLLM does.
At a glance
| What it reads | One line of handwritten Korean + English, from a photo |
| Model | ViT-Small encoder (ImageNet, Apache-2.0) + 8-layer TrOCR decoder trained from scratch |
| Vocabulary | 229 jamo tokens — not 11,172 syllables, so unseen characters are still writable |
| Parameters | 41M — one fifth of ko-trocr (213.7M) |
| Weights | 156 MB single model · 450 MB ensemble (3 models) (downloaded separately — GitHub Releases · Hugging Face) |
| Package | 119 KB — the code only |
| Plugs into | a server that speaks the OpenAI and Ollama APIs (ko-hand-ocr-serve) · an MCP tool for LM Studio (ko-hand-ocr-mcp) |
| Speed | 0.18 s per line · 1.04 s for a whole photo — CPU only, no GPU |
| Memory | 1.07 GB single model · 1.51 GB ensemble |
| Accuracy | 95.6% mean on six handwriting fonts never seen in training (95.6% ensemble) · worst font 90.4% (90.7% ensemble) · 11 handwriting photos 94.9% (96.6% ensemble) · 11 of 17 items over 95% (ensemble 12) |
| Against others | On the same sheets, ahead of the best of six public OCR models by +17.7 points on the six fonts (ko-trocr) and +15.7 points on the photos (ko-trocr) — single model |
| Training data | Synthetic — drawn on the fly from OFL fonts. Nothing is stored on disk |
| License | Apache-2.0, weights included — no dataset terms inherited |
| Python | 3.10+ · PyTorch 2.5+ |
Against six other OCR models, same sheets, same PC, same moment
| Rival | Kind | Size | How it was run |
|---|---|---|---|
ddobokki/ko-trocr |
Korean TrOCR (trained on AI Hub) | 214M | transformers, float32, beam 5, length cap raised 16 -> 64 |
| PaddleOCR PP-OCRv5 | line recogniser (korean_PP-OCRv5_mobile_rec) |
3.3M | paddleocr 3.7.0 / paddle 3.4.0, TextRecognition |
| EasyOCR | line recogniser (korean_g2) |
4.0M | easyocr 1.7.2, Reader(["ko", "en"]).recognize() |
| Tesseract 5 | line recogniser (LSTM, no parameter count published) | 6MB file | tesseract 5.5.3, --oem 1 --psm 7 -l kor+eng |
| PaddleOCR-VL-1.6 | OCR-specific VLM | 906M | transformers 5.16.1, bfloat16, prompt "OCR:" (model card) |
| Qwen3-VL-2B-Instruct | general VLM | 2.1B | transformers 5.16.1, bfloat16, Korean prompt — the best of three we tried |
Every rival runs exactly as its model card says, each in its own environment. Photos were cut once with our line cutter and the same cells went to everyone; font sheets were handed over raw.
Real handwriting photos · 34 lines, 209 characters
| Jamo similarity | Character error rate | Exact line match | Worst photo | |
|---|---|---|---|---|
| ko-hand-ocr ensemble (118M) | 96.52% | 8.61% | 67.6% | 90.28% |
| ko-hand-ocr single model (41M) | 94.24% | 13.40% | 61.8% | 83.03% |
| ddobokki/ko-trocr (214M) | 78.50% | 39.23% | 26.5% | 58.56% |
| PaddleOCR PP-OCRv5 (3.3M) | 71.75% | 39.71% | 8.8% | 34.81% |
| Qwen3-VL-2B (2.1B) | 69.86% | 38.28% | 8.8% | 46.86% |
| PaddleOCR-VL-1.6 (906M) | 61.71% | 44.02% | 2.9% | 18.52% |
| EasyOCR (4.0M) | 66.88% | 55.02% | 2.9% | 46.19% |
| Tesseract 5 (6MB) | 13.21% | 98.56% | 0.0% | 0.00% |
Six unseen handwriting fonts · 670 lines, 4708 characters
| Jamo similarity | Character error rate | Exact line match | Worst font | |
|---|---|---|---|---|
| ko-hand-ocr ensemble (118M) | 95.65% | 8.28% | 73.0% | 89.83% |
| ko-hand-ocr single model (41M) | 95.36% | 8.45% | 72.2% | 89.02% |
| ddobokki/ko-trocr (214M) | 77.62% | 41.67% | 29.4% | 65.24% |
| PaddleOCR PP-OCRv5 (3.3M) | 72.28% | 40.40% | 28.4% | 52.43% |
| Qwen3-VL-2B (2.1B) | 66.02% | 79.89% | 19.9% | 50.02% |
| PaddleOCR-VL-1.6 (906M) | 62.97% | 62.34% | 18.1% | 43.05% |
| EasyOCR (4.0M) | 56.35% | 64.21% | 9.7% | 36.95% |
| Tesseract 5 (6MB) | 17.42% | 98.13% | 0.9% | 7.29% |
Speed and size
| Parameters | Download | lines/s (GPU) | lines/s (CPU) | 30-line page (GPU) | 30-line page (CPU) | Peak VRAM | |
|---|---|---|---|---|---|---|---|
| ko-hand-ocr ensemble (118M) | 118M | 450MB | 6.9 | 0.66 | 4.76s | 45.8s | 1772MB |
| ko-hand-ocr single model (41M) | 41M | 156MB | 34.2 | 3.08 | 1.31s | 10.2s | 1052MB |
| ddobokki/ko-trocr (214M) | 214M | 408MB | 5.3 | 0.52 | 6.09s | 58.1s | 1798MB |
| PaddleOCR PP-OCRv5 (3.3M) | 3.3M | 13MB | 336.5 | 14.68 | 0.52s | 2.5s | 111MB |
| Qwen3-VL-2B (2.1B) | 2.1B | 4058MB | 6.5 | 0.47 | 5.02s | 64.1s | 5724MB |
| PaddleOCR-VL-1.6 (906M) | 906M | 1828MB | 0.9 | 0.04 | 32.80s | 680.7s | 1937MB |
| EasyOCR (4.0M) | 4.0M | 15MB | 32.5 | 14.88 | 1.36s | 2.5s | 70MB |
| Tesseract 5 (6MB) | — | 6MB | — | 22.25 | — | 1.8s | — |
- We are not the fastest. On CPU, Tesseract 5 (22.2 lines/s), EasyOCR (14.9 lines/s), PaddleOCR PP-OCRv5 (14.7 lines/s) beat our single model (3.1 lines/s). They read Tesseract 5 17.4%, EasyOCR 56.4%, PaddleOCR PP-OCRv5 72.3% of the six fonts. Read accuracy and speed together — on a form a wrong answer is worse than a blank.
- Against a model of similar purpose, ko-trocr (214M), we are 5.9× faster on CPU (a thirty-line page in 10.2s against 58.1s). Whether it fits on an office PC with no GPU is decided here.
- CPU figures ride hard on this laptop's state at the time. Same day, same cells, same
method: the single model ran at 4.6 lines/s in the morning, 3.1 in this run and 2.6 in
the evening, and its lead over ko-trocr moved through 6.3×, 5.9× and 4.4×. That is why
the 0.18 s per line under "Speed and footprint" below (
bench.py, measured separately) does not match. Read the order within one run rather than lines/s itself, and the lead as 4–6× — being faster than ko-trocr held all three times. - The large VLMs need a GPU. On CPU Qwen3-VL-2B does 0.47 lines/s and PaddleOCR-VL-1.6 0.04, so a thirty-line page takes 64s and 681s. Even on GPU they run at 6.5 and 0.9 lines/s, slower than our single model (34.2). On CPU they take 2s and 23s per line, so only the one-line batch was timed, once.
- Batch size is measured, not guessed. Each model reads at its own cheapest batch. A batch that does not fit this card is not timed at all — ko-trocr's 32-line batch is one, and the tool's own reason reads: "앞 묶음 4642MB(가중치 848MB)에서 활성값을 두 배 하면 8437MB 로 카드 8151MB 를 넘는다". A number that was not measured is not written down as if it were.
What is not equal — and which side it favours
| What | How | Favours |
|---|---|---|
| Test fonts | The six and the twenty-four were held out of our training only. What the rivals were trained on is unknown to us | rivals |
| Line cutting | Our cutter was lent to every rival (they all receive one line image) | rivals |
| Length cap | ko-trocr ships max_length: 16 (character tokenizer). Like the two VLMs it was given 64 tokens. Every line where Qwen3-VL hit the cap was a runaway repeat (11 11 11 …); no long answer was cut short (re-reading the 670 font lines, all 24 that hit the cap were like that) |
ko-trocr |
| Prompt | Qwen3-VL has no fixed OCR prompt; three were tried and the one it reads best with was used | Qwen3-VL |
| Precision | The two VLMs run in their released precision (bfloat16), the rest in float32. ko-trocr's float16 speed was also timed | — |
| VLM on CPU | Seconds to tens of seconds per line, so only the one-line batch was timed, once. An 8-line batch might be faster | ko-hand-ocr |
| All 480 fonts | PaddleOCR-VL takes about a second per line, so it was not run on all 480 (blank) | — |
| Design intent | The three line recognisers target mainly printed and scene text, the VLMs whole documents, ko-trocr AI Hub's handwriting and administrative documents. Only this model targets a single handwritten line | ko-hand-ocr |
| Number of models | Our ensemble is three models (one 36M + two 41M, 118M). For a size-matched view read the "single model" row | ko-hand-ocr |
| Training fonts | 450 of the 480 fonts are ones we have seen. That is why the herds are reported apart | ko-hand-ocr |
None of this says the rivals are bad models. Models built for printed documents, scene text or whole pages were handed a line of handwriting, and on that job this one does better. The other direction — whole printed documents, tables, scene text — was not measured, and if it were, this model would probably lose.
Measured on: NVIDIA GeForce RTX 5060 Laptop GPU · torch 2.14.0+cu130 · 2026-10-07 17:32. Accuracy for our two models and ko-trocr matches the same morning's run to the character — same models, same ruler. Speed rides on the machine's state, so compare only figures taken at the same moment.
- Downloads last month
- 26
Model tree for localdeel/ko-hand-ocr
Base model
facebook/deit-small-patch16-224