philschmid
/

donut-base-finetuned-cord-v2

vision-encoder-decoder

endpoints-template

Inference Endpoints

Model card Files Files and versions Community

donut-base-finetuned-cord-v2 / README.md

philschmid's picture

philschmid HF staff

Update README.md

941ee89 almost 2 years ago

|

1.09 kB

	---
	license: mit
	tags:
	- donut
	- image-to-text
	- vision
	---

	# Fork of [naver-clova-ix/donut-base-finetuned-cord-v2](https://huggingface.co/naver-clova-ix/donut-base-finetuned-cord-v2)

	> This is fork of [naver-clova-ix/donut-base-finetuned-cord-v2](https://huggingface.co/naver-clova-ix/donut-base-finetuned-cord-v2) implementing a custom `handler.py` as an example for how to use `flair` models with [inference-endpoints](https://hf.co/inference-endpoints)

	---

	# Donut (base-sized model, fine-tuned on CORD)

	Donut model fine-tuned on CORD. It was introduced in the paper [OCR-free Document Understanding Transformer](https://arxiv.org/abs/2111.15664) by Geewok et al. and first released in [this repository](https://github.com/clovaai/donut).

	Donut consists of a vision encoder (Swin Transformer) and a text decoder (BART). Given an image, the encoder first encodes the image into a tensor of embeddings (of shape batch_size, seq_len, hidden_size), after which the decoder autoregressively generates text, conditioned on the encoding of the encoder.

	# Use with Inference Endpoints