Instructions to use gonzalobosque/multimodal-document-intelligence with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Keras
How to use gonzalobosque/multimodal-document-intelligence with Keras:
# !pip install -U keras tensorflow huggingface_hub # Keras needs TensorFlow installed to read "hf://" paths, so the tensorflow backend is selected here; # "jax" and "torch" also work for computation once TensorFlow is installed. import os os.environ["KERAS_BACKEND"] = "tensorflow" import keras model = keras.saving.load_model("hf://gonzalobosque/multimodal-document-intelligence") - Notebooks
- Google Colab
- Kaggle
Multimodal Document Intelligence — Classifier Artifacts
This repository contains the trained runtime artifacts for the multimodal document classifier developed as part of the Multimodal Document Intelligence Master's Thesis project.
The classifier combines two independent branches:
- DistilBERT for OCR-derived document text;
- EfficientNetB0 for the document image.
Each branch produces a probability distribution over the 16 RVL-CDIP classes. The 16 text probabilities are concatenated with the 16 image probabilities and passed to a Logistic Regression late-fusion model, which produces the final prediction.
The complete application, Docker runtime and deployment code are maintained separately at:
https://github.com/gonzalobosque/multimodal-document-intelligence
Model overview
| Component | Role | Artifact |
|---|---|---|
| DistilBERT | Text classification from OCR text | text_model/ |
| EfficientNetB0 | Image classification | image_model/model.keras |
| Logistic Regression | Late fusion of both probability vectors | fusion_logreg.joblib |
| Class order | Canonical mapping between output indices and labels | class_order.json |
The text branch was fine-tuned from distilbert-base-uncased using its associated tokenizer.
The image branch was built from EfficientNetB0 initialized with ImageNet weights and fine-tuned for the 16 RVL-CDIP categories.
Final evaluation
The final multimodal classifier was evaluated once on the held-out test partition after model and fusion selection had been completed.
| Metric | Result |
|---|---|
| Macro-F1 | 0.9165 |
| Accuracy | 0.9166 |
| Valid test documents | 38,520 |
| Classes | 16 |
The held-out test partition was not used for model selection or final fusion fitting.
Inference contract
These artifacts are intended to reproduce the inference pipeline used during evaluation. The following contracts should be preserved:
- OCR text is expected in English;
- DistilBERT input is truncated to a maximum of 512 tokens;
- EfficientNetB0 receives RGB images resized to 384 × 384;
- the fusion model receives exactly 32 features;
- fusion feature order is fixed: 16 text probabilities first, followed by 16 image probabilities;
- class order must be loaded from
class_order.json; - changing preprocessing, probability order or class order requires revalidation.
The end-to-end application additionally performs OCR with Tesseract before running the text branch.
Class order
The canonical class order is:
[
"letter",
"form",
"email",
"handwritten",
"advertisement",
"scientific report",
"scientific publication",
"specification",
"file folder",
"news article",
"budget",
"invoice",
"presentation",
"questionnaire",
"resume",
"memo"
]
Artifact structure
.
├── README.md
├── checksums.sha256
├── class_order.json
├── fusion_logreg.joblib
├── image_model/
│ └── model.keras
└── text_model/
├── config.json
├── model.safetensors
├── special_tokens_map.json
├── tokenizer.json
├── tokenizer_config.json
└── vocab.txt
Only the artifacts required for classifier inference are included.
Training outputs, validation predictions, duplicate checkpoints and environment-specific files are intentionally excluded.
Loading the artifacts
The individual components can be loaded with the same libraries used by the application:
import json
import joblib
import tensorflow as tf
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("text_model")
text_model = AutoModelForSequenceClassification.from_pretrained("text_model")
image_model = tf.keras.models.load_model(
"image_model/model.keras",
compile=False,
)
fusion_model = joblib.load("fusion_logreg.joblib")
with open("class_order.json", "r", encoding="utf-8") as f:
class_order = json.load(f)
For the complete OCR, preprocessing, late-fusion and application workflow, use the companion GitHub repository rather than treating these files as a standalone inference package.
Data provenance
The models were developed using RVL-CDIP document images together with OCR transcriptions associated with that corpus.
No RVL-CDIP source images are distributed in this repository.
No QS-OCR corpus files are distributed in this repository.
The semantic retrieval catalogue used by the wider application, documents.parquet, is also not included, because it contains dataset-derived OCR content.
This repository therefore contains trained classifier artifacts and small configuration files only.
Scope and limitations
This is an academic prototype.
- Reported classification metrics apply to the RVL-CDIP domain.
- The classifier has not yet been systematically evaluated on documents from a different domain.
- OCR errors can affect the text branch.
- Some document classes remain more difficult to distinguish than others.
- The confidence returned by the late-fusion classifier should not be interpreted as a guarantee of correctness.
- These artifacts are not presented as a production-ready document-management system.
Licensing and third-party material
This model repository is marked license: other because the artifact bundle has mixed provenance and should not be represented as if a single permissive license automatically covered every upstream component and training source.
Third-party models, pretrained weights, datasets and document content remain subject to their respective original terms.
In particular:
- the DistilBERT branch was fine-tuned from
distilbert-base-uncased; - the EfficientNetB0 branch was initialized from ImageNet-pretrained weights;
- the training corpus is based on RVL-CDIP and OCR-derived text associated with that corpus;
- source dataset documents and the semantic catalogue are not redistributed here.
This repository does not grant rights over RVL-CDIP documents, QS-OCR corpus content or other third-party material.
The source code for the maintained application is distributed separately in the companion GitHub repository; its licensing is handled there.
SHA-256 checksums
The runtime artifacts in this release can be verified against checksums.sha256.
| File | SHA-256 |
|---|---|
text_model/model.safetensors |
0f816dce4aa40df26bdd528ffdd5beee076c0515099de27447aae52f952ef7e2 |
text_model/config.json |
274ccddb3b1c4841de147df99c1d82c5f678767dd96015f52babeb461bdddba8 |
text_model/tokenizer_config.json |
a3c9410f27554c6e26af779fd147e536c6c5f9e5cb0e6b8f38ddaa25e1b0f6af |
text_model/special_tokens_map.json |
b6d346be366a7d1d48332dbc9fdf3bf8960b5d879522b7799ddba59e76237ee3 |
text_model/tokenizer.json |
435667fab0c06c165b1283ecb422497c37124f2d6a35b2ac73dc876332fc9518 |
text_model/vocab.txt |
07eced375cec144d27c900241f3e339478dec958f92fddbc551f295c992038a3 |
image_model/model.keras |
f46c30bedc0761f270af41a3e260352787ac5327d0e3abd971c24416e5e98eb5 |
fusion_logreg.joblib |
fab7d8dc6cb13568c8b93d7fd4ddba1b77f86e7a5c625f83c7183579601991c0 |
class_order.json |
fc573a380d8a9cbb635a81edb60d20b3b0b42646d7293316b3bc1911cac005e6 |
Author
Gonzalo Bosque Rodríguez
Master's Thesis in Artificial Intelligence, 2026.
- Downloads last month
- 6