Multimodal Document Intelligence — Classifier Artifacts

This repository contains the trained runtime artifacts for the multimodal document classifier developed as part of the Multimodal Document Intelligence Master's Thesis project.

The classifier combines two independent branches:

  • DistilBERT for OCR-derived document text;
  • EfficientNetB0 for the document image.

Each branch produces a probability distribution over the 16 RVL-CDIP classes. The 16 text probabilities are concatenated with the 16 image probabilities and passed to a Logistic Regression late-fusion model, which produces the final prediction.

The complete application, Docker runtime and deployment code are maintained separately at:

https://github.com/gonzalobosque/multimodal-document-intelligence

Model overview

Component Role Artifact
DistilBERT Text classification from OCR text text_model/
EfficientNetB0 Image classification image_model/model.keras
Logistic Regression Late fusion of both probability vectors fusion_logreg.joblib
Class order Canonical mapping between output indices and labels class_order.json

The text branch was fine-tuned from distilbert-base-uncased using its associated tokenizer.

The image branch was built from EfficientNetB0 initialized with ImageNet weights and fine-tuned for the 16 RVL-CDIP categories.

Final evaluation

The final multimodal classifier was evaluated once on the held-out test partition after model and fusion selection had been completed.

Metric Result
Macro-F1 0.9165
Accuracy 0.9166
Valid test documents 38,520
Classes 16

The held-out test partition was not used for model selection or final fusion fitting.

Inference contract

These artifacts are intended to reproduce the inference pipeline used during evaluation. The following contracts should be preserved:

  • OCR text is expected in English;
  • DistilBERT input is truncated to a maximum of 512 tokens;
  • EfficientNetB0 receives RGB images resized to 384 × 384;
  • the fusion model receives exactly 32 features;
  • fusion feature order is fixed: 16 text probabilities first, followed by 16 image probabilities;
  • class order must be loaded from class_order.json;
  • changing preprocessing, probability order or class order requires revalidation.

The end-to-end application additionally performs OCR with Tesseract before running the text branch.

Class order

The canonical class order is:

[
  "letter",
  "form",
  "email",
  "handwritten",
  "advertisement",
  "scientific report",
  "scientific publication",
  "specification",
  "file folder",
  "news article",
  "budget",
  "invoice",
  "presentation",
  "questionnaire",
  "resume",
  "memo"
]

Artifact structure

.
├── README.md
├── checksums.sha256
├── class_order.json
├── fusion_logreg.joblib
├── image_model/
│   └── model.keras
└── text_model/
    ├── config.json
    ├── model.safetensors
    ├── special_tokens_map.json
    ├── tokenizer.json
    ├── tokenizer_config.json
    └── vocab.txt

Only the artifacts required for classifier inference are included.

Training outputs, validation predictions, duplicate checkpoints and environment-specific files are intentionally excluded.

Loading the artifacts

The individual components can be loaded with the same libraries used by the application:

import json
import joblib
import tensorflow as tf
from transformers import AutoModelForSequenceClassification, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("text_model")
text_model = AutoModelForSequenceClassification.from_pretrained("text_model")

image_model = tf.keras.models.load_model(
    "image_model/model.keras",
    compile=False,
)

fusion_model = joblib.load("fusion_logreg.joblib")

with open("class_order.json", "r", encoding="utf-8") as f:
    class_order = json.load(f)

For the complete OCR, preprocessing, late-fusion and application workflow, use the companion GitHub repository rather than treating these files as a standalone inference package.

Data provenance

The models were developed using RVL-CDIP document images together with OCR transcriptions associated with that corpus.

No RVL-CDIP source images are distributed in this repository.

No QS-OCR corpus files are distributed in this repository.

The semantic retrieval catalogue used by the wider application, documents.parquet, is also not included, because it contains dataset-derived OCR content.

This repository therefore contains trained classifier artifacts and small configuration files only.

Scope and limitations

This is an academic prototype.

  • Reported classification metrics apply to the RVL-CDIP domain.
  • The classifier has not yet been systematically evaluated on documents from a different domain.
  • OCR errors can affect the text branch.
  • Some document classes remain more difficult to distinguish than others.
  • The confidence returned by the late-fusion classifier should not be interpreted as a guarantee of correctness.
  • These artifacts are not presented as a production-ready document-management system.

Licensing and third-party material

This model repository is marked license: other because the artifact bundle has mixed provenance and should not be represented as if a single permissive license automatically covered every upstream component and training source.

Third-party models, pretrained weights, datasets and document content remain subject to their respective original terms.

In particular:

  • the DistilBERT branch was fine-tuned from distilbert-base-uncased;
  • the EfficientNetB0 branch was initialized from ImageNet-pretrained weights;
  • the training corpus is based on RVL-CDIP and OCR-derived text associated with that corpus;
  • source dataset documents and the semantic catalogue are not redistributed here.

This repository does not grant rights over RVL-CDIP documents, QS-OCR corpus content or other third-party material.

The source code for the maintained application is distributed separately in the companion GitHub repository; its licensing is handled there.

SHA-256 checksums

The runtime artifacts in this release can be verified against checksums.sha256.

File SHA-256
text_model/model.safetensors 0f816dce4aa40df26bdd528ffdd5beee076c0515099de27447aae52f952ef7e2
text_model/config.json 274ccddb3b1c4841de147df99c1d82c5f678767dd96015f52babeb461bdddba8
text_model/tokenizer_config.json a3c9410f27554c6e26af779fd147e536c6c5f9e5cb0e6b8f38ddaa25e1b0f6af
text_model/special_tokens_map.json b6d346be366a7d1d48332dbc9fdf3bf8960b5d879522b7799ddba59e76237ee3
text_model/tokenizer.json 435667fab0c06c165b1283ecb422497c37124f2d6a35b2ac73dc876332fc9518
text_model/vocab.txt 07eced375cec144d27c900241f3e339478dec958f92fddbc551f295c992038a3
image_model/model.keras f46c30bedc0761f270af41a3e260352787ac5327d0e3abd971c24416e5e98eb5
fusion_logreg.joblib fab7d8dc6cb13568c8b93d7fd4ddba1b77f86e7a5c625f83c7183579601991c0
class_order.json fc573a380d8a9cbb635a81edb60d20b3b0b42646d7293316b3bc1911cac005e6

Author

Gonzalo Bosque Rodríguez

Master's Thesis in Artificial Intelligence, 2026.

Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support