You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

ArcaAI Document Classifier

Status: in development. The pipeline, API contracts, label set, and endpoints may change without notice. Not yet recommended for production use.

A custom Hugging Face pipeline for classifying medical / health-related documents into one of six types:

Label Meaning
bill Hospital bill / invoice
discharge_summary Discharge summary
echs_card ECHS (Ex-Servicemen Contributory Health Scheme) card
lab_report Laboratory report
prescription_slip Prescription slip
referral_letter_by_polyclinic Referral letter from a polyclinic

Classification is performed by the ArcaAI hosted inference API (a fine-tuned document model served over Triton). This repository publishes a weightless custom pipeline: it loads configuration only and delegates the actual inference to the API. No model weights are downloaded.

Quickstart

pip install transformers torch requests

export DOCCLASSIFIER_API_KEY="dc_xxx"   # token from the ArcaAI admin API
from transformers import pipeline

classifier = pipeline(
    "document-classification",
    model="akbted/docclassifier",
    trust_remote_code=True,
)

print(classifier("invoice.pdf"))

print(classifier(["invoice.pdf", "lab_report.png", "prescription.pdf"]))

trust_remote_code=True is required: the repo's pipeline.py is executed locally to make the API call.

Supported inputs

  • PDFs and images (pdf, png, jpg, jpeg).
  • File paths (str), raw bytes, or a list of either.
  • Results are per page: [{"file_path", "page", "label", "confidence"}, ...].

Configuration

Env var Purpose
DOCCLASSIFIER_API_KEY API token (required).
DOCCLASSIFIER_API_URL Override the endpoint baked into the model config.

Threshold filtering per call:

classifier("invoice.pdf", confidence_threshold=0.9)

Try it

from transformers import pipeline

classifier = pipeline(
    "document-classification",
    model="akbted/docclassifier",
    trust_remote_code=True,
)

for page in classifier("invoice.pdf"):
    print(f"page={page['page']}  label={page['label']}  confidence={page['confidence']}")

Development status & limitations

  • In development; subject to breaking changes.
  • Requires network access to the ArcaAI API and a valid API token.
  • API requests are rate-limited per token.
  • The stand-in architecture ships no weights and is a loading vehicle only: AutoModel.from_pretrained("akbted/docclassifier", trust_remote_code=True) returns a weightless shell by design.

Getting an API token

Contact the ArcaAI team to provision a token for the public API.

License

See LICENSE (to be added).

Downloads last month
63
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support