docclassifier / README.md
akbted's picture
Upload folder using huggingface_hub
c4e0d2f verified
|
Raw
History Blame Contribute Delete
2.85 kB
metadata
pipeline_tag: other
tags:
  - document-classification
  - medical

ArcaAI Document Classifier

Status: in development. The pipeline, API contracts, label set, and endpoints may change without notice. Not yet recommended for production use.

A custom Hugging Face pipeline for classifying medical / health-related documents into one of six types:

Label Meaning
bill Hospital bill / invoice
discharge_summary Discharge summary
echs_card ECHS (Ex-Servicemen Contributory Health Scheme) card
lab_report Laboratory report
prescription_slip Prescription slip
referral_letter_by_polyclinic Referral letter from a polyclinic

Classification is performed by the ArcaAI hosted inference API (a fine-tuned document model served over Triton). This repository publishes a weightless custom pipeline: it loads configuration only and delegates the actual inference to the API. No model weights are downloaded.

Quickstart

pip install transformers torch requests

export DOCCLASSIFIER_API_KEY="dc_xxx"   # token from the ArcaAI admin API
from transformers import pipeline

classifier = pipeline(
    "document-classification",
    model="akbted/docclassifier",
    trust_remote_code=True,
)

print(classifier("invoice.pdf"))

print(classifier(["invoice.pdf", "lab_report.png", "prescription.pdf"]))

trust_remote_code=True is required: the repo's pipeline.py is executed locally to make the API call.

Supported inputs

  • PDFs and images (pdf, png, jpg, jpeg).
  • File paths (str), raw bytes, or a list of either.
  • Results are per page: [{"file_path", "page", "label", "confidence"}, ...].

Configuration

Env var Purpose
DOCCLASSIFIER_API_KEY API token (required).
DOCCLASSIFIER_API_URL Override the endpoint baked into the model config.

Threshold filtering per call:

classifier("invoice.pdf", confidence_threshold=0.9)

Try it

from transformers import pipeline

classifier = pipeline(
    "document-classification",
    model="akbted/docclassifier",
    trust_remote_code=True,
)

for page in classifier("invoice.pdf"):
    print(f"page={page['page']}  label={page['label']}  confidence={page['confidence']}")

Development status & limitations

  • In development; subject to breaking changes.
  • Requires network access to the ArcaAI API and a valid API token.
  • API requests are rate-limited per token.
  • The stand-in architecture ships no weights and is a loading vehicle only: AutoModel.from_pretrained("akbted/docclassifier", trust_remote_code=True) returns a weightless shell by design.

Getting an API token

Contact the ArcaAI team to provision a token for the public API.

License

See LICENSE (to be added).