dataset-schema-task-classifier

LiquidAI/LFM2.5-Encoder-350M fine-tuned for multi-label text classification on davanstrien/dataset-schemas-with-task-categories.

  • Labels (35): audio-classification, audio-to-audio, automatic-speech-recognition, feature-extraction, fill-mask, image-classification, image-feature-extraction, image-segmentation, image-text-to-text, image-to-3d, image-to-image, image-to-text, multiple-choice, object-detection, question-answering, reinforcement-learning, robotics, sentence-similarity, summarization, table-question-answering, tabular-classification, tabular-regression, text-classification, text-generation, text-retrieval, text-to-image, text-to-speech, text-to-video, time-series-forecasting, token-classification, … (35 total)
  • Date: 2026-07-29 13:00 UTC

This model uses a custom classification head (mean pooling over a backbone without a native sequence-classification class), so loading requires trust_remote_code=True. vLLM serving requires a standard architecture.

Evaluation

Metric Value
f1_micro @ 0.5 0.5898
f1_macro @ 0.5 0.3989
f1_micro @ tuned 0.6238
f1_macro @ tuned 0.5177

Per-label decision thresholds tuned on the eval split are stored in config.classifier_thresholds.

Choosing an operating point: the stored thresholds maximise per-label F1. For precision-first use (e.g. auto-applying labels), act only on predictions well above their threshold — sigmoid probabilities are a usable confidence signal, and filtering to high-confidence predictions trades coverage for precision. Route the rest to review.

Usage

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained("davanstrien/dataset-schema-task-classifier", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("davanstrien/dataset-schema-task-classifier", trust_remote_code=True)

inputs = tokenizer("your text here", return_tensors="pt", truncation=True)
probs = torch.sigmoid(model(**inputs).logits)[0]
thresholds = torch.tensor(model.config.classifier_thresholds)  # tuned on validation
labels = [model.config.id2label[i] for i in (probs >= thresholds).nonzero().flatten().tolist()]
print(labels)

Reproduction

Produced on Hugging Face Jobs (gpu) with the train-classifier.py recipe from uv-scripts. Run it yourself:

hf jobs uv run --flavor gpu --secrets HF_TOKEN \
    https://huggingface.co/datasets/uv-scripts/classification/raw/main/train-classifier.py \
    davanstrien/dataset-schemas-with-task-categories davanstrien/dataset-schema-task-classifier --label-column labels
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for davanstrien/dataset-schema-task-classifier

Finetuned
(15)
this model

Dataset used to train davanstrien/dataset-schema-task-classifier