Model Card for Model ID

PhysBERT-tag-recommender is a parameter-efficient fine-tuned (LoRA) adapter built on top of thellert/physbert_uncased. It is tailored for multi-label PhySH tag recommendation for condensed matter physics research articles, specifically relevant to APS Physical Review B (PRB).

Model Details

Model Description

  • Developed by: swetlanas
  • Model type: Transformer-based sequence classification with LoRA adapter (PEFT)
  • Language(s) (NLP): English
  • License: apache-2.0
  • Finetuned from model : thellert/physbert_uncased

Model Sources

Uses

Direct Use

This model is designed to recommend Physics subject heading (PhySH) tags for APS Physical Review B articles based on their titles and abstracts.

Out-of-Scope Use

  • General domain text classification (non-physics domains).
  • Processing full-length papers beyond standard sequence length limits without truncation.
  • Recommending tags outside of the PhySH scheme.

Bias, Risks, and Limitations

  • Domain Specificity: The model is heavily biased toward condensed matter, materials science, and solid-state physics terminology. Performance will degrade on astrophysics, high-energy physics, or non-physics fields.
  • Taxonomy Dependency: Label suggestions rely on the label set present during training and may omit newly introduced APS/PhySH terms past July 2026.

How to Get Started with the Model

Use the code snippet below to run inference on a title and abstract pair:

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel

MODEL_ID = "swetlanas/physbert-tag-recommender"

#Load Tokenizer & Model
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)
model.eval()

#Title + Abstract Input
input_title = "Three-dimensional zigzag correlations in the van der Waals Kitaev magnet RuBr3."
input_abstract = "Ruthenium trihalides RuX3 (X=Cl,Br,I) provide a tunable platform for Kitaev magnetism..."
input_text = input_title + " " + input_abstract

inputs = tokenizer(
    input_text,
    padding=True,
    truncation=True,
    max_length=512,
    return_tensors="pt"
)

#Compute Inference
threshold = 0.3  # Probability threshold for multi-label tag acceptance

with torch.no_grad():
    outputs = model(**inputs)
    probabilities = torch.sigmoid(outputs.logits)[0]

#Extract Predicted Tags
predicted_tags = [
    model.config.id2label[i] 
    for i, prob in enumerate(probabilities) 
    if prob > threshold
]

print("Recommended Tags:", predicted_tags)

Training Details

Training Data

  • Dataset: Abstract and title metadata from condensed matter physics publications (targeting APS Physical Review B categories / PhySH taxonomy). More info
  • Target Format: Multi-label binary indicator vectors.

Training Procedure

The model was fine-tuned using Hugging Face's Trainer and peft library with Low-Rank Adaptation (LoRA) on top of the base architecture thellert/physbert_uncased.

Preprocessing

  • Input texts were constructed by concatenating the paper Title and Abstract.
  • Tokenized with a maximum sequence length of 512 tokens.

Parameter-Efficient Fine-Tuning (PEFT/LoRA) Configuration

  • Task Type: Sequence Classification (SEQ_CLS)
  • LoRA Rank ($r$): 8
  • LoRA Alpha ($\alpha$): 16
  • LoRA Dropout: 0.1
  • Target Modules: ["query", "value"]
  • Modules to Save (Fully Fine-Tuned): ["classifier"]

trainable params: 2,156,661 || all params: 113,500,650 || trainable%: 1.9001

Training Hyperparameters

  • Objective: Multi-label classification with Binary Cross-Entropy Loss (BCEWithLogitsLoss).
  • Adapter Technique: Low-Rank Adaptation (LoRA / PEFT).
  • Base Architecture: PhysBERT Uncased.
  • Learning Rate: 3e-4
  • Learning Rate Scheduler: linear
  • Warmup Ratio: 0.1 (10% of total steps)
  • Weight Decay: 0.01
  • Gradient Accumulation Steps: 2
  • Gradient Checkpointing: Enabled (True)
  • Precision: bf16=True (bfloat16 mixed precision) with tf32=True (TensorFloat-32)
  • Evaluation & Save Strategy: Evaluated and checkpointed at the end of every epoch
  • Best Model Selection Metric: precision_at_5 (maximizing Precision@5)
  • Early Stopping: Triggered after 2 epochs of non-improving precision_at_5 (EarlyStoppingCallback(patience=2))
  • Load Best Model at End: Enabled (True)

Evaluation

Testing Data, Factors & Metrics

Testing Data

Held-out test split from the ~50,000 APS Physical Review B abstracts (2016–2026), using Hugging Face Dataset dataset.train_test_split(test_size=0.2, seed=0) for consistency across all model tracks.

Factors

Evaluation was run across the full label space (~3,000 PhySH tags) with no per-subfield breakdown; performance is expected to be stronger on condensed matter / materials science / solid-state topics, which dominate the training distribution.

Metrics

Precision@5 (P@5) was used as the primary metric and the early-stopping / best-checkpoint criterion, since the task is to recommend the top 5 most relevant tags per paper. Recall@5, NDCG@5, and Hit@5 were also tracked via a custom compute_metrics function in the HuggingFace Trainer.

Results

Model Precision@5
TF-IDF + MultinomialNB (baseline) 0.332
napkinXC (PLT) + frozen PhysBERT embeddings 0.371
LoRA fine-tuned PhysBERT (this model) 0.408

Summary

LoRA fine-tuning of PhysBERT outperforms both a classical TF-IDF baseline and a tree-based extreme multi-label classifier (napkinXC) using frozen PhysBERT embeddings, while updating only 2.16M parameters (1.9% of the base model's total).

Technical Specifications

Model Architecture and Objective

Base architecture: PhysBERT thellert/physbert_uncased, a BERT-style transformer encoder pretrained on physics (arXiv) text, with a sequence classification head for multi-label PhySH tag prediction. A LoRA (Low-Rank Adaptation) adapter (r=8, α=16, dropout=0.1) was applied to the query and value attention projection matrices, with the classification head fully fine-tuned. This updates 2.16M parameters (1.9% of the base model's total). The training objective is multi-label binary classification: each of the ~3,000 PhySH tags is treated as an independent binary target, optimized with BCEWithLogitsLoss (Binary Cross-Entropy with logits) over sigmoid outputs, rather than a single softmax over mutually exclusive classes.

Compute Infrastructure

Hardware

RTX 3060 12GB GPU

Software

PyTorch, Hugging Face Transformers, PEFT

Citation

BibTeX:

@misc{physh-tank,
  author = {Swetlana S},
  title = {PhySH-Tank: PhySH Tag Recommendation for PRB Abstracts},
  year = {2026},
  howpublished = {\url{https://github.com/swetlanas/PhySH-Tank}}
}

Model Card Authors

swetlanas

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for swetlanas/physbert-tag-recommender

Adapter
(1)
this model