TEM-LLM / README.md
LabSmart's picture
Update README.md
5f8ef08 verified
|
Raw
History Blame Contribute Delete
2.63 kB
metadata
license: llama2
language:
  - en
base_model:
  - liuhaotian/llava-v1.5-7b
pipeline_tag: visual-question-answering
tags:
  - tem
  - materials-science
  - multimodal

TEM-LLM

TEM-LLM is a domain-specific multimodal large language model for Transmission Electron Microscopy (TEM) image analysis, fine-tuned from LLaVA-v1.5-7B using a difficulty-aware curriculum learning strategy.

This model is associated with the following work (currently under review):

AI-Assisted Materials Characterization: A Curriculum-Guided Multimodal Framework for Transmission Electron Microscopy


Model Details

Base Model LLaVA-v1.5-7B (Vicuna-v1.5-7B)
Training Strategy Difficulty-Aware Curriculum Learning (4 stages)
Training Data ~216K QA pairs across 40K TEM images
Domain Transmission Electron Microscopy (TEM)
Modalities CTEM, HR-TEM, STEM, Diffraction
Fine-tuning Method LoRA (Low-Rank Adaptation)
License LLaMA 2

Intended Use

TEM-LLM is designed for:

  • Image Classification: Identifying TEM imaging modalities (CTEM, HRTEM, STEM, diffraction)
  • Microstructural Feature Detection: Localizing and counting structural elements such as nanoparticles and grain boundaries
  • Feature Analysis & Interpretation: Providing scientific explanations of observed microstructural phenomena
  • Comprehensive Image Description: Integrating multiple aspects of visual understanding into structured descriptions

Important: Inference Requirements

TEM-LLM is built on LLaVA and cannot be loaded directly via transformers. Inference requires the LLaVA repository.

Step 1 — Clone LLaVA:

git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
pip install -e .

Step 2 — Download weights:

from huggingface_hub import snapshot_download
snapshot_download(
    repo_id="LabSmart/TEM-LLM", 
    local_dir="./TEM-LLM"
)

Step 3 — Run inference:

python -m llava.serve.cli \
    --model-path "./TEM-LLM" \
    --image-file "path/to/your/tem_image.jpg" \
    --load-4bit

For the complete pipeline including data crawling, preprocessing, knowledge distillation, training, and evaluation scripts, please refer to the GitHub repository.


Limitations

  • TEM-LLM is optimized for TEM image analysis and may produce unreliable outputs when applied to out-of-scope questions or out-of-domain images.
  • The model provides approximate qualitative descriptions rather than exact quantitative measurements.