- HAR-Agent: Multilingual Multimodal Human Activity Recognition via Knowledge-Distilled LLM Reasoning Read Paper
HAR-Agent: Multilingual Multimodal Human Activity Recognition via Knowledge-Distilled LLM Reasoning Read Paper
Overview
HAR-Agent is a multimodal, multilingual human activity recognition (HAR) system that combines vision-language-audio perception with knowledge-distilled LLM reasoning. It achieves sophisticated activity understanding on consumer-grade hardware โ the best student model requires under 1 GB VRAM.
This repository contains all trained LoRA adapter weights for 12 fine-tuned models (2 teachers + 10 students) across two training paradigms, plus cached teacher logits and comprehensive evaluation results.
Key Results
| Model | Method | Params | Accuracy | F1 (Macro) | VRAM |
|---|---|---|---|---|---|
| IT-T-72B | IT | 72B | 60.3% | 0.621 | ~40 GB |
| IT-S-1.5B | IT | 1.5B | 41.4% | 0.405 | 0.8 GB |
| IT-S-3B | IT | 3B | 39.8% | 0.388 | 1.7 GB |
| SFT-T-72B | SFT | 72B | 42.6% | 0.401 | ~40 GB |
| Audio (multilingual) | IT-S-1.5B + Whisper | โ | 89.2% | 0.890 | ~6 GB |
Inverse scaling finding: The smallest IT student (1.5B) outperforms all larger students (3Bโ32B), making the most accessible model also the most capable.
Architecture
HAR-Agent uses a three-pathway multimodal architecture:
- Visual pathway: LLaVA-NeXT 7B generates natural language scene descriptions from video frames (17 uniformly sampled frames per clip), with Short-Term Memory (STM) and semantic deduplication (cosine similarity threshold 0.85โ0.92).
- Audio pathway: Whisper-medium transcribes spoken activity descriptions in any language, enabling multilingual HAR across 5+ languages.
- Text pathway: Direct text input for integration with existing systems.
All pathways converge to a unified text representation, which a single LLM reasoning module classifies into one of 14 activity classes.
Knowledge Distillation Pipeline
- Teacher: Qwen2.5-72B-Instruct fine-tuned with QLoRA (4-bit NormalFloat)
- Students: Qwen2.5-{32B, 14B, 7B, 3B, 1.5B}-Instruct distilled from the teacher
- IT (Instruction Tuning): Preserves the generative CausalLM paradigm; distillation via full-vocabulary KL divergence in shared token space
- SFT (Supervised Fine-Tuning): Adds a classification head; distillation via logit matching
IT outperforms SFT by 17โ28 percentage points at every model scale (all p < 0.0001).
Quick Start
Loading the Best Student (IT-S-1.5B)
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
# Load base model
base_model = "Qwen/Qwen2.5-1.5B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(base_model)
model = AutoModelForCausalLM.from_pretrained(
base_model,
torch_dtype=torch.float16,
device_map="auto",
)
# Load LoRA adapter
model = PeftModel.from_pretrained(model, "khashayargh/HAR-Agent/student_outputs_it_1_5b")
# Classify an activity from a scene description
prompt = """Based on the following scene description, classify the human activity.
Scene: A person is seen bending their knees and lowering their body onto a chair,
with their hands resting on the armrests for support.
Activity classes: Bending, CarryingObject, Cleaning, ClosingCan, Drinking,
LiftingObject, OpeningCan, PuttingDownObjects, Reaching, SittingDown,
StairsClimbingDown, StairsClimbingUp, StandingUp, Walking
The activity is:"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=10, temperature=0.3, top_p=0.95)
prediction = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(prediction.strip()) # โ SittingDown
Loading the Teacher (IT-T-72B)
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
# 4-bit quantisation for ~40 GB VRAM
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype="float16",
)
base_model = "Qwen/Qwen2.5-72B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
base_model,
quantization_config=bnb_config,
device_map="auto",
)
model = PeftModel.from_pretrained(model, "khashayargh/HAR-Agent/teacher_outputs_it_72b")
Activity Classes (14)
| Class | Description |
|---|---|
| Bending | Upper body forward flexion |
| CarryingObject | Transporting an object while walking |
| Cleaning | Wiping, sweeping, or tidying surfaces |
| ClosingCan | Sealing a container with a lid |
| Drinking | Raising a vessel to the mouth |
| LiftingObject | Picking up an object from a surface |
| OpeningCan | Removing a lid from a container |
| PuttingDownObjects | Placing held objects onto a surface |
| Reaching | Extending arm(s) toward a target |
| SittingDown | Transitioning from standing to seated |
| StairsClimbingDown | Descending a staircase |
| StairsClimbingUp | Ascending a staircase |
| StandingUp | Transitioning from seated to standing |
| Walking | Forward locomotion at normal pace |
Datasets
- RHM-HAR (Herts HAR RobotView): 6,701 video clips (5,354 train / 1,347 val), 14 activities, stratified 80/20 split by class label
- Toyota Smarthome: 186 samples, 8 mapped activity classes โ used for cross-domain generalisation evaluation
- Multilingual Audio: 250 TTS samples (5 activities ร 5 languages: English US, English UK, Chinese, Spanish, Farsi)
Training Details
All models use QLoRA (4-bit NormalFloat quantisation) with LoRA rank r=64, ฮฑ=128, dropout=0.05.
| Model | LR | Epochs | Batch | KD Temp | KD ฮฑ | GPUs |
|---|---|---|---|---|---|---|
| IT teachers | 5e-5 | 3 | 4 | โ | โ | 4รA100 80GB |
| IT students | 5e-5 | 3 | 4โ8 | 2.0 | 0.5 | 1โ4รA100 |
| SFT teachers | 2e-5 | 5 | 4 | โ | โ | 4รA100 80GB |
| SFT students | 2e-5 | 5 | 4โ8 | 2.0 | 0.5 | 1โ4รA100 |
Total training time for all 12 models: ~1 month on the UHHPC cluster (NVIDIA A100 80 GB).
Deployment Requirements
| Configuration | VRAM | Latency | Hardware |
|---|---|---|---|
| IT-T-72B + LLaVA | 48 GB | ~14s | 4รA100 (research only) |
| IT-S-1.5B + LLaVA | 9 GB | ~12s | RTX 3060 12GB |
| IT-S-1.5B + Whisper | 6 GB | ~0.6s | RTX 3060 (multilingual audio) |
| IT-S-1.5B only (text) | 0.8 GB | ~0.2s | Any GPU |
Knowledge Distillation Metrics
| Student | TSA | Cohen's ฮบ | Compression | CES |
|---|---|---|---|---|
| IT-S-1.5B | 0.457 | 0.417 | 48ร | 0.95 |
| IT-S-3B | 0.454 | 0.412 | 24ร | 1.89 |
| IT-S-14B | 0.434 | 0.377 | 5.1ร | 8.43 |
| SFT-S-7B (best SFT) | 0.078 | 0.042 | 10.3ร | 0.76 |
Cross-Domain Generalisation (Toyota Smarthome)
| Model | RHM-HAR F1 | Toyota F1 | Drop |
|---|---|---|---|
| IT-T-72B | 0.621 | 0.308 | 50.5% |
| IT-S-1.5B | 0.405 | 0.221 | 45.4% |
| IT-S-14B | 0.336 | 0.261 | 22.3% |
Citation
@article{ghamati2026har,
title={HAR-Agent: Multilingual Multimodal Activity Recognition via Knowledge-Distilled LLM Reasoning},
author={Ghamati, Khashayar and Alashti, Mohammad Reza Shahabian and Fallahirahmatabadi, Ali and Zaraki, Abolfazl},
year={2026},
publisher={Authorea}
}
License
Apache 2.0