Model Card

OneEmo is a unified multimodal reasoning model for emotion perception, understanding, and interaction in video scenarios. It simultaneously supports eight tasks, ranging from sentiment analysis and emotion/intention recognition to empathetic response generation and emotional support conversation within a single frameworks.

Overview

Support task

  • MER – Multimodal emotion recognition (single label)
  • MSA – Multimodal sentiment analysis (positive/neutral/negative)
  • OVMER – Open-vocabulary multimodal emotion recognition
  • MIR – Multimodal intention recognition (requires candidate labels)
  • MHD – Multimodal humour understanding (yes/no)
  • MSD – Multimodal sarcasm understanding (yes/no)
  • ERG – Empathetic response generation
  • ESC – Emotional support conversation

Software

The model was trained and evaluated in the following environment:

  • Python libraries:
    • torch: 2.10.0+cu128
    • torchaudio: 2.10.0+cu128
    • torchvision: 0.25.0+cu128
    • transformers: 5.2.0
    • vllm: 0.19.0
    • ms-swift: 4.1.3 (optional, for alternative inference engine)
  • Environment management: Conda (environment name: oneemo)

Model Sources

How to Get Started with the Model

Use the code below to get started with the model.

VLLM

vllm serve Jiaha0Hu4ng/OneEmo \
    --host 0.0.0.0 \
    --port 8000 \
    --served-model-name OneEmo \
    --max-model-len 32768 \
    --mm-encoder-tp-mode data \
    --media-io-kwargs '{"video": {"num_frames": 16}}' \
    --default-chat-template-kwargs '{"enable_thinking": true}' \
    --reasoning-parser qwen3

Swift

import torch
from swift import get_model_processor, get_template
from swift.infer_engine import TransformersEngine, InferRequest, RequestConfig

model_path = "Jiaha0Hu4ng/OneEmo"

model, processor = get_model_processor(
    model_path,
    model_type="qwen3_5",
    torch_dtype=torch.bfloat16,
    attn_impl="flash_attn",
)
template = get_template(processor, enable_thinking=True)
engine = TransformersEngine(model, template=template)
request_config = RequestConfig(max_tokens=2048, temperature=0.7)

infer_request = InferRequest(
    messages=[{
        "role": "user",
        "content": (
            "Based on the video, describe the range of emotions the character "
            "may be feeling."
        )
    }],
    videos=["path/to/your/video.mp4"]
)
resp_list = engine.infer([infer_request], request_config=request_config)
print(resp_list[0].choices[0].message.content)

Transformers

import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from qwen_vl_utils import process_vision_info

model_path = "Jiaha0Hu4ng/OneEmo"

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",  # fallback to "sdpa" if flash-attn is not installed
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_path)

# Example: open-vocabulary multimodal emotion recognition (OVMER)
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "video",
                "video": "path/to/your/video.mp4",
                "max_pixels": 360 * 420,
                "fps": 2.0,
            },
            {
                "type": "text",
                "text": "Based on the video, describe the range of emotions the character may be feeling.",
            },
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=2048,
        temperature=0.7,
        top_p=0.9,
        top_k=50,
    )
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])

For other tasks, replace the prompt text with the corresponding template (see the full model card on Hugging Face for the complete prompt table). The model also supports deployment with vLLM (OpenAI-compatible API) and inference via ms-swift; refer to the repository for those examples.

Training Details

Training Data

We use Emo-World-130K to train OneEmo.

Evaluation

Testing Data, Factors & Metrics

Testing Data

[To Be continued]

Factors

[To Be continued]

Metrics

[To Be continued]

Results

[To Be continued]

Technical Specifications [optional]

[To Be continued]

Citation [optional]

[To Be continued]

Uses

Direct Use

OneEmo can be used directly for video-based emotion analysis, intention recognition, sentiment classification, humour/sarcasm detection, and for generating empathetic or supportive textual responses. It is suitable for researchers and developers building affective computing applications, mental-health chatbots, or social signal processing tools.

Downstream Use [optional]

The model can be fine-tuned on domain-specific datasets (e.g., clinical interviews, customer service videos) or integrated into larger dialogue systems for emotionally aware conversation. When fine-tuning, users should follow the same prompt format and task prefix convention.

Note: The model is not intended for:

  • Clinical diagnosis of mental disorders without professional oversight.
  • Making high-stakes decisions (e.g., hiring, legal judgements) based on emotion predictions.
  • Use in surveillance or manipulative applications. It may underperform on languages or cultures not well represented in training data, and on videos longer than the maximum frame budget (16 frames by default).

Bias, Risks, and Limitations

  • The model inherits biases from its pre-training and fine-tuning data; emotion labels may reflect cultural stereotypes.
  • Generated empathetic or counselling responses can occasionally be unsafe, inappropriate, or not clinically sound.
  • Performance degrades with noisy audio transcripts, low-resolution video, or out-of-domain scenarios.
  • The chain-of-thought reasoning is not guaranteed to be factually correct or logically consistent.

Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. We recommend:

  • Not treating model outputs as professional mental-health advice.
  • Employing human review when used in sensitive contexts.
  • Evaluating the model on representative data before deployment.
  • Considering additional fairness and safety guardrails, especially for vulnerable populations.
Downloads last month
91
Safetensors
Model size
5B params
Tensor type
BF16
·
Video Preview
loading

Model tree for Jiaha0Hu4ng/OneEmo

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(1)
this model

Dataset used to train Jiaha0Hu4ng/OneEmo