You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Model Card for Fiafghan/humanoid-HDAR1: Qwen2.5 Fine-Tuned for Humanitarian Data Analytics

Model Details

Model Description

humanoid-HDAR1 is a causal language model based on Qwen2.5-3B, fine-tuned using LoRA (Low-Rank Adaptation) for tasks in humanitarian data analytics. The model is designed to understand and respond to queries about humanitarian datasets, summarize information, and provide insights, while politely refusing unrelated questions.

  • Developed by: Fardin Ibrahimi, CEO Of Humanoid International
  • Model type: Causal Language Model (CAUSAL_LM)
  • Language(s) (NLP): English
  • License: [Specify License]
  • Finetuned from model: Qwen/Qwen2.5-3B
  • Tags: LoRA, Humanitarian, NLP

Model Sources

Model Sources

Uses

Direct Use

from huggingface_hub import login import os import torch

os.environ["HF_TOKEN"] = "" # your token login(token=os.environ.get("HF_TOKEN"))

from transformers import AutoTokenizer, AutoModelForCausalLM from peft import PeftModel

BASE_MODEL = "Qwen/Qwen2.5-3B" LORA_REPO = "Fiafghan/humanoid-HDAR1"

tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, trust_remote_code=True) base_model = AutoModelForCausalLM.from_pretrained(BASE_MODEL, device_map="auto", torch_dtype="auto", trust_remote_code=True) model = PeftModel.from_pretrained(base_model, LORA_REPO) model.eval()

user_input = "Explain humanitarian disaster assessment (HDAR) in simple terms."

SYSTEM_PROMPT = ( "You are a humanitarian data analytics assistant.\n" "You ONLY answer questions related to humanitarian data analytics.\n" "If a question is NOT related, politely refuse." )

messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_input} ]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad(): output = model.generate( **inputs, max_new_tokens=256, do_sample=True, temperature=0.7, top_p=0.9 )

response = tokenizer.decode(output[0], skip_special_tokens=True) print("Assistant:", response)

Log in if private

os.environ["HF_TOKEN"] = "" login(token=os.environ.get("HF_TOKEN"))

BASE_MODEL = "Qwen/Qwen2.5-3B" LORA_REPO = "Fiafghan/humanoid-HDAR1"

Load base model

tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, trust_remote_code=True) base_model = AutoModelForCausalLM.from_pretrained( BASE_MODEL, torch_dtype=torch.float16, device_map="auto", trust_remote_code=True )

Attach LoRA adapter

model = PeftModel.from_pretrained(base_model, LORA_REPO) model.eval()

Downstream Use

The merged model can be fine-tuned further or used as a base for retrieval-augmented pipelines, dashboards, or other HDA tooling.

Out-of-Scope Use

  • Legal, medical, or life-critical decision-making without human oversight
  • Targeted surveillance or discriminatory profiling

Bias, Risks, and Limitations

The model inherits biases from both the base model (Qwen/Qwen2.5-3B) and the fine-tuning dataset (nlp-thedeep/humset). It can hallucinate, produce biased outputs, or be incorrect on out-of-domain queries.

Recommendations

  • Keep a human-in-the-loop for high-stakes decisions.
  • Evaluate on representative validation sets and run bias/safety checks.
  • Monitor outputs in production and log model responses for audit.

How to Get Started with the Model

Install the typical dependencies used in the script:

pip install -U pip
pip install transformers datasets trl peft accelerate huggingface_hub

Run the fine-tuning script (example):

# from repo root
python core/finetune-humanoid-hda-R1.py

Notes:

  • The script downloads train.jsonl, validation.jsonl, and test.jsonl from the nlp-thedeep/humset dataset.
  • The base model used is Qwen/Qwen2.5-3B.
  • Do NOT commit your Hugging Face token to the repository. Use environment variables or huggingface-cli login.

Example push workflow (secure):

export HF_TOKEN="hf_..."  # set locally, don't commit
python -c "from huggingface_hub import login; import os; login(token=os.environ['HF_TOKEN'])"

In Python, prefer to call huggingface_hub.login(token=os.environ.get('HF_TOKEN')) rather than embedding tokens in scripts.

Training Details

Training Data

  • Dataset: nlp-thedeep/humset (script downloads JSONL files directly)

Training Procedure

The script performs PEFT (LoRA) fine-tuning with trl.SFTTrainer and the following configuration derived from the code:

LoRA / PEFT configuration

  • r: 8
  • lora_alpha: 32
  • target_modules: ["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"]
  • lora_dropout: 0.05
  • bias: "none"
  • task_type: "CAUSAL_LM"

Preprocessing

Each example is formatted into a conversation-style prompt using this template:

<|system|>
{SYSTEM_PROMPT}
<|user|>
{example_text}
<|assistant|>

Where SYSTEM_PROMPT in the script is:

"You are a humanitarian data analytics assistant.\nYou ONLY answer questions related to humanitarian data analytics.\nIf a question is NOT related, politely refuse."

Training hyperparameters (from script)

  • model dtype: bfloat16 if supported else float16
  • output_dir: /content/humanoid-HDAR1 (training arg) and OUTPUT_DIR = /content/qwen2.5-humanitarian-analytics (script variable)
  • per_device_train_batch_size: 1
  • per_device_eval_batch_size: 1
  • gradient_accumulation_steps: 8
  • learning_rate: 2e-4
  • max_steps: 500
  • logging_steps: 10
  • save_steps: 100
  • fp16: True
  • save_total_limit: 2

The script uses get_peft_model(...) to attach LoRA to the base model and then PeftModel.from_pretrained(...).merge_and_unload() to create the final merged model for pushing.

Evaluation

The script does not produce evaluation metrics. Recommended evaluations before publishing:

  • Perplexity on held-out test split
  • Task-specific classification/regression metrics depending on the downstream HDA task
  • Safety checks (adversarial prompts, toxic output metrics)

Model Examination and Environmental Impact

Interpretability analyses and environmental accounting are not included in the script. If required, record hardware (GPU type), number of training hours, and region, then use the ML CO2 calculator.

  • Hardware Type: [More Information Needed]
  • Hours used: [More Information Needed]
  • Cloud Provider / Region: [More Information Needed]
  • Carbon Emitted: [More Information Needed]

Technical Specifications

  • Base model: Qwen/Qwen2.5-3B (decoder-only)
  • Fine-tuning method: supervised fine-tuning (SFT) with LoRA (PEFT)
  • Libraries: transformers, datasets, trl, peft, huggingface_hub, torch

Citation

If you publish results based on this fine-tuning run, cite the Qwen base model and the dataset used:

More Information, Authors & Contact

Before pushing to the Hub, update these fields with your chosen license and contact info.

  • Model Card Authors: Fardin Ibrahimi
  • Contact: update with Hugging Face username or org. (Script default HF_REPO_NAME = "Fiafghan/qwen2.5-humanitarian-analytics" โ€” change as needed.)

Security & publishing notes

  • Remove or rotate any secret tokens before pushing commits. The original script contains an example login(token=...) call; do not commit tokens.
  • Verify license compatibility with Qwen/Qwen2.5-3B and any dataset restrictions before publishing.

If you want, I can now:

  • update the script to remove hard-coded tokens and use os.environ for HF_TOKEN,
  • create a small inference example that loads the merged model and runs a sample prompt,
  • or generate a top-level README.md for the model repository that references this model card.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support