YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

TinyLLaMA-Indo-Property 🏠

Fine-tuned TinyLlama-1.1B-Chat-v1.0 model using the LoRA (Low-Rank Adaptation) method for the Indonesian-language property consultation domain. This model was developed as part of an undergraduate thesis research project with three experimental scenarios based on LoRA rank variations.


📋 Model Information

Attribute Detail
Base Model TinyLlama/TinyLlama-1.1B-Chat-v1.0
Fine-Tuning Method LoRA (Low-Rank Adaptation) via PEFT
Domain Indonesian-language property consultation
Language Indonesian 🇮🇩
Framework HuggingFace Transformers + TRL (SFTTrainer)
Training Environment Google Colab (GPU)

🧪 Experiment Scenarios

This research tested three scenarios with varying LoRA hyperparameters as follows:

Scenario LoRA Rank (r) LoRA Alpha (α) LoRA Dropout Learning Rate Epoch
Scenario 3 32 64 0 2e-4 12

🎬 Proof of Concept (PoC)

The fine-tuned model was integrated into an "AI Property Consultant" chatbot embedded on a property demo website (Opendoorz) as proof of the model's application in a real-world use case.

⚠️ Note: The source code for this demo/PoC (the Opendoorz website and chatbot UI) lives in a separate repository, which is currently private. Only the screenshots below are shown here for documentation purposes; the demo repository is not publicly accessible.

1. Landing Page — Property Website (Opendoorz)

The chatbot is integrated as a widget on the property website's homepage.

Landing Page Opendoorz

2. Conversation Demo — Mortgage (KPR) Questions

Example interaction where a user asks about additional costs beyond the down payment for a mortgage, as well as a topic guardrail test (the chatbot declines to answer questions outside the property domain, such as motorcycle prices).

Chatbot Demo - Mortgage

3. Conversation Demo — Credit Analysis & Document Legality (AJB)

Example follow-up interaction regarding mortgage eligibility based on salary & existing installments, as well as questions about the legality of a house sale between relatives (AJB / deed of sale).

Chatbot Demo - Credit and AJB

Note: This PoC demonstrates that the model can run end-to-end in a web-based chatbot scenario and is able to answer property-domain questions (mortgages, legal documents, transaction costs), including restricting its responses to questions outside the property topic.


🚀 Usage

Install dependencies

pip install -r requirements.txt

Load model and run inference

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

# Choose the scenario you want to use
SCENARIO = "skenario-3-new"  # change as needed

base_model_id = "TinyLlama/TinyLlama-1.1B-Chat-v1.0"
lora_path = f"RizalGo/TinyLLaMA-Indo-Property/{SCENARIO}/lora"

# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(base_model_id)

# Load base model
model = AutoModelForCausalLM.from_pretrained(
    base_model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# Load LoRA adapter
model = PeftModel.from_pretrained(model, lora_path)
model.eval()

# Inference
def chat(question: str) -> str:
    messages = [
        {
            "role": "system",
            "content": "You are a professional property consultant assistant that helps answer questions about property in Indonesia."
        },
        {
            "role": "user",
            "content": question
        }
    ]
    prompt = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True
    )
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            max_new_tokens=512,
            temperature=0.7,
            do_sample=True,
            pad_token_id=tokenizer.eos_token_id
        )

    response = tokenizer.decode(
        outputs[0][inputs["input_ids"].shape[-1]:],
        skip_special_tokens=True
    )
    return response

# Example usage
question = "What is the difference between SHM and HGB property ownership titles?"
print(chat(question))

📁 Repository Structure

TinyLLaMA-Indo-Property/
├── data-2/                  # Datasets
│   ├── merged/
        ├── data-real-estate-expert-reformat.json  
├── skenario-3-new/          # Scenario 3 (r=32)
│   ├── checkpoints/
│   ├── lora/
│   └── merged/             # the models .safetensor
|   └── skenario-3.gguf     # the models .gguf
├── images/                  # PoC documentation (demo screenshots)
└── README.md

⚙️ Training Details

# LoRA configuration (Scenario 3 example)
lora_config = LoraConfig(
    r=32,
    lora_alpha=64,
    lora_dropout=0.0,
    target_modules=["q_proj", "v_proj"],
    bias="none",
    task_type="CAUSAL_LM"
)

# Training configuration
training_args = SFTConfig(
    learning_rate=2e-4,
    num_train_epochs=12,
    per_device_train_batch_size=4,
    fp16=True,
)

📊 Evaluation

The model was evaluated using BLEU-4 and BERTScore metrics on an Indonesian-language property test dataset (300 samples), comparing the performance of the three LoRA rank scenarios.

Evaluation Results — Scenario 3 (r=32)

Metric Score
BLEU-4 (0–100) 14.54
BERTScore — Precision 0.7533
BERTScore — Recall 0.7442
BERTScore — F1 0.7487

BERTScore Comparison: Before vs After LoRA Fine-Tuning

Model BERTScore-F1
Base model (TinyLlama, before fine-tuning) ~0.65 (65%)
Fine-tuned (Scenario 3, LoRA r=32) 0.7487 (74.87%)

LoRA fine-tuning improved BERTScore-F1 by roughly +9.87 percentage points compared to the base model, showing that the fine-tuned model captures the semantic similarity of answers far better against the Indonesian-language property domain references.

BERTScore-F1 per Intent (Scenario 3)

Intent BERTScore-F1
biaya_dan_pajak (costs & taxes) 0.7568
spesifikasi_properti (property specifications) 0.7548
proses_transaksi (transaction process) 0.7534
legal_dokumen (legal documents) 0.7532
konsultasi_keputusan (decision consultation) 0.7527
risiko_dan_disclaimer (risk & disclaimer) 0.7496
kpr_simulasi (mortgage simulation) 0.7199

Note: The sizable gap between BERTScore-F1 (0.7487) and BLEU-4 (0.1454) (+0.60) indicates that the model produces answers with meaning/semantics that align with the reference even though the specific word choices (lexical level) differ — which is expected for a natural-language generative QA task.


📚 References

  • Hu, E. J., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
  • Biderman, D., et al. (2024). LoRA Learns Less and Forgets Less. Transactions on Machine Learning Research.
  • Zhou, C., et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023.
  • TinyLlama: https://github.com/jzhang38/TinyLlama

👤 Author

Rizal — Fachrizal Fazza Ashari


📄 License

Follows the license of the base model TinyLlama-1.1B-Chat-v1.0 (Apache 2.0).

Downloads last month
6
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for RizalGo/TinyLLaMA-Indo-Property