Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string

Model Card for teuken-fact20

teuken-fact20 is a finetune of openGPT-X/Teuken-7B-instruct-research-v0.4. Using the finetune it is possible to get much more correct answers on the question:

  • Which districts does a city have? Where city may be e.g. Berlin or Essen.

Only the districts of the 20 biggest cities in germany are considered. Language is german.

Model Details

Model Description

Teuken 7B (like other LLMs to) has problems with simple facts. E.g. if you ask it for the districts of the german city Dresden you will get varying results.

Using the dataset andy300/cities_bezirk_mp i finetuned openGPT-X/Teuken-7B-instruct-research-v0.4. The result is the LORA Adapter andy300/lora-teuken-fact20. Merging the LORA Adapter back into the base model: openGPT-X/Teuken-7B-instruct-research-v0.4 results in andy300/teuken-fact20

teuken-fact20 is able to correctly identify 229 out of 235 actual districts of the 20 biggest cities in Germany, and additionally generates 6 incorrect (fantasy) districts. Pure Teuken 7B on the other hand gives 97 correct districts.

  • Developed by: Andreas Wenzel
  • Model type: Transformer based decoder-only model
  • Language(s) (NLP): German
  • License: See License of openGPT-X/Teuken-7B-instruct-research-v0.4
  • Finetuned from model: openGPT-X/Teuken-7B-instruct-research-v0.4

Model Sources

Uses

teuken-fact20 has the same usage as the base model Teuken 7B. Additionally it improves Teuken 7B's ability to answer questions such as: "Which districts does the city of Berlin have?" Only the 20 largest cities in Germany are considered.

Bias, Risks, and Limitations

teuken-fact20 has the same Bias, Risks and Limitations like Teuken 7B.

How to Get Started with the Model

The model can be used the same way like Teuken 7B. Only change the model name from:

openGPT-X/Teuken-7B-instruct-research-v0.4 

to:

andy300/teuken-fact20 

Transformer

The model requires the same Python libraries as Teuken 7B.

python -m pip install numpy torch huggingface_hub transformers sentencepiece

Example usage of the model is translated from the corresponding Teuken 7B example:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model_name = "andy300/teuken-fact20"
tokenizer_name = "openGPT-X/Teuken-7B-instruct-research-v0.4"
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16
)
#model.load_adapter("adapter")
model = model.to(device).eval()
tokenizer = AutoTokenizer.from_pretrained(
    tokenizer_name,
    use_fast=False,
    trust_remote_code=True,
)


messages = [{"role": "User", "content": "Kannst du mir eine Liste aller Stadtbezirke in München geben?"}]
prompt_ids = tokenizer.apply_chat_template(messages, chat_template="DE", tokenize=True, add_generation_prompt=True, return_tensors="pt")

prediction = model.generate(
    prompt_ids.to(model.device),
    max_length=512,
    do_sample=True,
    top_k=50,
    top_p=0.95,
    temperature=0.7,
    num_return_sequences=1,
)
prediction_text = tokenizer.decode(prediction[0].tolist())
print(prediction_text)

Training Details

Training Data

Training is done using the dataset cities_bezirk_mp.

Training Procedure

The finetuned model is created merging the LORA Adapter lora-teuken-fact20 into base model Teuken-7B-instruct-research-v0.4.

import torch
from peft import AutoPeftModelForCausalLM

# Load PEFT model on CPU
model = AutoPeftModelForCausalLM.from_pretrained(
    pretrained_model_name_or_path="andy300/lora-teuken-fact20",
    torch_dtype=torch.float16,
    low_cpu_mem_usage=True,
)

# Merge LoRA and base model and save
merged_model = model.merge_and_unload()
merged_model.save_pretrained(
    'teuken-fact20', safe_serialization=True, max_shard_size="2GB"
)

Details of LORA adapter training can be found here lora-teuken-fact20.

Evaluation

Testing Data, Factors & Metrics

Testing Data

Dataset used for testing is: andy300/cities_bezirke.

Testing is done using the question:

question= "Bitte nenne mir alle Stadtbezirke der Stadt: " + city + ":"

Factors

Evaluation is using the same 20 largest cities like training. Only the question is different.

Metrics

The 20 largest cities in germany have 235 different areas. Asking for the areas of these cities me measure the total number of the correct areas for the given cities. Also the total number of areas given by teuken-fact20 is measured.

Results

Out of the 235 districts, teuken-fact20 can correctly name 231 — which corresponds to 98% correct answers and represents a clear improvement over the 50% of the original Teuken 7B model. The results even surpass those of GPT-5!

At the same time, teuken-fact20 behaves almost identically to the original Teuken 7B. My estimated perplexity for teuken-fact20 is 13, close to the original model’s value of 15. Interestingly, the fine‑tuned model’s perplexity is slightly lower than the original’s — so teuken‑fact20 has, in theory, been slightly improved by the fine‑tuning.

See my linkedin article for details.

Summary

Finetuning Teuken 7B for better at answering questions about 20 facts makes the fintuned model teuken-fact20 much better in this special field without changing the finetuned model too much.

Model Architecture and Objective

Model Architecture is exactly the same as Teuken 7B. The model adds 20 additional facts to the original version. This finetuning is an experimental approach to improve Teuken 7B for answering district-related questions about Germany's largest cities.

Compute Infrastructure

Finetune is done using colab.

Hardware

Tesla T4 used with colab

Software

colab with huggingface transformers.

Used colab notebooks and python scripts can be found here.

Model Card Contact

awenzel@net-haus.com

Downloads last month
17
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andy300/teuken-fact20

Dataset used to train andy300/teuken-fact20