Instructions to use seedboxai/Llama-3-KafkaLM-8B-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use seedboxai/Llama-3-KafkaLM-8B-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="seedboxai/Llama-3-KafkaLM-8B-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("seedboxai/Llama-3-KafkaLM-8B-v0.1") model = AutoModelForCausalLM.from_pretrained("seedboxai/Llama-3-KafkaLM-8B-v0.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use seedboxai/Llama-3-KafkaLM-8B-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "seedboxai/Llama-3-KafkaLM-8B-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "seedboxai/Llama-3-KafkaLM-8B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/seedboxai/Llama-3-KafkaLM-8B-v0.1
- SGLang
How to use seedboxai/Llama-3-KafkaLM-8B-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "seedboxai/Llama-3-KafkaLM-8B-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "seedboxai/Llama-3-KafkaLM-8B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "seedboxai/Llama-3-KafkaLM-8B-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "seedboxai/Llama-3-KafkaLM-8B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use seedboxai/Llama-3-KafkaLM-8B-v0.1 with Docker Model Runner:
docker model run hf.co/seedboxai/Llama-3-KafkaLM-8B-v0.1
Llama-3-KafkaLM-8B-v0.1
KafkaLM 8b is a Llama3 8b model which was finetuned on an ensemble of popular high-quality open-source instruction sets (translated from English to German).
Llama 3 KafkaLM 8b is a Seedbox project trained by Dennis Dickmann.
Why Kafka? The models are proficient, yet creative, and have some tendencies to linguistically push boundaries π
Model Details
The purpose of releasing the KafkaLM series is to contribute to the German AI community with a set of fine-tuned LLMs that are easy to use in everyday applications across a variety of tasks.
The main goal is to provide LLMs proficient in German, especially to be used in German-speaking business contexts where English alone is not sufficient.
Dataset
I used a 8k filtered version of the following seedboxai/multitask_german_examples_32k
Inference
Getting started with the model is straightforward
import transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "seedboxai/Llama-3-KafkaLM-8B-v0.1"
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
tokenizer = transformers.AutoTokenizer.from_pretrained(model_id)
pipeline = transformers.pipeline(
model=model, tokenizer=tokenizer,
return_full_text=True,
task='text-generation',
device="cuda",
)
messages = [
{"role": "system", "content": "Du bist ein hilfreicher KI-Assistent."},
{"role": "user", "content": "Wer ist eigentlich dieser Kafka?"},
]
prompt = pipeline.tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
terminators = [
pipeline.tokenizer.eos_token_id,
pipeline.tokenizer.convert_tokens_to_ids("<|eot_id|>")
]
outputs = pipeline(
prompt,
max_new_tokens=max_new_tokens,
num_beams=5,
num_return_sequences=1,
early_stopping=True,
eos_token_id=terminators,
do_sample=True,
temperature=temperature,
top_p=0.9,
)
print(outputs[0]["generated_text"][len(prompt):])
<|start_header_id|>system<|end_header_id|>Du bist ein hilfreicher KI-Assistent.<|eot_id|>
<|start_header_id|>user<|end_header_id|>Wer ist eigentlich dieser Kafka?<|eot_id|>
<|start_header_id|>assistant<|end_header_id|>Franz Kafka (1883-1924) war ein tschechisch-ΓΆsterreichischer Schriftsteller,
der als einer der bedeutendsten Autoren des 20. Jahrhunderts gilt. Seine Werke sind geprΓ€gt von existenzieller Angst, Isolation, Verwirrung und AbsurditΓ€t.
Kafka wurde in Prag, Γsterreich-Ungarn (heute Tschechische Republik), geboren und wuchs in einer jΓΌdischen Familie auf.
Er studierte Rechtswissenschaften und arbeitete als Versicherungsangestellter, bevor er sich dem Schreiben zuwandte.
Kafkas literarisches Werk umfasst Romane, ErzΓ€hlungen, Briefe und TagebΓΌcher. Zu seinen bekanntesten Werken gehΓΆren "Der Prozess" (1925),
"Das Schloss" (1926) und "Die Verwandlung" (1915). Seine Geschichten sind geprΓ€gt von surrealen, traumΓ€hnlichen Szenarien,
in denen die Charaktere oft in Situationen geraten, die sie nicht verstehen oder kontrollieren kΓΆnnen.
Kafkas Werk hat einen tiefgreifenden Einfluss auf die Literatur und Kultur des 20. Jahrhunderts gehabt.
Seine Werke wurden in viele Sprachen ΓΌbersetzt und inspirierten Autoren wie Albert Camus, Jean-Paul Sartre, Samuel Beckett und Thomas Mann.
Kafka starb 1924 im Alter von 40 Jahren an Tuberkulose. Trotz seines relativ kurzen Lebens hat er einen bleibenden Eindruck auf die Literatur und Kultur hinterlassen.
Disclaimer
The license on this model does not constitute legal advice. We are not responsible for the actions of third parties who use this model. This model should only be used for research purposes. The original Llama3 license and all restrictions of datasets used to train this model apply.
- Downloads last month
- 25
