Instructions to use vngrs/Kumru-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vngrs/Kumru-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vngrs/Kumru-2B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("vngrs/Kumru-2B") model = AutoModelForCausalLM.from_pretrained("vngrs/Kumru-2B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vngrs/Kumru-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vngrs/Kumru-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vngrs/Kumru-2B
- SGLang
How to use vngrs/Kumru-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vngrs/Kumru-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vngrs/Kumru-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vngrs/Kumru-2B with Docker Model Runner:
docker model run hf.co/vngrs/Kumru-2B
Tokenizer detayları
Selamlar,
Tebrikelr öncelikle. Biz firma olarak Türkçe TTS eğitimi üstüne calisisiyoruz. Hali hazırda eğitilmiş modellerimiz var. Acaba LLM tabanlı TTS tarafında Kumrunun tokenizeri kullanimi nasıl olur diye dusunduk?
Acaba Tokenizer ile ilgili daha fazla bilgi paylasmaniz mümkün mudur? Blog yazinizi okudum hali hazırda. Token azalmalari güzel duruyor. Sondan eklemeli dil olmasının bunun uzerinde etkisi vardır diye düşünüyorum. Ayni fikirde misiniz acaba?
TTS tarafında ihtiyacınızı bilmiyorum ama Kumru'nun tokenizer'ı kod, matematik ve web dahil olabildiğince büyük bir veride eğitilmiş modern LLM ihtiyaçlarına cevap veren bir tokenizer. O yüzden sizin işinizi de görür herhalde.
Türkçe'nin morfolojisi -özellikle sondan eklemeli oluşu- ve tokenizerlarla ilgili birkaç çalışma ve msc tezi vardı ama benim bildiğim net bir sonuç yok.
Tokenizer meselesi bazı requirement'ları (pretokenization regex, special tokens, math ve code desteği, chat template) sağladıktan sonra metni en verimli şekilde (fertility) represent etme ile ilgili. Bunun yolu da vocabulary'i veriye bakarak istatistiki olarak oluşturmak.
Bunları yaptıktan sonra token'ların birbirleri ile ilişkisini model kendi içinde öğreniyor ve diğer detayların önemi kalmıyor.
"Metinleri Türkçe yapım ve çekim eklerine ayırarak tokenize etmek" gibi romantik fikirler var. Ama günün sonunda model performansına ne kadar etki eder, buradan elde edilecek kazanım için metni tokenize ederken çalıştırılacak morphological analyzer gibi araçlara ne kadar ekstra compute harcanır, tüm bunlara değer mi? Çok şüpheliyim.
@meliksahturker bahsettiğiniz fikri biz gerçekleştirdik, morfolojiye duyarlı bir tokenizer fikrini, dediğiniz gibi zemberek'e ekstra computing ayırmak ilk başta mantıksız görünse de tamamen Pythonda yazılmış kütüphaneler ile fark yaratmadan gerçekleştirilebiliyor. Şu anda TR MMLU Benchmarkında 1. sıradayız, embedding modelleri üzerindeki eğitim ve testlerini de gerçekleştiriyoruz.
Sizin de görüşlerinizi ve eleştirilerinizi almak bizi mutlu eder, huggingface.co/Ethosoft/NedoTurkishTokenizer adresinden ulaşabilirsiniz ( şu anda JVM kullanıyor ama bayramın 1. günü güncelleyeceğiz )