Instructions to use vngrs/Kumru-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vngrs/Kumru-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vngrs/Kumru-2B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("vngrs/Kumru-2B") model = AutoModelForCausalLM.from_pretrained("vngrs/Kumru-2B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vngrs/Kumru-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vngrs/Kumru-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vngrs/Kumru-2B
- SGLang
How to use vngrs/Kumru-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vngrs/Kumru-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vngrs/Kumru-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vngrs/Kumru-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vngrs/Kumru-2B with Docker Model Runner:
docker model run hf.co/vngrs/Kumru-2B
Tokenizer efficiency hakkindaki kafa karisikligim
Merhabalar,
Cok guzel bir calisma, ellerinize saglik! Tarihe gectiniz, gurur verici olmali :)
Soyle bir kafa karsikligim var. Train data icinIt is pre-trained on a cleaned, deduplicated corpora of 500 GB for 300B tokens, and supervised fine-tuned on 1M examples.
demissiniz burdan en iyi ihtimalle 500B char'lik bir veri uzerinde, 300B token ile calistiginizi anliyorum. yani token basina 1.66 char gibi, bu gpt tokenizerina gore cok cok dusuk bir rakam. Ama plotlariniz tam tersi yonde.
Acaba neyi kaciriyorum diye danismak istedim.
Merhaba. 500GB verisetleri Huggingface Dataset formatında diske yazıldığındaki toplam boyut. Dolayısıyla 500GB sayısının token sayısı ile doğrudan bir ilişkisi yok.
Bu veri, Kumru tokenizer ile tokenize edildiğinde de toplam 100B token ediyor. Dolayısıyla 300B token'lık eğitim, 3 epoch'a tekabül ediyor.