Instructions to use FurkanNar/GPT-2_Instruct-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FurkanNar/GPT-2_Instruct-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FurkanNar/GPT-2_Instruct-v0.1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FurkanNar/GPT-2_Instruct-v0.1") model = AutoModelForCausalLM.from_pretrained("FurkanNar/GPT-2_Instruct-v0.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FurkanNar/GPT-2_Instruct-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FurkanNar/GPT-2_Instruct-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FurkanNar/GPT-2_Instruct-v0.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/FurkanNar/GPT-2_Instruct-v0.1
- SGLang
How to use FurkanNar/GPT-2_Instruct-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FurkanNar/GPT-2_Instruct-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FurkanNar/GPT-2_Instruct-v0.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FurkanNar/GPT-2_Instruct-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FurkanNar/GPT-2_Instruct-v0.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use FurkanNar/GPT-2_Instruct-v0.1 with Docker Model Runner:
docker model run hf.co/FurkanNar/GPT-2_Instruct-v0.1
GPT-2 Instruct
This project uses a GPT-2 model (124M parameters) fine-tuned on the SVAMP (Simple Variants of Arithmetic Math word Problems) model, available at FurkanNar/gpt-2_svamp.
Training Hyperparameters
- Base Model: GPT-2 (124M parameters)
- Dataset: Alpaca
- Max Sequence Length: 256 tokens
- Epochs: 3
- Batch Size: 4
- Learning Rate: 5e-5
- Optimizer: AdamW
- Loss Function: CrossEntropyLoss (tokens shifted by 1)
- Gradient Clipping Max Norm: 1.0
Training Progress
Epoch 1/3
- Average Train Loss: 0.7349
- Average Validation Loss: 0.6526
- Validation F1 Score (Macro): 0.2824
Epoch 2/3
- Average Train Loss: 0.6583
- Average Validation Loss: 0.6436
- Validation F1 Score (Macro): 0.2839
Epoch 3/3
- Average Train Loss: 0.6216
- Average Validation Loss: 0.6411
- Validation F1 Score (Macro): 0.2857
Training Metrics Visualization
The model showed consistent improvement across epochs, with training loss decreasing from 0.7349 to 0.6216, indicating effective learning of the instruction-following task.
Model Files
config.json- Model configurationgeneration_config.json- Generation parametersmodel.safetensors- Fine-tuned model weights (475MB)
Usage
The application uses the official Alpaca instruction format for inference:
Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
{user_input}
### Response:
Proof of Concept
Example interaction with the fine-tuned model:
You: Hello
AI: Hi there! How can I help you today?
Configuration
The model generation can be configured with the following parameters:
model_name: Hugging Face model identifier or local pathsystem_prompt: System prompt for the assistantmax_length: Maximum response lengthtemperature: Sampling temperature (default: 0.5)top_k: Top-k sampling parameter (default: 40)top_p: Nucleus sampling parameter (default: 0.9)repetition_penalty: Penalty for repeating tokens (default: 1.2)
Requirements
- PyTorch
- Transformers
- CUDA-capable GPU
Notes
- The model uses conversation history (last 10 messages) to maintain context
- Generation parameters are tuned to reduce data leakage and improve response quality
- CUDA GPU is required for inference
- Downloads last month
- 246
