Instructions to use IntegralPilot/gpt-oss-20b-bfloat16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IntegralPilot/gpt-oss-20b-bfloat16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IntegralPilot/gpt-oss-20b-bfloat16")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("IntegralPilot/gpt-oss-20b-bfloat16") model = AutoModelForCausalLM.from_pretrained("IntegralPilot/gpt-oss-20b-bfloat16", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IntegralPilot/gpt-oss-20b-bfloat16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IntegralPilot/gpt-oss-20b-bfloat16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IntegralPilot/gpt-oss-20b-bfloat16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/IntegralPilot/gpt-oss-20b-bfloat16
- SGLang
How to use IntegralPilot/gpt-oss-20b-bfloat16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IntegralPilot/gpt-oss-20b-bfloat16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IntegralPilot/gpt-oss-20b-bfloat16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IntegralPilot/gpt-oss-20b-bfloat16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IntegralPilot/gpt-oss-20b-bfloat16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use IntegralPilot/gpt-oss-20b-bfloat16 with Docker Model Runner:
docker model run hf.co/IntegralPilot/gpt-oss-20b-bfloat16
Model Card for GPT-OSS-20b-bfloat16
This is OpenAI's gpt-oss-20b, but repackaged in Bfloat16 format, as the MXFP4 format of the original is only supported on Hopper and later architectures.
It contains no differences or further training compared to OpenAI's model, and you should read their model card for the details of the model.
It was primarily created by me to reduce the need to for repeated manual upcasting from MXFP4 to Bfloat16, which I was doing for running on older GPUs (such as on Kaggle or Colab) and for my fine-tuning experiments. You are, of course, welcome to use it too if it helps!
- Downloads last month
- 10