Instructions to use 4fhct4sd/GPT-Redeaux with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use 4fhct4sd/GPT-Redeaux with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="4fhct4sd/GPT-Redeaux")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("4fhct4sd/GPT-Redeaux") model = AutoModelForCausalLM.from_pretrained("4fhct4sd/GPT-Redeaux", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 4fhct4sd/GPT-Redeaux with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 4fhct4sd/GPT-Redeaux:F16 # Run inference directly in the terminal: llama cli -hf 4fhct4sd/GPT-Redeaux:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 4fhct4sd/GPT-Redeaux:F16 # Run inference directly in the terminal: llama cli -hf 4fhct4sd/GPT-Redeaux:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 4fhct4sd/GPT-Redeaux:F16 # Run inference directly in the terminal: ./llama-cli -hf 4fhct4sd/GPT-Redeaux:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 4fhct4sd/GPT-Redeaux:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf 4fhct4sd/GPT-Redeaux:F16
Use Docker
docker model run hf.co/4fhct4sd/GPT-Redeaux:F16
- LM Studio
- Jan
- vLLM
How to use 4fhct4sd/GPT-Redeaux with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "4fhct4sd/GPT-Redeaux" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "4fhct4sd/GPT-Redeaux", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/4fhct4sd/GPT-Redeaux:F16
- SGLang
How to use 4fhct4sd/GPT-Redeaux with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "4fhct4sd/GPT-Redeaux" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "4fhct4sd/GPT-Redeaux", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "4fhct4sd/GPT-Redeaux" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "4fhct4sd/GPT-Redeaux", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Ollama
How to use 4fhct4sd/GPT-Redeaux with Ollama:
ollama run hf.co/4fhct4sd/GPT-Redeaux:F16
- Unsloth Desktop
- Docker Model Runner
How to use 4fhct4sd/GPT-Redeaux with Docker Model Runner:
docker model run hf.co/4fhct4sd/GPT-Redeaux:F16
- Lemonade
How to use 4fhct4sd/GPT-Redeaux with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 4fhct4sd/GPT-Redeaux:F16
Run and chat with the model
lemonade run user.GPT-Redeaux-F16
List all available models
lemonade list
- Atomic Chat
GPT-Redeaux
A GPT-2 XL fine-tune exported from Unsloth Studio. This revision contains complete BF16 model weights and F16 / Q8_0 GGUF exports. No separate base-model download or adapter merge is required.
Versions
| Branch | Source run | Training steps | Saved (UTC) |
|---|---|---|---|
main |
1789180799 |
133 | 2026-09-12 02:42:58 |
run-1789180737-step-7 |
1789180737 |
7 | 2026-09-12 02:39:23 |
This branch is main, with 133 training steps and 7 epochs recorded in its checkpoint. These are separate training runs. The latest run is the default; no comparative quality evaluation was performed.
Architecture: GPT-2 XL (48 layers, 1,600 hidden dimensions, 25 attention heads), approximately 1.56 billion parameters, a 1,024-token context, and 50,258 vocabulary entries. The original GPT-2 tokenizer includes a training-added padding token.
Files
model.safetensors: original, unquantized BF16 tensors, copied byte-for-byte from the completed run.config.json,generation_config.json, and tokenizer files: standalone Transformers loading configuration. Local machine references were removed and the inference KV cache enabled.gguf/GPT-Redeaux.F16.gguf: full floating-point GGUF converted from those weights.gguf/GPT-Redeaux.Q8_0.gguf: 8-bit quantized GGUF.export_info.json: source-run identifier, timestamps, weight checksum, and generation smoke test.SHA256SUMS: file checksums.
Transformers usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "4fhct4sd/GPT-Redeaux"
revision = "main"
# Authenticate with Hugging Face first if the repository is private.
tokenizer = AutoTokenizer.from_pretrained(repo, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
repo, revision=revision, dtype=torch.bfloat16, device_map="auto"
)
inputs = tokenizer("The meaning of life is", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=64, do_sample=False,
pad_token_id=tokenizer.pad_token_id)
print(tokenizer.decode(output[0], skip_special_tokens=True))
This is a text-completion model; no chat template is supplied.
llama.cpp usage
Download either GGUF file from this branch, then run:
llama-completion -m GPT-Redeaux.Q8_0.gguf -p "The meaning of life is" -n 64 --no-conversation
Provenance and validation
The source run identifier is openai-community_gpt2-xl__project-----gpt2-cpt-ttext_1789125890__project----gpt2-cpt-ttext_1789130121__project---gpt2-cpt-ttext_1789139360__project--gpt2-cpt-rosetta-n-ttext_1789154027__project-gpt-ttext_rosetta-soulfft001_1789180799. This is a GPT-2 XL lineage with preceding local continued-training stages; the metadata above identifies the original architecture, not a claim that this run started directly from untouched upstream weights.
All 580 saved tensors were checked for finite values. The BF16 model was loaded and used for short greedy generation in Transformers. Both GGUF files were loaded and used for short generation in the installed llama.cpp build. These are export integrity checks, not benchmark results. Dataset composition and intended-use claims were not inferred from run names.
- Downloads last month
- 68
Model tree for 4fhct4sd/GPT-Redeaux
Base model
openai-community/gpt2-xl