cu-llm Inference Endpoint
Custom handler for serving the credit-union fine-tuned Llama-3.1-8B
(cu-llm-merged) on HuggingFace Inference Endpoints.
Files
| File | Purpose |
|---|---|
handler.py |
EndpointHandler โ builds the training prompt, runs generation, returns the answer |
requirements.txt |
Runtime deps installed by the Endpoint builder |
Deploy
Push the merged model to a HF Hub repo (from Colab, reading
cu-llm-mergedoff Drive):from huggingface_hub import login from transformers import AutoModelForCausalLM, AutoTokenizer login(token=HF_TOKEN) p = "/content/drive/MyDrive/cu-llm-merged" AutoModelForCausalLM.from_pretrained(p).push_to_hub("rdmurugan007/cu-llm", private=True) AutoTokenizer.from_pretrained(p).push_to_hub("rdmurugan007/cu-llm", private=True)Add the handler to the same repo. Upload
handler.pyandrequirements.txtto the root ofrdmurugan007/cu-llm:from huggingface_hub import upload_file upload_file(path_or_fileobj="handler.py", path_in_repo="handler.py", repo_id="rdmurugan007/cu-llm") upload_file(path_or_fileobj="requirements.txt", path_in_repo="requirements.txt", repo_id="rdmurugan007/cu-llm")Create the Endpoint. huggingface.co/new-endpoint โ select
rdmurugan007/cu-llm.- Task: Custom (HF auto-detects
handler.py) - GPU: Nvidia T4 (smallest; 4-bit load fits) or A10G for fp16 (set
LOAD_IN_4BIT=Falseinhandler.py) - Set to scale-to-zero if you want to avoid idle cost.
- Task: Custom (HF auto-detects
Call it
curl https://<your-endpoint>.endpoints.huggingface.cloud \
-H "Authorization: Bearer $HF_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"inputs": "What is a CUSO and how do credit unions use them?",
"parameters": {"max_new_tokens": 300, "temperature": 0.0}
}'
Response:
{"answer": "A CUSO (Credit Union Service Organization) is ...", "raw": "..."}
Wiring into fin360ai
The fin360ai router calls this Endpoint first for CU-domain questions and only
escalates to Claude Sonnet on low confidence. From the Next.js /api/cu-chat
route:
const res = await fetch(process.env.CU_LLM_ENDPOINT_URL!, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.HF_TOKEN}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ inputs: question, parameters: { max_new_tokens: 300 } }),
});
const { answer } = await res.json();
- Downloads last month
- 2
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support