cu-llm Inference Endpoint

Custom handler for serving the credit-union fine-tuned Llama-3.1-8B (cu-llm-merged) on HuggingFace Inference Endpoints.

Files

File Purpose
handler.py EndpointHandler โ€” builds the training prompt, runs generation, returns the answer
requirements.txt Runtime deps installed by the Endpoint builder

Deploy

  1. Push the merged model to a HF Hub repo (from Colab, reading cu-llm-merged off Drive):

    from huggingface_hub import login
    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    login(token=HF_TOKEN)
    p = "/content/drive/MyDrive/cu-llm-merged"
    AutoModelForCausalLM.from_pretrained(p).push_to_hub("rdmurugan007/cu-llm", private=True)
    AutoTokenizer.from_pretrained(p).push_to_hub("rdmurugan007/cu-llm", private=True)
    
  2. Add the handler to the same repo. Upload handler.py and requirements.txt to the root of rdmurugan007/cu-llm:

    from huggingface_hub import upload_file
    upload_file(path_or_fileobj="handler.py", path_in_repo="handler.py", repo_id="rdmurugan007/cu-llm")
    upload_file(path_or_fileobj="requirements.txt", path_in_repo="requirements.txt", repo_id="rdmurugan007/cu-llm")
    
  3. Create the Endpoint. huggingface.co/new-endpoint โ†’ select rdmurugan007/cu-llm.

    • Task: Custom (HF auto-detects handler.py)
    • GPU: Nvidia T4 (smallest; 4-bit load fits) or A10G for fp16 (set LOAD_IN_4BIT=False in handler.py)
    • Set to scale-to-zero if you want to avoid idle cost.

Call it

curl https://<your-endpoint>.endpoints.huggingface.cloud \
  -H "Authorization: Bearer $HF_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": "What is a CUSO and how do credit unions use them?",
    "parameters": {"max_new_tokens": 300, "temperature": 0.0}
  }'

Response:

{"answer": "A CUSO (Credit Union Service Organization) is ...", "raw": "..."}

Wiring into fin360ai

The fin360ai router calls this Endpoint first for CU-domain questions and only escalates to Claude Sonnet on low confidence. From the Next.js /api/cu-chat route:

const res = await fetch(process.env.CU_LLM_ENDPOINT_URL!, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.HF_TOKEN}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ inputs: question, parameters: { max_new_tokens: 300 } }),
});
const { answer } = await res.json();
Downloads last month
2
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support