YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Kronos-Base Inference Endpoint

HuggingFace Inference Endpoint handler for Kronos-Base time-series model on GPU.

Files

  • handler.py β€” Custom endpoint handler (loads model, serves predictions)
  • Dockerfile β€” Custom container image (optional, for custom deployments)
  • model/ β€” Symlink or copy of the Kronos model code from https://github.com/shiyu-coder/Kronos

Deploy to HuggingFace Inference Endpoint

Option A: Custom Handler (recommended for T4 GPU)

  1. Create a repo on HuggingFace (e.g. your-username/kronos-gold-endpoint)
  2. Copy this directory to the repo root, including the model/ folder from Kronos:
    cp -r /path/to/Kronos/model ./
    git add . && git commit -m "Add Kronos handler" && git push
    
  3. Create an Inference Endpoint at https://ui.endpoints.huggingface.co/
    • Repository: your-username/kronos-gold-endpoint
    • Type: Custom β†’ "Custom image" or select the repo directly
    • GPU: NVIDIA T4 (16GB) β€” cheapest GPU tier (~$0.06/hr)
    • Region: eu-west-1 or us-east-1 (whichever is closer to your VPS)
    • Min instances: 1 (keeps warm, no cold starts) β€” OR 0 (scale to zero, ~20s cold start, much cheaper)
    • Max instances: 1
    • Environment variables:
      • KRONOS_MODEL_ID=NeoQuasar/Kronos-base
      • KRONOS_TOKENIZER_ID=NeoQuasar/Kronos-Tokenizer-base
      • KRONOS_DEVICE=cuda
  4. Wait for endpoint to build + deploy (~5 min first time)
  5. Copy the endpoint URL β€” it'll be something like: https://xxxxx.aws.endpoints.huggingface.cloud/

Option B: Use the Pre-built Model Directly

If HF adds serverless support for Kronos, you can deploy the model directly:

  • Repository: NeoQuasar/Kronos-base
  • Task: Custom

Configure Gold Bot

Set these in your .env file:

KRONOS_ENDPOINT_URL=https://xxxxx.aws.endpoints.huggingface.cloud/
KRONOS_HF_TOKEN=hf_xxxxxxxxxxxx
KRONOS_DEVICE=cpu  # ignored when endpoint is set

If both KRONOS_ENDPOINT_URL and KRONOS_HF_TOKEN are set, the bot will use the remote GPU endpoint. Otherwise it falls back to local CPU (slow).

Expected Performance on T4

Kronos-base (102M params) on T4:

  • 1 sample: ~1-2 seconds
  • 8 samples: ~3-6 seconds
  • Model loading: ~3 seconds (one-time)

This is well within the 15-minute bar timeframe.

API Contract

Request:

{
  "inputs": {
    "ohlcv": [[open, high, low, close, volume], ...],
    "x_timestamps": ["2025-01-01T00:00:00", ...],
    "y_timestamps": ["2025-01-01T16:00:00", ...],
    "pred_len": 8,
    "sample_count": 8,
    "temperature": 1.0,
    "top_p": 0.9
  }
}

Response:

{
  "predictions": [
    {"open": ..., "high": ..., "low": ..., "close": ..., "volume": ..., "amount": ...},
    ...
  ],
  "timestamps": ["2025-01-01T16:00:00", ...],
  "inference_seconds": 3.45
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support