YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Kronos-Base Inference Endpoint
HuggingFace Inference Endpoint handler for Kronos-Base time-series model on GPU.
Files
handler.pyβ Custom endpoint handler (loads model, serves predictions)Dockerfileβ Custom container image (optional, for custom deployments)model/β Symlink or copy of the Kronos model code from https://github.com/shiyu-coder/Kronos
Deploy to HuggingFace Inference Endpoint
Option A: Custom Handler (recommended for T4 GPU)
- Create a repo on HuggingFace (e.g.
your-username/kronos-gold-endpoint) - Copy this directory to the repo root, including the
model/folder from Kronos:cp -r /path/to/Kronos/model ./ git add . && git commit -m "Add Kronos handler" && git push - Create an Inference Endpoint at https://ui.endpoints.huggingface.co/
- Repository: your-username/kronos-gold-endpoint
- Type: Custom β "Custom image" or select the repo directly
- GPU: NVIDIA T4 (16GB) β cheapest GPU tier (~$0.06/hr)
- Region: eu-west-1 or us-east-1 (whichever is closer to your VPS)
- Min instances: 1 (keeps warm, no cold starts) β OR 0 (scale to zero, ~20s cold start, much cheaper)
- Max instances: 1
- Environment variables:
KRONOS_MODEL_ID=NeoQuasar/Kronos-baseKRONOS_TOKENIZER_ID=NeoQuasar/Kronos-Tokenizer-baseKRONOS_DEVICE=cuda
- Wait for endpoint to build + deploy (~5 min first time)
- Copy the endpoint URL β it'll be something like:
https://xxxxx.aws.endpoints.huggingface.cloud/
Option B: Use the Pre-built Model Directly
If HF adds serverless support for Kronos, you can deploy the model directly:
- Repository:
NeoQuasar/Kronos-base - Task: Custom
Configure Gold Bot
Set these in your .env file:
KRONOS_ENDPOINT_URL=https://xxxxx.aws.endpoints.huggingface.cloud/
KRONOS_HF_TOKEN=hf_xxxxxxxxxxxx
KRONOS_DEVICE=cpu # ignored when endpoint is set
If both KRONOS_ENDPOINT_URL and KRONOS_HF_TOKEN are set, the bot will
use the remote GPU endpoint. Otherwise it falls back to local CPU (slow).
Expected Performance on T4
Kronos-base (102M params) on T4:
- 1 sample: ~1-2 seconds
- 8 samples: ~3-6 seconds
- Model loading: ~3 seconds (one-time)
This is well within the 15-minute bar timeframe.
API Contract
Request:
{
"inputs": {
"ohlcv": [[open, high, low, close, volume], ...],
"x_timestamps": ["2025-01-01T00:00:00", ...],
"y_timestamps": ["2025-01-01T16:00:00", ...],
"pred_len": 8,
"sample_count": 8,
"temperature": 1.0,
"top_p": 0.9
}
}
Response:
{
"predictions": [
{"open": ..., "high": ..., "low": ..., "close": ..., "volume": ..., "amount": ...},
...
],
"timestamps": ["2025-01-01T16:00:00", ...],
"inference_seconds": 3.45
}
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support