AI & ML interests
LLM inference, model serving, low-latency and high-throughput inference, OpenAI-compatible APIs, GPU optimization, model routing, and open-source AI infrastructure.
Recent Activity
Organization Card
VeloInfer
VeloInfer is building fast, reliable, and cost-efficient inference infrastructure for open-source AI models.
What we focus on
- Low-latency and high-throughput LLM inference
- OpenAI-compatible APIs
- Streaming chat completions
- GPU inference optimization
- Model routing and scalable serving
- Simple and transparent usage-based pricing
Hugging Face integration
We are preparing VeloInfer for integration with Hugging Face Inference Providers.
Our public API endpoints, supported model catalog, pricing, benchmarks, and documentation will be published after staging validation is complete.
Links
- GitHub: https://github.com/Zhihuauk
- Hugging Face: https://huggingface.co/VeloInfer
Status
Infrastructure and provider integration are currently under development.
models 0
None public yet
datasets 0
None public yet