AI & ML interests

LLM inference, model serving, low-latency and high-throughput inference, OpenAI-compatible APIs, GPU optimization, model routing, and open-source AI infrastructure.

Recent Activity

zhihuauk  updated a Space 2 days ago
VeloInfer/README
zhihuauk  published a Space 2 days ago
VeloInfer/README
View all activity

Organization Card

VeloInfer

VeloInfer is building fast, reliable, and cost-efficient inference infrastructure for open-source AI models.

What we focus on

  • Low-latency and high-throughput LLM inference
  • OpenAI-compatible APIs
  • Streaming chat completions
  • GPU inference optimization
  • Model routing and scalable serving
  • Simple and transparent usage-based pricing

Hugging Face integration

We are preparing VeloInfer for integration with Hugging Face Inference Providers.

Our public API endpoints, supported model catalog, pricing, benchmarks, and documentation will be published after staging validation is complete.

Links

Status

Infrastructure and provider integration are currently under development.

models 0

None public yet

datasets 0

None public yet