Support / Roadmap for serving Clef-Flash via vLLM?

#3
by Vishva007 - opened

Hi @Cloudflare team,

Thanks for releasing Clef-Flash! The joint schema evaluation approach and single forward-pass probability output design are exceptional for high-throughput decision and classification pipelines.

Since standard deployment in production environments relies heavily on vLLM (for PagedAttention, continuous batching, and high concurrency), I wanted to check on vLLM compatibility:

  1. Custom Architecture Support: Since clef-flash relies on custom code (joint_schema_model.py / systemone) and evaluates typed questions via direct hidden-state readouts rather than token-by-token autoregressive generation, is there any current pathway or plan to register ClefForConditionalGeneration / its joint classification heads as an out-of-the-box model architecture in vLLM?
  2. Forward-Pass / Embedding API: In vLLM, classification and embedding models typically use the pooling/forward-pass endpoints rather than /v1/completions. Has the team explored integrating systemone schema scoring on top of vLLM’s forward/pooling engine?
  3. Current Serving Recommendations: Outside of running native Hugging Face Transformers + FastAPI/Triton, what is your recommended high-throughput serving stack for serving Clef-Flash at scale?

Appreciate any guidance or pointers from the team or community!

@Vishva007 If stock vLLM support is the immediate requirement, there is another implementation approach: score the next-token logits restricted to the decision's option keys, rather than serving a custom joint head. I built Seb-9B around that approach: https://huggingface.co/ironbcc/seb-9b

The card includes a vLLM chat API example with constrained choices and option logprobs. It handles one question per request, with up to 20 choice options; that differs from Clef's joint multi-question design. Released-checkpoint GPU throughput measurements are still pending, so I can't claim it solves your high-concurrency requirement yet. It may be useful as a stock-vLLM baseline while the Clef integration question is worked through.

Sign up or log in to comment