Qwen3-Reranker-0.6B β€” ExecuTorch

A reranker that is not a cross-encoder. The other five on this shelf are BERTs with a regression head; this one is a causal language model asked a yes/no question, and the score is how much more it would answer "yes" than "no". Same job β€” read a query and one document together and score the pair β€” with a different machine underneath.

input_ids      (1, 512) int64   the prompt, LEFT-padded
attention_mask (1, 512) int64
    -> score   (1, 1)   fp32    log-odds: logit("yes") - logit("no")
  • Source: Qwen/Qwen3-Reranker-0.6B β€” 595.8M parameters, 28 Qwen3 layers, hidden 1024, 151,669 vocabulary
  • License: apache-2.0
  • Output: one number per pair. Higher is more relevant; sigmoid(x) maps it to 0..1 and does not change the ordering.

Variants

build file size (MB) worst score error vs eager Mac ms* backend takes
fp32 rerank_qwen3_0_6b_xnnpack_fp32.pte 2383.7 0.0001 logits 320.5 80.8%
fp16 rerank_qwen3_0_6b_xnnpack_fp16.pte 1192.4 0.0268 logits 1188.8 70.4%
Core ML (fp16, iOS) rerank_qwen3_0_6b_coreml_all.pte 1196.2 0.0886 logits 83.4 100%

*Mac arm64, one 512-token pair, fastest of five medians of ten β€” a reference point for relative cost, not a device number. The host shares its cores with other work and a single median does not survive that; contention only ever adds time, so the fastest repetition is the one that means something. Torch eager fp32, measured the same way, is 320.1 ms. Core ML is the one to use where it exists: 3.8x faster than XNNPACK fp32, the whole graph in one delegated subgraph where XNNPACK takes 80.8% across 172. A reranker earns its keep over a list of fifty candidates, so that factor is the whole story.

Correlation cannot judge this model β€” the output is a single number, and the correlation of a one-element vector is undefined. The gate is the score error in the model's own units together with whether the ranking survives. Over six real pairs, every shipped build reproduces eager's order, and the narrowest adjacent gap in that ranking is 0.7433 logits.

Running it

1. Build the prompt. The model was trained to read one specific frame, and it is not a chat model you can prompt freely:

PREFIX = ('<|im_start|>system\nJudge whether the Document meets the requirements '
          'based on the Query and the Instruct provided. Note that the answer can '
          'only be "yes" or "no".<|im_end|>\n<|im_start|>user\n')
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCT = "Given a web search query, retrieve relevant passages that answer the query"

text = (f"{PREFIX}<Instruct>: {INSTRUCT}\n<Query>: {query}\n"
        f"<Document>: {document}{SUFFIX}")

INSTRUCT is a real input, not decoration: it names the retrieval task, and Qwen reports it is worth a few points to write your own instead of the generic default.

2. Tokenise with LEFT padding to exactly 512.

tokenizer = AutoTokenizer.from_pretrained(repo, padding_side="left")
batch = tokenizer([text], padding="max_length", truncation=True, max_length=512,
                  return_tensors="pt")
assert batch["attention_mask"][0, -1] == 1

This is the one that will bite. The graph reads position βˆ’1 unconditionally, which is the last real token only when the padding is on the left. Right-padded input scores a pad token and returns a confident-looking number that means nothing β€” nothing raises. The assertion above is one line and catches it.

3. Run it once per candidate and sort by the score, descending.

The output projection is folded into one vector

The score needs two of the 151,669 logits, and the difference of two dot products is one dot product against the difference of two rows:

logit_yes - logit_no = h . W[yes] - h . W[no] = h . (W[yes] - W[no])

So lm_head never runs. A 151,669-wide matmul per pair does not happen, and the delegate does not take its own copy of a 621 MB table β€” which is exactly why this shelf's Whisper decoders come out bigger than their weights.

The two vocabulary ids are read from the repo's own 1_LogitScore/config.json (9693 for "yes", 2152 for "no") rather than looked up through the tokenizer, where a leading-space variant is a different token.

Checked against the reference, at matching precision. sentence-transformers' CrossEncoder scores these pairs at +6.6875 / βˆ’6.0625 / βˆ’7.2812 and the folded head at +6.7271 / βˆ’6.1107 / βˆ’7.3159 β€” a gap of 4.8e-02 that has nothing to do with the fold. The reference loads the checkpoint in its own bfloat16; run both arms in fp32 and the folded head matches it to 4.673e-05, while the reference's bf16 and fp32 arms differ from each other by the full 4.820e-02.

It ranks across languages

The six test pairs include a Japanese passage that says what the query asks. It comes out first of six at +7.470, above the English document carrying the actual number (+6.727). The shelf's ms-marco-MiniLM rerankers are English-only and put it far lower β€” which is the right answer for them, not a defect.

The attention is eager, and the left padding is why that needed checking

F.scaled_dot_product_attention does not survive export as one operation. The edge dialect lowers it through _safe_softmax, whose guard against a row with no unmasked key at all leaves eleven operations XNNPACK cannot take, in every attention block β€” scalar_tensor, where, mul.Scalar, logical_not, eq, full_like, any.dim. Over 28 layers that is 308 operations, each one cutting the subgraph in two. Exporting with attn_implementation="eager" removes them: 72.0% to 80.8% delegated, fp32 from 381.5 to 320.5 ms.

This is the one model on the shelf where that guard could plausibly have been doing work. Its prompt is left-padded, so the leading rows of a causal mask have no unmasked key at all β€” exactly the case the guard exists for. Every other model here is right-padded, where even an all-padding row still sees the real tokens and the question never arises.

So it was measured rather than assumed. With 109 real tokens of 512, sdpa against eager: no NaN appears on either arm, and the scored last position reads 54.7696 against 54.7695. The reason is that the guard fires only on -inf, and eager masks with torch.finfo(dtype).min, a large finite number β€” a fully-masked row comes back uniform instead of zeroed. Both answers are defined; they differ only about padding rows, and this graph reads position βˆ’1, which is a real token by construction.

Not shipped: int8

rerank_qwen3_0_6b_xnnpack_int8.pte is 1063.9 MB and runs at 265.4 ms β€” smaller and faster than the fp32 build, and 85.3% of it delegates, the most of any build here. It is withheld on the number that decides.

Dynamic int8 moves a pair score by 0.6336 logits, against a narrowest adjacent gap of 0.7433 in the ranking it has to preserve. It happens to keep the order on these six pairs, but an error that is 85% of the distance between two neighbours is not a build that ranks reliably β€” a slightly different candidate list would reorder.

The size was predictable before it was built: the token embedding table is 151,669 x 1024 x 4 = 621 MB of the 2383.7 MB model, a 26.1% share, and this shelf's rule int8/fp16 = 0.5 + 1.5 x (table share) puts the file at 0.891 of fp16. It came out at 0.892 β€” the same figure as Qwen3-Embedding-0.6B, which shares this backbone.

Going eager (below) made this build 13% faster and took it from 78.1% to 85.3% delegated. It did not move the verdict: speed and quality are separate questions, and what withholds this file is the score error, which barely changed.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlboydaisuke/Qwen3-Reranker-0.6B-ExecuTorch

Quantized
(76)
this model