ProactiveInquirer-Qwen3-8B-Merged

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM   2Weizmann Institute of Science

Project page Paper Code License

The trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents, with its LoRA adapter merged into Qwen3-8B. It is a standard full-weight model: it loads without PEFT and serves with vLLM, SGLang or TGI like any Qwen3-8B.

This is training seed 1, the adapter at the root of the adapter repository. The merge ran in float32 and the weights are stored in bfloat16. On the adapter card's two-turn example, greedy decoding with this model returns the adapter's output character for character.

How to use it

The questioner reads the prompt template it was trained on, in prompts/, and replies with one JSON action per step: {"action": "ask", "question": ...} or {"action": "stop", ...}. Keep Qwen3's thinking off, as in training.

import re

import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "dolev31/ProactiveInquirer-Qwen3-8B-Merged"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")
template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()


def next_action(**state):
    fields = dict(state, user_channel=placebo.strip())
    prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
    ids = tok.apply_chat_template(
        [{"role": "user", "content": prompt}],
        add_generation_prompt=True,
        enable_thinking=False,
        return_tensors="pt",
        return_dict=True,
    ).to(model.device)
    out = model.generate(**ids, max_new_tokens=200, do_sample=False)
    return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)


print(next_action(
    question="Who was the spouse of the director of the film The Great Flamarion?",
    instructions="Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
    "queries against that pool before answering; several paragraphs are distractors, and the answer "
    "usually requires composing facts from more than one of them.",
    evidence="(nothing retrieved yet)", draft="(no draft yet)", history="(nothing asked yet)",
))
# {"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}

With vLLM, serve it and send the filled template as the user message, with thinking off:

vllm serve dolev31/ProactiveInquirer-Qwen3-8B-Merged
# request body: {"messages": [{"role": "user", "content": "<the filled template>"}],
#                "chat_template_kwargs": {"enable_thinking": false}, "temperature": 0}

Citation

@article{levy2026asking,
  title   = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
  author  = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
  journal = {arXiv preprint},
  year    = {2026}
}

License

Apache-2.0, like the base model Qwen3-8B.

Downloads last month
337
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dolev31/ProactiveInquirer-Qwen3-8B-Merged

Finetuned
Qwen/Qwen3-8B
Finetuned
(2128)
this model
Quantizations
1 model

Datasets used to train dolev31/ProactiveInquirer-Qwen3-8B-Merged