Instructions to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dolev31/ProactiveInquirer-Qwen3-8B-Merged") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dolev31/ProactiveInquirer-Qwen3-8B-Merged") model = AutoModelForCausalLM.from_pretrained("dolev31/ProactiveInquirer-Qwen3-8B-Merged", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dolev31/ProactiveInquirer-Qwen3-8B-Merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-Merged
- SGLang
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dolev31/ProactiveInquirer-Qwen3-8B-Merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dolev31/ProactiveInquirer-Qwen3-8B-Merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with Docker Model Runner:
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-Merged
ProactiveInquirer-Qwen3-8B-Merged
Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM 2Weizmann Institute of Science
The trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents, with its LoRA adapter merged into Qwen3-8B. It is a standard full-weight model: it loads without PEFT and serves with vLLM, SGLang or TGI like any Qwen3-8B.
- The adapter, with the results, the training details, the limitations and a complete two-turn example: dolev31/ProactiveInquirer-Qwen3-8B.
- Quantized for llama.cpp, Ollama and LM Studio: dolev31/ProactiveInquirer-Qwen3-8B-GGUF.
This is training seed 1, the adapter at the root of the adapter repository. The merge ran in float32 and the weights are stored in bfloat16. On the adapter card's two-turn example, greedy decoding with this model returns the adapter's output character for character.
How to use it
The questioner reads the prompt template it was trained on, in prompts/, and replies
with one JSON action per step: {"action": "ask", "question": ...} or {"action": "stop", ...}.
Keep Qwen3's thinking off, as in training.
import re
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "dolev31/ProactiveInquirer-Qwen3-8B-Merged"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")
template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()
def next_action(**state):
fields = dict(state, user_channel=placebo.strip())
prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
ids = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
out = model.generate(**ids, max_new_tokens=200, do_sample=False)
return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)
print(next_action(
question="Who was the spouse of the director of the film The Great Flamarion?",
instructions="Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
"queries against that pool before answering; several paragraphs are distractors, and the answer "
"usually requires composing facts from more than one of them.",
evidence="(nothing retrieved yet)", draft="(no draft yet)", history="(nothing asked yet)",
))
# {"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
With vLLM, serve it and send the filled template as the user message, with thinking off:
vllm serve dolev31/ProactiveInquirer-Qwen3-8B-Merged
# request body: {"messages": [{"role": "user", "content": "<the filled template>"}],
# "chat_template_kwargs": {"enable_thinking": false}, "temperature": 0}
Citation
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint},
year = {2026}
}
License
Apache-2.0, like the base model Qwen3-8B.
- Downloads last month
- 337