Instructions to use bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct") model = AutoModelForCausalLM.from_pretrained("bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct
- SGLang
How to use bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct with Docker Model Runner:
docker model run hf.co/bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct
Bactrainus HotpotQA Rationale Reader — Llama 3 8B Instruct
Artifact identity
- Status: complete merged causal-language-model checkpoint
- Base model:
meta-llama/Meta-Llama-3-8B-Instruct - Audited Hub revision:
852277e5b9534ff51a66adbad1ad43b7a3ef4457 - Public artifact date: August 2024
- Role: rationale-plus-answer generation from supplied evidence
This is a historical Llama 3 artifact. It must not be represented as either revised Llama 3.1 rationale-reader variant described in the updated manuscript.
Model summary
This reader is adapted to generate an intermediate natural-language rationale followed by an answer. The rationale is process supervision generated for task adaptation; it is not a hidden trace recovered from the base model and is not a gold supporting-fact annotation.
Intended use
- Studying natural-language rationale supervision for HotpotQA readers.
- Qualitative inspection of an evidence-conditioned answer path.
- Reader-stage comparisons where evidence is supplied independently.
Out-of-scope use
- Treating generated rationales as faithful explanations or verified proofs.
- Using rationale text as a substitute for HotpotQA supporting-fact labels.
- Open-domain retrieval, safety-critical decisions, or factual verification.
- Associating revised Llama 3.1 rationale results with this legacy checkpoint.
Input and output contract
Input should contain a question and selected, title-preserving evidence. Output is expected to contain rationale text and a final answer. Downstream code must parse the final answer explicitly and must keep rationale evaluation separate from answer/evidence metrics.
The public legacy configuration does not preserve a complete prompt-version manifest or an independently verified rationale delimiter. Do not assume that a newly invented delimiter exactly matches historical training.
Loading
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct"
REVISION = "852277e5b9534ff51a66adbad1ad43b7a3ef4457"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, revision=REVISION)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
revision=REVISION,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
Training data and lineage
The checkpoint derives from Meta Llama 3 8B Instruct and HotpotQA-based reader/rationale supervision. Its matching canonical training view is cot-reader-sft, pinned at revision 7f3a1d4d21f22aad7262d8ffd6520f31186b284d. The view provides rationale-plus-answer targets for all 90,447 training source IDs and joins to the original examples and every other view through source_id.
from datasets import load_dataset
train = load_dataset(
"bactrianus/bactrainus-hotpotqa",
"cot-reader-sft",
split="train",
revision="7f3a1d4d21f22aad7262d8ffd6520f31186b284d",
)
The canonical view is a deterministic, evidence-grounded serialization. It is not claimed to be byte-identical to every historical 2024 training file or to establish faithfulness of generated rationales.
Likewise, the revised reader_8b_rationale_8b.yaml and reader_8b_rationale_70b.yaml files describe Llama 3.1 experiments, not this Llama 3 weight artifact.
Evaluation boundary
No predictions or evaluation results are bundled with this card. The Bactrainus paper reports rationale-supervision experiments with explicit recipe caveats. A correct final answer does not establish rationale faithfulness.
Limitations
- Generated rationales can be post-hoc, incomplete, contradictory, or unsupported.
- Longer outputs increase parsing and truncation risk.
- Evidence omissions propagate to both rationale and answer.
- The model is specialized for English HotpotQA-style inputs.
- Wikipedia-derived data carries temporal and representational biases.
- Historical prompt, generator, and environment details are incomplete.
License and attribution
The weights remain subject to the Meta Llama 3 Community License and Acceptable Use Policy.
Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright Meta Platforms, Inc. All Rights Reserved.
Built with Meta Llama 3.
HotpotQA-derived data is licensed under CC BY-SA 4.0. Bactrainus code is Apache-2.0 licensed.
Citation
@article{barati2025bactrainus,
title = {Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks},
author = {Barati, Iman and Ghafouri, Arash and Minaei-Bidgoli, Behrouz},
journal = {arXiv preprint arXiv:2501.06286},
year = {2025},
doi = {10.48550/arXiv.2501.06286},
url = {https://arxiv.org/abs/2501.06286}
}
- Downloads last month
- 2
Model tree for bactrianus/HotpotQA-Reader-CoT-Llama-3-8B-Instruct
Base model
meta-llama/Meta-Llama-3-8B-Instruct