ArcANE-32B-SFT

Paper Dataset Model Model

🏆 Accepted to EMNLP 2026 Main Conference

ArcANE-32B-SFT is a Qwen3-32B model fine-tuned with LoRA on the supervised split of ArcANE. It is designed for English role-playing responses that reflect a character's behavioral state at a specified point in a narrative, rather than treating the character as a fixed persona.

ArcANE character-arc construction and probe-generation pipeline

Model details

Field Value
Base model Qwen/Qwen3-32B
Parameters 32B class
Training stage Supervised fine-tuning (SFT)
Parameter update LoRA, rank 64 and alpha 128
Recommended mode Qwen3 non-thinking mode

Intended use

ArcANE-32B-SFT is intended for research on:

  • point-in-time character role-play;
  • character responses conditioned on a chapter-truncated Character Arc;
  • behavioral and value changes across narrative phases;

The strongest evaluated setup supplies the relevant Character Arc only up to the queried chapter. Future phases must not be exposed to the model.

Training data

The SFT split is derived from 12 training novels in the ArcANE corpus, covering 55 characters and 339 character axes. For each (probe, phase) pair under Arc context, three teacher completions were sampled from gpt-5.4-mini. The teacher privately received the phase reference, but each stored training row retained only the character system prompt, scenario-question user prompt, and answer. Reference text was therefore not included in the model input.

The main validated evaluation slice is held out at the novel, character, arc, and probe levels from the training pool.

Training parameters

Hyperparameter Value
Epochs 1
Learning rate 1e-4
Effective batch size 32
Maximum sequence length 8,192 tokens

Reproducibility

The released recipe is training/sft/configs/lora/sft.yaml in the ArcANE repository. From training/sft, run bash scripts/train_sft_lora.sh.

Evaluation

The paper evaluates free-form role-playing responses on a held-out five-novel slice containing 25 principal characters, 205 arcs, and 1,754 probes. A separate DeepSeek-V4-Flash judge scores four 1 to 100 metrics:

  • APF: Action Phase-Fidelity;
  • RPF: Reasoning Phase-Fidelity;
  • RAE: Reasoning-Action Entailment;
  • PTF: Phase Trajectory Fidelity.

Scores are pooled within each novel, novels receive equal weight, and Overall is the mean of the 12 probe-category by metric cells. Higher is better.

Held-out results with Arc context

Probe category APF RPF RAE PTF
In-Scenario 62.6 61.3 55.9 56.7
In-World 61.4 60.6 54.7 51.8
Out-of-World 63.4 62.3 57.9 52.2
Comparison Overall
ArcANE-32B-SFT, Arc context 58.4
ArcANE-32B-SFT, strongest non-Arc context 53.7
Qwen3-32B, Arc context 50.1

The Arc-context score is 4.7 points above this checkpoint's strongest non-Arc context and 8.3 points above Qwen3-32B under the same Arc context.

Usage

Use the Qwen3 chat template with thinking disabled. The example below is illustrative; replace the compact context with a valid chapter-truncated ArcANE Character Arc.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "holi-lab/ArcANE-32B-SFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {
        "role": "system",
        "content": (
            'You are <character>, from "<novel>". You are at the point in '
            "the story corresponding to chapter <query_chapter>.\n\n"
            "Background you have access to:\n"
            "<context>\n<chapter-truncated Character Arc JSON>\n</context>"
        ),
    },
    {
        "role": "user",
        "content": "Scenario:\n<scenario>\n\nQuestion:\n<question>",
    },
]

inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    enable_thinking=False,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(inputs, max_new_tokens=1024, do_sample=True, temperature=1.0)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))

For faithful point-in-time conditioning, remove all phases later than the query chapter. Always remove literary_validation and evidence_summary; if any later phase is hidden, also remove pole_end and arc_direction. The paper's evaluation drew one sample with backend-default sampling, effectively temperature 1.0, and capped generation at 8,192 tokens.

Citation

@misc{song2026arcaneroleplayinglanguageagents,
      title={ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?}, 
      author={Woojung Song and Nalim Kim and Sangjun Song and Chaewon Heo and Jongwon Lim and Yohan Jo},
      year={2026},
      eprint={2606.05553},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.05553}, 
}
Downloads last month
242
Safetensors
Model size
33B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for holi-lab/ArcANE-32B-SFT

Base model

Qwen/Qwen3-32B
Finetuned
(538)
this model
Finetunes
1 model
Quantizations
2 models

Dataset used to train holi-lab/ArcANE-32B-SFT

Collection including holi-lab/ArcANE-32B-SFT

Paper for holi-lab/ArcANE-32B-SFT