AutoIndexer-8B-LoRA-Instruct

This is an AutoIndexer checkpoint: a Qwen3-8B backbone fine-tuned (LoRA, merged into the base weights) with AutoIndexer's chain-of-edits objective and custom attention machinery (edit/return/EOS markers, a start/end "index head" for cursor placement, and a marker head for edit-vs-normal-token classification).

Run name: autoindexer_8B_lora_p5_20260918-220528_focal-loss-top-p-1.0.

See the sibling ADSKAILab/AutoIndexer-8B-LoRA-Base repo for the base-model-trained counterpart from the same training sweep.

Architecture

  • Backbone: Qwen3-8B (hidden_size=4096, num_hidden_layers=36, num_attention_heads=32, num_key_value_heads=8).

  • Custom model_type: autoindexer_qwen3, custom class AutoIndexerQwen3Model (defined in this repo's modeling_autoindexer.py).

  • Custom attention backend, selected via config.autoindexer_attn_implementation:

    • autoindexer_cutedsl (default in this config) โ€” fastest, requires CuteDSL/flash_attn built for Hopper+ GPUs (sm 9.0+).
    • autoindexer_triton โ€” used automatically if CuteDSL isn't available but Triton + CUDA are.
    • autoindexer_eager โ€” pure PyTorch fallback, always available (CPU-compatible, slower, more memory).

    The model automatically falls back down this list at load time depending on what's installed/available in your environment โ€” no action needed unless you want to force a specific backend (pass attn_implementation="autoindexer_eager" to from_pretrained, for example).

Usage

import torch
from transformers import AutoModel, AutoTokenizer

repo_id = "ADSKAILab/AutoIndexer-8B-LoRA-Instruct"

model = AutoModel.from_pretrained(repo_id, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained(repo_id)

trust_remote_code=True is required: this architecture's modeling/attention code ships as custom code inside this repo (there is no AutoIndexerQwen3Model in the transformers package itself).

Notes

  • This model's forward expects labels for its training objective (chain-of-edits perturbation + token/marker/index losses). Calling it with only input_ids (no labels) runs a plain causal-LM scoring pass instead (useful for likelihood/eval harnesses).
  • generate() uses a custom sampling loop (AutoIndexerModelBase._sample) that also samples the marker head (edit/return/EOS) and the index head (cursor start/end) at every step; see the docstring on that method in modeling_autoindexer.py for the available generation_config knobs (marker_calibration_weight, greedy_mid_edit, max_delete_span, exclude_prompt_from_edits, editable_start, marker_top_p, marker_temperature).
Downloads last month
17
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ADSKAILab/AutoIndexer-8B-LoRA-Instruct

Adapter
(146)
this model

Collection including ADSKAILab/AutoIndexer-8B-LoRA-Instruct