apertus-8b-financial-reasoner-v2

A fine-tune of swiss-ai/Apertus-8B-Instruct-2509 for two-stage financial analysis of a stock ticker:

  • Task A – classify how the market reacted to a news item given the price move (good, bad, neutral, overreaction_down, overreaction_up).
  • Task B – given that reaction, a valuation gap and fundamentals, write the reasoning to find a recommendation (BUY/SELL/HOLD) and answer as JSON.

Format

Full merged checkpoint in bfloat16 (~16 GB, 4 safetensors shards). It is not pre-quantized — load it in 4-bit at runtime if you need to fit a small GPU:

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    "gaparecido/apertus-8b-financial-reasoner-v2",
    max_seq_length=2048,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
model.config.use_cache = True
model.generation_config.use_cache = True

Serving notes

- Needs a bf16-capable GPU (L4 / A10G / A100 or better). Apertus is bf16-trained; in fp16 on a T4 the logits overflow to NaN and every generated token is <unk>.
- The upstream Apertus config ships use_cache: false and this repo has no generation_config.json to override it, so pin use_cache=True (as above) or generation recomputes attention over the full sequence per token.
- The CUDA-fused xIELU not available warning is harmless — the Python fallback costs ~10% throughput.

Training

Trained with Unsloth (https://github.com/unslothai/unsloth) (QLoRA, 4-bit) and Hugging Face TRL, loaded from Unsloth's unsloth/apertus-8b-instruct-2509-unsloth-bnb-4bit mirror of the base model, on an NVIDIA L4.
Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gaparecido/apertus-8b-financial-reasoner-v2

Finetuned
(19)
this model