You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

LFM2.5-1.2B-RLCD (Certa)

A System-One calibrated decision model: given a state and typed questions, it returns typed answers with calibrated probabilities in a single forward pass β€” no text generation, no parsing. Runs with the open-source certa library (pip install certa); training recipe summarized below.

  • Base: LiquidAI/LFM2.5-1.2B-Base
  • Task: Choice (pick one of ≀26 options), Score (ordered level), Noul (yes/no probability)
  • In-distribution test: 0.766 accuracy Β· 0.319 Brier Β· 0.043 ECE Β· 0.079 AURC
  • Latency: β‰ˆ 14 ms per multi-question decision (warm, L40S)
  • API: compatible with TypeSafe.ai's Jev POST /v1/systemone

Usage

pip install certa
from certa.engine import ModelEngine
from certa.client import choice, score, verify

engine = ModelEngine("goutam/LFM2.5-1.2B-RLCD")

resp = engine.decide(
    "I was double-charged and support has ignored me for three days.",
    {
        "team":   choice("Route this ticket", {"billing": "charges", "technical": "bugs", "sales": "pricing"}),
        "anger":  score("How angry is the customer?", ["calm", "annoyed", "furious"]),
        "urgent": verify("Does the message convey urgency?"),
    },
)
print(resp.answers["team"].choice, resp.answers["team"].confidence)   # billing 0.79
print(resp.answers["anger"].score)                                    # 1.44 (0–2)
print(resp.answers["urgent"].noul)                                    # 0.78

To serve the Jev-compatible HTTP API:

CERTA_CHECKPOINT=goutam/LFM2.5-1.2B-RLCD pip install 'certa[serve]' && python -m certa --port 8000

Training

  1. Supervised warm start on 12k labelled rows from goutam/rlcd-decision-atlas: focal (Ξ³=2) + multiclass Brier over the option-token logits.
  2. RLCD refinement (Reinforcement Learning for Calibrated Decisions) on the remaining ~519k rows: REINFORCE in a single-step bandit setup (sample one option, one scalar reward), reward r = correct βˆ’ p(action) (RPS partial-credit for ordinal questions), RLOO (leave-one-out) baseline, light KL to the supervised reference.

Evaluation

Split Accuracy Brier ↓ ECE ↓ AURC ↓
Test (in-distribution) 0.766 0.319 0.043 0.079
Validation 0.772 0.333 0.054 β€”
Holdout (unseen sources) 0.458 0.678 0.157 0.395

Multiclass-sum Brier (range 0–2), top-label 15-bin ECE. Selective prediction: acting only when the top probability β‰₯ 0.9 covers β‰ˆ 39% of cases at β‰ˆ 2.9% error β€” a calibrated confidence gate for act-vs-escalate routing.

Intended use & limitations

Use for: fast calibrated routing, triage, classification, and grading over text where you need a probability you can threshold β€” and to abstain/escalate on low-confidence cases.

Limitations:

  • Out-of-distribution: weak (~0.46 accuracy) on entirely unseen source types β€” broaden training coverage for new domains, and re-check calibration on yours.
  • ≀ 26 options per Choice (single-letter decode); larger sets are rejected.
  • Trained on English-leaning text sources; non-text / non-English inputs are unsupported.

License

Weights are released under the LFM Open License v1.0 (inherited from the base model). Accompanying code (certa, certa-rlcd) is Apache-2.0. Not affiliated with TypeSafe.ai.

Downloads last month
4
Safetensors
Model size
1B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for goutam/LFM2.5-1.2B-RLCD

Finetuned
(59)
this model

Dataset used to train goutam/LFM2.5-1.2B-RLCD