AFM-D Encoder (afm_de)

On-device System 1 decisions over a typed answer space.
Site: ariacompute.com · Org: ariacompute · Hub: ariacompute/afm-de

AFM-D Encoder is the non-autoregressive track of AFM-D: a Laya-style ModernBERT-large DecisionModel with a MASK option head, RLCD-fine-tuned from convaiinnovations/laya. It scores a text/JSON state against user-supplied options and returns a distribution — not open-ended chat, not TypeSafe Jev.

Primitive Answer space Result
Choice 1–255 options choice, probabilities, confidence
Score 2–10 ordered levels expected score, distribution, confidence
Noul false / true noul = P(true)

Companion Decoder (PEFT LoRA on MiniCPM5-2B): ariacompute/afm-dd.

Model details

Subject id afm_de
Init / base convaiinnovations/laya (ModernBERT-large DecisionModel)
Encoder backbone answerdotai/ModernBERT-large
Context max_len=1024; shared head_max_len=512 for question + option text
Option text OPTION_DESC_MAX=96 tokens; long states keep the tail (truncate_left)
High-cardinality Choice embedding shortlist → one forward pass
Decision temperature fixed 1.0
Confidence default normalized entropy; product calib uses noul=max / choice=max / score=max; per-bucket confidence temperatures fitted by ECE grid search
Training RLCD on option distribution (log + spherical + RPS for Score); LABEL_SMOOTHING=0.02 on hard one-hots
Export Laya-layout safetensors

Checkpoint layout

model.safetensors
rl_agent_config.json
tokenizer/
encoder/

Intended use

  • Local / on-device typed decision scoring (Choice / Score / Noul) over a supplied state.
  • Offline eval and JevBench comparison via the AFM-D harness.

How to use

Load through Aria Engine (native AFM runtime) or Ollaya (Ollama-style local decision daemon).

Aria Engine

Pure Rust inference for AFM Encoder (afm_de): CLI, POST /v1/systemone, FFI, and language SDKs. Docs: engine README.

# setup (writes ~/.ariacompute/engine.yml; .com → HF token, .cn → ModelScope)
aria-engine setup
aria-engine download afm-de          # → ~/.ariacompute/models/afm-de

# one-shot decide (AFM-D record or System One body on stdin / --file)
aria-engine decide --track encoder --model-name afm-de --file record.json

# local HTTP server (default http://127.0.0.1:8010)
aria-engine serve --model-name afm-de

System One body (POST /v1/systemone or decide):

curl -s http://127.0.0.1:8010/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "user wants a refund",
    "questions": {
      "q1": {
        "type": "choice",
        "instructions": "Pick the best action",
        "criteria": {
          "refund": "issue a full refund",
          "deny": "deny the request"
        }
      }
    }
  }'
type criteria
choice object: option name → description
score array of 2–10 ordered level strings
noul object with true / false (optional)

Python SDK (pip install aria-engine; needs libaria-engine_ffi or ARIA_FFI_LIB):

from aria_engine import AriaEngine

eng = AriaEngine("/path/to/afm-de", "encoder")  # or ~/.ariacompute/models/afm-de
out = eng.systemone({
    "state": "user wants a refund",
    "questions": {
        "q1": {
            "type": "choice",
            "instructions": "Pick the best action",
            "criteria": {
                "refund": "issue a full refund",
                "deny": "deny the request",
            },
        }
    },
})
print(out["answers"]["q1"])
eng.destroy()

Also: TypeScript @ariacompute/engine-ts, Rust ariacompute-engine, Go / Flutter / Swift / Kotlin — see engine bindings/.

Ollaya

Ollaya pulls AFM-D by name and serves TypeSafe-compatible /v1/systemone. Use the AFM-D-enabled builds from ariacompute/ollaya releases (not the default ollaya-dev/ollaya channel). Weights stay on Hugging Face / ModelScope (Ollaya does not re-host).

# install from https://github.com/ariacompute/ollaya/releases (latest AFM-D build)
curl -fsSL https://raw.githubusercontent.com/ariacompute/ollaya/main/scripts/install.sh \
  | OLLAYA_REPO=ariacompute/ollaya sh
# pin a release: OLLAYA_REPO=ariacompute/ollaya OLLAYA_VERSION=0.7.5+afm-d.1.0.0
# Windows (PowerShell):
#   $env:OLLAYA_REPO='ariacompute/ollaya'; irm https://raw.githubusercontent.com/ariacompute/ollaya/main/scripts/install.ps1 | iex

# weights: HF (ariacompute/afm-de) or ModelScope (AriaCompute/afm-de)
# OLLAYA_HUB=huggingface|modelscope|auto  (auto → ModelScope when LANG looks Chinese)
ollaya pull afm-de
ollaya run afm-de --preset triage \
  "Third time this year you've double-charged me. Refund it today or I'm cancelling."

Or download a platform asset from the releases page (e.g. ollaya-linux-amd64.tar.zst, ollaya-darwin-arm64.tgz, ollaya-windows-amd64.zip), unpack, put ollaya on PATH, then pull / run as above.

HTTP (daemon default http://localhost:11435):

curl http://localhost:11435/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "afm-de",
    "state": "user wants a refund",
    "questions": {
      "q1": {
        "type": "choice",
        "instructions": "Pick the best action",
        "criteria": {
          "refund": "issue a full refund",
          "deny": "deny the request"
        }
      }
    }
  }'

Same question schema as Engine. Native route: POST /api/decide. Family notes: Ollaya repo docs/families/afm-de.md.

Eval

Decision T=1.0; calib auto picked noul=max / choice=max / score=max.

AFM-D Encoder Base Laya
Agreement 71.1% 55.7%
ECE 0.033 0.141
Brier 0.386 0.605
Task / bucket n Agree ECE
choice (all) 2036 73.8% 0.049
choice:3-5 1360 71.6% 0.059
score:3-5 1019 52.6% 0.095
noul:2 1255 81.8% 0.052

Confidence temperatures (approx.): choice:2 ≈1.9, noul:2 ≈1.75, choice:3-5 ≈1.7, choice:6-10/11+ ≈1.1, score:3-5 =1.0. High-conf errors (conf≥0.7 among wrongs) ≈25%.

JevBench

# System Score Intel. Calib. Speed Acc. Hard
1 SemIf 83.9 74.9 87.1 90.6 81.0% 61.3%
2 Bespoke Nimble-9B 79.2 72.5 76.2 89.8 79.7% 61.3%
3 NeoHorse-Jev-4B 78.9 63.1 84.6 91.9 72.3% 45.0%
4 Kev-4B 78.3 67.2 76.7 92.9 75.8% 54.1%
5 AFM-D Encoder 45.9 41.0 82.7 94.1 59.3% 38.7%
6 AgentJev-0.6B 41.8 40.0 79.8 88.2 58.0% 36.0%
7 Laya 30.9 36.4 57.7 93.9 53.2% 27.9%

AFM-D Encoder tiers: easy 100%, standard 63.9%, hard 38.7%.

Limitations

  • ModernBERT encoder track: strong calibration / speed relative to base Laya, weaker hard-tier Intelligence than larger causal peers on JevBench public-proxy.
  • Score buckets remain the weakest local eval slice; prefer targeted data over blind extra epochs when ECE stays high.
  • English-centric product corpus; peer hard-label imports are best-effort and may be skipped if missing.
  • Must be loaded through AFM-D / Laya DecisionModel code — not a drop-in AutoModelForCausalLM chat checkpoint.

License

MIT

Citation / links

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ariacompute/afm-de

Finetuned
(89)
this model

Datasets used to train ariacompute/afm-de