BetterWright Encoder 350M

A task-conditioned accessibility-tree relevance encoder for reducing browser agent context while preserving action targets and evidence. It is trained from LiquidAI/LFM2.5-Encoder-350M in bfloat16.

All 355,011,331 parameters are trained; this checkpoint does not use LoRA, adapters, or a frozen backbone. The model has three outputs: chunk relevance, token/node relevance, and a confidence signal for deterministic fallback. BetterWright keeps its original snapshot whenever the confidence gate is not met.

The training corpus contains 500,000 deduplicated structural examples grounded in 5,256 independently generated and audited tasks from 141 real-site domains. Domains—not rows—define the split: 110 train, 8 validation, and 23 held-out test domains. The corpus is primarily BetterWright accessibility trees, with a smaller share of alternate accessibility serializations for robustness.

Input

[BETTERWRIGHT_TASK]
<the current browser task or observation query>
[ACCESSIBILITY_SUBTREE]
<BetterWright aria tree chunk>

The released checkpoint is progressively trained at 8K, 16K, 32K, and 64K sequence lengths. Normal inference should still use hierarchical chunks for lower latency; 64K is intended for unusually large dumps and packed tabs.

Quick start

import torch
from transformers import AutoModel, AutoTokenizer

model_id = "ProCreations/betterwright-encoder-350m"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).eval()
batch = tokenizer(
    "[BETTERWRIGHT_TASK]\nOpen pricing\n[ACCESSIBILITY_SUBTREE]\n"
    '- link "Pricing" [ref=e12]',
    return_tensors="pt",
)
with torch.inference_mode():
    output = model(**batch)
relevance = torch.sigmoid(output.logits)
confidence = torch.sigmoid(output.uncertainty_logits)
token_relevance = torch.sigmoid(output.token_logits)

Production pruning should use the validated thresholds in relevance_config.json, deterministic must-retain rules, and exact full-tree fallback. A raw relevance score alone is not a safe deletion decision.

Training

  • 500,000 unique pairs: 77.7% BetterWright trees, 13.1% alternate ARIA trees, and 9.1% compact accessibility outlines.
  • 391,777 train, 25,891 validation, and 82,332 untouched-test pairs, split by domain before structural expansion.
  • Full-parameter BF16 AdamW training with global batch 128, a cosine schedule, and a class weight derived from the exact training distribution.
  • A balanced positive/negative packed-tree curriculum at 8K, 16K, 32K, and 64K tokens, with exact required-node span supervision.

Evaluation policy

Fallback thresholds are selected on validation domains. The production policy uses eight ranked 1,800-character structural regions, preserves referenced nodes, bounded child context, ancestry, and deterministic must-retain states, and returns the exact original tree on low confidence, errors, insufficient benefit, or an over-limit result.

Domain split Tasks Target/evidence recall Perfect-task recall Pruning coverage Mean token savings when active Token-weighted savings
Validation (8 domains) 288 100.000% 100.000% 12.50% 33.38% 3.42%
Held-out test (23 domains) 877 99.584% 99.544% 22.46% 44.09% 9.69%
Post-training reserve (17 additional domains) 161 100.000% 100.000% 33.54% 16.78% 2.42%

The exact reports and frozen policy are distributed as eval.json, reserve_eval.json, and relevance_config.json. The 99.5% target is evaluated on whole tasks and required accessibility refs, not inferred from training loss. A nine-site live navigational replay retained and clicked 6/9 current targets; two large pages exceeded the runtime ceiling after fallback and one site did not expose the requested label. This limitation is intentionally not hidden by substituting easier sites.

Safety and limitations

This is a relevance model, not an autonomous browser agent. It must not be the only path for security decisions. Snapshot text is untrusted data. Consumers should keep deterministic must-retain rules and fall back to the full tree on model errors, timeouts, low confidence, or insufficient retained context. Savings are substantial when pruning activates, but coverage is deliberately conservative and very large pages can still fall back above the caller's size ceiling. This checkpoint is therefore an alpha integration, not a claim that every page becomes smaller.

Downloads last month
174
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/betterwright-encoder-350m

Finetuned
(27)
this model