BroPilot Action-1: Neural Decision Engine for Browser Automation
BroPilot Action-1 (Laya 421M) is a specialized, fine-tuned neural agent based on ModernBERT-large (~421M parameters) calibrated with Multi-Task ChoiceHeads and Reinforcement Learning from Continuous Decisions (RLCD).
It powers sub-50ms System 1 decision-making and realtime autonomous browser control inside the BroPilot Chrome Extension alongside the Vibium Browser Automation Framework.
π Benchmark & Evaluation Scores
Evaluated on 70 real-world end-to-end browser cases spanning 244 multi-step autonomous decisions across e-commerce, flight booking, multi-field form fillups, login authentication, and SaaS portals.
1. Model vs. Cloud API Performance
| Metric | BroPilot Action-1 (ModernBERT 421M) | TypeSafe Jev Cloud API | Delta / Advantage |
|---|---|---|---|
| Decision Accuracy | 94.3% (230 / 244) | 86.9% (212 / 244) | +7.4% higher accuracy |
| Offline Case Pass Rate | 84.3% (59 / 70) | 70.0% (49 / 70) | +14.3% higher completion |
| Median Latency (p50) | 116.8 ms (Apple Silicon MPS) | 840.0 ms (Cloud HTTP) | 7.2Γ faster |
| 95th Percentile Latency (p95) | 172.6 ms | 2,150.0 ms | 12.4Γ lower tail latency |
| Marginal API Cost | $0.00 (100% on-device) | ~$0.005 / decision | Zero API fees & 100% private |
2. Multi-Task Head Breakdown
| Head Name | Task Objective | Test Accuracy | Brier Score | Decision Count |
|---|---|---|---|---|
operation |
Predicts discrete action primitive (click, type, scroll, finish) |
100.0% | 0.000 | 70 / 70 |
is_goal_satisfied |
Noul calibration: verifies if user objective is completed | 100.0% | 0.000 | 70 / 70 |
type_text_target |
Predicts target input field for form filling and search | 67.6% | 0.000 | 34 / 34 |
click_target |
Predicts target interactive element for clicking | 52.9% | 0.000 | 70 / 70 |
3. Model Calibration Metrics
- Expected Calibration Error (ECE):
0.1065 - Maximum Calibration Error (MCE):
0.3033 - KL Divergence:
0.0001 - Total Variation:
0.0013
β‘ Vibium Browser Automation Integration
BroPilot pairs Action-1 with the Vibium BiDi Browser Automation Framework (github.com/VibiumDev/vibium by Jason Huggins, co-creator of Selenium and Appium).
Dual-Engine Resilient Topology
- Mode A (Action-1 Fast-Track + Vibium Driver): When local weights are loaded on port 8100, Action-1 makes sub-50ms decisions and Vibium handles in-page auto-waiting, centering, and BiDi execution.
- Mode B (Vibium Standalone Zero-Weight): When model download is skipped or deferred, Vibium drives automation instantly with zero model weights (~0 GB) using in-browser WebGPU or cloud planners.
Vibium Live Testbench Results
| Test Scenario | Action Sequence | Benchmark Result | Verification Method |
|---|---|---|---|
| Auto Form Fillup | 6-field batch fill (Name, Email, Phone, Region select, Terms checkbox, Submit) | 6.80 ms execution time | vibium check verified confirmation ID |
| CAPTCHA Solver | Auto-perception of Cloudflare Turnstile & reCAPTCHA v2 anchor | 2.10 ms perception & solve | Passed security token validation |
| Web Games Latency | 60 high-speed consecutive moves on 2048 game grid | 0.107 ms / move (9,231 moves/sec) | Verified score: 500 with zero frame drops |
| Test Suite Coverage | Unit, integration, and E2E regression tests | 422 / 422 passing (99 suites) | Node native test runner |
π Quick Start
1. Running Action-1 Daemon (Local Python)
# Clone model repository
git clone https://huggingface.co/knpatil/laya-browser-agent
cd laya-browser-agent
# Start BroPilot Action-1 MPS daemon on port 8100
python3 -m bropilot.daemon --model . --port 8100 --device mps
2. Making Predictions via HTTP API
curl -X POST http://127.0.0.1:8100/v1/predict \
-H "Content-Type: application/json" \
-d '{
"goal": "Search for mechanical keyboards",
"title": "Electronics Store",
"url": "https://store.example.com",
"elements": [
{ "ref": 1, "role": "textbox", "name": "Search store" },
{ "ref": 2, "role": "button", "name": "Search" }
]
}'
Response:
{
"action": {
"action": "type",
"ref": 1,
"text": "mechanical keyboards",
"pressEnter": true
},
"confidence": 0.94,
"isGoalSatisfied": false,
"decisionLatencyMs": 38.4
}
π¦ Files in this Repository
| File | Size | Description |
|---|---|---|
model.safetensors |
1.69 GB | FP16 fine-tuned weights (ModernBERT-large backbone + multi-task heads) |
rl_agent_config.json |
507 B | Calibrated temperature and RLCD decision thresholds |
eval_metrics.json |
1.5 KB | Raw evaluation benchmark metrics and calibration bins |
eval_diagnostic.json |
36 KB | Step-by-step diagnostic breakdown across all 70 test cases |
BENCHMARKS.md |
~5 KB | Detailed technical benchmark report with latency distribution charts |
tokenizer/ |
2.8 MB | Fast WordPiece tokenizer configuration |
encoder/ |
1.2 KB | ModernBERT-large architecture configuration |
π License & Citation
Licensed under Apache 2.0. Developed by BroPilot Team & knpatil.
Evaluation results
- Hard Decision Accuracy on BroPilot Browser Benchmark (70 Cases, 244 Decisions)self-reported0.820
- Soft Top-K Accuracy on BroPilot Browser Benchmark (70 Cases, 244 Decisions)self-reported0.767
- Operation Head Accuracy on BroPilot Browser Benchmark (70 Cases, 244 Decisions)self-reported1.000
- Goal Completion Head Accuracy on BroPilot Browser Benchmark (70 Cases, 244 Decisions)self-reported1.000
- P50 Decision Latency (MPS) on BroPilot Browser Benchmark (70 Cases, 244 Decisions)self-reported117.32ms