Falco Sovereign Decision Engine (Experimental Prototype)
Falco is an experimental, non-autoregressive System 1 routing model exploring multi-query sequence packing over a shared context prefix. It evaluates multiple categorical classification queries concurrently in a single encoder forward pass without inter-query token interference.
It was designed to test whether 4D block-diagonal prefix attention masking and dynamic position ID resets can eliminate slot permutation drift in multi-query encoder setups while keeping computational footprints under 300 MB.
- Primary Backbones: Fine-tuned DistilBERT-base-uncased (66M) for English; INT8-quantized IndicBERT-v2 (278M) for 22 Indic scripts.
- Format: Exported directly to ONNX Runtime (CPU and DirectML/CUDA compatible).
- Scope: Experimental research artifact. Designed for structured triage, policy gating, and fast classification—not open-ended generation.
Intended Use & Boundaries
Recommended Use Cases
- High-Frequency Policy Pre-Routing: Screening inbound text (support tickets, logs, prompts) for binary or low-cardinality ($K \le 4$) routing categories before invoking expensive LLMs.
- Air-Gapped & Edge Deployments: Environments requiring local, CPU-based classification with deterministic sub-30ms execution times and zero marginal API fees.
- Cross-Query Invariance Research: Applications requiring mathematical guarantees that query order does not leak information between concurrent classification heads.
Out-of-Scope & Unsupported Uses
- Do NOT use for open-ended text generation: Falco is an encoder with discrete linear heads. It cannot generate rationales, summaries, or conversational prose.
- Do NOT use as a replacement for high-capacity language models: DistilBERT (66M) has limited linguistic capacity. It will fail on complex, multi-sentence reasoning chains and subtle legal/medical nuances.
- Do NOT use for high-cardinality classification ($K > 4$): The output heads are structurally bounded to 4 logits. Passing queries with 5 or more options will result in clipped classes.
- Do NOT deploy in high-stakes clinical or safety-critical decisions without human verification: The model's training distribution is bounded, and out-of-distribution performance degrades sharply.
How It Works: Architectural Mechanics
Contiguous Input Sequence:
[CLS] State Prefix Context Tokens... [SEP] Q1 Tokens... [SEP] Q2 Tokens... [SEP] Q3 Tokens... [SEP]
│ │ │
4D Attention Mask: ▼ ▼ ▼
Each slot attends to: [State Only + Q1] [State Only + Q2] [State Only + Q3]
Position IDs: [0 ... L_state-1] [L_state ... +N1] [L_state ... +N2]
│ │ │
Head Pooling & Clamping: ▼ ▼ ▼
Logit Clamping: Clamp(K=2) Clamp(K=3) Clamp(K=4)