---

language: en

license: apache-2.0

tags:

- biology

- multi-agent

---

huggingface-model-card.md

language:

en tags:

biology

bioinformatics

tokenizers

systems-biology

biosecurity

secure-by-design

multi-agent

clinical-to-lab license: apache-2.0 datasets:

synthetic-nsclc-10k

Model Card: scalar-clinical-metadata-tokenizer

An open-source, biosecurity-compliant, secure-by-design metadata tokenizer designed to normalize and parse volatile, multi-scale clinical and biological datasets. This tokenizer serves as the foundational data-ingestion layer for the Bio-Stochastic Continuum (BSC) Engine open-source foundation, preparing raw inputs for real-time multi-agent choreography.

Model Details

Developed by: Department of Biology & Statistical Computer Science, Scalar Logic Group

Namespace: scalarlogicgroup

Model Type: Biological/Clinical Metadata Tokenizer

Language(s): Python / Tokenization Scripts

License: Apache 2.0 (Open-Source Foundation)

Repository: scalarlogicgroup/scalar-bio-choreography

Preprint: Non-Markovian Trajectory Modeling in Multi-Agent Biological Information Systems (submitted to bioRxiv)

Intended Use

This tokenizer is designed to ingest unstructured electronic health records (EHR) and volatile biological data, mapping raw inputs into structured state-space representations. It is optimized to:

Support zero-friction ingestion pipelines for multi-agent biological networks (reducing clinical click-rate friction by up to 84%).

Prepare non-Markovian trajectory data for real-time uncertainty quantification via Non-Parametric Bayesian Swarms.

Establish a standard, structured formatting layer for competitive "Clinician Proxy Agent" and "Biochemist Proxy Agent" feedback systems.

Out of Scope Uses

This tokenizer must NOT be utilized to process, represent, or design hazardous biological materials, regulated pathogens, select agents, or toxins. It should not be used in physical synthesis pipelines without active screening active.

Biosecurity & Safe-by-Design Curation

Consistent with global biosecurity frameworks—such as the 2024 White House Office of Science and Technology Policy (OSTP) Framework for Nucleic Acid Synthesis Screening, the 2023 HHS Screening Framework Guidance, and the Nuclear Threat Initiative (NTI) guidelines on AI biodesign tools (BDTs)—the training and data-ingestion pipeline of this tokenizer features a permanent, hard-coded safety filter:

Pathogen & Toxin Filtering: The tokenizer has been trained on a sanitized corpus that systematically excludes eukaryotic viral sequences and pathogenic genomic constructs.

Model-Level Safeguard Precedent: This structural exclusion mirrors safe-by-design precedents set by leading biological foundation models (e.g., ESM3-open, which omitted eukaryotic viral hosts to successfully limit pathogenic generation capabilities). Exposure to restricted sequences results in high perplexity, preventing the tokenizer from accurately parsing or normalizing hazardous genetic sequences of concern (SOCs).

Compliance: This active, pre-training dataset filtering layer ensures that downstream generative systems remain securely isolated from dual-use biological threat vectors while preserving high-performance modeling capabilities in translational oncology and immunology.

Performance Benchmarks

When deployed within the modified BSC Engine pipeline alongside the Non-Parametric Bayesian Throttling Gate (acting as a ≥95% confidence filter), the clinical-metadata tokenizer demonstrated the following baseline advantages under a simulated 400% stress surge:

Metric Traditional Pipeline BSC Modified Engine Net System Advantage

Data Ingestion Friction Manual parsing Automated multi-agent mapping 84% reduction in clinical clicks

Processing Latency Static matching / <120 ms <45 ms continuous execution 62.5% faster execution speed

Compute Expenditure Exponential escalation Logarithmic stabilization 41% reduction in total compute costs

System Uptime Telemetry Cascade Failure 100% Uptime under 4x surge Zero cascade failures

Citation

If you use this tokenizer or the associated biological routing protocol in your research, please cite our preprint:

@article{scalarlogicgroup2026nonmarkovian,

title={Non-Markovian Trajectory Modeling in Multi-Agent Biological Information Systems},

author={Office of the Director, Department of Biology \& Statistical Computer Science, Scalar Logic Group},

journal={bioRxiv},

year={2026},

publisher={Scalar Logic Group}

}

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train scalarlogicgroup/scalar-clinical-metadata-tokenizer