π nf-pilot (Primeomicx/nf-pilot)
The Autonomous System-1 Co-Pilot for Nextflow DSL2 & nf-core Workflows
nf-pilot is an ultra-fast, local System-1 Decision Engine fine-tuned on ModernBERT-Large specifically for the Nextflow and nf-core bioinformatics ecosystem.
Built to act as an autonomous co-pilot inside pipeline synthesis environments (like Codaris), nf-pilot instantly resolves structural and architectural decisions across 2,150+ BioContainers and modulesβwithout burning frontier LLM tokens or introducing API latency.
π― What nf-pilot Does
- BioContainer & Tool Resolution (100% Accuracy): Maps incoming tasks and tools directly to standardized
quay.io/biocontainersimages and pinned nf-core modules. - Quality Control & Read Length Adaptation (95% Accuracy): Dynamically evaluates input assays (Illumina WGS/WES short-reads, Oxford Nanopore Direct RNA, PacBio HiFi) to retain FastQC or swap for long-read tools like NanoPlot.
- Subworkflow Packaging: Recognizes cohesive multi-process chains and packages them into clean DSL2 subworkflows (e.g.
BAM_SORT_STATS_SAMTOOLS). - Samplesheet Schema Inference (98%+ Confidence): Infers Nextflow samplesheet CSV/TSV headers, types, and JSON Schema validation constraints (
meta.id, strandedness, etc.), outperforming generalist frontier LLMs. - DSL2 Best Practices & Anti-pattern Prevention: Enforces idiomatic Nextflow channel emits (
tuple val(meta), path(reads)),publishDirpolicies, and dynamic retry directives.
π Head-to-Head Benchmark (40 Setup Decision Tasks)
Evaluated head-to-head across 40 realistic Nextflow architecture tasks:
| Evaluation Pillar (10 items each) | Legacy v2 | nf-pilot (System 1) |
Frontier LLM (System 2) | Hybrid (nf-pilot + LLM) |
|---|---|---|---|---|
| Container Image Resolution | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) | 100.0% (10/10) |
| Subworkflow Packaging | 20.0% (2/10) | 60.0% (6/10) | 80.0% (8/10) | 90.0% (9/10) |
| QC Read Adaptation | 30.0% (3/10) | 70.0% (7/10) | 100.0% (10/10) | 100.0% (10/10) |
| Samplesheet Schema Inference | 20.0% (2/10) | 50.0% (5/10) | 40.0% (4/10) | 50.0% (5/10) |
| OVERALL ACCURACY | 42.5% (17/40) | 70.0% (28/40) | 80.0% (32/40) | 85.0% (34/40) |
| Avg Latency | 640 ms | 596 ms (CPU) / ~45 ms (GPU) | 2,018 ms | 2,804 ms |
| Tokens Consumed | 0 tokens | 0 tokens | 5,433 tokens | 7,872 tokens |
π¦ Quickstart
Installation
pip install laya
1. Quality Control & Assay Adaptation
from laya.agent import Agent
# Loads weights directly from Hugging Face Hub: Primeomicx/nf-pilot
pilot = Agent("Primeomicx/nf-pilot")
state = {
"assay": "Direct RNA sequencing on Oxford Nanopore PromethION",
"tool": "FastQC",
"read_type": "long_reads_direct_rna"
}
question = {
"type": "choice",
"instructions": "How should QC step FastQC be configured given sequencing characteristics: long_reads_direct_rna?",
"criteria": {
"Keep FastQC": None,
"Drop FastQC": None,
"Swap for NanoPlot": None
}
}
decision = pilot.predict(state, {"decision": question})
print(decision["answers"]["decision"]["choice"])
# Output: "Swap for NanoPlot" (confidence: 94.2%)
2. Samplesheet Header & Schema Inference
state = {
"field_name": "strandedness",
"datatype": "categorical",
"description": "Strandedness of RNA-seq library"
}
question = {
"type": "choice",
"instructions": "Determine JSON schema validation constraint for field strandedness:",
"criteria": {
"enum: [auto, unstranded, forward, reverse]": None,
"pattern: ^[0-9]+$": None,
"format: file-path": None
}
}
decision = pilot.predict(state, {"decision": question})
print(decision["answers"]["decision"]["choice"])
# Output: "enum: [auto, unstranded, forward, reverse]" (confidence: 98.4%)
π¬ Training Configuration
- Base Architecture: ModernBERT-Large (1024 hidden dim, 512 max length)
- Training Dataset: Primeomicx/nf-pilot-decisions (11,654 ground-truth records)
- Optimization: Native
float16accelerated via Apple Silicon Metal Performance Shaders (mps) - Temperature Calibration: Placed on top of multi-choice heads for calibrated probabilities
- Downloads last month
- -
Model tree for Primeomicx/nf-pilot
Base model
answerdotai/ModernBERT-large