AI & ML interests
AI reliability, legal AI, evaluation benchmarks, citation verification, evidence infrastructure, reproducibility, provenance systems, retrieval evaluation, small language models, AI safety.
Recent Activity
Dali
The open verification layer for AI
Dali is the open verification layer for AI: it creates, scores, and preserves evidence so AI-assisted outputs can be independently verified, exchanged, and replayed. Legal AI is the proving ground.
Open Artifacts
5 hand-curated cases, 14 authorities — methodology review sample (not the full run).
524 citations · 3 models · 5 jurisdiction tracks — on GitHub data/results/.
Open verification engine, methodologies, and reproducible evaluation workflows.
Quick Start
from datasets import load_dataset
dataset = load_dataset("yenklabs/open-evidence-corpus")
GitHub · Contribute · yenklabs.com
Research models (roadmap)
Lightweight, reproducible research models — not foundation models. All planned.
| Release | Model | Purpose | Status |
|---|---|---|---|
| v0.1 | Dali Verification Taxonomy Classifier | Predict standardized verification outcome labels | planned |
| v0.2 | Dali Citation Risk Classifier | Estimate citation verification risk from evidence metadata | planned |
| v0.3 | Dali Authority Matching Baseline | Reproducible baseline for authority matching experiments | planned |
| v0.4 | Dali Proposition Support Classifier | Classify proposition support relationships for legal authorities | planned |
Ecosystem
- Datasets — Open Evidence Corpus, Citation Benchmark seed sample, Verification Taxonomy · available
- Full evaluation run — 524 citations on github.com/yenklabs/Dali/data/results · available
- Models — Taxonomy Classifier, Citation Risk, Authority Matching, Proposition Support · planned
- Spaces — Evidence Explorer, Benchmark Dashboard, Citation Verification Demo · planned
- Future datasets — Evaluation Prompts, Replay Corpus, Evidence Artifacts · planned
Dali is the open verification layer for AI: it creates, scores, and preserves evidence so AI-assisted outputs can be independently verified, exchanged, and replayed.