I work at the intersection of applied research and production LLM systems. My focus is on the patterns that turn impressive demos into reliable, governable products:
Grounded generation — citation-aware RAG, confidence refusal, hybrid retrieval, and hallucination control
Agent observability — multi-step pipelines with full trajectory logging, critique–recovery loops, and evaluation harnesses
Parameter-efficient domain adaptation — QLoRA / PASTA-style workflows that specialise small models for healthcare, finance, legal, manufacturing, and more
LLMOps & FinOps — live latency, throughput, cache, and cost metrics that teams actually monitor
Security & guardrails — prompt-injection defence, red-team surfaces, and policy-driven access
Multi-tenant platforms — quotas, rate limits, audit logs, and shared-backend control planes
I build CPU-friendly, open, and inspectable reference implementations (Gradio + HF Spaces) so others can learn the patterns and adapt them. More projects in the pipeline: evaluation gates, continuous red-teaming, adapter routing, and tighter FinOps feedback loops.
Currently exploring: production RAG reliability, agent trajectory evaluation, and low-resource domain specialisation.