A 15+ minute manual process reduced to under 90 seconds. RAG-based document intelligence with confidence scoring and human-in-the-loop review for exceptions.
Projects
Agentic AI systems I've shipped to production
A selection of AI systems I've taken end-to-end — from problem framing and architecture through to production, adoption, and measurable business impact — across regulated finance, legal, and consumer domains. Ownership spans multi-agent orchestration, RAG, LLM-as-judge and eval-driven quality, explainability, and human-in-the-loop design.
A 2–3 day proposal process collapsed to under 60 minutes. A multi-agent pipeline with LLM-as-judge quality scoring and approval workflows.
Natural language to SQL with enterprise-grade accuracy. LIME/SHAPLEY explainability overlays meet compliance and adoption requirements in regulated environments.
Executives ask business questions in plain English and get SQL-powered answers, charts, and recommendations. A LangGraph pipeline classifies intent, retrieves schema and query patterns via RAG, plans and generates PostgreSQL, self-validates before executing live, then layers on insights and follow-ups.
Banks receive legal notices across email, portal, and SFTP. LEANM ingests and normalises them, extracts structured data with PII redaction, then routes each notice to the right team with priority and SLA — including multi-directive notices carrying independent deadlines.
Upload a contract and Claude classifies it, extracts key terms, and checks every clause against a configurable rubric — returning a risk level, recommendation, and clause-by-clause gap analysis in seconds. Structured tool-use output, backed by a ground-truth eval harness.
A consumer travel app built around a conversational planner (Atlas), a "Travel DNA" personalisation engine, smart destination recommendations, and proactive nudges — all reshaped in real time by a four-stage travel-lifecycle state machine from exploring to in-destination.
An enrichment API that turns sparse transaction data into structured merchant intelligence — categories, tags, and locations — by combining LLM classification against a curated taxonomy with web-scraped signals.
A vision pipeline that scores brand imagery for Sharia'a compliance — detecting alcohol, gambling, tobacco, and other flagged categories with per-object confidence — and extracts structured product data straight from images.
Field and support teams lose time hunting for answers buried in dense equipment manuals. This RAG assistant ingests technical and repair documentation — iFixit guides plus Toshiba and Otis equipment PDFs — and answers natural-language repair and troubleshooting questions grounded in the source documents.
The data-as-a-service layer beneath the analytics agents: it auto-maps and reconciles schemas across sources using embedding similarity and the Valentine matching library, so downstream agents query clean, unified data.
Want the deeper story on any of these — architecture, trade-offs, or outcomes? Get in touch.
