Professional Overview
I build and verify the knowledge bases behind production conversational AI. Over the past year I have authored more than 1,500 knowledge base entries across 12 client deployments in 7 industries and 8 markets, covering hospitality, retail, FMCG, healthcare, entertainment, real estate and MICE venues, and designed the test suites that verify them.
My work centres on one question: can this answer be traced back to a source? I design control-based test methodology that separates genuine model failure from grading error, write mechanical validators that refuse to ship a file when a rule is violated, and treat an unsourced claim as a defect regardless of how plausible it reads.
I was trained in archaeology, which is source criticism as a discipline: date every source, rank its authority, and never let inference pass as evidence. That is the same method applied to AI content.
Core Expertise
Knowledge Base Engineering
Turning unstructured client source material into retrieval-ready entries that an assistant can actually answer from.
- Taxonomy and schema design for multi-entry knowledge bases
- Coverage mapping from sitemaps, so no source sub-page is silently missed
- Deduplication and keyword collision resolution across hundreds of entries
- Provenance tracking, so every entry records where its facts came from
- Largest single builds: a 533-entry retail catalogue and a 500-entry product knowledge base
Evaluation and Test Design
Designing the tests that decide whether a knowledge base is fit to ship.
- Control-row methodology that separates genuine model failure from grading error
- Edge-case suites probing the boundary of the knowledge base: questions with no answer in source, cross-entry retrieval, non-English input
- A 100-question suite for a theme park deployment returned 100 of 100 answers traceable to source with zero fabricated facts
- A 15-question edge suite found three defects that a 100-question suite of answerable questions had missed, and removed two planned prompt rules that proved unnecessary
- Audited my own test design after finding 24 of 30 recorded failures were errors in my expected answers, not assistant errors