Artificial Intelligence
Artificial Intelligence & Technology Systems
I work on the difficult part of AI: determining whether systems reason reliably when the subject matter demands scientific precision.
Professional AI Work
AI Systems Evaluator, Scientific Benchmarking
I evaluate advanced AI systems on complex scientific reasoning and computational problem-solving tasks. Project-specific details remain confidential under NDA.
What the work involves
- Designing and validating scientific benchmarks, reference solutions and automated evaluation frameworks.
- Building multi-step scientific-computing tasks that draw on research literature, quantitative reasoning, data analysis and reproducible code.
- Analysing model failures in reasoning, interpretation, implementation and adherence to scientific specifications.
- Quality assurance, independent validation and robustness testing so tasks are sound, reproducible and appropriately hard.
- Working with AI agents and large language models to produce structured evidence for model evaluation and improvement.
Why a clinician
Medicine is a discipline of reasoning under uncertainty, checking assumptions, and catching the confident wrong answer before it causes harm. That is exactly what good AI evaluation demands. I bring the habits of differential diagnosis to model outputs: what else could explain this, what was missed, and what would a careful expert have done differently.
Scientific Computing
I work in technical and scientific problem settings that require careful specification, reproducible testing, Python-based workflows, and disciplined interpretation of model behaviour.