Artificial Intelligence

Artificial Intelligence & Technology Systems

I work on the difficult part of AI: determining whether systems reason reliably when the subject matter demands scientific precision.

Professional AI Work

AI Systems Evaluator, Scientific Benchmarking

Confidential AI project (NDA) · Remote · to present

I evaluate advanced AI systems on complex scientific reasoning and computational problem-solving tasks. Project-specific details remain confidential under NDA.

What the work involves

  • Designing and validating scientific benchmarks, reference solutions and automated evaluation frameworks.
  • Building multi-step scientific-computing tasks that draw on research literature, quantitative reasoning, data analysis and reproducible code.
  • Analysing model failures in reasoning, interpretation, implementation and adherence to scientific specifications.
  • Quality assurance, independent validation and robustness testing so tasks are sound, reproducible and appropriately hard.
  • Working with AI agents and large language models to produce structured evidence for model evaluation and improvement.

Why a clinician

Medicine is a discipline of reasoning under uncertainty, checking assumptions, and catching the confident wrong answer before it causes harm. That is exactly what good AI evaluation demands. I bring the habits of differential diagnosis to model outputs: what else could explain this, what was missed, and what would a careful expert have done differently.

Scientific Computing

I work in technical and scientific problem settings that require careful specification, reproducible testing, Python-based workflows, and disciplined interpretation of model behaviour.