Data Scientist · ML Researcher · LLM/NLP · Liverpool, UK
I'm a data scientist and machine learning researcher who spends most days elbow deep in messy real world data clinical time series, environmental exposure records, behavioural cohorts and increasingly, in language itself: building NLP pipelines and exploring what LLMs actually understand versus what they're good at sounding like they understand. I care about models that are accurate, explainable, and a little kinder to the world they touch. Currently applying to PhD programmes.
the (tidy) corner where I keep my models honest
I am a data scientist and machine learning researcher with practical experience building models on complex, multi source datasets including environmental exposure data, clinical time-series, and large scale behavioural cohorts.
Across four independent research preprints, I have developed strong instincts for making models not just accurate but interpretable and actionable. My work spans spatial epidemiology, clinical risk prediction, and ICU patient monitoring and increasingly, NLP and large language models, from retrieval augmented pipelines over clinical literature to auditing bias in real hiring language.
I am drawn to doctoral research that applies these skills to sustainability and systems level problems building tools that help researchers and policymakers see complex interactions in real time, not just in reports.
output, newest first — tap a card to open the record
SHAP-driven clinical risk stratification with class imbalance handling.
A mathematical index for personalised behavioural guidance.
Explainable early prediction using routine vital signs.
Local bounded imputation for ordered numerical sequences.
some more successful than others — tap a card for the full story
99.5% accuracy — but was it looking at the right thing?
VGG16 vs EfficientNetB0 on 923 images.
Scratch CNNs vs ResNet50 transfer learning.
123,842 postings, one lexicon, some uncomfortable truths.
Teaching a model to actually look at the animal.
A classroom study on copying vs understanding.
Parkinson's & Alzheimer's literature, made queryable.
13 urinary biomarkers, cross-validated AUC 0.985.
Spatial epidemiology across English regions.
Atmospheric data, binary prediction.
XGBoost, SHAP, and a clean AUC-ROC.
100,000+ users, 88% drop-off, one MSc thesis.
the paper trail
what's actually in the toolbox
Python (Advanced), R, SQL
XGBoost, Random Forest, SVM, LSTM, Autoencoders, SHAP, SMOTE/ADASYN, Survival Analysis, MICE/KNN Imputation
GeoPandas, Spatial Data Linkage, Epidemiological Methods, Exposure Analysis
TensorFlow, PyTorch, Scikit-learn, Hugging Face Transformers, Retrieval-Augmented Generation (RAG), NumPy, Pandas, SciPy
Matplotlib, Seaborn, Power BI, Amplitude
Jupyter, Git/GitHub, Google Colab, Docker, Excel (Advanced)