Data Scientist · Causal & Interpretable ML Researcher · Liverpool, UK
Shweta Debjit Sarkar
I'm a data scientist and machine learning researcher who works across a wide range of real world data, clinical time series, environmental exposure records, behavioural cohorts, and increasingly, language itself, building NLP pipelines and exploring what LLMs actually understand versus what they're good at sounding like they understand. My most active current line of work is causal ML for climate prediction, causal discovery, structural break detection, and graph and diffusion based models applied to Arctic sea ice, and I bring the same instincts to whatever data I'm working with. I care about models that are accurate, explainable, and a little kinder to the world they touch. Currently applying to PhD programmes.
I am a data scientist and machine learning researcher who works across a range of complex, real world datasets, environmental exposure data, clinical time-series, and large scale behavioural cohorts, tied together by the same two questions: is this relationship causal, and can someone actually explain why the model made the call it did. My most developed current line of work applies these questions to causal ML for climate prediction, causal discovery, structural break detection, and graph and diffusion based models for Arctic sea ice.
Across four independent research preprints, I have developed strong instincts for making models not just accurate but interpretable and actionable. My work spans spatial epidemiology, clinical risk prediction, and ICU patient monitoring and increasingly, NLP and large language models, from retrieval augmented pipelines over clinical literature to auditing bias in real hiring language.
I am drawn to doctoral research that puts these methods to work on real, systems level problems, climate very much included, building tools that help researchers and policymakers see complex interactions in real time, not just in reports.
Tests whether winter Arctic Oscillation predicts September sea ice extent beyond trend and persistence. A reduced three-feature Ridge model, evaluated with leave-one-out cross validation, beat a persistence baseline by 23%, with a 1,000-run permutation test putting the finding at p = 0.004.
A Chow test finds a structural break in Arctic ice decline at 2007 (p = 0.0001) that a single linear trend badly misses. Correcting for the break alone cuts holdout RMSE by 37%, beating every climate variable tested as an additional predictor.
Change Point DetectionChow TestStructural BreakTime Series
Causal Feature Selection — Arctic Sea Ice
2026 · Independent Research
PCMCI causal discovery run under three specifications, raw, detrended, and forcing-adjusted, to separate genuine seasonal drivers from shared secular trend. Causally selected predictors beat a full feature set, but no model, causal or otherwise, beats trend plus persistence, consistent with the spring predictability barrier reported in the seasonal forecasting literature.
Rechecks the AO project's own finding with a different, more standard test. Once ice extent is made properly stationary, using a break-aware detrending step that independently reconfirms the 2007 break, Granger causality finds no significant AO link on its own. A follow-up check shows AO's original significance holds once persistence is controlled for, resolving rather than hiding the disagreement.
A graph neural network built from scratch in PyTorch over four Arctic regions, testing whether real spatial structure beats the regional averaging used elsewhere in this series. Two real training bugs were caught and fixed, an oversmoothing hypothesis was tested and ruled out, and the honest result, beaten by trend and persistence, mirrors the rest of this project series.
A denoising diffusion model (DDPM) built from scratch to generate a calibrated distribution of next-month global temperature anomalies, rather than a single point forecast. Calibration lands close to an 80% coverage target, though the model's central estimate overreacts to its conditioning input compared to a simple linear baseline, a finding investigated directly rather than hidden.
A larger, real-world graph neural network test, 30 major cities connected by real geographic distance. The graph initially made predictions worse than removing it entirely, isolated and confirmed with an ablation test, until a skip connection fixed the underlying signal dilution problem and beat every baseline, persistence included.
Two paired wildfire detection models exposing a critical ML lesson: a satellite classifier achieved 99.5% accuracy, but Grad-CAM revealed it learned land-use patterns, not fire damage — a textbook shortcut learning failure. A second model on real fire/smoke imagery achieved genuine 100% accuracy, confirmed by interpretability analysis.
EfficientNetB0Grad-CAMPyTorchShortcut Learning
Coral Reef Health Classification
2026
Binary classifier (healthy vs bleached coral) comparing VGG16 and EfficientNetB0 transfer learning on 923 images. VGG16 overfit severely with 119M trainable parameters, while fine-tuned EfficientNetB0 achieved 79.5% accuracy with only a 2.1% train/val gap. Grad-CAM confirms ecologically valid feature attention.
EfficientNetB0VGG16Grad-CAMTransfer Learning
Skin Lesion Classifier — CNN from Scratch
2026
7-class skin lesion classification on HAM10000, training two custom CNNs from scratch (39.5% / 42.1% balanced accuracy) and benchmarking against ResNet50 transfer learning (83.8%). Dermatofibroma recall was 0% for both scratch models vs 96% for ResNet50 — a clear case for transfer learning in medical imaging.
PyTorchResNet50HAM10000Medical Imaging
Job Posting Bias Analyser
2026
NLP pipeline detecting gender-coded, age-biased, and exclusionary language in 123,842 real LinkedIn job postings, grounded in the Gaucher et al. (2011) lexicon. Found 46% of postings lean masculine, with Tech and Venture Capital among the most biased industries.
NLPFairness in AIPythonText Analysis
Wildlife Camera Trap Detector
2025
Binary classifier (blank vs animal-present) for conservation camera trap images using fine-tuned ResNet18 on the Serengeti2 dataset. Data augmentation reduced the generalisation gap from 6.93% to 1.03%, achieving 85.11% test accuracy. Grad-CAM confirms the model attends to animal regions, not background.
ResNet18Grad-CAMPyTorchConservation AI
AI Agency in Student Learning
2026
Pilot classroom observation study (N=73, Years 7–11) across two Liverpool secondary schools. Key finding: 32% of students claimed ownership of an answer they could not explain — a gap invisible to current assessment methods, consistent across both SEN and non-SEN populations.
Education ResearchAI EthicsMixed MethodsSEN
Clinical RAG — Parkinson's & Alzheimer's
2026
Retrieval-Augmented Generation pipeline for question-answering over Parkinson's and Alzheimer's research literature, with retrieval evaluation metrics and query similarity analysis to assess answer grounding quality.
RAGNLPClinical AILLM
T2DM Urinary Metabolomics Analysis
2026
End-to-end metabolomics pipeline identifying urinary biomarkers of Type 2 Diabetes Mellitus from NMR spectroscopy data (N=132). Random Forest classification (cross-validated AUC = 0.985) surfaced 13 metabolites confirmed by both methods, consistent with published T2DM literature.
MetabolomicsExWASRandom ForestBiomarker Discovery
Air Pollution & Breast Cancer Incidence in England
2026
Spatial epidemiological analysis linking environmental exposure data to cancer outcomes across English regions using ML and geospatial data linkage.
GeoPandasSpatial AnalysisEpidemiologyPython
Rainfall Classification
2025
Binary classification of daily rainfall occurrence from multi-variable atmospheric observations — humidity, pressure, wind direction, and temperature.
ClassificationAtmospheric DataScikit-learn
Breast Cancer Wisconsin Prediction
2026
Clinical ML classification of malignant vs benign tumours using Logistic Regression, XGBoost, and Random Forest, with SHAP explainability for diagnostic feature identification.
XGBoostSHAPAUC-ROCClinical ML
MoovBuddy Cohort Analysis — MSc Thesis
2025
Longitudinal cohort analysis on 100,000+ anonymised user records across 10+ countries. Identified 88% install-to-subscription drop-off and modelled country-level variation.