Research & Preprints
Our recommendations are grounded in active research on AI evaluation, reliability, and safety, with a focus on clinical AI. Our work is publicly available as preprints and reflects ongoing doctoral and applied research.
Why this matters for your AI
Deploying AI in healthcare, finance, and other regulated industries demands more than a working demo: it demands evidence that the system is reliable, calibrated, and safe. Because our practice is built on active research into how AI systems are evaluated, shared openly as preprints, the assessments and validations we deliver are rooted in methods we develop and document, not vendor marketing.
Preprints & Publications
Preprints are posted openly on arXiv and medRxiv ahead of peer review, so you can read the methods and judge them for yourself.
Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors
M. O. Eniolade · Preprint, arXiv · 2026
Tests whether frontier AI agents can audit the security of clinical prediction models on their own: implementing attacks from pseudocode, scoring security posture, and reporting findings through a bash interface alone. Across three models, two datasets, and three architectures, the strongest agents completed every audit while a weaker model succeeded in 61% of attempts.
Underpins our AI Security work.
arXiv:2607.13411Demographic Calibration Gaps in Breast Cancer Risk Prediction
M. O. Eniolade · Preprint, medRxiv · 2026
Examines how breast cancer risk models can appear accurate in aggregate while remaining poorly calibrated for specific demographic groups: the kind of gap that surfaces only when fairness is measured directly rather than assumed.
Underpins our AI Trust & Validation work.
Calibration, Uncertainty Communication, and Deployment Readiness in CKD Risk Prediction: A Framework Evaluation Study
M. O. Eniolade · Preprint, arXiv · 2026
Five chronic kidney disease risk classifiers reached near-perfect discrimination on internal test data, then degraded sharply on an external MIMIC-IV cohort, with calibration error rising alongside. The study argues that calibration stability and conformal coverage must be checked on external data before any clinical model moves toward deployment.
Underpins our AI Trust & Validation work.
arXiv:2605.21566Put this research to work
Want an evidence-based assessment of whether your AI is ready for production? Let's talk.
Book a Discovery Call