Theosyntra Solutions Logo

Research & Preprints

Our recommendations are grounded in active research on AI evaluation, reliability, and safety, with a focus on clinical AI. Our work is publicly available as preprints and reflects ongoing doctoral and applied research.

Why this matters for your AI

Deploying AI in healthcare, finance, and other regulated industries demands more than a working demo: it demands evidence that the system is reliable, calibrated, and safe. Because our practice is built on active research into how AI systems are evaluated, shared openly as preprints, the assessments and validations we deliver are rooted in methods we develop and document, not vendor marketing.

Preprints & Publications

Preprints are posted openly on arXiv and medRxiv ahead of peer review, so you can read the methods and judge them for yourself.

Preprint

Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors

M. O. Eniolade · Preprint, arXiv · 2026

Tests whether frontier AI agents can audit the security of clinical prediction models on their own: implementing attacks from pseudocode, scoring security posture, and reporting findings through a bash interface alone. Across three models, two datasets, and three architectures, the strongest agents completed every audit while a weaker model succeeded in 61% of attempts.

Underpins our AI Security work.

arXiv:2607.13411
Preprint

Demographic Calibration Gaps in Breast Cancer Risk Prediction

M. O. Eniolade · Preprint, medRxiv · 2026

Examines how breast cancer risk models can appear accurate in aggregate while remaining poorly calibrated for specific demographic groups: the kind of gap that surfaces only when fairness is measured directly rather than assumed.

Underpins our AI Trust & Validation work.

Preprint

Calibration, Uncertainty Communication, and Deployment Readiness in CKD Risk Prediction: A Framework Evaluation Study

M. O. Eniolade · Preprint, arXiv · 2026

Five chronic kidney disease risk classifiers reached near-perfect discrimination on internal test data, then degraded sharply on an external MIMIC-IV cohort, with calibration error rising alongside. The study argues that calibration stability and conformal coverage must be checked on external data before any clinical model moves toward deployment.

Underpins our AI Trust & Validation work.

arXiv:2605.21566

Put this research to work

Want an evidence-based assessment of whether your AI is ready for production? Let's talk.

Book a Discovery Call