Author(s):
Shaymaa Al-tharwane and Mohammed S Al Abadie
Background: The skin prick test (SPT) is the first-line in-vivo investigation for immunoglobulin E-mediated sensitisation, but its reading is constrained by interobserver variability, subjectivity in delineating irregular reactions, and the time burden of measuring multiple wheals. Artificial intelligence (AI), particularly deep learning and computer vision, offers a potential means of standardising measurement and expanding capacity. Current gaps in validation, calibration and reporting, however, confine these systems to exploratory roles rather than deployable clinical instruments.
Aim: This review aims to critically evaluate the application of AI to skin prick testing, focusing on diagnostic and measurement performance, methodological strengths, and current limitations regarding clinical integration.
Materials and Methods: Following PRISMA 2020 guidelines, MEDLINE (PubMed) was searched for records from inception to July 2026. Eligible studies included original research involving human participants that utilised recognised AI or computational techniques, such as convolutional neural networks (CNNs), fully convolutional networks or classical image-processing pipelines, for the detection, segmentation, measurement or interpretation of SPT reactions, and reported quantifiable performance metrics.
Results: Nine reports, representing eight unique studies, met the inclusion criteria Methodologies spanned automated wheal segmentation, smartphone and three-dimensional imaging, thermography-based classification, and one deviceintegrated AI-assisted readout workflow. Segmentation performance was moderate, with Dice similarity coefficients ranging from 0.66 to 0.81, while lesion-detection sensitivity was frequently low, from 0.56 to 0.70, despite reported accuracies exceeding 0.96 that reflect class imbalance rather than clinical performance. The largest study reported 85.0% sensitivity and 98.4% specificity against physician measurement, falling to 77.4% sensitivity for allergen wheals, with a 3.7-fold faster readout. Critical limitations included the complete absence of external validation and calibration, small single-centre datasets, weak or single-reader reference standards, and no quantification of performance across skin phototypes.
Conclusion: AI demonstrates credible technical capability in wheal detection and measurement but is currently constrained by dataset heterogeneity, internal-only validation and inadequate reporting. Consequently, AI currently functions best as a decision-support system rather than an autonomous diagnostic tool, and no included system is ready for routine clinical implementation.