Early Detection of Oral Cancer with Artificial Intelligence: A Systematic and Critical Review of Imaging Modalities, Learning Architectures, Evaluation Practice and Clinical Translation
Keywords:
Oral cancer, oral squamous cell carcinoma, oral potentially malignant disorders, early detection, systematic review, deep learning, histopathology, explainable artificial intelligence, external validation, screening prevalence, PRISMA.Abstract
Oral cancer is one of the few common malignancies whose precursor lesions are directly visible and directly biopsiable, and yet the majority of cases worldwide are still diagnosed at an advanced stage. Artificial intelligence has therefore been positioned as the technology that will convert a visible but unrecognised lesion into an actionable referral. This article presents a systematic and deliberately critical review of 153 records spanning 1988-2026, assembled under a PRISMA-conformant protocol and appraised with an explicit Diagnostic Evidence and Transparency Evaluation for Cancer Triage (DETECT) rubric. The corpus comprises 50 primary oral-cancer studies, 16 reviews and meta-analyses, 10 epidemiological sources, 51 methodological foundations and 26 reporting standards and metrics, and 37% of it was published in 2023-2026. We organise the field along four axes, namely input modality, learning architecture, learning regime and clinical task, and resolve the evidence density across a modality-by-task matrix that exposes a specific and consequential imbalance: histopathological binary discrimination attracts 14 records while the referral-triage task that actually governs early detection attracts far fewer, and prospective evidence is essentially absent from every cell. Our central findings are three. First, an evaluation deficit: only 34% of primary studies state a patient-wise partitioning strategy, so a large fraction of reported performance may estimate within-patient similarity rather than diagnostic generalisation. Second, a translation deficit: 16% report any external or multi-site validation, 12% compare against a clinician reading the same images, and 14% release data or code; without a clinician comparator, a reported accuracy cannot establish that a model adds anything to current practice. Third, a prevalence deficit: most studies report metrics computed on artificially balanced case-control samples, whereas screening operates at a prevalence two to three orders of magnitude lower, at which the same sensitivity and specificity yield a positive predictive value that the literature does not report. We formalise eight recurrent failure modes, derive the prevalence-corrected predictive-value relations and a screening-yield model that make the third deficit quantitative, and contribute OSCAR-Eval (Oral Screening and Cancer AI Reporting Evaluation), an eleven-stage evaluation and reporting protocol with six invariants that a reviewer can check in minutes. A three-horizon roadmap separates reporting reforms achievable now from the multi-centre benchmark and prospective-trial infrastructure that is not. All 153 records were retained only after their bibliographic identity was verified against a retrievable source.





