Advances in Non-Invasive Dermatologic Imaging and Large Language Models for Skin Cancer and Inflammatory Disease Diagnosis
Mehdi Boostani
KÁROLY RÁCZ CONSERVATIVE MEDICINE PROGRAM
Dr. Fekete Andrea
Semmelweis Egyetem, Klinikai Kórélettani Intézet, könyvtár
2026-09-22 14:00:00
Dermatology and Venereology
Dr. Sárdy Miklós
Dr. Kiss Norbert Ferenc
Dr. Budai András
Dr. Szabó Imre Lőrinc
Dr. Hamar Péter
Dr. Perge Pál
Dr. Gellén Emese
Background and aims. Skin diseases are among the most prevalent human disorders, yet many clinically important lesions remain difficult to diagnose on visual inspection alone, particularly in early stages or when benign and malignant conditions overlap in appearance. The consequences are direct. Delayed recognition of malignancy may postpone life-saving treatment, overdiagnosis leads to unnecessary biopsies and scarring, and misclassified inflammatory disease delays appropriate therapy. This thesis evaluates two complementary technologies against distinct clinical problems. 1) Non-invasive imaging for oncologic margin assessment. 2) Multimodal large language models (LLMs) for image-based diagnosis and staging across melanocytic, keratinocyte, and inflammatory conditions, comparing model behaviors across tasks to define their strengths, limitations, and potential complementary role in dermatologic practice.
Methods and findings.
•
Dermoscopy-guided high-frequency ultrasound for basal cell carcinoma margins. In a prospective single-center study, adults with suspected basal cell carcinoma had dermoscopic margins drawn 2 mm beyond visible borders, and four oriented margins were assessed with a Dermus SkinScanner against blinded histopathology. Dermoscopy-guided high-frequency ultrasound detected tumor extension within 2 mm of the excision line with 94.4% sensitivity and 93% specificity, supporting its use as a practical adjunct for tissue-sparing surgical planning.
•
Melanoma vs nevi. Across 807 histopathology-verified images, GPT-4o achieved very high sensitivity (96.5% dermoscopic, 90.6% clinical) while Gemini 2.0 Flash favored specificity (98.8% dermoscopic, 90.6% clinical); at least one model identified melanoma in 96.2% of cases. Their contrasting error profiles indicate that performance must be interpreted against the tool's clinical purpose rather than by a single summary metric.
•
Actinic keratosis vs SCC. Across 129 lesions, GPT-4o outperformed Gemini overall (accuracy 80.5% vs 66.9%; specificity 88.6%, PPV 91.5%), while Gemini reached 100% specificity but only 40% sensitivity. Both had insufficient negative predictive value (57.5% GPT-4o, 50.0% Gemini), so a non-malignant LLM output cannot safely exclude squamous cell carcinoma in real clinical practice.
•
Acne vs rosacea. Against dermatologist consensus, GPT-4o achieved 93% overall accuracy (sensitivity 90.9% acne, 100% rosacea, with near-perfect specificities), whereas Gemini responded in only 21% of cases. However, subtype classification was substantially weaker (overall sensitivity 54.6% acne, 50% rosacea), limiting reliability for therapy selection without specialist oversight.
•
Hidradenitis suppurativa. Using GPT-4o, Gemini 2.0 Flash, and Claude Sonnet 3.7, GPT-4o led on recognition (87.3% vs 32.4% Gemini, 77.5% Claude), Hurley staging (79.0% accuracy), treatment recommendations (88.7%), and reproducibility (>85%), yet staging errors persisted, notably between stages II and III. LLMs may thus assist but not replace clinical severity assessment in complex inflammatory disease.
Conclusions. The studies reveal two complementary innovation pathways in dermatology. Advanced imaging offers structural precision and objective lesion characterization, while large language models offer accessible, scalable interpretive support that is triage-oriented but not yet safe for autonomous exclusion of malignancy or definitive severity grading. Their future clinical value lies in integration, supported by broader validation datasets, disease-specific benchmarking, and careful implementation within supervised clinical workflows