232 studies, and 1.3% recorded skin type
A systematic review found AI detecting skin cancer at 90% accuracy across 232 studies. Almost none of those studies recorded who the patients were, and the ones that checked found performance dropping to chance.
TL;DR. A 2023 systematic review of 232 studies put AI accuracy for detecting skin cancer at 90%, sensitivity 87%, specificity 91%. Only 1.3% of those studies described Fitzpatrick skin type, and only 3.2% of images were type IV to VI. The headline figure is therefore uninterpretable, because for 98.7% of the literature nobody recorded the variable that would tell you who it applies to. Where anyone did measure, the drops are large. On the Diverse Dermatology Images dataset, one widely cited model's sensitivity fell from 0.69 on lighter skin to 0.23 on darker, another from 0.41 to 0.12, and one model's AUROC fell to 0.50, which is chance. A 2025 meta-analysis puts the pooled gap at 0.89 against 0.82. And training on diverse data narrows it, which makes this a recorded choice rather than a limitation.
---
Status: strong primary literature, and the central finding is an absence. Sources are peer-reviewed: a systematic review of 232 studies, the Diverse Dermatology Images benchmark work, a narrative review covering 2020 to 2025, and published evaluations of general-purpose models. The most important number here is what was not recorded rather than what was measured.
---
The figure and what it omits
A 2023 systematic review of 232 studies reported that AI detection of cutaneous malignancy averaged 90% accuracy, 87% sensitivity and 91% specificity.
Those are the numbers that circulate, and as summaries of the literature they are accurate.
The same review reports that only 1.3% of the included studies described Fitzpatrick skin type, and that only 3.2% of images were Fitzpatrick type IV to VI.
Which means the 90% describes a population nobody characterised. For roughly 229 of 232 studies, the variable that determines whether the finding transfers was not recorded at all.
This is not a claim that those studies were biased. It is a claim that they are uninterpretable on this dimension, and an aggregate assembled from them inherits the gap rather than averaging it away.
It is also the strongest instance in this corpus of the aggregate evidence gap: each study passed peer review, the systematic review was conducted properly, and the resulting number cannot be applied to any specific patient because the composition is unknown.
What happens when somebody measures
The Diverse Dermatology Images dataset was built to benchmark algorithm performance across skin types, and evaluating established models against it produced the field's clearest results.
One widely cited model showed sensitivity of 0.69 on lighter skin and 0.23 on darker, a nearly threefold difference.
Another fell from 0.41 to 0.12.
In AUROC terms, one model dropped from 0.72 on Fitzpatrick I to II to 0.57 on Fitzpatrick V to VI. Another dropped to 0.50, which is chance.
A 2025 meta-analysis puts the pooled figures at 0.89 for Fitzpatrick I to III against 0.82 for IV to VI.
The meta-analytic gap is much smaller than the individual-model gaps, which is what pooling does, and the seven-point aggregate conceals cases where a specific deployed model performs no better than guessing.
A patient is treated by one model, not by a pooled estimate.
Where it comes from
The cause is documented and unglamorous: the training data.
The International Skin Imaging Collaboration dataset and other major benchmarks are heavily skewed toward fair skin, in places exceeding 70% representation, with particular underrepresentation of Fitzpatrick types V and VI among malignant lesions, which is the category that matters most.
Fitzpatrick 17k, a widely used clinical image dataset with 114 skin conditions, has significantly fewer dark skin images than light, and an imbalance of condition labels across skin types on top of that.
One commercial tool was reported to include 2.7% Fitzpatrick type V and a single instance of type VI.
Models trained predominantly on lighter skin learn the features that distinguish disease in lighter skin. That is not a defect in the training method. It is the training method working correctly on the data it was given.
And the same physical difficulty affects human clinicians, who report harder diagnosis on darker skin due to pigmentation contrast, which means the AI gap partly reflects a pre-existing clinical gap rather than creating one.
The general-purpose models are worse
Published evaluations of general models are more concerning than the specialist ones, and matter more because patients use them directly.
GPT-4 evaluated on 50 Fitzpatrick 17k images gave the correct diagnosis for 44% of lighter-skin images and 12% of darker-skin images, a difference reaching statistical significance in a small sample. With each unit increase on the Fitzpatrick scale, accuracy fell 11.4% for differential diagnosis and 7.1% for correct diagnosis.
Overall accuracy was 28%, with the correct answer appearing among the top three differentials in 48% of cases.
A separate evaluation of ChatGPT-4o against the Diverse Dermatology Images dataset found significantly lower sensitivity, specificity and accuracy for melanoma in darker skin tones, with the gradient appearing in melanoma detection within the top three differentials.
Both studies gave the model images without clinical context, which is a limitation the authors state and which makes these lower-bound rather than realistic estimates.
The relevant comparison is not to a dermatologist. These models are freely available, patients reach them before they reach a clinic, and the alternative for many is a web search.
And the generated images repeat it
A 2026 study generated 4,000 images across the 20 most prevalent dermatologic conditions using four text-to-image models, with a standardised prompt, rated by two independent raters against US Census distributions.
89.8% of the images showed lighter skin. 10.2% showed darker skin.
Three of the four models significantly underrepresented darker skin: 3.9%, 6.0% and 8.7%, all at P below .001. One model produced 38.1%, aligned with census data and not significantly different from expected.
That one model met the bar demonstrates the others could have.
And the second finding is worse than the first. In a blinded review of a 200-image subset by dermatology residents, only 15% of images were correctly identified as the intended condition.
So the images are unrepresentative and mostly wrong, which matters because generated medical images are already used in teaching material, patient information and training data.
The fix is known and it works
Worth stating clearly, because the finding is otherwise dispiriting.
When researchers enhanced datasets with diverse skin images, the accuracy gaps narrowed significantly. A narrative review covering eight relevant papers from 2020 to 2025 reports that training with diverse datasets produced overall improvement in recognising pathology in Fitzpatrick IV to VI.
Which changes the character of the problem. This is not a limitation of the technology or an unsolved research question. It is a data collection decision, and the decision has been made the same way repeatedly.
And the reporting fix is cheaper still. Recording Fitzpatrick skin type in a study costs a column, and 98.7% of 232 studies did not.
A field that cannot state who its 90% applies to has a documentation problem before it has a fairness problem, and the documentation problem is the one that could be solved this year.
The numbers, arranged by what was measured
The literature makes more sense once the studies are sorted by whether skin type was recorded at all.
| Evidence | Skin type recorded | Finding |
|---|---|---|
| Systematic review, 232 studies | 1.3% of studies | 90% accuracy, population unknown |
| DDI benchmark, specialist models | Yes, by design | 0.69 to 0.23, and 0.41 to 0.12 sensitivity |
| Meta-analysis, 2025 | Yes | 0.89 against 0.82 pooled AUROC |
| GPT-4, Fitzpatrick 17k | Yes | 44% against 12% correct diagnosis |
| Generated images, 4 models | Yes, by design | 89.8% lighter skin, 15% correct condition |
The first row is the one that circulates and the only one that cannot be interpreted.
Every study that looked found a gap. The size varies by model and by metric, and no evaluation designed to detect a difference failed to find one.
Which makes the 1.3% figure the load-bearing one in this subject. If measuring reliably reveals a gap, then a literature that overwhelmingly did not measure is not neutral evidence. It is evidence with a known direction and an unknown magnitude.
And the honest reading of the 90% is narrower than it appears. It is a well-supported statement about AI performance on the population those 232 studies happened to contain, which was 3.2% Fitzpatrick IV to VI by image count, and that is a statement about lighter skin with a small unmeasured remainder.
What this adds to the territory
Territory 11 has found five failures that good evidence infrastructure did not prevent, and this is the sixth and the simplest.
Setting effects, transmission chains, missing category registers, inherited benchmarks and comparator choice are all subtle. Each requires understanding something about study design to see.
This one requires reading a column that is not there.
No methodological sophistication is needed to notice that 1.3% of studies recorded a variable. The systematic review reported it plainly. It sat in the same paper as the 90%, and it did not travel, which is the transmission failure from the mammography article arriving in a different form.
And it connects the territory's two halves. The evidence problem is that skin type was not recorded. The transmission problem is that when somebody did record it and reported both numbers, only one moved.
Which suggests the corrective is not more sophisticated methodology. It is the habit of asking, of any accuracy figure, on whom.
That question costs nothing, requires no statistical training, and would have caught every instance in this article. It is also, on the evidence of how the 90% circulates, not being asked.
Three things this establishes
An aggregate figure inherits an absence rather than averaging it. 90% across 232 studies, with 1.3% recording skin type, is not a 90% with unknown error bars. It is a number about an uncharacterised population, and no amount of pooling recovers the missing variable.
Pooled gaps conceal individual failures. A meta-analytic 0.89 against 0.82 sounds tolerable. A deployed model at 0.50 AUROC on darker skin is not, and patients meet models rather than meta-analyses.
And the cause is a recorded choice. Training data composition is documented, the effect of fixing it is documented, and one image model out of four met census representation, which establishes the bar was reachable.
What it does not establish
That dermatology AI does not work. The performance on lighter skin is real, and skin cancer detection at those rates is clinically valuable.
That the gap is unique to AI. Human clinicians report the same difficulty from pigmentation contrast, and the technology partly inherits a pre-existing clinical disparity.
That general-model results transfer to practice. Both LLM studies used images without clinical history, which the authors state, and real use would supply more.
And nothing about any specific product's current performance. The DDI evaluations date from 2022 and models have changed.
What is unresolved
Whether current commercial products have closed the gap. The clearest measurements are several years old and no comparable independent evaluation of current systems has been published.
Whether reporting has improved. The 1.3% figure comes from a 2023 review, and no subsequent audit establishes whether Fitzpatrick reporting has become standard.
How much of the gap is data and how much is physics. Pigmentation genuinely reduces contrast for some lesion features, and no study separates the data effect from the optical one.
And whether regulators will require it. The FDA's credibility framework turns on context of use, and a population characteristic is exactly what a context of use should specify, with no requirement currently in place.
Why this variable and not others
A reasonable objection is that no study records everything, so singling out one omission requires a reason. There is one, and it generalises.
A stratifier is worth recording when three conditions hold.
It plausibly affects the mechanism. Skin pigmentation changes the optical properties of the image a model reads, which is a physical pathway rather than a statistical association. That is different from a variable that might correlate with outcome for unknown reasons.
The population is known to be unevenly sampled. Major dermatology benchmarks exceed 70% fair skin, which was documented before most of these studies were run. An imbalance nobody disputes is a reason to stratify, not a reason to hope it averages out.
And the cost of recording is near zero. Fitzpatrick type is a six-point scale assigned by a rater in seconds. The 1.3% figure is not explained by expense.
Where all three hold, an unrecorded stratifier is a choice, and the choice was made the same way across 229 studies.
Contrast that with variables where the case is weaker. Recording every patient characteristic that might matter is impossible and would produce underpowered subgroup analyses that mislead in their own way. The argument here is not for exhaustive stratification. It is that a physically implicated, known-imbalanced, cheaply recorded variable is the specific case where the omission cannot be defended on cost or on principle.
And the test transfers. In any AI evaluation, ask whether a variable plausibly affects the mechanism, whether the sample is known to be skewed on it, and whether recording it is cheap. Where the answer is yes three times and the column is missing, the aggregate is not neutral evidence.
The counter-argument
The 1.3% figure describes historical practice, not current standards. Reporting expectations in medical AI have tightened considerably since 2023, several reporting guidelines now specify demographic disclosure, and criticising a literature for lacking a convention that arrived after most of it was written is unfair.
The DDI results are dated and were the point. Those evaluations were published precisely to force the field to improve, models have been retrained since, and citing 2022 sensitivity figures as though they describe current products is exactly the staleness this corpus criticises elsewhere.
Pooling is the right approach and the article rejects it selectively. A meta-analytic estimate exists to summarise heterogeneous evidence, and dismissing 0.89 against 0.82 in favour of the worst individual result selects the most alarming datapoint and calls it the real one.
And the human comparison cuts deeper than acknowledged. If dermatologists also perform worse on darker skin, then a model matching human performance has not created a disparity, and the appropriate comparator is the care actually available rather than an idealised standard, which is the same argument this corpus made about mental health tools and applies here.
The short version
A 2023 systematic review of 232 studies reported AI detecting skin cancer at 90% accuracy, 87% sensitivity and 91% specificity. Only 1.3% of those studies described Fitzpatrick skin type, and only 3.2% of images were type IV to VI. The headline figure describes a population nobody characterised.
Where anyone measured, the drops are large. On the Diverse Dermatology Images benchmark, one model's sensitivity fell from 0.69 to 0.23 and another's from 0.41 to 0.12. In AUROC, one fell from 0.72 to 0.57 and another to 0.50, which is chance. A 2025 meta-analysis pools the gap at 0.89 against 0.82, and a patient is treated by a model rather than by a pooled estimate.
The cause is training data. Major benchmarks exceed 70% fair-skin representation, with type V and VI particularly scarce among malignant lesions, and one commercial tool included 2.7% type V and a single type VI.
General-purpose models perform worse and matter more, because patients reach them first. GPT-4 was correct on 44% of lighter-skin images and 12% of darker, with accuracy falling 11.4% per Fitzpatrick unit for differential diagnosis.
And the generated images repeat the pattern. Across 4,000 images from four models, 89.8% showed lighter skin, three models produced 3.9%, 6.0% and 8.7% darker-skin images, one produced 38.1% and matched census data, and only 15% of images correctly depicted the intended condition.
The fix is documented and it works. Diverse training data narrows the gaps, which makes this a data collection decision rather than a research problem, and recording skin type in a study costs a column.
Common questions
What is the headline accuracy figure for AI skin cancer detection? A 2023 systematic review of 232 studies reported average accuracy of 90%, sensitivity of 87% and specificity of 91%. The same review reports that only 1.3% of those studies described Fitzpatrick skin type and only 3.2% of images were Fitzpatrick type IV to VI, which means the aggregate describes a population that was not characterised in almost the entire underlying literature.
How large is the measured performance gap? Large where anyone measured it. On the Diverse Dermatology Images dataset, built specifically to benchmark across skin types, one widely cited model showed sensitivity of 0.69 on lighter skin against 0.23 on darker, and another fell from 0.41 to 0.12. In AUROC terms one model dropped from 0.72 on Fitzpatrick I to II to 0.57 on V to VI, and another dropped to 0.50, which is chance performance. A 2025 meta-analysis pools the gap at 0.89 for Fitzpatrick I to III against 0.82 for IV to VI.
Why does the pooled figure look smaller? Because pooling averages across models and studies, which is what it is for. The consequence is that a seven-point aggregate gap conceals individual deployed models performing at chance on darker skin. A patient is treated by one model rather than by a meta-analytic estimate, so the distribution matters more than the average.
What causes it? Training data composition. Major benchmarks including the International Skin Imaging Collaboration dataset are heavily skewed toward fair skin, in places exceeding 70% representation, with particular scarcity of Fitzpatrick types V and VI among malignant lesions. One commercial tool was reported to include 2.7% type V and a single instance of type VI. Models trained predominantly on lighter skin learn the features that distinguish disease in lighter skin, which is the training process working correctly on the data supplied.
How do general-purpose models perform? Worse, and they matter because patients reach them before they reach a clinic. GPT-4 evaluated on 50 Fitzpatrick 17k images gave the correct diagnosis for 44% of lighter-skin images and 12% of darker-skin images, with accuracy falling 11.4% per Fitzpatrick unit for differential diagnosis and 7.1% for correct diagnosis, against an overall accuracy of 28%. A separate evaluation of ChatGPT-4o on the Diverse Dermatology Images dataset found significantly lower sensitivity, specificity and accuracy for melanoma in darker skin tones. Both studies supplied images without clinical context, which the authors note.
What about AI-generated dermatology images? A 2026 study generated 4,000 images across the 20 most common conditions using four text-to-image models with a standardised prompt. 89.8% featured lighter skin and 10.2% darker. Three models produced 3.9%, 6.0% and 8.7% darker-skin images, all significantly below expected proportions, while one produced 38.1% and aligned with US Census data. In a blinded review of a 200-image subset by dermatology residents, only 15% of images correctly depicted the intended condition.
Can it be fixed? Yes, and this is the most important part. Research training models on enhanced diverse datasets found the accuracy gaps narrowed significantly, and a narrative review covering 2020 to 2025 reports overall improvement in recognising pathology in Fitzpatrick IV to VI after diverse training. That makes the gap a data collection decision rather than a technological limit. The reporting fix is cheaper still: recording Fitzpatrick skin type in a study costs a column, and 98.7% of 232 studies did not.
What is the strongest objection to this article? That it cites dated measurements. The clearest performance figures come from 2022 evaluations published precisely to push the field to improve, models have been retrained since, and treating those numbers as descriptions of current products is the staleness this corpus criticises elsewhere. A second objection is that the 1.3% reporting figure describes practice before demographic disclosure guidelines tightened, which makes it a criticism of a literature for lacking a convention that arrived after most of it was written.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Exploring the Diagnostic Capability of Artificial Intelligence in Dermatology for Darker Skin Tones: A Narrative Review PMC, covering 2020 to 2025 The systematic review figures: 232 studies averaging 90% accuracy, 87% sensitivity and 91% specificity, with only 1.3% describing Fitzpatrick skin type and only 3.2% of images type IV to VI, and the finding that diverse training data improves recognition in darker skin.
- Assessing GPT-4's Diagnostic Accuracy with Darker Skin Tones medRxiv preprint The Fitzpatrick 17k evaluation: 44% correct diagnosis on lighter skin against 12% on darker, overall accuracy of 28%, and accuracy falling 11.4% per Fitzpatrick unit for differential diagnosis and 7.1% for correct diagnosis.
- Performance Evaluation of ChatGPT-4o in Dermatological Diagnoses Across Fitzpatrick Skin Types PMC The Diverse Dermatology Images evaluation finding significantly lower sensitivity, specificity and accuracy for melanoma in darker skin tones, with images supplied without clinical context.
- AI-Generated Dermatologic Images Show Poor Diagnostic Accuracy, Underrepresent Skin of Color Patient Care Online, on Joerg et al. The 4,000-image study across four models: 89.8% lighter skin against 10.2% darker, three models at 3.9%, 6.0% and 8.7%, one at 38.1% matching census data, and only 15% of images correctly depicting the intended condition.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a widely deployed system whose aggregate performance required an independent study nobody was obliged to run. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Aggregate Evidence Gap
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how a category's composition determines what its aggregate can support. :: https://arxiv.org/abs/2111.15366 Aggregate Evidence Gap
Related articles
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.
- One trial, a waitlist control, and a letterThe best evidence for AI mental health support is a single randomised trial of a purpose-built clinical tool. Its own journal published three methodological objections, and almost nobody uses the thing that was tested.
- The framework that exists produced 1.6%The FDA has two AI tracks. One is final, has authorised over 1,350 devices, and is the regime under which almost none of them cite a trial. The other missed its own deadline five weeks ago.
- The flags land on lower prior attainmentDetection tools misclassify human writing in 10 to 20% of cases, and an analysis of 10,725 assessments found the flags falling disproportionately on younger students, male students and those with weaker prior results.