Skip to content
← All work

Dermaid

Archived

Dermaid is a nine-class skin condition classifier built as a KNUST final-year project in 2024 and rebuilt as a browser demo in 2026. It is on this site less for the classifier itself than for what happened after it shipped: an audit of its own test set found that 24 percent of it was perceptually identical to training data, the reported accuracy was corrected downward once that leakage was removed, a second explanation was hypothesised for the model's strongest result, and that hypothesis was then tested and publicly retracted when the evidence didn't support it.

Role
Sole author
Period
2024 (training), 2026 (browser redeployment and audit)
Status note
Original training work from 2024. Redeployed as a static client-side demo in 2026.

Try the classifier

This runs the actual model in your browser through TensorFlow.js. Nothing you upload leaves your device. It is not a diagnostic tool, it has not been clinically validated or reviewed by a dermatologist, and public dermatology datasets skew toward lighter skin tones, so it is expected to underperform on darker skin. See a clinician if you’re worried about a skin condition.

Open the demo on Hugging Face →

The problem

A classifier trained on public dermatology datasets is only as trustworthy as its evaluation. Public dermatology data also skews toward lighter skin tones, which matters directly for a tool meant to be useful in Ghana.

The constraint

The training data was severely imbalanced, 1,751 images across nine classes with clear skin alone making up 47 percent, and it had been assembled from public sources and passed through Roboflow before Ron ever saw it. Nothing in the pipeline guaranteed the train and test splits were actually independent.

Architecture

MobileNetV2 with ImageNet weights frozen, a head of GlobalAveragePooling2D, Dense(128, ReLU), and Dense(9, softmax). Inputs are 128x128 RGB, rescaled to [0,1]. Trained with sparse categorical cross-entropy, Adam, 20 epochs, batch size 32.

It was originally served through a Flask backend. In 2026 it was converted to TensorFlow.js and redeployed as a static Hugging Face Space that runs the model entirely in the browser, so no uploaded image ever leaves the user's device.

Decisions, and what they cost

Audit the test set before trusting it

Rather than accept the shipped accuracy figure, Ron perceptually hashed all 2,076 images across both the training and test splits using a 16x16 average hash and compared them by Hamming distance. The result: 24 percent of the test set, 78 of 325 images, was perceptually identical to a training image (Hamming distance 0), 45 percent was near-identical, and 19 images were byte-identical. The worst-affected classes were chickenpox at 59 percent and shingles at 39 percent. This is an artefact of how the dataset was assembled and passed through Roboflow, not a defect in the training code.

Report the deduplicated number, not the flattering one

Re-evaluating on the 247-image remainder after removing every perceptually-identical duplicate gives 92.7 percent accuracy against a 34.0 percent majority-class baseline (always predicting clear skin), and a macro F1 of 0.895. Accuracy fell only 1.8 points once the duplicates were removed, from 94.5 to 92.7, which is evidence the model is actually generalising rather than memorising: a model that were mostly memorising would have fallen much further. Confidence also separates correct from incorrect predictions cleanly, a mean softmax confidence of 0.984 when correct against 0.771 when wrong, which means a confidence threshold could plausibly gate low-certainty predictions in a future version.

Zero of the 163 images with a genuine condition were classified as clear skin. That is the failure mode that matters most in a screening tool, since every other confusion still sends someone to a clinician, and it held at zero even after deduplication.

Test the second hypothesis before publishing it

Every clear-skin image in the dataset came from a different source than every condition-class image, at a different resolution (Roboflow, 640x640, versus 224x224 or smaller), with no exceptions in either split. That is exactly the kind of shortcut a model can learn instead of the clinical signal, and it looked like a plausible second explanation for why clear skin classifies perfectly.

Ron tested it rather than publishing it as a finding: clear-skin photographs taken from a phone and pulled from the web, carrying none of that telltale resolution, still classify correctly. The confound is real and remains uncontrolled in this dataset, but it is not the demonstrated cause of the clear-skin result, and the honest thing to do was retract the hypothesis rather than let it stand.

What didn't work

An unexplained misclassification under warm light

A forearm close-up under warm indoor lighting was confidently misclassified as cellulitis, while other clear-skin photographs classified correctly. Cellulitis presents as erythema, and warm indoor light adds a red cast to skin, so illumination sensitivity is a plausible explanation. It has not been tested, and it is left here as an open question rather than a finding.

More

Errors are clinically coherent

The confusions that remain are not random: chickenpox with shingles share a virus family, athlete's foot with cutaneous larva migrans both present on the feet, ringworm with cellulitis. That pattern of error is itself evidence the model learned something resembling the real clinical signal, not just dataset artefacts. Ringworm has the weakest recall of any class at 0.737.

Caveats that apply wherever this demo appears

This is not a diagnostic tool. It has not been clinically validated and has not been reviewed by a dermatologist. Public dermatology datasets skew toward lighter skin tones, so it is expected to underperform on darker skin, which is directly relevant to a tool aimed at a Ghanaian audience. If you are worried about a skin condition, see a clinician.

The numbers

test images perceptually identical to a training image
24%Measured. 16x16 average-hash comparison, evaluate_heldout.py

78 of 325 images, Hamming distance 0. Chickenpox and shingles were worst affected, at 59% and 39%.

accuracy, deduplicated (247 images)
92.7%Measured. evaluate_heldout.py

34.0% majority-class baseline. Macro F1 0.895.

accuracy, uncleaned test set (contaminated)
94.5%Measured. evaluate_heldout.py

The originally reported figure. Included for comparison, not used as a headline number.

genuine conditions classified as clear skin
0 of 163Measured. evaluate_heldout.py, deduplicated set
mean confidence, correct vs. incorrect
0.984 vs. 0.771Measured. evaluate_heldout.py, deduplicated set
weakest class recall (ringworm)
0.737Measured. evaluate_heldout.py, deduplicated set

Screenshots

1 / 2

Demo walkthrough

A silent, 56-second recording of the browser interface, for anyone who’d rather watch than upload an image.

Stack

  • TensorFlow / Keras
  • MobileNetV2
  • TensorFlow.js
  • Flask (original serving)
  • Hugging Face Spaces