The Cyprus Garbage Project - Pitsilia Region

The deepTerra research tool has been used to download satellite images for the main part of the Pitsilia region. These images consist of 362,880 20m x 20m image patches, each around 200x200 pixels, and represent 280 square kilometers land mass. The machine learning models already developed (see here) have been applied to these satellite images to generate predictions of where garbage exists. The results of these have been captured in various forms: json, csv and visually in google maps html. WARNING - this is quite a big file (70MB) so will take a few moments to a few tens of seconds to load, depending on connection bandwidth. It also consumes around 2GB of memory, so may not work on older or low specced devices. Alternatively, the following three screenshots are of the google maps visualisation, at three different zoom levels.

Zoomed google maps screenshot
Zoomed out garbage map of Pitsilia region

Zoomed google maps screenshot
Partially zoomed in garbage map of Pitsilia region

Zoomed google maps screenshot
Fully zoomed in garbage map of Pitsilia region

Summary of Results

Garbage was predicted to occur in 94,288 images patches, or 26% of those surveyed. Whilst this seems alarming, the following caveats apply.
  1. The CNN architectures used (such as ResNet50) once trained on a suitable dataset, have been found to be very good at detecting any rubbish visually present in the images, including items as small as 50cm x 50cm in area (for example) a microwave or small fridge. Whilst this shows how powerful the predictive power of the approach can be, it should be noted that this is probably counter productive in terms of offering a practical starting point of addressing garbage tipping. It would be more useful if only image patches were reported where garbage dumps above a certain size were present.
  2. The deepTerra toolset supports use of random sampling and the labeling of subsets of a dataset of predictions, which gives an estimate of the predictive accuracy of the results. In this case, it was found that whilst the Recall of the results was good (above 80%), the Precision was less good (around 60%). This indicates that the model developed is very good at detecting garbage if it exists in an image patch, but also has a tendency towards False Positives: predicting garbage where none exists.
  3. The approach results in predictions of the presence/absence of garbage within a 20m x 20m image, it gives no indication of where in the 400m^2 area of the image the garbage occurs.
There are several possible ways in which the above issues might be addressed:
  • Optimizing tuning of hyperparameters such as: learning rate, batch size, number of epochs, optimizer, L1/L2 regularization, early stopping, dropout, etc.
  • Expanding on the training data and/or improving the data augmentation capabilities used to generate training data.
  • Adjusting the decision threshold to modify what is considered to be "garbage".
  • Applying heatmap algorithms such as Grad-CAM to determine how much of each image contributes to a garbage classification, and from this infer how big the area of an image consists of garbage.
  • Use of deep ensemble learning to combine several different architectures.
All of the above offer interesting areas for further research.