Please use this identifier to cite or link to this item: http://hdl.handle.net/1893/37221
Appears in Collections:Computing Science and Mathematics Journal Articles
Peer Review Status: Refereed
Title: Detecting Human Actions in Drone Images Using YoloV5 and Stochastic Gradient Boosting
Author(s): Ahmad, Tasweer
Cavazza, Marc
Matsuo, Yutaka
Prendinger, Helmut
Contact Email: marc.cavazza@stir.ac.uk
Keywords: action detection
YoloV5
gradient boosting classifier
Issue Date: 16-Sep-2022
Date Deposited: 10-Jul-2025
Citation: Ahmad T, Cavazza M, Matsuo Y & Prendinger H (2022) Detecting Human Actions in Drone Images Using YoloV5 and Stochastic Gradient Boosting. <i>Sensors</i>, 22 (18), Art. No.: 7020. https://doi.org/10.3390/s22187020
Abstract: Human action recognition and detection from unmanned aerial vehicles (UAVs), or drones, has emerged as a popular technical challenge in recent years, since it is related to many use case scenarios from environmental monitoring to search and rescue. It faces a number of difficulties mainly due to image acquisition and contents, and processing constraints. Since drones’ flying conditions constrain image acquisition, human subjects may appear in images at variable scales, orientations, and occlusion, which makes action recognition more difficult. We explore low-resource methods for ML (machine learning)-based action recognition using a previously collected real-world dataset (the “Okutama-Action” dataset). This dataset contains representative situations for action recognition, yet is controlled for image acquisition parameters such as camera angle or flight altitude. We investigate a combination of object recognition and classifier techniques to support single-image action identification. Our architecture integrates YoloV5 with a gradient boosting classifier; the rationale is to use a scalable and efficient object recognition system coupled with a classifier that is able to incorporate samples of variable difficulty. In an ablation study, we test different architectures of YoloV5 and evaluate the performance of our method on Okutama-Action dataset. Our approach outperformed previous architectures applied to the Okutama dataset, which differed by their object identification and classification pipeline: we hypothesize that this is a consequence of both YoloV5 performance and the overall adequacy of our pipeline to the specificities of the Okutama dataset in terms of bias–variance tradeoff.
DOI Link: 10.3390/s22187020
Rights: © 2022 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Licence URL(s): http://creativecommons.org/licenses/by/4.0/

Files in This Item:
File Description SizeFormat 
sensors-22-07020.pdfFulltext - Published Version7.59 MBAdobe PDFView/Open



This item is protected by original copyright



A file in this item is licensed under a Creative Commons License Creative Commons

Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.

The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/

If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.