Supervised versus unsupervised approaches to classification of accelerometry data

Maitreyi Sur; Jonathan C. Hall; Joseph Brandt; Molly Astell; Sharon Poessel; Todd E. Katzner

doi:10.1002/ece3.10035

Supervised versus unsupervised approaches to classification of accelerometry data

Ecology and Evolution

By: Maitreyi Sur, Jonathan C. Hall, Joseph Brandt, Molly Astell, Sharon Poessel, and Todd E. Katzner

https://doi.org/10.1002/ece3.10035

Links

More information: Publisher Index Page (via DOI)
Data Release: USGS data release - Tri-axial acceleration data from California condors (Gymnogyps californianus), California, USA
Open Access Version: Publisher Index Page
Download citation as: RIS | Dublin Core

Abstract

Sophisticated animal-borne sensor systems are increasingly providing novel insight into how animals behave and move. Despite their widespread use in ecology, the diversity and expanding quality and quantity of data they produce have created a need for robust analytical methods for biological interpretation. Machine learning tools are often used to meet this need. However, their relative effectiveness is not well known and, in the case of unsupervised tools, given that they do not use validation data, their accuracy can be difficult to assess. We evaluated the effectiveness of supervised (n = 6), semi-supervised (n = 1), and unsupervised (n = 2) approaches to analyzing accelerometry data collected from critically endangered California condors (Gymnogyps californianus). Unsupervised K-means and EM (expectation–maximization) clustering approaches performed poorly, with adequate classification accuracies of <0.8 but very low values for kappa statistics (range: −0.02 to 0.06). The semi-supervised nearest mean classifier was moderately effective at classification, with an overall classification accuracy of 0.61 but effective classification only of two of the four behavioral classes. Supervised random forest (RF) and k-nearest neighbor (kNN) machine learning models were most effective at classification across all behavior types, with overall accuracies >0.81. Kappa statistics were also highest for RF and kNN, in most cases substantially greater than for other modeling approaches. Unsupervised modeling, which is commonly used for the classification of a priori-defined behaviors in telemetry data, can provide useful information but likely is instead better suited to post hoc definition of generalized behavioral states. This work also shows the potential for substantial variation in classification accuracy among different machine learning approaches and among different metrics of accuracy. As such, when analyzing biotelemetry data, best practices appear to call for the evaluation of several machine learning techniques and several measures of accuracy for each dataset under consideration.

Additional publication details
Publication type	Article
Publication Subtype	Journal Article
Title	Supervised versus unsupervised approaches to classification of accelerometry data
Series title	Ecology and Evolution
DOI	10.1002/ece3.10035
Volume	13
Issue	5
Publication Date	May 17, 2023
Year Published	2023
Language	English
Publisher	Wiley
Contributing office(s)	Forest and Rangeland Ecosystem Science Center
Description	e10035, 11 p.
Google Analytic Metrics	Metrics page