Multi-class, unsupervised detection and classification of biological and anthropogenic sounds in coral reefs
Plos.org·July 20, 2026
AI Summary
Researchers have developed an unsupervised AI method to detect and classify biological and human-made sounds in coral reefs without requiring large labeled datasets. This approach enables more practical monitoring of coral reef restoration progress by reducing the manual annotation burden typically needed for sound classification systems.
Analyzing the complex and diverse soundscapes of ecosystems such as coral reefs remains a challenge for understanding environmental dynamics and processes. While machine learning techniques can significantly improve detection and classification capabilities, applications of traditional supervised learning to underwater acoustics are limited by the size and class-coverage of labeled datasets. Unsupervised machine learning offers the potential to detect and classify sounds without the guidance of human labels, including signals that were unknown to the human analyst. However, the majority of previously developed unsupervised approaches characterize reef soundscapes from correlative metrics without identifying specific sounds, and the few that detect individual signals have been trained on limited data (<10 days), which constrains the potential to generalize across datasets and geographical localities. Here, a convolutional autoencoder was built and trained on year-long acoustic datasets from four Hawaiian coral reefs, and latent embeddings were clustered using Gaussian mixture modeling. A total of 29 classes were automatically generated, and a manual review of samples in each class determined that nine of the classes corresponded to distinct biological and anthropogenic sounds. The classes were identified to be two call types from the damselfish, Dascyllus albisella, parrotfish feeding sounds, holocentrid calls, an unidentified fish sound, three humpback whale song units, and ship noise. The classifier was found to be robust against an independently-collected test dataset with D. albisella calls (AUC = 0.9) with no extra training on the labels. Diel, lunar, and seasonal trends were observed for all nine classes, including previously-unidentified responses of the holocentrid and unknown fish groups to lunar illumination. This work demonstrates the capability of unsupervised algorithms to cluster acoustic signals into identifiable biological and anthropogenic categories in order to examine and characterize ecological trends.
Full story reconstructed from Plos.org. Formatting and media may differ from the original.
Detection and classification of biological sounds is essential for monitoring the restoration progress of coral reefs, yet most artificial intelligence (AI) methods require large, hand‑labeled training datasets, which are difficult to obtain in acoustically complex reef environments. Here we present an unsupervised AI algorithm that automatically detects and clusters sounds from acoustic recordings with no manual annotation. We applied the algorithm to four concurrent, year‑long recordings from Hawaiian coral reefs, and it automatically grouped detections into nine distinct signal types: two call types from Domino damselfish, feeding sounds from parrotfish, calls from holocentrids (squirrelfishes and soldierfishes), an unidentified fish sound, three separate humpback whale song units, and ship noise. When tested on an independent dataset with labeled Domino damselfish calls, the model correctly identified more than 82% of damselfish calls while producing false positives on fewer than 15% of non‑damselfish sounds. All nine classes exhibited patterns linked to sunlight or the seasons, and we discovered previously‑unreported responses of the holocentrid and unknown fish groups to moonrise and lunar phase. This work demonstrates that unsupervised AI can detect and cluster sound types and uncover ecological trends without the guidance of a human analyst.
Citation: Duane D, Duggan MT, Berlik E, Dantzker MS, Rice AN, Freeman LA (2026) Multi-class, unsupervised detection and classification of biological and anthropogenic sounds in coral reefs. PLoS Comput Biol 22(7): e1014516. https://doi.org/10.1371/journal.pcbi.1014516
Editor: Aidan D. Meade, Technological University Dublin, IRELAND
Received: September 5, 2025; Accepted: June 30, 2026; Published: July 20, 2026
This is an open access article, free of all copyright, and may be freely reproduced, distributed, transmitted, modified, built upon, or otherwise used by anyone for any lawful purpose. The work is made available under the Creative Commons CC0 public domain dedication.
Data Availability: All code underlying these findings, including the trained autoencoder and gaussian mixture model, is publicly available at https://doi.org/10.34740/kaggle/m/442941. The minimal data set required to replicate the study’s conclusions, consisting of the processed spectrogram samples used to train the algorithm, is publicly available at https://doi.org/10.34740/kaggle/dsv/12972413. Video clips used to verify Dascyllus albisella sounds can be found at https://www.fisheyecollaborative.org/fish-sounds/dascyllus-albisella. The raw acoustic data supporting this research were collected by the Naval Undersea Warfare Center (NUWC) Division Newport. These raw data are subject to regulatory access restrictions imposed by the U.S. Department of Defense and cannot be distributed publicly by the authors. Interested researchers may apply for access to these raw datasets by contacting the Naval Undersea Warfare Center Division Newport directly at https://www.navsea.navy.mil/Home/Warfare-Centers/NUWC-Newport/Contact-Us/.
Funding: This work was supported by the Naval Undersea Warfare Center Division Newport, Code 00X, via Internal Investment (DD, LAF), Oceankind, LLC (ANR, MSD), Schmidt Marine Technology Partners, the Schmidt Family Foundation (MSD, ANR), and the National Science Foundation Graduate Research Fellowship Program under Grant No. DGE – 2139899 (MTD). MSD received salaries from Schmidt Marine Technology Partners and Oceankind. EB received a salary from Schmidt Marine Technology Partners. MTD received a salary from the National Science Foundation Graduate Research Fellowship Program. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Coral reefs are acoustically dynamic environments, with hundreds of diverse biological sounds often contained within a single minute of acoustic data [e.g., 1 –3] alongside physical and anthropogenic noise [4,5]. These soundscapes serve as important indicators of reef health, with acoustic diversity and activity shown to increase recruitment of pelagic fish, coral, and crustacean larva [2,6–11], and to directly correlate with more resilient reefs [2,12,13]. However, to date, reef ecosystem function as inferred through soundscape metrics has largely been correlative, based on unidentified sound classes [12,14,15], band-limited sound levels [5,16–19], or acoustic indices [2,20–24], all of which lack taxonomic identification. By parsing multiple individual biological sounds and identifying them to a more precise taxonomic level, coral reef soundscapes can be interpreted to provide improved ecological context and actionable metrics for reef monitoring, conservation, and restoration. This is critical because natural resource management is primarily focused on individual species rather than sound production [3].
The application of machine learning detectors and classifiers allows for the identification of individual signals that would be missed by traditional intensity-based metrics and acoustic indices. For automated detection of fish sounds, supervised machine learning has been typically applied to large, labeled datasets focusing on a single call type, including calls from groupers [25,26], damselfishes [27], and toadfishes [28]. While effective for identifying spatial or temporal call patterns, these techniques are typically not transferable to other environments/locations, other sensors, or other call-types. Other supervised detectors have been trained on fish calls more broadly, without differentiating call-type or species [14,15].
In contrast to supervised learning, unsupervised techniques enable rapid and inexpensive signal labeling without manual annotation. Unsupervised learning reduces inter- and intra-observer variability in event classification by circumventing subjective human interpretation, which can be particularly useful for acoustically diverse datasets where individual sounds are unknown or difficult to disentangle. Previous unsupervised algorithms have been trained to identify biological choruses in long-term spectral averages with minute-scale temporal resolution [29–33], or to characterize soundscape recordings without identifying individual signals [23,24]. While a number of studies have used unsupervised clustering to detect and identify sounds associated with marine mammals [34–39], fewer studies have used unsupervised techniques to classify individual fish sounds. In a study by Noble et al. [40], handpicked spectral and temporal features were extracted from manually labeled fish calls in order to cluster them into 55 unidentified sub-groups. A study by Ozanich et al. [41] differentiated fish and whale vocalizations using two unsupervised approaches: 1) clustering of handpicked spectral and temporal features similar to Noble et al. [40], and 2) clustering of deep latent features generated by convolutional autoencoders, where the deep learning approach was found to yield significant improvements in classification accuracy. The algorithms developed by Noble et al. [40] and Ozanich et al. [41] only provided broad categorical labels and were both applied to less than ten days of data, which constrained the ability to interpret soundscape composition and resolve diel or seasonal ecological trends.
Here, an unsupervised deep clustering algorithm was trained on four concurrent, year-long acoustic datasets from Hawaiian coral reefs, using an autoencoder-Gaussian mixture model framework similar to Ozanich et al. [41]. A simple manual review of automatically-generated clusters identified nine classes corresponding to distinct acoustic signals, including two call types from the Domino damselfish, Dascyllus albisella, soniferous signatures from holocentrids (squirrelfishes and soldierfishes), an unidentified fish call, scraping from parrotfish grazing, three song units from humpback whales, and noise from ship traffic. The clustering algorithm was verified against an independently collected, human-annotated dataset containing labeled D. albisella calls with synchronized video-audio array verification (AUC = 0.9). Seasonal and diel characteristics of well-studied signals (such as humpback whale and damselfish sounds) are consistent with known behaviors, while the temporal patterns of less-studied signals reveal previously-unidentified ecological trends, including the response of the holocentrid group to lunar illumination. This work demonstrates the capability of unsupervised algorithms to derive biologically meaningful acoustic classes from unlabeled reef soundscapes at an event-level resolution and provide ecological insights through long-term analysis of species- and family-level acoustic signatures.
Long-term passive acoustic measurements were collected at four survey sites with sensor depths ranging from 20-30 m off the western coast of Hawai’i Island (Fig 1). Survey sites were selected based on prior local knowledge to capture a wide range of coastal reef settings. HTI-96-min hydrophones were deployed with Loggerhead LS1 (Survey Sites 1 and 4) or Loggerhead LS1X (Survey Sites 2 and 3) recording packages. Acoustic recorders were set to record at a sampling rate of 96 kHz, recording for 1 minute every 15 minutes for the LS1 recorders and 1 minute every 10 minutes for the LS1X recorders to accommodate the 1-year deployment duration. The sensitivity of the HTI 96-min hydrophones used here was -170 dB re V/μPa, and the frequency range was 2 Hz to 30 kHz. A synchronized video-audio array was independently deployed at a fifth location (“Test Site”, Fig 1) in order to provide ground-truth verification of calls from Dascyllus albisella (Domino damselfish). This is a well-documented vocalizing species in Hawaiian coral reefs [42–45], but few of their sounds are publicly available for model training. This passive acoustic camera combines an Insta-360 X2 360° camera within a tetrahedral hydrophone array and is described further in Dantzker et al [3]. The passive acoustic camera was deployed near nests of Dascyllus albisella and allowed for the opportunity to match sounds from specific focal fish species and their associated behaviors with vocalizations.
Red, yellow, green, and blue dots correspond to long-term single-sensor survey sites, and the black dot corresponds to the test site with a synchronized video-audio array. Bathymetric data is from the Main Hawaiian Islands Multibeam Bathymetry Synthesis (https://www.soest.hawaii.edu/hmrg/multibeam).
https://doi.org/10.1371/journal.pcbi.1014516.g001
The data analyzed from the four survey sites spans from May 1, 2020 at 00:00 local time to May 1, 2021 at 00:00 local time. Each one-minute audio sample was downsampled to a sampling frequency of 1600 Hz and run through a 4th-order highpass Butterworth filter with critical frequency 30 Hz to reduce low-frequency seismic noise. Spectrograms were generated with 32-sample fast-Fourier transforms with an overlap of 28. Frequencies between 150 and 750 Hz were isolated in order to target a frequency band of interest where diverse biological sounds are present [46].
Detections were made using the Wang & Willett power-law detector, which was chosen for its suitability for transient signals of unknown structure. The power law statistic is defined [47] as
where is the power spectral density in linear scale at time and frequency , is the background noise level, and =1.7 [47]. Here was calculated for each frequency band as the 3-second median of centered at time . The power law statistic was then smoothed using a time-domain Gaussian filter with a standard deviation of 10. This filter width was chosen to focus the detector on longer-duration sounds (on the order of the 0.36 s sample window used in subsequent analysis) and to help ensure detections remained centered within that window. Peaks were then identified at a threshold of 1 standard deviation above the one-minute mean of the filtered output. This low detection threshold was intentionally chosen so that the automated clustering algorithm would be trained to distinguish relevant signals from background or low signal to noise ratio (SNR) samples. For each detection, a spectrogram sample () centered at the detection time was generated with size 12x144, corresponding to frequency range 150–750 Hz and duration 0.36 s. Spectrogram samples were standardized according to
where and are the mean and standard deviation of across both time and frequency bins. This standardization step ensures that signals with similar temporal/spectral characteristics but different intensities will be clustered similarly. Normalized samples were then obtained by clipping according to
effectively setting the lower threshold at one standard deviation above the within-sample mean and the upper threshold at two standard deviations above the within-sample mean. Samples were clipped in this way in order to zero out background fluctuations and emphasize higher-SNR features. The detector was run on the full year of data in all four survey sites. In order to equalize detection rates between sites with different duty cycles, we discarded every third detection at Survey Sites 2 and 3, which had 1.5 times greater temporal coverage than Sites 1 and 4.
A total of 7,767,943 detections were made across the four survey sites (1,514,293 in Site 1, 1,933,200 in Site 2, 1,931,952 in Site 3, and 2,388,498 in Site 4). The detections were randomly split into a training set (90%) and a held-out validation set (10%) and then fed into a convolutional autoencoder which compresses inputs into a 16-dimensional latent space before reconstruction (Fig 2, Table 1). This dimensionality was determined by testing several configurations, selecting the smallest latent space that ensured the reconstructed outputs captured the essential spectral and temporal features of the input spectrograms. Training utilized the Adam optimizer with a learning rate of 0.0001 and a batch size of 64, minimizing the mean squared error (MSE) between input and output tensors. Training concluded automatically when the epoch-to-epoch MSE reduction fell below 1 × 10 ⁻ ⁵, requiring 35 epochs to reach convergence on the training set (Fig 3). On a machine equipped with 128 GB of RAM and two NVIDIA TITAN V GPUs, this training process was completed in less than 26 hours.
https://doi.org/10.1371/journal.pcbi.1014516.t001
The encoder compresses a 12x144 input into a 16 element latent embedding (top row). The decoder constructs a recreation of the input using only the latent embedding (bottom row).
https://doi.org/10.1371/journal.pcbi.1014516.g002
Training and validation loss (left) is nearly identical after training ends, indicating that the autoencoder is not overfitting to the training set. Randomly-selected sample input spectrograms (middle) and reconstructions (right) show the autoencoder is retaining the basic spectral and temporal features in the latent embeddings.
https://doi.org/10.1371/journal.pcbi.1014516.g003
After training, the encoder reprocessed the entirety of the dataset to extract latent feature representations. These representations were clustered using Gaussian mixture models (GMMs) with k-means++ initialization and full covariance matrices in order to account for potential correlations between latent features. To ensure clustering stability and mitigate the risk of the expectation-maximization algorithm converging to suboptimal local extrema, we employed a standard multi-start protocol, selecting the final model based on the maximum log-likelihood achieved across ten random restarts. To assess sensitivity to the number of clusters (k), we computed standard model-selection metrics including the Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Silhouette score, and Calinski–Harabasz index, across this range of candidate cluster counts. We observed that these traditional metrics for evaluating cluster count did not converge to an optimal k for this dataset (S1 Fig). This is likely a consequence of a high proportion of noisy or ambiguous samples, which are prevalent in complex coral reef soundscapes. Therefore, we focused our model selection on the portions of the data that were well-modeled by the GMM. For each candidate cluster count k, we identified “well-defined” clusters—those where >10% of assigned samples had likelihoods >0.99—and computed the Akaike Information Criterion exclusively using samples within these well-defined clusters. This approach evaluated the model’s ability to describe well-structured data, while remaining robust to poorly-fit, noisy outliers. The optimal model (k = 29) minimized this modified AIC.
To interpret the acoustic content of these unsupervised groupings, we manually inspected 100 high-likelihood samples (likelihood > 0.99) per cluster by: 1) viewing their normalized spectrograms, and 2) listening to their raw audio. This manual review step is necessary to assign biological and ecological meaning to the automatically generated clusters. We focused this review on high-likelihood samples to establish the dominant acoustic signatures of each soft cluster, acknowledging that low-likelihood samples at the cluster peripheries would likely contain more ambiguous sounds or noise, a natural consequence of the complex acoustic environment. When possible, clusters were associated to source-specific identified labels (e.g., species or vessel) based on sounds documented in the literature [e.g., 46, 48, 49] or focal observations confirmed with video [3].
Manual inspection of automatically generated clusters revealed that nine of the 29 clusters were associated with identifiable manmade and biological signals. These included two distinct damselfish sounds produced by Dascyllus albisella (“damselfish1,2”), scraping sounds from parrotfish feeding on coral (“parrotfish”), sounds from holocentrids (“holocentrid”), an unidentified fish sound (“unknown fish”), three distinct humpback whale song units (“humpback1,2,3”), and ship noise (“ship”). Sounds from holocentrid fishes were inferred based on the similarity to sounds produced by Myripristis berndti and other members of the family [50–52], but we did not feel confident about assigning species-level labels to the holocentrid class. While the identified humpback song unit clusters represent acoustically distinct signals, we did not assess whether they are functionally distinct within the context of songs they are part of. Example spectrograms from very high-likelihood (>0.999) samples in each of the characterized classes are shown in Fig 4, and example spectrograms with variable thresholds for inclusion ( > 0.9, 0.5, 0.25, and 0) are shown in S2 Fig. While consistent spectral and temporal features are generally preserved as the detection threshold is lowered, this adjustment leads to the inclusion of noisy or misclassified samples.
Audio files containing sample sounds from each class are in S1 Audio.
https://doi.org/10.1371/journal.pcbi.1014516.g004
Spectrogram samples for the uncharacterized clusters (numbered 1–20) are shown in S3 Fig. Manual review confirmed that many of these uncharacterized clusters contain a mixture of unidentified biological pulses and non-biological transient sounds, or low-SNR signals with few visible features in the spectrograms. One of these uncharacterized clusters (“18”) was found to contain a mixture of humpback whale sounds and ship tonals, indicating some confusion between these signal types.
To evaluate the statistical properties of the 29 clusters automatically generated by the Gaussian mixture model, we analyzed their internal variance, inter-cluster similarity, and classification confidence. The generalized variance for each cluster is shown in S4 Fig, where the comparatively high variance of the damselfish2 and holocentrid classes suggests they encompass a diverse range of vocalizations, and the lower variance of the parrotfish class indicates a narrower range of sound types. The probability distributions for these clusters were found to be mathematically distinct, as the Bhattacharyya coefficients for nearly all cluster pairs were below 0.36, indicating minimal overlap (S5 Fig). An analysis of likelihood scores revealed that the nine characterized sound classes are dominated by high-confidence detections, with prominent modes above a 0.99 likelihood (S6 Fig). In contrast, many of the uncharacterized clusters lack a high-confidence peak, suggesting ambiguous or noisy samples with uncertain assignments.
The distribution of detections for the characterized classes across all survey sites are shown in Fig 5. Significant variations in detection rates for each class are seen across the survey sites. Survey Site 1 saw a majority of damselfish2 call detections (51%) and a plurality of detections across all three humpback classes (40%, 39%, and 52%, respectively), as well as a strikingly low number of unknown fish detections (<0.5%). Site 2 saw a majority of unknown fish detections (67%) and a plurality of damselfish1 detections (46%). Site 3 saw a plurality of holocentrid (48%) and parrotfish (35%) detections. Site 4 was dominated by anthropogenic noise (47% of ship detections) with comparatively few biological detections (3.4% of humpback detections, 5.1% of holocentrid detections, and <15% for all other biological classes). The number of detections per class ranged from 20,000–50,000, with the exception of the damselfish1 class, where there were approximately 5,000.
https://doi.org/10.1371/journal.pcbi.1014516.g005
Detection performance for the damselfish1 and 2 classes was evaluated using the labeled dataset of verified Dascyllus albisella calls collected at the Test Site. The passive acoustic camera was deployed near a nest of Dascyllus albisella on July 27, 2022, and manual detections were made over a total of 57 minutes spanning 11:32–11:59 and 12:57–13:27 local time. 85 manual labels of the Dascyllus calls were verified by collocating beamformed acoustic detections with images of the individual from the 360° video imagery (Fig 6a). All of these sounds were associated with Dascyllus courtship/territorial displays, and video samples for four of the calls can be found at https://www.fisheyecollaborative.org/fish-sounds/dascyllus-albisella. The entire 57-minute dataset was run through the detection mechanism outlined in Section I, where 6,098 initial detections were made, 82 of which overlapped with the manual labels. The 6,098 detections were classified using the pre-trained encoder and GMM clustering algorithms, with no extra training on the labeled dataset. Detections were classified as damselfish when where and are the likelihood scores for the damselfish1 and damselfish2 classes assigned by the Gaussian mixture model, and is an adjustable detection threshold. A ROC curve is generated by varying from 0 to 1 (Fig 6b). True positive rates of >82.5% are possible with false positive rates of <15%, and the area under the curve is 0.90. To validate the choice of this clustering framework, we also implemented standard k-means clustering using the same learned latent representations to serve as a baseline comparison (S7 Fig). The Gaussian mixture model framework demonstrated improved discrimination performance relative to k-means, achieving a higher ROC AUC (0.90 vs. 0.85) and Precision-Recall AUC (0.32 vs. 0.23). This supports the use of GMM clustering for this application.