USGS ScienceSearch

USGS · 70227319

Automated detection of clipping in broadband earthquake records

Abstract

Because the amount of available ground‐motion data has increased over the last decades, the need for automated processing algorithms has also increased. One difficulty with automated processing is to screen clipped records. Clipping occurs when the ground‐motion amplitude exceeds the dynamic range of the linear response of the instrument. Clipped records in which the amplitude exceeds the dynamic range are relatively easy to identify visually yet challenging for automated algorithms. In this article, we seek to identify a reliable and fully automated clipping detection algorithm tailored to near‐real‐time earthquake response needs. We consider multiple alternative algorithms, including (1) an algorithm based on the percentage difference in adjacent data points, (2) the standard deviation of the data within a moving window, (3) the shape of the histogram of the recorded amplitudes, (4) the second derivative of the data, and (5) the amplitude of the data. To quantitatively compare these algorithms, we construct development and holdout datasets from earthquakes across a range of geographic regions, tectonic environments, and instrument types. We manually classify each record for the presence of clipping and use the classified records. We then develop an artificial neural network model that combines all the individual algorithms. Testing on the holdout dataset, the standard deviation and histogram approaches are the most accurate individual algorithms, with an overall accuracy of about 93%. The combined artificial neural network method yields an overall accuracy of 95%, and the choice of classification threshold can balance precision and recall.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

James Kael Kleckner, Kyle B. Withers, Eric M. Thompson, J.M. Rekoske, Emily Wolin, Morgan P. Moschetti. 2021-12-22. Automated detection of clipping in broadband earthquake records. https://doi.org/10.1785/0220210028

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Digitizer Suite: The Albuquerque Seismological Laboratory Digitizer Testing Suite

Laboratory testing of digitizers and seismometers helps ensure that prior to deployment the instrumentation can produce high quality data and is operating within specifications. In this work we detail the software package called: the Albuquerque Seismological Laboratory (ASL) Digitizer Test Suite. This Java software package provides several algorithms to verify various performance parameters of digitizers commonly used for recording analog seismic instruments. The goal of these tests is not to be exhaustive, but to identify common failures that could compromise the integrity of seismic data being recorded on the digitizer. For example, Sandia National Laboratories (e.g., Slad and Merchant, 2018) routinely do comprehensive testing of digitizers for various monitoring missions. While these tests reports are valuable for comprehensively characterizing a recording system, it would be resource intensive to conduct such tests on every seismic recorder used in a network. We focus on tests that include ways to estimate the sensitivity, timing, self-noise, and clip-level of the digitizer, as well as the fidelity of the signal being recorded. The software is publicly available and provides a way for the community to verify the integrity of a digitizer using a minimum amount of outside equipment.

Seismological Research Letters

The digital archivist: Automating legacy macroseismic data processing using large language models

Macroseismic data are a key resource to investigate shaking and damage from preinstrumental and early instrumental eras. However, data are often stored as inconsistently formatted reports describing observed shaking and damage, making manually parsing and interpreting accounts labor‐intensive. We introduce a novel workflow using Google’s Gemini 2.5 Pro large language model (LLM) to automate the extraction and structuring of macroseismic observations from summary reports. We apply this workflow to the 22 March 1957 M 5.3 Daly City, California, earthquake as a case study. We used Gemini to extract addresses, originally assigned modified Mercalli intensity values, and descriptions from each report. To address coordinate precision limits, addresses were geocoded via Google’s Geocoding application programming interface. This workflow yielded over 2300 geocoded intensity reports for the Daly City earthquake. We use the geocoded accounts, with the original report intensity assignments, to develop a shaking intensity map that in some respects rivals modern Did You Feel It? Maps. We also extract and present data for the 9 February 1971 M L 6.7 Sylmar, California, earthquake. Our results demonstrate the potential of LLMs for reliably extracting and analyzing large, unstructured macroseismic datasets. LLMs offer a scalable solution for rapidly digitizing macroseismic archives, enabling their broader use to constrain ground‐motion models in modern seismic hazard analysis and to improve our understanding of site effects in urban areas. The concepts explored here may also be applied to the handling of other legacy seismological and earth science data.

Seismological Research Letters