USGS ScienceSearch

USGS · sir20255064

Grammar to graph—An approach for semantic transformation of annotations to triples

Abstract

Data annotation is the process of labeling data to show the outcome that a related data model should predict. In this study, annotation data were transformed into semantic graph triples, mainly for use with the Resource Description Framework (RDF), a type of entity-relationship-attribute data model for graph databases. The transformation of annotation data to semantic graph triples provides complex linguistic meaning with data handling advantages such as reduced data storage needs, improved logical specification of relations between objects, and reusable classes and properties that support logic and inference. A grammar-based framework in graph form supports user questions and queries. The words defining approximately 334 topographic feature types compiled by the U.S. Geological Survey were tokenized as units of analysis and grouped by part of speech. Their dependency relations were identified for this study using natural language processing libraries. Dependency concepts are used as structured semantic relations among part-of-speech classes. Tokens, units equivalent to words, form instances of classes and were quantified within a tabular output format using PostgreSQL data storage software. Table data were logically aligned as triples following a mapping file and stored with an ontology file using Ontop virtual triplestore software. A grammar ontology schema for the data was synchronized to match queries whose results validated the graph’s structure. The text analysis produced 8 part-of-speech classes of content words for object representations and 4 classes of function words for operational applications. Dependency relations formed 27 ontology properties for topographic subgraph structures. Token occurrences shaped overall ontology salience and formed a lexicon of syntactic terms for subgraph objects and properties. The schema ontology of class and property population shapes formed the lexicon of English terms. SPARQL Protocol and RDF Query Language (SPARQL) was used with the lexicon to conform data to RDF guidelines. This study confirms the hypothesis that although linguistic logic varies from description logic, its approximation applies to ontology design. Property and query use case patterns extracted from the analysis support queries concerning complex topographic relations and patterns normally embedded within text definitions. The method used in this study could be applied to text forms in other domains, such as survey notes.

Explore related subjects

Keep this discovery

BibTeXRIS

Dalia E. Varanka, Emily Abbott. 2025-09-02. Grammar to graph—An approach for semantic transformation of annotations to triples. https://doi.org/10.3133/sir20255064

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

Methodology for quantitative spatial sensitivity analysis of volcanic geodetic networks

Introduction This report introduces a methodology for assessing the state of the U.S. Geological Survey Volcano Observatories’ geodetic monitoring networks that measure how volcanoes deform or change shape. This new method uses a model-based approach that considers the uniqueness of the instrument environments at each volcano. This report focuses on simplified volcanic sources, is independent of the shape or size of the volcano, or the network geometry, and thus highlights the strengths and potential vulnerabilities of each volcano’s geodetic network in an actionable visual format. This analysis can help observatories to make informed decisions about whether volcanoes have an adequate level of geodetic monitoring and indicate where improvements are needed.

Lassen Peak, Mount Shasta

Assessment of water chemistry of the Coconino aquifer in northeastern Arizona

The Coconino aquifer was investigated as a potential groundwater resource for the Hopi Tribe and Navajo Nation in northeastern Arizona. Basic groundwater chemistry, including major ions, total dissolved solids, and selected trace metal concentrations, are presented and analyzed to characterize the Coconino aquifer. The geochemical compositions of groundwater are associated with changes in geology and groundwater movement and are compared to drinking-water standards to determine suitable areas for potential groundwater resource development. Dissolved-solids concentrations in much of the Coconino aquifer water were higher than the U.S. Environmental Protection Agency’s secondary drinking-water standard of 500 milligrams per liter (mg/L) due to a buried halite body in the southeastern part of the study area. However, trace metal concentrations were generally low. Groundwater may need to be treated for high dissolved-solids concentrations before it is suitable for use as a resource for the Hopi Tribe and Navajo Nation.

Arizona

Water-resources inventory and assessment at Katahdin Woods and Waters National Monument

The U.S. Geological Survey, in cooperation with the National Park Service, prepared a water-resources inventory and assessment for Katahdin Woods and Waters National Monument (KAWW). This compilation includes published and publicly accessible hydrologic data and resource assessments of streams, rivers, ponds, lakes, wetlands, vernal pools, and groundwater in and near KAWW. It also includes reports and datasets summarizing attributes of KAWW’s hydrologic infrastructure, such as stream crossings, dams, wastewater discharge plants, groundwater monitoring wells, and U.S. Geological Survey streamflow-gaging stations. Descriptions of data and details of current limitations in available datasets are included. Wetland, groundwater, streamflow, and water-quality information are all limited. Hydrography data are available; however, there are limited ground-truth data. Accurate streamlines within KAWW were developed from light detection and ranging (lidar) as a part of this work. Hydrologic infrastructure information is available from multiple sources; however, differences exist among the datasets. Datasets are summarized in appendix 1.

Maine