Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Terminology-driven mining of biomedical literature

Nenadic, Goran, Spasic, Irena ORCID: https://orcid.org/0000-0002-8132-3885 and Ananiadou, Sophia 2003. Terminology-driven mining of biomedical literature. Bioinformatics 19 (8) , pp. 938-943. 10.1093/bioinformatics/btg105

Full text not available from this repository.

Abstract

MOTIVATION: With an overwhelming amount of textual information in molecular biology and biomedicine, there is a need for effective literature mining techniques that can help biologists to gather and make use of the knowledge encoded in text documents. Although the knowledge is organized around sets of domain-specific terms, few literature mining systems incorporate deep and dynamic terminology processing. RESULTS: In this paper, we present an overview of an integrated framework for terminology-driven mining from biomedical literature. The framework integrates the following components: automatic term recognition, term variation handling, acronym acquisition, automatic discovery of term similarities and term clustering. The term variant recognition is incorporated into terminology recognition process by taking into account orthographical, morphological, syntactic, lexico-semantic and pragmatic term variations. In particular, we address acronyms as a common way of introducing term variants in biomedical papers. Term clustering is based on the automatic discovery of term similarities. We use a hybrid similarity measure, where terms are compared by using both internal and external evidence. The measure combines lexical, syntactical and contextual similarity. Experiments on terminology recognition and clustering performed on a corpus of MEDLINE abstracts recorded the precision of 98 and 71% respectively. AVAILABILITY: software for the terminology management is available upon request.

Item Type: Article
Status: Published
Schools: Computer Science & Informatics
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Q Science > QH Natural history > QH301 Biology
Uncontrolled Keywords: Spasic; Text Mining; Biologists; Literature Mining; Biology; Biomedical Sciences; Biomedicine
Publisher: Oxford University Press
ISSN: 1460-2059
Related URLs:
Last Modified: 17 Oct 2022 09:56
URI: https://orca.cardiff.ac.uk/id/eprint/6223

Citation Data

Cited 32 times in Scopus. View in Scopus. Powered By Scopus® Data

Actions (repository staff only)

Edit Item Edit Item