MapReduce in the Cloud: A Use Case Study for Efficient Co-Occurrence Processing of MEDLINE Annotations with MeSH

Kreuzthaler, Markus; Mi&#241;arro-Gim&#233;nez, Jose Antonio; Schulz, Stefan

doi:10.3233/978-1-61499-678-1-582

MapReduce in the Cloud: A Use Case Study for Efficient Co-Occurrence Processing of MEDLINE Annotations with MeSH

Authors

Markus Kreuzthaler, Jose Antonio Miñarro-Giménez, Stefan Schulz

Pages

582 - 586

DOI

10.3233/978-1-61499-678-1-582

Series

Studies in Health Technology and Informatics

Ebook

Volume 228: Exploring Complexity in Health: An Interdisciplinary Systems Approach

Abstract

Big data resources are difficult to process without a scaled hardware environment that is specifically adapted to the problem. The emergence of flexible cloud-based virtualization techniques promises solutions to this problem. This paper demonstrates how a billion of lines can be processed in a reasonable amount of time in a cloud-based environment. Our use case addresses the accumulation of concept co-occurrence data in MEDLINE annotation as a series of MapReduce jobs, which can be scaled and executed in the cloud. Besides showing an efficient way solving this problem, we generated an additional resource for the scientific community to be used for advanced text mining approaches.

This website uses cookies

This website uses cookies