Efficient Identification of Duplicate Bibliographical References

de Melo, Vin&#237;cius Veloso; de Andrade Lopes, Alneu

Abstract

In this work we present an approach to extract and to structure bibliographical references from BibTex files, allowing the identification of the duplicate ones, which can appear slightly different in different files. To deal with this problem, existing systems use classifiers, clustering or others algorithms, allied with an Edit Distance metric, to distinguish between duplicate and nonduplicate records. The main challenge is to identify the duplicate records in database where the volume of the references can reach millions, in an efficient computational time. The technique proposed constructs a key (string) with information from each reference and stores them in a metric data structure called Slim-Tree. The Slim-Tree structure allows the minimization of the comparisons between references (being close to O(n log (n))), considering only the most similar keys to a given one.

Contact

IOS Press Copyright 2024

Contact

IOS Press Copyright 2024

This website uses cookies

This website uses cookies