German English

Distributed Holistic Clustering on Linked Data

PDF

Google Scholar

Nentwig, Markus; Groß, Anika; Möller, Maximilian; Rahm, Erhard
Distributed Holistic Clustering on Linked Data
Proc. OTM 2017 - Confederated International Conferences: CoopIS, C&TC, and ODBASE 2017, LNCS 10574, pp 371-382
2017-10

Description

Link discovery is an active field of research to support data integration in the Web of Data. Due to the huge size and number of available data sources, efficient and effective link discovery is a very challenging task. Common pairwise link discovery approaches do not scale to many sources with very large entity sets. We propose a distributed holistic approach to link many data sources based on a clustering of entities that represent the same real-world object. Our approach provides a compact and fused representation of entities, and can identify errors in existing links as well as many new links. We support distributed execution, show scalability for large real-world data sets and evaluate our methods with respect to effectiveness and efficiency for two domains.

BibTex

@inproceedings{DBLP:conf/otm/NentwigGMR17,
  author    = {Markus Nentwig and
               Anika Gro{\ss} and
               Maximilian M{\"{o}}ller and
               Erhard Rahm},
  title     = {Distributed Holistic Clustering on Linked Data},
  booktitle = {On the Move to Meaningful Internet Systems. {OTM} 2017 Conferences
               - Confederated International Conferences: CoopIS, C{\&}TC, and
               {ODBASE} 2017, Rhodes, Greece, October 23-27, 2017, Proceedings, Part
               {II}},
  pages     = {371--382},
  year      = {2017},
  url       = {https://doi.org/10.1007/978-3-319-69459-7_25},
  doi       = {10.1007/978-3-319-69459-7_25},
  timestamp = {Wed, 25 Oct 2017 18:09:20 +0200},
  biburl    = {http://dblp.org/rec/bib/conf/otm/NentwigGMR17},
  bibsource = {dblp computer science bibliography, http://dblp.org}
}