Search
2006 Volume 21
Article Contents
RESEARCH ARTICLE   Open Access    

Partitioning strategies for distributed association rule mining

More Information
  • In this paper a number of alternative strategies for distributed/parallel association rule mining are investigated. The methods examined make use of a data structure, the T-tree, introduced previously by the authors as a structure for organizing sets of attributes for which support is being counted. We consider six different approaches, representing different ways of parallelizing the basic Apriori-T algorithm that we use. The methods focus on different mechanisms for partitioning the data between processes, and for reducing the message-passing overhead. Both ‘horizontal’ (data distribution) and ‘vertical’ (candidate distribution) partitioning strategies are considered, including a vertical partitioning algorithm (DATA-VP) which we have developed to exploit the structure of the T-tree. We present experimental results examining the performance of the methods in implementations using JavaSpaces. We conclude that in a JavaSpaces environment, candidate distribution strategies offer better performance than those that distribute the original dataset, because of the lower messaging overhead, and the DATA-VP algorithm produced results that are especially encouraging.
  • 加载中
  • Cite this article

    FRANS COENEN, PAUL LENG. 2006. Partitioning strategies for distributed association rule mining. The Knowledge Engineering Review. 21:786 doi: 10.1017/S0269888906000786
    FRANS COENEN, PAUL LENG. 2006. Partitioning strategies for distributed association rule mining. The Knowledge Engineering Review. 21:786 doi: 10.1017/S0269888906000786

Article Metrics

Article views(17) PDF downloads(107)

Other Articles By Authors

RESEARCH ARTICLE   Open Access    

Partitioning strategies for distributed association rule mining

The Knowledge Engineering Review  21 Article number: 10.1017/S0269888906000786  (2006)  |  Cite this article

Abstract: In this paper a number of alternative strategies for distributed/parallel association rule mining are investigated. The methods examined make use of a data structure, the T-tree, introduced previously by the authors as a structure for organizing sets of attributes for which support is being counted. We consider six different approaches, representing different ways of parallelizing the basic Apriori-T algorithm that we use. The methods focus on different mechanisms for partitioning the data between processes, and for reducing the message-passing overhead. Both ‘horizontal’ (data distribution) and ‘vertical’ (candidate distribution) partitioning strategies are considered, including a vertical partitioning algorithm (DATA-VP) which we have developed to exploit the structure of the T-tree. We present experimental results examining the performance of the methods in implementations using JavaSpaces. We conclude that in a JavaSpaces environment, candidate distribution strategies offer better performance than those that distribute the original dataset, because of the lower messaging overhead, and the DATA-VP algorithm produced results that are especially encouraging.

    • © 2006 Cambridge University Press
  • About this article
    Cite this article
    FRANS COENEN, PAUL LENG. 2006. Partitioning strategies for distributed association rule mining. The Knowledge Engineering Review. 21:786 doi: 10.1017/S0269888906000786
    FRANS COENEN, PAUL LENG. 2006. Partitioning strategies for distributed association rule mining. The Knowledge Engineering Review. 21:786 doi: 10.1017/S0269888906000786
  • Catalog

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return