A data structure perspective to the RDD-based Apriori algorithm on Spark

Pankaj Singh; Sudhakar Singh; P.K. Mishra; Rakhi Garg

Title:
A data structure perspective to the RDD-based Apriori algorithm on Spark

dc.contributor.author	Pankaj Singh
dc.contributor.author	Sudhakar Singh
dc.contributor.author	P.K. Mishra
dc.contributor.author	Rakhi Garg
dc.date.accessioned	2026-02-07T11:02:45Z
dc.date.issued	2022
dc.description.abstract	During the recent years, a number of efficient and scalable frequent itemset mining algorithms for big data analytics have been proposed by many researchers. Initially, MapReduce-based frequent itemset mining algorithms on Hadoop cluster were proposed. Although, Hadoop has been developed as a cluster computing system for handling and processing big data, but the performance of Hadoop does not meet the expectation for the iterative algorithms of data mining, due to its high I/O, and writing and then reading intermediate results in the disk. Consequently, Spark has been developed as another cluster computing infrastructure which is much faster than Hadoop due to its in-memory computation. It is highly suitable for iterative algorithms and supports batch, interactive, iterative, and stream processing of data. Many frequent itemset mining algorithms have been re-designed on the Spark, and most of them are Apriori-based. All these Spark-based Apriori algorithms use Hash Tree as the underlying data structure. This paper investigates the efficiency of various data structures for the Spark-based Apriori. Although, the data structure perspective has been investigated previously, but for the MapReduce-based Apriori, and it must be re-investigated in the distributed computing environment of Spark. The considered underlying data structures are Hash Tree, Trie, and Hash Table Trie. The experimental results on the benchmark datasets show that the performance of Spark-based Apriori with Trie and Hash Table Trie are almost similar but both perform many times better than Hash Tree in the distributed computing environment of Spark. © 2019, Bharati Vidyapeeth's Institute of Computer Applications and Management.
dc.identifier.doi	10.1007/s41870-019-00337-3
dc.identifier.issn	25112104
dc.identifier.uri	https://doi.org/10.1007/s41870-019-00337-3
dc.identifier.uri	https://dl.bhu.ac.in/bhuir/handle/123456789/41505
dc.publisher	Springer Science and Business Media B.V.
dc.subject	Apriori
dc.subject	Big data analytics
dc.subject	Frequent itemset mining
dc.subject	Parallel and distributed algorithms
dc.subject	RDD
dc.subject	Spark
dc.title	A data structure perspective to the RDD-based Apriori algorithm on Spark
dc.type	Publication
dspace.entity.type	Article

Collections

2022

Title: A data structure perspective to the RDD-based Apriori algorithm on Spark

Files

Collections

Title:
A data structure perspective to the RDD-based Apriori algorithm on Spark