期刊文献+

一种稳定的并行分布式频繁集挖掘算法及其应用

A STABLE PARALLEL DISTRIBUTED FREQUENT ITEMSET MINING ALGORITHM AND ITS APPLICATION
下载PDF
导出
摘要 为解决大规模医药数据分析中的频繁集挖掘问题,提出一种稳定且具有良好扩展性的并行分布式算法P-FIM。该算法将挖掘任务分割成无相互依赖关系的同构子任务,实现有效的并行计算;并且充分利用Map/Reduce框架和集群环境的优势提高自身的鲁棒性和负载均衡能力。采用最大规模为512万条记录的中医药方剂数据进行算法性能分析实验,其结果表明,该算法在分布式集群环境中表现稳定,而且随着集群规模的增加其加速比接近线性。以P-FIM算法为基础设计实现的中医药数据相关性分析方案,可有效地从大规模临床数据中获得全面、可靠的病、症、药间相关性的信息。 This paper proposes P-FIM,a stable parallel distributed algorithm with good scalability,to deal with frequent itemset mining issue in large scale medicine data analysis.It divides the mining task into independent isomorphic subtasks to achieve effective parallel computation,and takes full advantage of Map/Reduce infrastructure as well as computing cluster to improve its own robustness and load balance capability.In this paper we carry out analytical experiment on performance of the P-FIM algorithm based on TCM prescription data that contain largest records up to 51.2million.The result shows that the algorithm performs stably in distributed clustering condition,and approaches linear speedup along with the augment of clustering scale.The correlation analysis scheme of traditional Chinese medicine designed and implemented based on P-FIM algorithm can effectively gain comprehensive and reliable information correlating with the disease,symptoms and medicine from large scale clinical data.
出处 《计算机应用与软件》 CSCD 2011年第3期83-85,124,共4页 Computer Applications and Software
基金 国家高技术研究发展计划项目(2006AA01A123) 杰出青年基金(NSFC60525202)
关键词 数据挖掘 频繁集挖掘 Map/Reduce并行框架 医药数据分析 Data mining Frequent itemset mining Map/Reduce parallel infrastructure Analysis of medicine data
  • 相关文献

参考文献6

  • 1Agrawal R ,Srikant R. Fast algorithms for mining association rules. Santiago Chile : Very Large Data Bases ( VLDB' 94) :487 - 499. 被引量:1
  • 2Dean J, Ghemawat S. MapReduee:Simplified data processing on large clusters [ C ]//Proc. of the 6th OSDI ( Dec. 2004 ) : 137 - 150. 被引量:1
  • 3Agrawal R, Sharer J C. Parallel mining of association rules. IEEE Transaction On Knowledge And Data Engineering. 1996(8) :962 -969. 被引量:1
  • 4Ye Y, Chiang C C. A parallel apriori algorithm for frequent itemsets mining[ C ]//SERA ,2006. 被引量:1
  • 5Liu Li, Li Eric, Zhang Yimin. Optimization of frequent itemset mining on multiple-core processor[ C]//VLDB ,2007. 被引量:1
  • 6LI H, Wang Y,Zhang D, et al. PFP: Parallel FP-Growth for Query Recommendation. ACM Recommender Systems,2008. 被引量:1

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部