用于重复充电运营记录的基于块采样的高效聚集查询算法

Efficient block-based sampling algorithm for aggregation query processing on duplicate charged records

下载PDF

导出

摘要现有查询分析方法通常将实体识别作为线下预处理过程清洗整个数据集,然而,随着数据规模的不断增大,这种高计算复杂性的线下清洗模式已经很难满足实时性分析应用的需求。针对重复充电运营记录上的聚集查询问题,提出一种将近似聚集查询处理与实体识别相结合的方法。首先,通过基于块的采样策略采集样本;然后,在采集到的样本上利用实体识别方法识别出重复的实体;最后,根据实体识别的结果重构得到聚集结果的无偏估计。所提方法避免了识别全部实体的时间代价,通过识别少量样本数据即可返回满足用户需求的查询结果。真实数据集和合成数据集上的实验结果验证了所提方法的高效性和可靠性。 The existing query analysis methods usually treat the entity resolution as an offline preprocessing process to clean the whole data set. However, with the continuous increasing of data size, such offline cleaning mode with high computing complexity has been difficult to meet the needs of real-time analysis in most applications. In order to solve the problem of aggregation query on duplicate charged records, a new method integrating entity resolution with approximate aggregation query processing was proposed. Firstly, a block-based sampling strategy was adopted to collect samples. Then, an entity recognition method was used to identify the duplicate entities on the sampled samples. Finally, the unbiased estimation of aggregated results was reconstructed according to the results of entity recognition. The proposed method avoids the time cost of identifying all entities, and returns the query results that satisfy user needs by identifying only a small number of sample data. The experimental results on both real dataset and synthetic dataset demonstrate the efficiency and reliability of the proposed method.

作者潘鸣宇张禄龙国标李香龙马冬雪徐亮 PAN Mingyu 1 , ZHANG Lu 1, LONG Guobiao 1, LI Xianglong 1, MA Dongxue 1, XU Liang 2(1. State Grid Beijing Electric Power Company, Beijing 100075, China ; 2. NARI Group, Beijing 102299, Chin)

机构地区国网北京市电力公司南瑞集团

出处《计算机应用》 CSCD 北大核心 2018年第6期1596-1600,1607,共6页 journal of Computer Applications

基金国家电网公司总部科技项目(52020116000j)~~

关键词大数据实体识别聚集查询块采样分布式计算 big data entity resolution aggregation query block sampling distributed computing

分类号 TP311 [自动化与计算机技术—计算机软件与理论]

引文网络
相关文献

参考文献3

1王宏志,樊文飞.复杂数据上的实体识别技术研究[J].计算机学报,2011,34(10):1843-1852. 被引量：19
2孙琛琛,申德荣,寇月,聂铁铮,于戈.面向关联数据的联合式实体识别方法[J].计算机学报,2015,38(9):1739-1754. 被引量：9
3寇月,申德荣,刘恒,王泰明,聂铁铮,于戈.异构网络中关联实体识别模型及增量式验证算法研究[J].计算机学报,2013,36(10):2096-2108. 被引量：6

二级参考文献80

1Nikki S. Gartner warns firms of "dirty data". Information Management Journal, 2007, 41 (3). http://www, allbusi ness. com/company-activities-management/operations quality-control/8901885-1. html. 被引量：1
2Kohn L T, Corrigan J M, Donaldson M S. To err is human, building a safer health system. Washington, D. C. , USA: National Academies Press, 2000. 被引量：1
3Eckerson W. Data quality and the bottom line: Achieving business success through a commitment to high quality data. The Data Warehousing Institute: Technical Report, 2002. http://download. 101com. com/pub/tdwi/Files/DQReport. pdf. 被引量：1
4Weis M, Naumann F. DogmatiX tracks down duplicates in XML//Proceedings of the ACM S1GMOD International Con ference on Management of Data. Baltimore, Maryland, USA, 2005:431 -442. 被引量：1
5Augsten N, Bohlen M H, Gamper J. Approximate matching of hierarchical data using pq-grams//Proceedings of the 31st International Conference on Very Large Data Bases. Trondheim, 2005:301-312. 被引量：1
6Ananthakrishna R, Chaudhuri S, Ganti V. Eliminating fuzzy duplicates in data warehouses//Proceedings of the 28th International Conference on Very Large Data Bases. Hong Kong, China, 2002: 586-597. 被引量：1
7Weis M, Naumann F. Detecting duplicates in complex XML data//Proceedings of the 22nd International Conference on Data Engineering. Atlanta, GA, USA, 2006:109. 被引量：1
8Weis M, Naumann F. Detecting duplicate objects in XML documents//Proceedings of the International Workshop on Information Quality in Information Systems. Paris, France, 2004:10-19. 被引量：1
9Zeng Z, Tung A K H, Wang J, Feng J, Zhou L. Comparing stars: on approximating graph edit distance. PVLDB, 2009, 2(1) : 25 -36. 被引量：1
10Cho J, Shivakumar N, Garcia-Molina H. Finding replicated Web collections//Proceedings of the 2000 ACM SIGMOD In ternational Conference on Management of Data. Dallas: Texas, USA, 2000:355-366. 被引量：1

共引文献30

1王东,牛军钰.基于多角度关联模型的实体检索方法[J].计算机工程,2013,39(1):71-75. 被引量：1
2陈爽,刁兴春,宋金玉,曹建军,丁晨路.基于伸缩窗口和等级调整的SNM改进方法[J].计算机应用研究,2013,30(9):2736-2739. 被引量：14
3寇月,申德荣,刘恒,王泰明,聂铁铮,于戈.异构网络中关联实体识别模型及增量式验证算法研究[J].计算机学报,2013,36(10):2096-2108. 被引量：6
4宋金玉,陈爽,郭大鹏,王内蒙.数据质量及数据清洗方法[J].指挥信息系统与技术,2013,4(5):63-70. 被引量：31
5祝永健,陈琛.一种新的异构网络中基于上下文相关的推荐模型[J].硅谷,2013,6(23):44-45.
6谭明超,刁兴春,曹建军.实体分辨研究综述[J].计算机科学,2014,41(4):9-12. 被引量：10
7朱灿,曹健.实体解析技术综述与展望[J].计算机科学,2015,42(3):8-12. 被引量：5
8刘文奇.中国公共数据库数据质量控制模型体系及实证[J].中国科学：信息科学,2014,44(7):836-856. 被引量：18
9官思发,孟玺,李宗洁,刘扬.大数据分析研究现状、问题与对策[J].情报杂志,2015,34(5):98-104. 被引量：73
10沈忱,曾卫明,吴爱华.融合修复代价的不一致关系数据中相似重复记录识别[J].现代计算机（中旬刊）,2015(6):3-9. 被引量：1

1刘雪莉,李建中.不一致弱可用数据近似计算可行性判定问题[J].智能计算机与应用,2018,8(2):1-6.
2迟梦园,罗朝军,董丽薇.基于O2O汽车清洗营销战略的分析与研究[J].现代商业,2017(33):21-22.
3孙玉芳.消毒供应中心实施PDCA循环对手术室腔镜器械清洗效果的影响[J].医学食疗与健康,2018,0(6):167-167. 被引量：3
4王鹏,付成龙.马鞍山公交新型洗车机顺利投用[J].人民公交,2017,0(9):15-15.
5陈听枢.谈谈对总体的估计[J].工业工程,1983(1):38-41.
6吴国发,徐哲,张燕征.具有模型检验功能的一元回归分析程序[J].计算机应用,1986,8(4):39-48.
7刘羽,王朝元,施正香,李保明.储奶罐电解水清洗除菌效果与清洗模式优选[J].农业工程学报,2017,33(20):300-306. 被引量：7
8李国安,李建峰,李穆真.二元Marshall-Olkin型指数分布的矩估计及最大似然估计[J].高等数学研究,2018,21(3):48-51.
9王子龙,陈伟杰,付强,姜秋香,印玉明,常广义.基于优先级指数的土壤采样设计方法研究[J].农业机械学报,2018,49(7):244-251. 被引量：7
10林春土.回归系统的一种有效估计方法[J].系统工程理论与实践,1985,5(3):20-24.

计算机应用

2018年第6期

浏览历史

内容加载中请稍等...

用于重复充电运营记录的基于块采样的高效聚集查询算法

参考文献3

二级参考文献80

共引文献30

相关作者

相关机构

相关主题

浏览历史