基于最大频繁等价类的Web信息自动抽取

Automatic Web Information Extraction Based on Maximal and Frenquent Equivalence Classes

下载PDF

导出

摘要在定义模板的基础上,提出了页面创建模型。该模型描述了如何使用模板将来自于后台数据库的值编码生成页面。基于这个模型,设计了一个基于最大频繁等价类的抽取算法EBMFEC,通过分析给定的数据导向型页面的终端符号的出现情况,找出最大频繁等价类,并推导出用于生成页面的未知模板。然后使用推导出的模板,从输入页面中提取出相关信息。在大量实际HTML页面上的实验证明,EBMFEC在大部分情况下都可以从给定页面中推导出模板,并正确抽取出数据信息。 A novel approach based on MFEC （Maximal and Frenquent Equivalence Classes）is proposed to solve the problem of automatically extracting data from data-intensive Web pages. A template is defined and a model of page creation is proposed to describe how values are encoded into pages using the defined template. We present an algorithm, EBMFEC that takes,as input, a set of template-generated pages, analyzes the page-tokens of given pages to discover MFEC, deduces the unknown template used to generate the pages and extracts, as output, the values encoded in the pages. Experiments on a large number of HTML pages indicate that our algorithm correctly extracts data in most cases and the results are also provided.

作者陈华昌薛永生任仲晟张东站

机构地区厦门大学计算机科学系福建生态工程职业技术学校福州

出处《计算机科学》 CSCD 北大核心 2006年第12期169-173,202,共6页 Computer Science

基金国家自然科学基金(50474033) 福建省自然科学基金(A0310008) 福建省重点科技项目(2003H043)。

关键词等价类信息抽取模式模板 Equivalence classes, Information extraction,Schema,Template

分类号 TP391.2 [自动化与计算机技术—计算机应用技术]

引文网络
相关文献

参考文献10

1Haas L M,Kossmann D, Wimmers E L,et al. Optimizing queries across diverse data sources. In: Proc of the 23th VLDB Conf.Athens, 1997. 276-285 被引量：1
2Levy A,Rajaraman A,Ordille J J. Querying heterogeneous information sources using source descriptions. In: Proc. of the 22th VLDB Conf. Bombay,1996. 251-262 被引量：1
3Kushmerik N. Wrapper induction: Efficiency and expressiveness.Journal of Artificial intelligence,2000,118(1-2): 15-68 被引量：1
4Soderland S, Learning information extraction rules for semi- structured and free text. Journal of Machine learning, 1999, 34(1-3):233-272 被引量：1
5Embley D W,Campbell D M. A conceptual-modeling approach to extracting data from the web. In:Proc. of the 17th Intl. Conf on Conceptual Modeling. Singapore,1998. 78-91 被引量：1
6Chang C, Lui S. IEPAD: Information extraction based on pattern discovery. In:Proc of 10th WWW Conf. Hong Kong, 2001. 681-688 被引量：1
7Crescenzi V, Mecca G, Merialdo P. ROADRUNNER: Towards automatic data extraction from large web sites. In: Proc of the 27th VLDB Conf. Roma,2001. 109-118 被引量：1
8Crescenzl V, Mecca G. Automatic Information Extraetlon from Large Websites. Journal of the ACM, 2004,51 (5): 731-779 被引量：1
9Sarawagi S. Automation in InformationExtraction and Data Integration (tutorial). VLDB, 2002 被引量：1
10Myllymaki J. Effective Web data extraction with standard XML technologies. In:Proc of 10th WWW Conf. Hong Kong, 2001.689-696 被引量：1

1周宝刚,刘杰,李成.基于Struts的WEB页面构建系统[J].电脑知识与技术,2008(2):695-698. 被引量：2
2舒红平,吴自恒,蒋建民.基于JAVA的WEB页面自动生成系统[J].成都信息工程学院学报,2003,18(4):343-348. 被引量：1
3任庆丽,杨曙年.远程教学系统中基于数据库的动态页面设计方法研究[J].计算机与现代化,2002(5):47-50.
4陈杰.CAD/CAM网站中动态页面的设计[J].煤矿机械,2004,25(12):84-85.
5李磊.P4P浅谈[J].新课程研究（职业教育）,2008(12):83-83. 被引量：2
6贾洋洋,蒋泽军,王丽芳.基于XML的组件系统的设计与实现[J].科学技术与工程,2009,9(2):441-445. 被引量：3
7彭珍瑞,殷红,刘颖.基于Web和Matlab的控制系统仿真实验平台的开发[J].科技信息,2009(27). 被引量：2
8服务园地·问答与讨论信箱[J].现代铸铁,2013,33(3):102-105.
9张如云.计算机常见故障排除的九大原则[J].金融科技时代,2012,20(4):85-85.
10李晓,邱玉辉.基于MAS的Web用户数据预处理[J].广西师范大学学报（自然科学版）,2003,21(A01):160-163. 被引量：3

计算机科学

2006年第12期

浏览历史

内容加载中请稍等...

基于最大频繁等价类的Web信息自动抽取

参考文献10

相关作者

相关机构

相关主题

浏览历史