基于页面实体空间关系的Web对象抽取

Object Extraction Based on Spatial-Relation of Entities from the World Wide Web

  • 摘要: 针对Web同一对象内部信息组件之间的空间距离小于不同对象之间信息组件之间的距离这一显示特征. 提出一种新的Web对象抽取方法. 通过分析给定页面中不同实体间的空间位置关系来判断哪些信息成分属于同一对象,与Web文档的表示无关. 通过Web页的文档对象模型(DOM)获得不同信息成分之间的位置关系,进而判断这些信息组件是否属于同一对象. 实验结果表明,该方法对于多个领域中不同结构的Web文档具有很好的适应性. 对于设计结构规则,含有多个数据对象的页面,抽取结果的准确率可以达到100%.

     

    Abstract: The spatial distance between components within one object is always less than that between different objects in Web pages. A novel method of object extraction from the World Wide Web is reported. This proposed method considers mainly the layout characteristic of Web contents and is independent of underlying documentation representation such as HTML code. The method is based on document object model (DOM) to obtain the bounding-box of various kinds of Web information such as image, text or link. Then the distance of adjacent components is computed to get the spatial relation. Finally, all the Web information components of the same object can be integrated. Experiments showed that the proposed method could work well even when the HTML structure was far different from layout structure, and the experimental results are quite encouraging.

     

/

返回文章
返回
Baidu
map