QA-Pagelet: Data preparation techniques for large-scale data analysis of the Deep Web

abstract

This paper presents the QA-Pagelet as a fundamental data preparation technique for large-scale data analysis of the Deep Web. To support QA-Pagelet extraction, we present the Thor framework for sampling, locating, and partioning the QA-Pagelets from the Deep Web. Two unique features of the Thor framework are 1) the novel page clustering for grouping pages from a Deep Web source into distinct clusters of control-flow dependent pages and 2) the novel subtree filtering algorithm that exploits the structural and content similarity at subtree level to identify the QA-Pagelets within highly ranked page clusters. We evaluate the effectiveness of the Thor framework through experiments using both simulation and real data sets. We show that Thor performs well over millions of Deep Web pages and over a wide range of sources, including e-Commerce sites, general and specialized search engines, corporate Web sites, medical and legal resources, and several others. Our experiments also show that the proposed page clustering algorithm achieves low-entropy clusters, and the subtree filtering algorithm identifies QA-Pagelets with excellent precision and recall. 2005 IEEE.

authors

Caverlee, James

published proceedings

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

author list (cited authors)

Caverlee, J., & Liu, L.

citation count

8

complete list of authors

Caverlee, J||Liu, L

publication date

September 2005

publisher

Institute of Electrical and Electronics Engineers (IEEE) Publisher

published in

IEEE Transactions on Knowledge and Data Engineering Journal

keywords

Clustering
Data Extraction
Data Preparation
Deep Web
Pagelets

Digital Object Identifier (DOI)

10.1109/TKDE.2005.151

start page

1247

end page

1262

volume

17

issue

9

URL

http://dx.doi.org/10.1109/tkde.2005.151

QA-Pagelet: Data preparation techniques for large-scale data analysis of the Deep Web Conference Paper

Overview

abstract

authors

published proceedings

author list (cited authors)

citation count

complete list of authors

publication date

publisher

published in

Research

keywords

Identity

Digital Object Identifier (DOI)

Additional Document Info

start page

end page

volume

issue

Other

URL