Computer Science and Engineering, Department of
Citation
Procs. International Conference on Document Recognition (ICDAR'11), Beijing, September 2011
Abstract
We present a method based on header paths for efficient and complete extraction of labeled data from tables meant for humans. Although many table configurations yield to the proposed syntactic analysis, some require access to semantic knowledge. Clicking on one or two critical cells per table, through a simple interface, is sufficient to resolve most of these problem tables. Header paths, a purely syntactic representation of visual tables, can be transformed (“factored”) into existing representations of structured data such as category trees, relational tables, and RDF triples. From a random sample of 200 web tables from ten large statistical web sites, we generated 376 relational tables and 34,110 subject-predicate-object RDF triples.
Included in
Computer Engineering Commons, Electrical and Computer Engineering Commons, Other Computer Sciences Commons