Computing, School of

School of Computing: Faculty Publications

Syntactic Segmentation and Labeling of Digitized Pages from Technical Journals

Mukkai Krishnamoorthy, Rensselaer Polytechnic Institute
George Nagy, Rensselaer Polytechnic InstituteFollow
Sharad C. Seth, University of Nebraska-LincolnFollow
Mahesh Viswanathan, IBM Pennant Systems, Boulder, CO

Document Type

Article

Date of this Version

1993

Comments

Abstract

Alternating horizontal and vertical projection profiles are extracted from nested sub-blocks of scanned page images of technical documents. The thresholded profile strings are parsed using the compiler utilities Lex and Yacc. The significant document components are demarcated and identified by the recursive application of block grammars. Backtracking for error recovery and branch and bound for maximum-area labeling are implemented with Unix Shell programs. Results of the segmentation and labeling process are stored in a labeled X-Y tree. It is shown that families of technical documents that share the same layout conventions can be readily analyzed. More than 20 types of document entities can be identified in sample pages from the IBM Journal of Research and Development and IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE. Potential applications include preprocessors for optical character recognition, document archival, and digital reprographics.

Download

Included in

Computer Sciences Commons

COinS

Computing, School of

School of Computing: Faculty Publications

Syntactic Segmentation and Labeling of Digitized Pages from Technical Journals

Document Type

Date of this Version

Comments

Abstract

Included in

Search

Browse

Author Corner

Links

Computing, School of

School of Computing: Faculty Publications

Syntactic Segmentation and Labeling of Digitized Pages from Technical Journals

Authors

Document Type

Date of this Version

Comments

Abstract

Included in

Share

Search

Browse

Author Corner

Links