Document segmentation using textural features summarization and feedforward neural network.

Saved in:
Bibliographic Details
Title: Document segmentation using textural features summarization and feedforward neural network.
Authors: Oyedotun, Oyebade1 oyebade.oyedotun.k@ieee.org, Khashman, Adnan adnan.khashman@bun.edu.tr
Source: Applied Intelligence. Jul2016, Vol. 45 Issue 1, p198-212. 15p.
Subjects: Document selection, Texture analysis (Image processing), Information retrieval, Artificial neural networks, Parameters (Statistics)
Abstract: Document Segmentation is a process that aims to filter documents while identifying certain regions of interest. Generally, the regions of interest include texts, graphics (image occupied regions) and the background. This paper presents a novel top-bottom approach to perform document segmentation using texture features that are extracted from the specified/selected documents. A mask of suitable size is used to summarize textural features, and statistical parameters are captured as blocks in document images. Four textural features that are extracted from masks using the gray level co-occurrence matrix (glcm) include entropy, contrast, energy and homogeneity. Furthermore, two statistical parameters extracted from corresponding masks are the modal and median pixel values. The extracted attributes allow the classification of each mask or block as text, graphics, and background. A feedforward network is trained on the 6 extracted attributes, using documents obtained from a public database ; an error rate of 15.77 % is achieved. Furthermore, it is shown that this novel approach produces promising performance in segmenting documents and is expected to be significantly efficient for content-based information retrieval systems. Detection of duplicate documents within large databases is another potential area of application. [ABSTRACT FROM AUTHOR]
Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 117358861
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Document segmentation using textural features summarization and feedforward neural network.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Oyedotun%2C+Oyebade%22">Oyedotun, Oyebade</searchLink><relatesTo>1</relatesTo><i> oyebade.oyedotun.k@ieee.org</i><br /><searchLink fieldCode="AR" term="%22Khashman%2C+Adnan%22">Khashman, Adnan</searchLink><i> adnan.khashman@bun.edu.tr</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Applied+Intelligence%22">Applied Intelligence</searchLink>. Jul2016, Vol. 45 Issue 1, p198-212. 15p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Document+selection%22">Document selection</searchLink><br /><searchLink fieldCode="DE" term="%22Texture+analysis+%28Image+processing%29%22">Texture analysis (Image processing)</searchLink><br /><searchLink fieldCode="DE" term="%22Information+retrieval%22">Information retrieval</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Parameters+%28Statistics%29%22">Parameters (Statistics)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Document Segmentation is a process that aims to filter documents while identifying certain regions of interest. Generally, the regions of interest include texts, graphics (image occupied regions) and the background. This paper presents a novel top-bottom approach to perform document segmentation using texture features that are extracted from the specified/selected documents. A mask of suitable size is used to summarize textural features, and statistical parameters are captured as blocks in document images. Four textural features that are extracted from masks using the gray level co-occurrence matrix (glcm) include entropy, contrast, energy and homogeneity. Furthermore, two statistical parameters extracted from corresponding masks are the modal and median pixel values. The extracted attributes allow the classification of each mask or block as text, graphics, and background. A feedforward network is trained on the 6 extracted attributes, using documents obtained from a public database ; an error rate of 15.77 % is achieved. Furthermore, it is shown that this novel approach produces promising performance in segmenting documents and is expected to be significantly efficient for content-based information retrieval systems. Detection of duplicate documents within large databases is another potential area of application. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=117358861
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10489-015-0753-z
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 198
    Subjects:
      – SubjectFull: Document selection
        Type: general
      – SubjectFull: Texture analysis (Image processing)
        Type: general
      – SubjectFull: Information retrieval
        Type: general
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Parameters (Statistics)
        Type: general
    Titles:
      – TitleFull: Document segmentation using textural features summarization and feedforward neural network.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Oyedotun, Oyebade
      – PersonEntity:
          Name:
            NameFull: Khashman, Adnan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2016
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-print
              Value: 0924669X
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Applied Intelligence
              Type: main
ResultId 1