Using N-Grams for Arabic Text Searching.

Saved in:
Bibliographic Details
Title: Using N-Grams for Arabic Text Searching.
Authors: Mustafa, Suleiman H.1 smustafa@yu.edu.jo, Al-Radaideh, Qasem A.1 qasemr@yu.edu.jo
Source: Journal of the American Society for Information Science & Technology. Sep2004, Vol. 55 Issue 11, p1002-1007. 6p.
Subjects: Orthography & spelling, Spelling errors, Affixes (Grammar), Semantics, Information theory, Vocabulary
Abstract: Word variation is one of the major challenges involved in free text searching. The most common types of variation that are encountered in textual databases are affixes, multiword concepts, spelling errors, alternative spellings, transliteration, and abbreviations. Several conflation techniques have been devised to handle these variations. As defined in the literature, conflation is the act of bringing together nonidentical textual words that are semantically related and reducing them to a controlled or single form for retrieval purposes. Conflation techniques can be classified as being one of two major approaches: traditional and nontraditional. Conflation has traditionally been performed by means of a comprehensive thesaurus. A thesaurus provides a precise and controlled vocabulary specifying the words and concepts of a given subject domain together with their various conceptual and morphological relationships that are to be used for indexing and searching. Modern algorithmic conflation approaches, on the other hand, have emerged in response to the need for reducing the labor and cost involved in the manual generation of a carefully designed, reliable thesaurus and in response to the skepticism raised over the possibility of fully automating this process.
Database: Engineering Source
Description
Abstract:Word variation is one of the major challenges involved in free text searching. The most common types of variation that are encountered in textual databases are affixes, multiword concepts, spelling errors, alternative spellings, transliteration, and abbreviations. Several conflation techniques have been devised to handle these variations. As defined in the literature, conflation is the act of bringing together nonidentical textual words that are semantically related and reducing them to a controlled or single form for retrieval purposes. Conflation techniques can be classified as being one of two major approaches: traditional and nontraditional. Conflation has traditionally been performed by means of a comprehensive thesaurus. A thesaurus provides a precise and controlled vocabulary specifying the words and concepts of a given subject domain together with their various conceptual and morphological relationships that are to be used for indexing and searching. Modern algorithmic conflation approaches, on the other hand, have emerged in response to the need for reducing the labor and cost involved in the manual generation of a carefully designed, reliable thesaurus and in response to the skepticism raised over the possibility of fully automating this process.
ISSN:15322882
DOI:10.1002/asi.20051