Using N-Grams for Arabic Text Searching.
Saved in:
| Title: | Using N-Grams for Arabic Text Searching. |
|---|---|
| Authors: | Mustafa, Suleiman H.1 smustafa@yu.edu.jo, Al-Radaideh, Qasem A.1 qasemr@yu.edu.jo |
| Source: | Journal of the American Society for Information Science & Technology. Sep2004, Vol. 55 Issue 11, p1002-1007. 6p. |
| Subjects: | Orthography & spelling, Spelling errors, Affixes (Grammar), Semantics, Information theory, Vocabulary |
| Abstract: | Word variation is one of the major challenges involved in free text searching. The most common types of variation that are encountered in textual databases are affixes, multiword concepts, spelling errors, alternative spellings, transliteration, and abbreviations. Several conflation techniques have been devised to handle these variations. As defined in the literature, conflation is the act of bringing together nonidentical textual words that are semantically related and reducing them to a controlled or single form for retrieval purposes. Conflation techniques can be classified as being one of two major approaches: traditional and nontraditional. Conflation has traditionally been performed by means of a comprehensive thesaurus. A thesaurus provides a precise and controlled vocabulary specifying the words and concepts of a given subject domain together with their various conceptual and morphological relationships that are to be used for indexing and searching. Modern algorithmic conflation approaches, on the other hand, have emerged in response to the need for reducing the labor and cost involved in the manual generation of a carefully designed, reliable thesaurus and in response to the skepticism raised over the possibility of fully automating this process. |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 14256755 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Using N-Grams for Arabic Text Searching. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Mustafa%2C+Suleiman+H%2E%22">Mustafa, Suleiman H.</searchLink><relatesTo>1</relatesTo><i> smustafa@yu.edu.jo</i><br /><searchLink fieldCode="AR" term="%22Al-Radaideh%2C+Qasem+A%2E%22">Al-Radaideh, Qasem A.</searchLink><relatesTo>1</relatesTo><i> qasemr@yu.edu.jo</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+the+American+Society+for+Information+Science+%26+Technology%22">Journal of the American Society for Information Science & Technology</searchLink>. Sep2004, Vol. 55 Issue 11, p1002-1007. 6p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Orthography+%26+spelling%22">Orthography & spelling</searchLink><br /><searchLink fieldCode="DE" term="%22Spelling+errors%22">Spelling errors</searchLink><br /><searchLink fieldCode="DE" term="%22Affixes+%28Grammar%29%22">Affixes (Grammar)</searchLink><br /><searchLink fieldCode="DE" term="%22Semantics%22">Semantics</searchLink><br /><searchLink fieldCode="DE" term="%22Information+theory%22">Information theory</searchLink><br /><searchLink fieldCode="DE" term="%22Vocabulary%22">Vocabulary</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Word variation is one of the major challenges involved in free text searching. The most common types of variation that are encountered in textual databases are affixes, multiword concepts, spelling errors, alternative spellings, transliteration, and abbreviations. Several conflation techniques have been devised to handle these variations. As defined in the literature, conflation is the act of bringing together nonidentical textual words that are semantically related and reducing them to a controlled or single form for retrieval purposes. Conflation techniques can be classified as being one of two major approaches: traditional and nontraditional. Conflation has traditionally been performed by means of a comprehensive thesaurus. A thesaurus provides a precise and controlled vocabulary specifying the words and concepts of a given subject domain together with their various conceptual and morphological relationships that are to be used for indexing and searching. Modern algorithmic conflation approaches, on the other hand, have emerged in response to the need for reducing the labor and cost involved in the manual generation of a carefully designed, reliable thesaurus and in response to the skepticism raised over the possibility of fully automating this process. |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=14256755 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/asi.20051 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 6 StartPage: 1002 Subjects: – SubjectFull: Orthography & spelling Type: general – SubjectFull: Spelling errors Type: general – SubjectFull: Affixes (Grammar) Type: general – SubjectFull: Semantics Type: general – SubjectFull: Information theory Type: general – SubjectFull: Vocabulary Type: general Titles: – TitleFull: Using N-Grams for Arabic Text Searching. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Mustafa, Suleiman H. – PersonEntity: Name: NameFull: Al-Radaideh, Qasem A. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: Sep2004 Type: published Y: 2004 Identifiers: – Type: issn-print Value: 15322882 Numbering: – Type: volume Value: 55 – Type: issue Value: 11 Titles: – TitleFull: Journal of the American Society for Information Science & Technology Type: main |
| ResultId | 1 |