Identification of related multilingual documents using ant clustering algorithms.

Saved in:
Bibliographic Details
Title: Identification of related multilingual documents using ant clustering algorithms.
Alternate Title: Identificación de documentos multilingües relacionados mediante algoritmos de clustering de hormigas.
Authors: Cobo, Ángel1 acobo@unican.es, Rocha, Rocío2 rochar@unican.es
Source: INGENIARE - Revista Chilena de Ingeniería. Dec2011, Vol. 19 Issue 3, p351-358. 8p. 4 Color Photographs, 3 Charts.
Subjects: Document clustering, Multilingual computing, Ant algorithms, Computational linguistics, Text mining, Records management
Abstract (English): This paper presents a document representation strategy and a bio-inspired algorithm to cluster multilingual collections of documents in the field of economics and business. The proposed approach allows the user to identify groups of related economics documents written in Spanish and English using techniques inspired on clustering and sorting behaviours observed in some types of ants. In order to obtain a language independent vector representation of each document two multilingual resources are used: an economic glossary and a thesaurus. Each document is represented using four feature vectors: words, proper names, economic terms in the glossary and thesaurus descriptors. The proper name identification, word extraction and lemmatization are performed using specific tools. The tf-idf scheme is used to measure the importance of each feature in the document, and a convex linear combination of angular separations between feature vectors is used as similarity measure of documents. The paper shows experimental results of the application of the proposed algorithm in a Spanish-English corpus of research papers in economics and management areas. The results demonstrate the usefulness and effectiveness of the ant clustering algorithm and the proposed representation scheme. [ABSTRACT FROM AUTHOR]
Abstract (Spanish): Este artículo presenta una estrategia de representación documental y un algoritmo bioinspirado para realizar procesos de agrupamiento en colecciones multilingües de documentos en las áreas de la economía y la empresa. El enfoque propuesto permite al usuario identificar grupos de documentos económicos relacionados escritos en español o inglés usando técnicas inspiradas en comportamientos de organización y agrupamiento de objetos observados en algunos tipos de hormigas. Para conseguir una representación vectorial de cada documento independiente del idioma, se han utilizado dos recursos lingüísticos: un glosario económico y un tesauro. Cada documento es representado usando cuatro vectores de rasgos: palabras, nombres propios, términos económicos del glosario y descriptores del tesauro. La identificación de los nombres propios y la extracción y lematización de palabras se realizan usando herramientas específicas. El esquema tf-idf es utilizado para medir la importancia de cada rasgo en el documento, y se utiliza una combinación lineal convexa de separaciones angulares de los vectores de rasgos como medida de similitud de documentos. El trabajo muestra resultados experimentales de aplicación del algoritmo propuesto sobre un corpus español-inglés de documentos científicos de áreas económica y de gestión empresarial. Los resultados demuestran la utilidad y efectividad de las técnicas de ant clustering y del esquema de representación propuesto. [ABSTRACT FROM AUTHOR]
Copyright of INGENIARE - Revista Chilena de Ingeniería is the property of Universidad de Tarapaca and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 72370736
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Identification of related multilingual documents using ant clustering algorithms.
– Name: TitleAlt
  Label: Alternate Title
  Group: TiAlt
  Data: Identificación de documentos multilingües relacionados mediante algoritmos de clustering de hormigas.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Cobo%2C+Ángel%22">Cobo, Ángel</searchLink><relatesTo>1</relatesTo><i> acobo@unican.es</i><br /><searchLink fieldCode="AR" term="%22Rocha%2C+Rocío%22">Rocha, Rocío</searchLink><relatesTo>2</relatesTo><i> rochar@unican.es</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22INGENIARE+-+Revista+Chilena+de+Ingeniería%22">INGENIARE - Revista Chilena de Ingeniería</searchLink>. Dec2011, Vol. 19 Issue 3, p351-358. 8p. 4 Color Photographs, 3 Charts.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Document+clustering%22">Document clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Multilingual+computing%22">Multilingual computing</searchLink><br /><searchLink fieldCode="DE" term="%22Ant+algorithms%22">Ant algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Computational+linguistics%22">Computational linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Text+mining%22">Text mining</searchLink><br /><searchLink fieldCode="DE" term="%22Records+management%22">Records management</searchLink>
– Name: Abstract
  Label: Abstract (English)
  Group: Ab
  Data: This paper presents a document representation strategy and a bio-inspired algorithm to cluster multilingual collections of documents in the field of economics and business. The proposed approach allows the user to identify groups of related economics documents written in Spanish and English using techniques inspired on clustering and sorting behaviours observed in some types of ants. In order to obtain a language independent vector representation of each document two multilingual resources are used: an economic glossary and a thesaurus. Each document is represented using four feature vectors: words, proper names, economic terms in the glossary and thesaurus descriptors. The proper name identification, word extraction and lemmatization are performed using specific tools. The tf-idf scheme is used to measure the importance of each feature in the document, and a convex linear combination of angular separations between feature vectors is used as similarity measure of documents. The paper shows experimental results of the application of the proposed algorithm in a Spanish-English corpus of research papers in economics and management areas. The results demonstrate the usefulness and effectiveness of the ant clustering algorithm and the proposed representation scheme. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label: Abstract (Spanish)
  Group: Ab
  Data: Este artículo presenta una estrategia de representación documental y un algoritmo bioinspirado para realizar procesos de agrupamiento en colecciones multilingües de documentos en las áreas de la economía y la empresa. El enfoque propuesto permite al usuario identificar grupos de documentos económicos relacionados escritos en español o inglés usando técnicas inspiradas en comportamientos de organización y agrupamiento de objetos observados en algunos tipos de hormigas. Para conseguir una representación vectorial de cada documento independiente del idioma, se han utilizado dos recursos lingüísticos: un glosario económico y un tesauro. Cada documento es representado usando cuatro vectores de rasgos: palabras, nombres propios, términos económicos del glosario y descriptores del tesauro. La identificación de los nombres propios y la extracción y lematización de palabras se realizan usando herramientas específicas. El esquema tf-idf es utilizado para medir la importancia de cada rasgo en el documento, y se utiliza una combinación lineal convexa de separaciones angulares de los vectores de rasgos como medida de similitud de documentos. El trabajo muestra resultados experimentales de aplicación del algoritmo propuesto sobre un corpus español-inglés de documentos científicos de áreas económica y de gestión empresarial. Los resultados demuestran la utilidad y efectividad de las técnicas de ant clustering y del esquema de representación propuesto. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of INGENIARE - Revista Chilena de Ingeniería is the property of Universidad de Tarapaca and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=72370736
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 8
        StartPage: 351
    Subjects:
      – SubjectFull: Document clustering
        Type: general
      – SubjectFull: Multilingual computing
        Type: general
      – SubjectFull: Ant algorithms
        Type: general
      – SubjectFull: Computational linguistics
        Type: general
      – SubjectFull: Text mining
        Type: general
      – SubjectFull: Records management
        Type: general
    Titles:
      – TitleFull: Identification of related multilingual documents using ant clustering algorithms.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Cobo, Ángel
      – PersonEntity:
          Name:
            NameFull: Rocha, Rocío
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Text: Dec2011
              Type: published
              Y: 2011
          Identifiers:
            – Type: issn-print
              Value: 07183291
          Numbering:
            – Type: volume
              Value: 19
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: INGENIARE - Revista Chilena de Ingeniería
              Type: main
ResultId 1