Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.

Saved in:
Bibliographic Details
Title: Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.
Authors: Sagi, Tomer1 (AUTHOR) tsagi@cs.aau.dk, Zaga, Moran2 (AUTHOR) mzaga@staff.haifa.ac.il, Rusinek, Sinai2 (AUTHOR) sinai.rusinek@gmail.com, Fekete, Marcell R.1 (AUTHOR) mrfe@cs.aau.dk, Bjerva, Johannes1 (AUTHOR) jbjerva@cs.aau.dk, Hose, Katja1,3 (AUTHOR) katja.hose@tuwien.ac.at
Source: Language Resources & Evaluation. Sep2025, Vol. 59 Issue 3, p2427-2451. 25p.
Subjects: Hebrew language, Arabic language, Transliteration, Geographic names, Bilingualism, Phonology, Historical source material
Abstract: The writings of one ancient civilization often overlap in time and space with others. Many of these sources comprise unstructured text in ancient languages, causing scholars studying these civilizations to be siloed, often relying on sources in specific languages. Most recent efforts to extract structured information from historical scripts into place (toponym) and people databases (prospographies) have followed this pattern, focusing on one civilization and selected sources. The path to creating a common database runs through aligning names or toponyms between sources from disparate languages utilizing different scripts. Existing multi-lingual orthographic (string-based) comparison often relies on transliteration to a common script (Latin/English). Transliteration often creates multiple options and even more confusion. However, when integrating sources that overlap in space and time, the languages often share a common phonetic background. This commonality may prove beneficial. In this work, we present a benchmark for comparing toponyms from two linguistically and culturally related languages, namely Hebrew and Arabic. We provide a benchmark comprised of a set of dataset pairs created from historical sources written in Medieval variants of these languages, later historical Gazetteers and a modern dataset curated from Wikidata. We empirically evaluate several toponym comparison approaches over the benchmark: transliteration to a common script, direct transliteration, and phonetic comparison using a common phonetic representation. We discuss the results and the limitations of the various methods and outline future work. [ABSTRACT FROM AUTHOR]
Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 186909071
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Sagi%2C+Tomer%22">Sagi, Tomer</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> tsagi@cs.aau.dk</i><br /><searchLink fieldCode="AR" term="%22Zaga%2C+Moran%22">Zaga, Moran</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> mzaga@staff.haifa.ac.il</i><br /><searchLink fieldCode="AR" term="%22Rusinek%2C+Sinai%22">Rusinek, Sinai</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> sinai.rusinek@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Fekete%2C+Marcell+R%2E%22">Fekete, Marcell R.</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> mrfe@cs.aau.dk</i><br /><searchLink fieldCode="AR" term="%22Bjerva%2C+Johannes%22">Bjerva, Johannes</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jbjerva@cs.aau.dk</i><br /><searchLink fieldCode="AR" term="%22Hose%2C+Katja%22">Hose, Katja</searchLink><relatesTo>1,3</relatesTo> (AUTHOR)<i> katja.hose@tuwien.ac.at</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Language+Resources+%26+Evaluation%22">Language Resources & Evaluation</searchLink>. Sep2025, Vol. 59 Issue 3, p2427-2451. 25p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Hebrew+language%22">Hebrew language</searchLink><br /><searchLink fieldCode="DE" term="%22Arabic+language%22">Arabic language</searchLink><br /><searchLink fieldCode="DE" term="%22Transliteration%22">Transliteration</searchLink><br /><searchLink fieldCode="DE" term="%22Geographic+names%22">Geographic names</searchLink><br /><searchLink fieldCode="DE" term="%22Bilingualism%22">Bilingualism</searchLink><br /><searchLink fieldCode="DE" term="%22Phonology%22">Phonology</searchLink><br /><searchLink fieldCode="DE" term="%22Historical+source+material%22">Historical source material</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The writings of one ancient civilization often overlap in time and space with others. Many of these sources comprise unstructured text in ancient languages, causing scholars studying these civilizations to be siloed, often relying on sources in specific languages. Most recent efforts to extract structured information from historical scripts into place (toponym) and people databases (prospographies) have followed this pattern, focusing on one civilization and selected sources. The path to creating a common database runs through aligning names or toponyms between sources from disparate languages utilizing different scripts. Existing multi-lingual orthographic (string-based) comparison often relies on transliteration to a common script (Latin/English). Transliteration often creates multiple options and even more confusion. However, when integrating sources that overlap in space and time, the languages often share a common phonetic background. This commonality may prove beneficial. In this work, we present a benchmark for comparing toponyms from two linguistically and culturally related languages, namely Hebrew and Arabic. We provide a benchmark comprised of a set of dataset pairs created from historical sources written in Medieval variants of these languages, later historical Gazetteers and a modern dataset curated from Wikidata. We empirically evaluate several toponym comparison approaches over the benchmark: transliteration to a common script, direct transliteration, and phonetic comparison using a common phonetic representation. We discuss the results and the limitations of the various methods and outline future work. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=186909071
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10579-025-09812-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
        StartPage: 2427
    Subjects:
      – SubjectFull: Hebrew language
        Type: general
      – SubjectFull: Arabic language
        Type: general
      – SubjectFull: Transliteration
        Type: general
      – SubjectFull: Geographic names
        Type: general
      – SubjectFull: Bilingualism
        Type: general
      – SubjectFull: Phonology
        Type: general
      – SubjectFull: Historical source material
        Type: general
    Titles:
      – TitleFull: Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Sagi, Tomer
      – PersonEntity:
          Name:
            NameFull: Zaga, Moran
      – PersonEntity:
          Name:
            NameFull: Rusinek, Sinai
      – PersonEntity:
          Name:
            NameFull: Fekete, Marcell R.
      – PersonEntity:
          Name:
            NameFull: Bjerva, Johannes
      – PersonEntity:
          Name:
            NameFull: Hose, Katja
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1574020X
          Numbering:
            – Type: volume
              Value: 59
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Language Resources & Evaluation
              Type: main
ResultId 1