Semantic search as extractive paraphrase span detection: Semantic search as extractive...: J. Kanerva et al.

Saved in:
Bibliographic Details
Title: Semantic search as extractive paraphrase span detection: Semantic search as extractive...: J. Kanerva et al.
Authors: Kanerva, Jenna1 (AUTHOR) jmnybl@utu.fi, Kitti, Hanna1 (AUTHOR), Chang, Li-Hsin1 (AUTHOR), Vahtola, Teemu2 (AUTHOR), Creutz, Mathias2 (AUTHOR), Ginter, Filip1 (AUTHOR)
Source: Language Resources & Evaluation. Mar2025, Vol. 59 Issue 1, p257-276. 20p.
Subjects: Heather, Paraphrase, Corpora, Terms & phrases, Possibility
Abstract: In this paper, we approach the problem of semantic search by introducing a task of paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to identify its paraphrase in a given document, the same modelling setup as typically used in extractive question answering. While current work in paraphrasing has almost uniquely focused on sentence-level approaches, the novel span detection approach gives a possibility to retrieve a segment of arbitrary length. On the Turku Paraphrase Corpus of 100,000 manually extracted Finnish paraphrase pairs including their original document context, we find that by achieving an exact match of 88.73 our paraphrase span detection approach outperforms widely adopted sentence-level retrieval baselines (lexical similarity as well as BERT and SBERT sentence embeddings) by more than 20pp in terms of exact match, and 11pp in terms of token-level F-score. This demonstrates a strong advantage of modelling the paraphrase retrieval in terms of span extraction rather than commonly used sentence similarity, the sentence-level approaches being clearly suboptimal for applications where the retrieval targets are not guaranteed to be full sentences. Even when limiting the evaluation to sentence-level retrieval targets only, the span detection model still outperforms the sentence-level baselines by more than 4 pp in terms of exact match, and almost 6pp F-score. Additionally, we introduce a method for creating artificial paraphrase data through back-translation, suitable for languages where manually annotated paraphrase resources for training the span detection model are not available. [ABSTRACT FROM AUTHOR]
Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 183750667
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Semantic search as extractive paraphrase span detection: Semantic search as extractive...: J. Kanerva et al.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kanerva%2C+Jenna%22">Kanerva, Jenna</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> jmnybl@utu.fi</i><br /><searchLink fieldCode="AR" term="%22Kitti%2C+Hanna%22">Kitti, Hanna</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Chang%2C+Li-Hsin%22">Chang, Li-Hsin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Vahtola%2C+Teemu%22">Vahtola, Teemu</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Creutz%2C+Mathias%22">Creutz, Mathias</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Ginter%2C+Filip%22">Ginter, Filip</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Language+Resources+%26+Evaluation%22">Language Resources & Evaluation</searchLink>. Mar2025, Vol. 59 Issue 1, p257-276. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Heather%22">Heather</searchLink><br /><searchLink fieldCode="DE" term="%22Paraphrase%22">Paraphrase</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora%22">Corpora</searchLink><br /><searchLink fieldCode="DE" term="%22Terms+%26+phrases%22">Terms & phrases</searchLink><br /><searchLink fieldCode="DE" term="%22Possibility%22">Possibility</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In this paper, we approach the problem of semantic search by introducing a task of paraphrase span detection, i.e. given a segment of text as a query phrase, the task is to identify its paraphrase in a given document, the same modelling setup as typically used in extractive question answering. While current work in paraphrasing has almost uniquely focused on sentence-level approaches, the novel span detection approach gives a possibility to retrieve a segment of arbitrary length. On the Turku Paraphrase Corpus of 100,000 manually extracted Finnish paraphrase pairs including their original document context, we find that by achieving an exact match of 88.73 our paraphrase span detection approach outperforms widely adopted sentence-level retrieval baselines (lexical similarity as well as BERT and SBERT sentence embeddings) by more than 20pp in terms of exact match, and 11pp in terms of token-level F-score. This demonstrates a strong advantage of modelling the paraphrase retrieval in terms of span extraction rather than commonly used sentence similarity, the sentence-level approaches being clearly suboptimal for applications where the retrieval targets are not guaranteed to be full sentences. Even when limiting the evaluation to sentence-level retrieval targets only, the span detection model still outperforms the sentence-level baselines by more than 4 pp in terms of exact match, and almost 6pp F-score. Additionally, we introduce a method for creating artificial paraphrase data through back-translation, suitable for languages where manually annotated paraphrase resources for training the span detection model are not available. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=183750667
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10579-023-09715-7
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 257
    Subjects:
      – SubjectFull: Heather
        Type: general
      – SubjectFull: Paraphrase
        Type: general
      – SubjectFull: Corpora
        Type: general
      – SubjectFull: Terms & phrases
        Type: general
      – SubjectFull: Possibility
        Type: general
    Titles:
      – TitleFull: Semantic search as extractive paraphrase span detection: Semantic search as extractive...: J. Kanerva et al.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kanerva, Jenna
      – PersonEntity:
          Name:
            NameFull: Kitti, Hanna
      – PersonEntity:
          Name:
            NameFull: Chang, Li-Hsin
      – PersonEntity:
          Name:
            NameFull: Vahtola, Teemu
      – PersonEntity:
          Name:
            NameFull: Creutz, Mathias
      – PersonEntity:
          Name:
            NameFull: Ginter, Filip
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1574020X
          Numbering:
            – Type: volume
              Value: 59
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Language Resources & Evaluation
              Type: main
ResultId 1