MPI framework for parallel searching in large biological databases

Saved in:
Bibliographic Details
Title: MPI framework for parallel searching in large biological databases
Authors: Battré, Dominic dominic@battre.de, Angulo, David Sigfredo1 dangulo@cti.depaul.edu
Source: Journal of Parallel & Distributed Computing. Dec2006, Vol. 66 Issue 12, p1503-1511. 9p.
Subjects: Information storage & retrieval systems, Databases, Amino acid sequence, Mass spectrometry
Abstract: Abstract: In this paper, we address the problem of searching huge biological databases on the scale of at least several gigabytes by utilizing parallel processing. Biological databases storing DNA sequences, protein sequences, or mass spectra are growing exponentially. Searches through these databases consume exponentially growing computational resources as well. We demonstrate herein a general use, MPI based, framework for generically splitting databases amongst several computational nodes. The combined RAM of the nodes working in tandem is often sufficient to keep the entire database in memory, and therefore to search it efficiently without paging to disk. The framework runs as a persistent service, processing all submitted queries. This allows for query reordering and better utilization of the memory. Thereby, we achieve superlinear speedups compared to single processor implementations. We demonstrate the utility and speedup of the framework using a real biological database and an actual searching algorithm for mass spectrometry. [Copyright &y& Elsevier]
Copyright of Journal of Parallel & Distributed Computing is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 23161000
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: MPI framework for parallel searching in large biological databases
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Battré%2C+Dominic%22">Battré, Dominic</searchLink><i> dominic@battre.de</i><br /><searchLink fieldCode="AR" term="%22Angulo%2C+David+Sigfredo%22">Angulo, David Sigfredo</searchLink><relatesTo>1</relatesTo><i> dangulo@cti.depaul.edu</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Parallel+%26+Distributed+Computing%22">Journal of Parallel & Distributed Computing</searchLink>. Dec2006, Vol. 66 Issue 12, p1503-1511. 9p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Information+storage+%26+retrieval+systems%22">Information storage & retrieval systems</searchLink><br /><searchLink fieldCode="DE" term="%22Databases%22">Databases</searchLink><br /><searchLink fieldCode="DE" term="%22Amino+acid+sequence%22">Amino acid sequence</searchLink><br /><searchLink fieldCode="DE" term="%22Mass+spectrometry%22">Mass spectrometry</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Abstract: In this paper, we address the problem of searching huge biological databases on the scale of at least several gigabytes by utilizing parallel processing. Biological databases storing DNA sequences, protein sequences, or mass spectra are growing exponentially. Searches through these databases consume exponentially growing computational resources as well. We demonstrate herein a general use, MPI based, framework for generically splitting databases amongst several computational nodes. The combined RAM of the nodes working in tandem is often sufficient to keep the entire database in memory, and therefore to search it efficiently without paging to disk. The framework runs as a persistent service, processing all submitted queries. This allows for query reordering and better utilization of the memory. Thereby, we achieve superlinear speedups compared to single processor implementations. We demonstrate the utility and speedup of the framework using a real biological database and an actual searching algorithm for mass spectrometry. [Copyright &y& Elsevier]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Parallel & Distributed Computing is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=23161000
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.jpdc.2006.08.003
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 1503
    Subjects:
      – SubjectFull: Information storage & retrieval systems
        Type: general
      – SubjectFull: Databases
        Type: general
      – SubjectFull: Amino acid sequence
        Type: general
      – SubjectFull: Mass spectrometry
        Type: general
    Titles:
      – TitleFull: MPI framework for parallel searching in large biological databases
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Battré, Dominic
      – PersonEntity:
          Name:
            NameFull: Angulo, David Sigfredo
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Text: Dec2006
              Type: published
              Y: 2006
          Identifiers:
            – Type: issn-print
              Value: 07437315
          Numbering:
            – Type: volume
              Value: 66
            – Type: issue
              Value: 12
          Titles:
            – TitleFull: Journal of Parallel & Distributed Computing
              Type: main
ResultId 1