A deep learning mechanism to detect phishing URLs using the permutation importance method and SMOTE-Tomek link.
Saved in:
| Title: | A deep learning mechanism to detect phishing URLs using the permutation importance method and SMOTE-Tomek link. |
|---|---|
| Authors: | Zaimi, Rania1 (AUTHOR) rania.zaimi@univ-annaba.org, Hafidi, Mohamed1 (AUTHOR), Lamia, Mahnane1 (AUTHOR) |
| Source: | Journal of Supercomputing. Aug2024, Vol. 80 Issue 12, p17159-17191. 33p. |
| Subjects: | Deep learning, Uniform Resource Locators, Phishing, Feature selection, Machine learning, Permutations, Web-based user interfaces |
| Abstract: | In contemporary times, the proliferation of phishing attacks presents a substantial and growing challenge to cybersecurity. This fraudulent tactic is designed to deceive unsuspecting individuals, enticing them to access malicious websites and disclose sensitive personal information such as usernames, passwords, and financial details. As a result, malevolent actors exploit this data for illicit purposes. As the sophistication and maliciousness of phishing continue to evolve, researchers are earnestly developing multiple anti-phishing solutions in the literature. Among these solutions, those based on machine learning and deep learning models have gained substantial attention in recent years. This study proposes an intelligent mechanism to detect phishing URLs. The proposed system is based on the permutation importance method to select the most relevant URL features and the SMOTE-Tomek link method to solve the problem of an unbalanced dataset. In addition, the XGBoost classifier and four deep learning models—CNN, LSTM, and two hybrid models (CNN-LSTM and LSTM-CNN)—are employed to classify URLs as phishing or legitimate and to compare their performance. The experimental results demonstrate the successful functioning of the proposed phishing detection mechanism. It is observed that the proposed mechanism achieved an accuracy ranging from 93.36 to 97.05% without feature selection and data balance across two variants of datasets and different classifiers. It also achieved an accuracy ranging from 94.12 to 97.82% with feature selection and data balance. Finally, our phishing detection mechanism is implemented as a web application to enhance its usability for web users. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 178339410 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A deep learning mechanism to detect phishing URLs using the permutation importance method and SMOTE-Tomek link. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Zaimi%2C+Rania%22">Zaimi, Rania</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> rania.zaimi@univ-annaba.org</i><br /><searchLink fieldCode="AR" term="%22Hafidi%2C+Mohamed%22">Hafidi, Mohamed</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Lamia%2C+Mahnane%22">Lamia, Mahnane</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Supercomputing%22">Journal of Supercomputing</searchLink>. Aug2024, Vol. 80 Issue 12, p17159-17191. 33p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Uniform+Resource+Locators%22">Uniform Resource Locators</searchLink><br /><searchLink fieldCode="DE" term="%22Phishing%22">Phishing</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+selection%22">Feature selection</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Permutations%22">Permutations</searchLink><br /><searchLink fieldCode="DE" term="%22Web-based+user+interfaces%22">Web-based user interfaces</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In contemporary times, the proliferation of phishing attacks presents a substantial and growing challenge to cybersecurity. This fraudulent tactic is designed to deceive unsuspecting individuals, enticing them to access malicious websites and disclose sensitive personal information such as usernames, passwords, and financial details. As a result, malevolent actors exploit this data for illicit purposes. As the sophistication and maliciousness of phishing continue to evolve, researchers are earnestly developing multiple anti-phishing solutions in the literature. Among these solutions, those based on machine learning and deep learning models have gained substantial attention in recent years. This study proposes an intelligent mechanism to detect phishing URLs. The proposed system is based on the permutation importance method to select the most relevant URL features and the SMOTE-Tomek link method to solve the problem of an unbalanced dataset. In addition, the XGBoost classifier and four deep learning models—CNN, LSTM, and two hybrid models (CNN-LSTM and LSTM-CNN)—are employed to classify URLs as phishing or legitimate and to compare their performance. The experimental results demonstrate the successful functioning of the proposed phishing detection mechanism. It is observed that the proposed mechanism achieved an accuracy ranging from 93.36 to 97.05% without feature selection and data balance across two variants of datasets and different classifiers. It also achieved an accuracy ranging from 94.12 to 97.82% with feature selection and data balance. Finally, our phishing detection mechanism is implemented as a web application to enhance its usability for web users. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=178339410 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11227-024-06124-7 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 33 StartPage: 17159 Subjects: – SubjectFull: Deep learning Type: general – SubjectFull: Uniform Resource Locators Type: general – SubjectFull: Phishing Type: general – SubjectFull: Feature selection Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Permutations Type: general – SubjectFull: Web-based user interfaces Type: general Titles: – TitleFull: A deep learning mechanism to detect phishing URLs using the permutation importance method and SMOTE-Tomek link. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Zaimi, Rania – PersonEntity: Name: NameFull: Hafidi, Mohamed – PersonEntity: Name: NameFull: Lamia, Mahnane IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 08 Text: Aug2024 Type: published Y: 2024 Identifiers: – Type: issn-print Value: 09208542 Numbering: – Type: volume Value: 80 – Type: issue Value: 12 Titles: – TitleFull: Journal of Supercomputing Type: main |
| ResultId | 1 |