SVG-CNN: A shallow CNN based on VGGNet applied to intra prediction partition block in HEVC.
Saved in:
| Title: | SVG-CNN: A shallow CNN based on VGGNet applied to intra prediction partition block in HEVC. |
|---|---|
| Authors: | Linck, Iris1 (AUTHOR) iris.linck@ucdenver.edu, Gómez, Arthur Tórgo2 (AUTHOR), Alaghband, Gita1 (AUTHOR) |
| Source: | Multimedia Tools & Applications. Sep2024, Vol. 83 Issue 30, p73983-74001. 19p. |
| Subjects: | Convolutional neural networks, Recursive partitioning, Video coding, Probability theory, Forecasting, Video compression |
| Abstract: | High Efficiency Video Coding (HEVC) offers superior compression rates, but its adoption introduces increased coding complexity due to its reliance on a recursive quad-tree for partitioning frames into varying block sizes. This quad-tree process is a central feature in upcoming video coding standards. Our paper presents a novel framework, SVG-CNN, which integrates three shallow Convolutional Neural Networks (CNNs) inspired by VGGNet. Each CNN is specifically designed for individual quad-tree levels to predict the Code Unit (CU) partition in HEVC, leading to reduced intra-frame coding time. SVG-CNN has an inherent capability for early terminations, leveraging sequential CNN feeding based on quad-tree level probabilities. This provides a mechanism to halt processes when further refinement is seemed unlikely. Enhancing the model's efficacy, we have crafted three specialized datasets, each focusing on distinct quad-tree levels and quantization parameter (QP) contexts. This allows each CNN within our framework to undergo targeted training, establishing a cutting-edge training methodology. Our study shows that performance, in terms of accuracy and F1 metrics, is highly dependent on QP settings, with lower QPs yielding better results, and higher QPs diminishing performance due to potential loss of critical features. To enhance our model, we tackled hyperparameter selection and CU split threshold determination for HEVC prediction. We utilized Grid Search Cross-Validation for the former and assessed multiple thresholds across selected videos for the latter. The model has a moderate complexity with over 328,000 parameters across 18 layers, which ensures memory efficiency. It boasts a swift prediction time of 0.05 ms and reduces HEVC encoding time by 61.64%, while slightly improving the bitrate-distortion performance by -0.24% BDBR, indicating better compression without notable PSNR loss. Significantly, our approach outperforms other CNN-based quad-tree partitioning methods that reduce HEVC coding complexity but sacrifice compression performance. [ABSTRACT FROM AUTHOR] |
| Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 179395179 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: SVG-CNN: A shallow CNN based on VGGNet applied to intra prediction partition block in HEVC. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Linck%2C+Iris%22">Linck, Iris</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> iris.linck@ucdenver.edu</i><br /><searchLink fieldCode="AR" term="%22Gómez%2C+Arthur+Tórgo%22">Gómez, Arthur Tórgo</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Alaghband%2C+Gita%22">Alaghband, Gita</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Multimedia+Tools+%26+Applications%22">Multimedia Tools & Applications</searchLink>. Sep2024, Vol. 83 Issue 30, p73983-74001. 19p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Convolutional+neural+networks%22">Convolutional neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Recursive+partitioning%22">Recursive partitioning</searchLink><br /><searchLink fieldCode="DE" term="%22Video+coding%22">Video coding</searchLink><br /><searchLink fieldCode="DE" term="%22Probability+theory%22">Probability theory</searchLink><br /><searchLink fieldCode="DE" term="%22Forecasting%22">Forecasting</searchLink><br /><searchLink fieldCode="DE" term="%22Video+compression%22">Video compression</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: High Efficiency Video Coding (HEVC) offers superior compression rates, but its adoption introduces increased coding complexity due to its reliance on a recursive quad-tree for partitioning frames into varying block sizes. This quad-tree process is a central feature in upcoming video coding standards. Our paper presents a novel framework, SVG-CNN, which integrates three shallow Convolutional Neural Networks (CNNs) inspired by VGGNet. Each CNN is specifically designed for individual quad-tree levels to predict the Code Unit (CU) partition in HEVC, leading to reduced intra-frame coding time. SVG-CNN has an inherent capability for early terminations, leveraging sequential CNN feeding based on quad-tree level probabilities. This provides a mechanism to halt processes when further refinement is seemed unlikely. Enhancing the model's efficacy, we have crafted three specialized datasets, each focusing on distinct quad-tree levels and quantization parameter (QP) contexts. This allows each CNN within our framework to undergo targeted training, establishing a cutting-edge training methodology. Our study shows that performance, in terms of accuracy and F1 metrics, is highly dependent on QP settings, with lower QPs yielding better results, and higher QPs diminishing performance due to potential loss of critical features. To enhance our model, we tackled hyperparameter selection and CU split threshold determination for HEVC prediction. We utilized Grid Search Cross-Validation for the former and assessed multiple thresholds across selected videos for the latter. The model has a moderate complexity with over 328,000 parameters across 18 layers, which ensures memory efficiency. It boasts a swift prediction time of 0.05 ms and reduces HEVC encoding time by 61.64%, while slightly improving the bitrate-distortion performance by -0.24% BDBR, indicating better compression without notable PSNR loss. Significantly, our approach outperforms other CNN-based quad-tree partitioning methods that reduce HEVC coding complexity but sacrifice compression performance. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=179395179 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11042-024-18412-8 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 19 StartPage: 73983 Subjects: – SubjectFull: Convolutional neural networks Type: general – SubjectFull: Recursive partitioning Type: general – SubjectFull: Video coding Type: general – SubjectFull: Probability theory Type: general – SubjectFull: Forecasting Type: general – SubjectFull: Video compression Type: general Titles: – TitleFull: SVG-CNN: A shallow CNN based on VGGNet applied to intra prediction partition block in HEVC. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Linck, Iris – PersonEntity: Name: NameFull: Gómez, Arthur Tórgo – PersonEntity: Name: NameFull: Alaghband, Gita IsPartOfRelationships: – BibEntity: Dates: – D: 21 M: 09 Text: Sep2024 Type: published Y: 2024 Identifiers: – Type: issn-print Value: 13807501 Numbering: – Type: volume Value: 83 – Type: issue Value: 30 Titles: – TitleFull: Multimedia Tools & Applications Type: main |
| ResultId | 1 |