Large Language Model Selection for Test-Driven Prompt Android iOS Development.
Saved in:
| Title: | Large Language Model Selection for Test-Driven Prompt Android iOS Development. |
|---|---|
| Authors: | Rizqullah, Muhammad1 mrizqullah@stu.kau.edu.sa, Albassam, Emad1 |
| Source: | International Journal of Interactive Mobile Technologies. 2026, Vol. 20 Issue 3, p71-82. 12p. |
| Subjects: | Mobile app development, iOS (Operating system), Empirical research, Code generators, Android (Operating system), Prompt engineering, Language models |
| Abstract: | Large language model (LLM) code generation research predominantly focuses on Python, with test-driven prompt engineering exclusively targeting this language. This study presents a comprehensive LLM selection framework for mobile development through rigorous empirical analysis. We conducted 8,704 evaluations across 544 programming tasks (HumanEval and MBPP datasets) on Android (Java) and iOS (Swift) platforms using four state-of-the-art LLMs (GPT-4o, GPT-4o-mini, Qwen 14B, and Qwen 32B), two prompting strategies (base and test-driven), and two metrics (accuracy and remediation accuracy). Systematic analysis of platform-specific patterns yielded a decision tree incorporating first-attempt correctness, budget constraints, and self-hosting requirements, validated through three industry-relevant use cases. Results show test-driven prompting (TDP) achieves a +2.22 pp average accuracy improvement over baseline (95% CI [1.22-3.23 pp], p < 0.001, d = 0.3974). However, LLMs consistently underperform in mobile development (66.85%--88.87%) compared to Pythonbased code generation (86.90%-91.30%) regardless of model size or type. This framework establishes groundwork for platform-specific optimizations while providing practitioners with actionable guidance for model selection in mobile development contexts. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Interactive Mobile Technologies is the property of International Journal of Interactive Mobile Technologies and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 191627704 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Large Language Model Selection for Test-Driven Prompt Android iOS Development. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Rizqullah%2C+Muhammad%22">Rizqullah, Muhammad</searchLink><relatesTo>1</relatesTo><i> mrizqullah@stu.kau.edu.sa</i><br /><searchLink fieldCode="AR" term="%22Albassam%2C+Emad%22">Albassam, Emad</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Interactive+Mobile+Technologies%22">International Journal of Interactive Mobile Technologies</searchLink>. 2026, Vol. 20 Issue 3, p71-82. 12p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Mobile+app+development%22">Mobile app development</searchLink><br /><searchLink fieldCode="DE" term="%22iOS+%28Operating+system%29%22">iOS (Operating system)</searchLink><br /><searchLink fieldCode="DE" term="%22Empirical+research%22">Empirical research</searchLink><br /><searchLink fieldCode="DE" term="%22Code+generators%22">Code generators</searchLink><br /><searchLink fieldCode="DE" term="%22Android+%28Operating+system%29%22">Android (Operating system)</searchLink><br /><searchLink fieldCode="DE" term="%22Prompt+engineering%22">Prompt engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Large language model (LLM) code generation research predominantly focuses on Python, with test-driven prompt engineering exclusively targeting this language. This study presents a comprehensive LLM selection framework for mobile development through rigorous empirical analysis. We conducted 8,704 evaluations across 544 programming tasks (HumanEval and MBPP datasets) on Android (Java) and iOS (Swift) platforms using four state-of-the-art LLMs (GPT-4o, GPT-4o-mini, Qwen 14B, and Qwen 32B), two prompting strategies (base and test-driven), and two metrics (accuracy and remediation accuracy). Systematic analysis of platform-specific patterns yielded a decision tree incorporating first-attempt correctness, budget constraints, and self-hosting requirements, validated through three industry-relevant use cases. Results show test-driven prompting (TDP) achieves a +2.22 pp average accuracy improvement over baseline (95% CI [1.22-3.23 pp], p < 0.001, d = 0.3974). However, LLMs consistently underperform in mobile development (66.85%--88.87%) compared to Pythonbased code generation (86.90%-91.30%) regardless of model size or type. This framework establishes groundwork for platform-specific optimizations while providing practitioners with actionable guidance for model selection in mobile development contexts. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of International Journal of Interactive Mobile Technologies is the property of International Journal of Interactive Mobile Technologies and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191627704 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.3991/ijim.v20i03.59861 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 71 Subjects: – SubjectFull: Mobile app development Type: general – SubjectFull: iOS (Operating system) Type: general – SubjectFull: Empirical research Type: general – SubjectFull: Code generators Type: general – SubjectFull: Android (Operating system) Type: general – SubjectFull: Prompt engineering Type: general – SubjectFull: Language models Type: general Titles: – TitleFull: Large Language Model Selection for Test-Driven Prompt Android iOS Development. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Rizqullah, Muhammad – PersonEntity: Name: NameFull: Albassam, Emad IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: 2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 18657923 Numbering: – Type: volume Value: 20 – Type: issue Value: 3 Titles: – TitleFull: International Journal of Interactive Mobile Technologies Type: main |
| ResultId | 1 |