Secondary Students' Reading of Socio-Scientific Image-Texts on Climate Change in a GPT-4 Scenario
Saved in:
| Title: | Secondary Students' Reading of Socio-Scientific Image-Texts on Climate Change in a GPT-4 Scenario |
|---|---|
| Language: | English |
| Authors: | Jack Pun (ORCID |
| Source: | Research in Science Education. 2026 56(1):183-202. |
| Availability: | Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ |
| Peer Reviewed: | Y |
| Page Count: | 20 |
| Publication Date: | 2026 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Secondary Education |
| Descriptors: | Secondary School Students, Science and Society, Climate, Reader Text Relationship, Artificial Intelligence, Reading Comprehension, Internet, Search Engines, Imagery, Multimedia Materials |
| DOI: | 10.1007/s11165-025-10258-w |
| ISSN: | 0157-244X 1573-1898 |
| Abstract: | The prominence of multimodal generative artificial intelligence (GenAI) facilitates students' comprehension of scientific knowledge through linguistic and visual modes. However, there is a lack of research that investigates how students read image-text outputs created in GenAI. We conceptualize a model of image-text reading of GenAI scientific texts that comprises the interpretation, exchange, and evaluation domains. Based on this theoretical model, we explored how 68 junior secondary students read two image-text socio-scientific texts created by GPT-4 with DALL.E plugins, one focusing on cognitive-epistemic aspects and another focusing on social-institutional aspects of climate change. Our findings indicated that these domains did not exhibit a hierarchical structure, while students' performance in the evaluation domain in the cognitive-epistemic text was better than that in the social-institutional text. More importantly, students expressed a range of uninformed ideas regarding the nature of GenAI when they read the two texts, including equating GenAI to an Internet search engine, picture creators, and human. We discussed how teaching and learning can foster students' "image-text and epistemic" reading by targeting the three domains of our theoretical model. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Accession Number: | EJ1504722 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwG2yrDtJJ-NFS-8VWJMMnAhAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDJOmqJ-HWKt3ZZU2VgIBEICBmy97kskASs4LCTgJDc1AGqmMYxOI0_UoQb6usFVt2wtHjsNAi_Z0HRjiDmn5rPgqcJsNrE6lDCMc2xAhTG6uf_ZceLbW63veZ8ba1TRDa6e4b9j0NIZcVL-xyDFIQXIe1y7PMegPqPYVCVKzE21BvoqO-3ecOVa5r50ubLny1IC_HT__xSsp0N4SJIEtCSOq1-YWc32dJ28VtTQx Text: Availability: 1 Value: <anid>AN0191207466;g7201feb.26;2026Feb02.06:20;v2.2.500</anid> <title id="AN0191207466-1">Secondary Students' Reading of Socio-Scientific Image-Texts on Climate Change in a GPT-4 Scenario </title> <p>The prominence of multimodal generative artificial intelligence (GenAI) facilitates students' comprehension of scientific knowledge through linguistic and visual modes. However, there is a lack of research that investigates how students read image-text outputs created in GenAI. We conceptualize a model of image-text reading of GenAI scientific texts that comprises the interpretation, exchange, and evaluation domains. Based on this theoretical model, we explored how 68 junior secondary students read two image-text socio-scientific texts created by GPT-4 with DALL.E plugins, one focusing on cognitive-epistemic aspects and another focusing on social-institutional aspects of climate change. Our findings indicated that these domains did not exhibit a hierarchical structure, while students' performance in the evaluation domain in the cognitive-epistemic text was better than that in the social-institutional text. More importantly, students expressed a range of uninformed ideas regarding the nature of GenAI when they read the two texts, including equating GenAI to an Internet search engine, picture creators, and human. We discussed how teaching and learning can foster students' image-text and epistemic reading by targeting the three domains of our theoretical model.</p> <p>Keywords: Generative artificial intelligence; GPT-4; ChatGPT; Reading of science; Multimodality; Education Specialist Studies In Education Psychology and Cognitive Sciences Psychology</p> <p>Jack Pun will handle correspondence at all stages of refereeing and publication. If there are any queries pertaining to this manuscript, please do not hesitate to contact Jack Pun.</p> <hd id="AN0191207466-2">Introduction</hd> <p>Generative artificial intelligence (GenAI) applications, exemplified by DALL.E, possess the capability to generate diverse modes of representation beyond linguistic texts, including images, videos, and music (Ferrara, [<reflink idref="bib24" id="ref1">24</reflink>]). For instance, GPT-4 has a DALL.E plugin so that GPT-4 can generate content containing both an image and written text (OpenAI, [<reflink idref="bib52" id="ref2">52</reflink>]). Owing to GenAI's capacity to create and interpret different kinds of representations, science educators have been exploring the pedagogical potential of GenAI that can create both images and texts (Alasadi &amp; Baiz, [<reflink idref="bib2" id="ref3">2</reflink>]; Bewersdorff et al., [<reflink idref="bib5" id="ref4">5</reflink>]; Cooper &amp; Tang, [<reflink idref="bib15" id="ref5">15</reflink>]; Tang, [<reflink idref="bib63" id="ref6">63</reflink>]).</p> <p>As GenAI can create "nonsensical" scientific images that can be published in authoritative scientific journals (Franzen, [<reflink idref="bib26" id="ref7">26</reflink>]), students need to critically interpret and evaluate these images, as well as exchanging with GenAI cautiously. Especially at present, GenAI can create images of fake experts and activists that manipulate and shape scientific misinformation (Harris, [<reflink idref="bib31" id="ref8">31</reflink>]; Marchal et al., [<reflink idref="bib46" id="ref9">46</reflink>]). Using large language models, GPT-4 recognises and generates linguistic texts according to pre-trained data (Kasneci et al., [<reflink idref="bib35" id="ref10">35</reflink>]); using simple diffusion models, DALL·E encodes texts and maps them to a representational space, followed by image encoding and decoding to create an image based on the probability of datasets (Derevyanko &amp; Zalevska, [<reflink idref="bib16" id="ref11">16</reflink>]). As GenAI does not directly perform any scientific investigate to generate scientific information, images and texts created by GenAI may also have the potential to communicate misinformation found in socio-scientific texts, such as climate change. As shown in previous research, ChatGPT fails to address how the social structure, including economic activities and social actors, impacts climate change (Sommer &amp; von Querfurth, [<reflink idref="bib60" id="ref12">60</reflink>]). Yet, there is not any theoretical model that sets student's performance of reading scientific image-texts created by GenAI that guides teaching and learning, despite the popularity of GenAI in creating image-texts of scientific information.</p> <p>In this regard, this paper aims to theorise a model of reading of socio-scientific image-texts and explore students' performance in relation to the components of this model. This model's primary purpose is to provide a framework for understanding how students engage with socio-scientific image-texts generated by GenAI, rather than solely for evaluating performance or drafting texts. The model is designed to assess students' comprehension and critical engagement with different elements (linguistic and visual) of GenAI-generated texts, focusing on three key domains: Interpretation, Exchange, and Evaluation. Such a model can guide teachers' and researchers' evaluation of students' reading socio-scientific image-texts. More importantly, this model can be incorporated to designing discipline-specific pedagogy for incorporating GenAI in science classrooms. Grounded in a language and literacy perspective, Tang ([<reflink idref="bib63" id="ref13">63</reflink>]) suggested that an interactive-constructive model is needed for science educators to understand how students <emph>read</emph> images and written scientific texts in GPT-4. However, not only does a reading model of scientific texts consider how students interpret and interact with the scientific image-texts in GenAI, but such a reading model also needs to consider the <emph>epistemic</emph> nature of science and nature of GenAI (Cheung et al., [<reflink idref="bib12" id="ref14">12</reflink>], [<reflink idref="bib11" id="ref15">11</reflink>]). Hence, the GenAI reading model of socio-scientific image-texts consists of three major domains: (<reflink idref="bib1" id="ref16">1</reflink>) the <emph>interpretation domain</emph> refers to how students interpret content of images and linguistic texts in GenAI; (<reflink idref="bib2" id="ref17">2</reflink>) the <emph>exchange domain</emph> refers to how students reason and act on socio-scientific image-texts in GenAI; (<reflink idref="bib3" id="ref18">3</reflink>) the <emph>evaluation domain</emph> refers to how students critically judge claims portrayed in images and linguistic texts in GenAI in relation to nature of science and nature of GenAI.</p> <p>To explore students' performance in these three domains, we used a construct-driven approach (Wilson, [<reflink idref="bib69" id="ref19">69</reflink>]) to design a reading instrument that simulates the interaction between students and GPT-4 with DALL.E plugin in the context of verifying climate information. As GenAI generates new texts based on the students' inputs, consistently measuring students' reading of image-texts in real-time poses difficulty. Therefore, we situated students in simulated conversations with GPT-4, where we screenshotted two conversation records with image-text outputs regarding verifying common climate claims. One conversation is situated within the cognitive-epistemic dimension of climate science, whereas the other one is situated within the social-institutional science (Kaya &amp; Erduran, [<reflink idref="bib36" id="ref20">36</reflink>]). As students' grade levels (Diakidoy et al., [<reflink idref="bib17" id="ref21">17</reflink>]) and prior exposure to digital applications (Masataka, [<reflink idref="bib47" id="ref22">47</reflink>]) can be related to reading performance, we hypothesize that students' performance in the three domains might be different in various grade levels and exposure to ChatGPT. The following research questions guide the present study:</p> <p> <bold> <emph>RQ1</emph> </bold>. What is students' performance in reading socio-scientific image-texts on climate change in a ChatGPT scenario? Specifically, what are their reading performances in the three domains: the <emph>Interpretation</emph>, <emph>Exchange</emph>, and <emph>Evaluation</emph> domains?</p> <p> <bold> <emph>RQ2</emph> </bold>. Is there any difference in students' performance in reading of socio-scientific image-texts regarding their different grade levels and prior exposure to ChatGPT?</p> <hd id="AN0191207466-3">Framing the Study</hd> <p></p> <hd id="AN0191207466-4">Reading of Socio-Scientific Texts on Climate Change</hd> <p>Climate change is considered a type of socio-scientific text since it generates both interest and controversy and is undermined in nature (Sadler, [<reflink idref="bib54" id="ref23">54</reflink>]). Also, climate change is qualified as a kind of socio-scientific issues owing to its connection to societal, political and economic implications (Sadler, [<reflink idref="bib54" id="ref24">54</reflink>]; Sadler &amp; Dawson, [<reflink idref="bib55" id="ref25">55</reflink>]). Reading socio-scientific texts is crucial in enhancing students' understanding and reasoning of climate change, which subsequently shapes their attitudes and behaviors towards climate change (Clark et al., [<reflink idref="bib13" id="ref26">13</reflink>]; Cheung et al., [<reflink idref="bib11" id="ref27">11</reflink>]). Despite the pivotal role of reading in this context, socio-scientific texts on climate change often contain varying degrees of misinformation (Samantray &amp; Pin, [<reflink idref="bib56" id="ref28">56</reflink>]; Treen et al., [<reflink idref="bib66" id="ref29">66</reflink>]) that requires students' <emph>critical</emph> reading. Such misinformation can be classified into two types: canonical climate misinformation and epistemic climate misinformation. Canonical climate misinformation includes false claims regarding causes, processes, and effects of climate change, such as falsifying human-caused climate change (Cook et al., [<reflink idref="bib14" id="ref30">14</reflink>]); epistemic climate misinformation includes incorrect claims about how science works to generate knowledge claims regarding climate change, such as climate research being independent of funding networks (Farrell et al., [<reflink idref="bib20" id="ref31">20</reflink>]).</p> <p>Reading climate socio-scientific texts requires students to interpret canonical knowledge as well as epistemic knowledge presented. In the past, reading in science has been positioned by science educators as decoding semantics and content knowledge of written texts; while at present, reading scientific texts is positioned as <emph>epistemic</emph> in nature that involves image-texts (Tang et al., [<reflink idref="bib64" id="ref32">64</reflink>]; Yore &amp; Tang, [<reflink idref="bib71" id="ref33">71</reflink>]). By <emph>epistemic</emph>, learners need to draw on their understanding about the characteristics of science, such as source and justification of scientific claims, to interpret knowledge claims (Cheung, Pun and Fu, [<reflink idref="bib9" id="ref34">9</reflink>]). This is also evidenced by the correlation between conceptual gains and scientific epistemic beliefs (Ferguson et al., [<reflink idref="bib23" id="ref35">23</reflink>]; Yang et al., [<reflink idref="bib70" id="ref36">70</reflink>]).</p> <hd id="AN0191207466-5">Reading of Scientific Image-Texts in GenAI</hd> <p>Reading scientific texts has been a long tradition of scientific literacy (Glynn &amp; Muth, [<reflink idref="bib28" id="ref37">28</reflink>]). More recently, scientific literacy was defined as "understanding of scientific terminology and concepts, scientific enquiry and practice, and the interactions of science, technology, and society" (Jarman &amp; McClune, [<reflink idref="bib34" id="ref38">34</reflink>], p. 3). In a more recent version of "Vision III scientific literacy", the intersection between science, technology and society, as well as science and decisions has been a critical part of scientific literacy (Sjöström, [<reflink idref="bib59" id="ref39">59</reflink>]). Yet, in the era of GenAI, there is a need for empirical studies to conceptualise what scientific literacy means. Specifically, as part of scientific literacy influenced by GenAI, Tang ([<reflink idref="bib63" id="ref40">63</reflink>]) calls for a reading model in this area.</p> <p>The advent of GenAI in interpreting and creating image-texts calls for an update of a reading model within science education. Although there was a theorization of a GenAI-science reading model that comprises the content-interpretation, genre-reasoning and epistemic-evaluation domains (Authors, 2024), the reading model only focused on linguistic mode of GenAI instead of reading image-text outputs. Different GenAI applications, such as GPT-4 with DALL.E plugin, can handle and create pictures other than linguistic outputs (Gill et al., [<reflink idref="bib27" id="ref41">27</reflink>]; OpenAI, [<reflink idref="bib52" id="ref42">52</reflink>]). In the current state of educational research, the affordances of GenAI in relation to image-texts have been explored from a composing perspective (e.g., Liu et al., [<reflink idref="bib45" id="ref43">45</reflink>]), rather than a reading perspective. Reading image-texts is a prerequisite for interacting with GenAI sensibly. Other than interpreting the image-text outputs, students also need to acquire the competence to give further commands to prompt for improved outputs (Liu et al., [<reflink idref="bib45" id="ref44">45</reflink>]).</p> <p>Importantly, low-skilled comprehends require extensive scaffolding for reading image-texts (Meneses et al., [<reflink idref="bib48" id="ref45">48</reflink>]). These scaffolds involves that the image complements with the meaning with the linguistic texts (Meneses et al., [<reflink idref="bib48" id="ref46">48</reflink>]). In science education, despite the potential value of image-texts in shaping students' understanding of scientific issues, studies are still at an empirical level without developing a set of performance expectation (e.g., Fazio et al., [<reflink idref="bib22" id="ref47">22</reflink>]). In addition, for the disciplinary application of GenAI tools, there is a research gap on to what extent the generated image is scaffolded by the linguistic texts. There is an increasing interest in realising versions of GPT, such as GPT-4 V, to realise answering questions involving image-text (Lee &amp; Zhai, [<reflink idref="bib42" id="ref48">42</reflink>]). Still, there is reading model that does not characterise how students read scientific image-texts, particularly socio-scientific texts.</p> <p>While GenAI applications such as GPT-4 are capable of analyzing the molecular structure of a chemical (Alasadi &amp; Baiz, [<reflink idref="bib2" id="ref49">2</reflink>]), it is essential to underscore their epistemic foundation in the development of a reading model of image-texts in science. Based on enormous datasets trained with millions of images, GPT-4 utilises a mathematical processing system to interpret scientific images (Hatakeyama-Sato et al., [<reflink idref="bib32" id="ref50">32</reflink>]), such as the type and frequency of bonding of a chemical compound (Alasadi &amp; Baiz, [<reflink idref="bib2" id="ref51">2</reflink>]). To create scientific images, DALL.E decodes textual descriptions, segments them into smaller inputs, and then converts them to low latent representations that generate images (Singh et al., [<reflink idref="bib58" id="ref52">58</reflink>]). Owing to its reliance on training models, GenAI sometimes display stereotypic images of science classrooms, such as portraying a male teacher in a white lab coat (Cooper &amp; Tang, [<reflink idref="bib15" id="ref53">15</reflink>]). Researchers in science education have cautioned about biases in the training models (Avraamidou, [<reflink idref="bib3" id="ref54">3</reflink>]; Krist &amp; Kubsch, [<reflink idref="bib39" id="ref55">39</reflink>]), while students need to develop an informed epistemic understanding of how these GenAI systems create images and texts. It contrasts with evidence-based scientific representations that images in GenAI can lack direct empirical evidence, so they might generate scientific images like an incorrect rat male reproductive system with wrong labels and specification (Guo et al., [<reflink idref="bib30" id="ref56">30</reflink>]). Therefore, it is important for students to draw on their epistemic understanding of science and GenAI to critically <emph>read</emph> these image-texts in GenAI outputs.</p> <hd id="AN0191207466-6">Construct Map</hd> <p>To explore students' reading of socio-scientific image-texts in GenAI, we create a conceptual framework (construct map) that targets the expected performance set. Reading of scientific image-texts comprises three domains: the Interpretation (IN), Exchange (EX), and Evaluation (EV) (Fig. 1). Descriptions of each domain will be further elaborated below.</p> <p>Graph: Fig. 1 Conceptual framework of reading of socio-scientific image-texts in GenAI</p> <hd id="AN0191207466-7">The Interpretation (IN) Domain</hd> <p>This domain concerns interpretation, elaboration, and inferring elements socio-scientific image-texts in GenAI. These elements refer to both linguistic texts and images in GPT-4 outputs. In a previously conceptualized model of reading socio-scientific texts in GenAI linguistic outputs, students at a lower level can only detect fragments of ideas from linguistic texts (Oliveras et al., [<reflink idref="bib50" id="ref57">50</reflink>], [<reflink idref="bib51" id="ref58">51</reflink>]), whereas students at a higher level can provide a complete explanation of the linguistic scientific outputs (Cheung, Pun and Li, [<reflink idref="bib10" id="ref59">10</reflink>]). In the context of reading image-texts, students need to selectively interpret visual or written elements in GenAI outputs to make sense of the controversies and debates of the socio-scientific texts. More specifically, in reading texts related to climate science, students need to identify linguistic evidence from the outputs or imagery of how scientists examine evidence regarding the causes of climate science.</p> <hd id="AN0191207466-8">The Exchange (EX) Domain</hd> <p>This domain describes how students take actions to follow up outputs in GenAI using different actions, such as inputting linguistic texts or uploading files (e.g., images and pdf). In existing educational research, students use GenAI to create different modes of representation such as images and videos to facilitate their learning of content knowledge (Berg et al., [<reflink idref="bib4" id="ref60">4</reflink>]; Tilak et al., [<reflink idref="bib65" id="ref61">65</reflink>]). For example, the potential of GenAI has been discussed in language learning (Law, [<reflink idref="bib40" id="ref62">40</reflink>]) instead of science education. GenAI can motivate students to develop their reading and writing skills, but its disciplinary application in reading of scientific texts is limited. Yet, GenAI in social media might have deep-fake images in socio-scientific texts (Fatima, Kinger, &amp; Kumar, [<reflink idref="bib21" id="ref63">21</reflink>]). While generating socio-scientific image-texts, students need to critically interpret the image-text outputs (Rowsell et al., [<reflink idref="bib53" id="ref64">53</reflink>]) and determine if the current outputs satisfy users' needs. If the current outputs do not satisfy the user's needs, students also need to further enter prompts (Liu et al., [<reflink idref="bib45" id="ref65">45</reflink>]) to get more information. In the context of reading GenAI socio-scientific texts on climate change, if GPT-4 does not provide satisfying evidence or images related to climate science, students can prompt for more evidence as well as request a more "authentic" scientific image that truly represents the consensus among climate scientists. Such prompts need to be specific to science and grounded in scientific reasoning; otherwise, GPT-4 will give unspecific outputs.</p> <hd id="AN0191207466-9">The Evaluation (EV) Domain</hd> <p>This domain refers to how students evaluate images and linguistic texts by drawing on their understanding of the epistemologies of science and GenAI. When students encounter scientific claims, they <emph>adhere</emph> to the epistemic resources (e.g., understanding of characteristics of scientific knowledge and GenAI knowledge) to make such evaluations (Leung, [<reflink idref="bib43" id="ref66">43</reflink>], [<reflink idref="bib44" id="ref67">44</reflink>]). For example, in the context of climate change, ChatGPT give general answers without taking a stance on questions like limiting global warming to 1.5<sups>o</sups>C per year (Vaghefi et al., [<reflink idref="bib67" id="ref68">67</reflink>]). Although GenAI systems like Midjourney can create fake scientific images (Guo et al., [<reflink idref="bib30" id="ref69">30</reflink>]), they can also generate future-oriented images that visualize potential future hazards and solutions to mitigate the negative consequences of global warming (LC et al., [<reflink idref="bib41" id="ref70">41</reflink>]).</p> <p>To evaluate GenAI image-texts on climate science, students at least need to identify the mechanisms for GenAI to create linguistic outputs and image outputs. More specifically, GPT-4 generates linguistic outputs through large language models (Ferrara, [<reflink idref="bib24" id="ref71">24</reflink>]; Kasneci et al., [<reflink idref="bib35" id="ref72">35</reflink>]), while it creates images using diffusion models that encode linguistic texts as input (Derevyanko &amp; Zalevska, [<reflink idref="bib16" id="ref73">16</reflink>]). According to Cheung et al., ([<reflink idref="bib11" id="ref74">11</reflink>]), students' performance in this domain was considered as lower compared to other domains as students are not given explicit opportunities to discuss the validity of GenAI scientific information in lessons. Specifically, students equated ChatGPT to Google search without an informed understanding of large language models (Cheung et al., [<reflink idref="bib11" id="ref75">11</reflink>]).</p> <hd id="AN0191207466-10">Methodology</hd> <p>A construct-driven methodology (Wilson, [<reflink idref="bib69" id="ref76">69</reflink>]) was used to examine students' reading of socio-scientific image-texts in GenAI. Such a methodology helps align each item with the description of domains set out in the constructs (Schäpers et al., [<reflink idref="bib57" id="ref77">57</reflink>]), as well as setting levels expected for items of each domain. This further upholds construct validity of the instrument developed. Based on this approach, an instrument was designed following the three stages: (<reflink idref="bib1" id="ref78">1</reflink>) <emph>developing a conversation about climate science in GenAI</emph>; (<reflink idref="bib2" id="ref79">2</reflink>) <emph>designing items according to the construct map</emph>; (<reflink idref="bib3" id="ref80">3</reflink>) <emph>specifying the outcome spaces</emph>. The instrument was then administered to 68 junior science students to explore students' performance of reading of socio-scientific image-texts in GenAI.</p> <hd id="AN0191207466-11">Participants and Contexts</hd> <p>This study is part of a large-scale research project that focuses on enhancing junior secondary students' reading of scientific texts through the utilization of diverse digital tools, including AI and online platforms. The research team sent an invitation letter to all mainstream secondary schools in Hong Kong for inviting participation for an extra-curricular program. In this program, students took after-school lessons reading scientific texts on digital tools including GenAI and other sources. 68 junior secondary students from three subsidized secondary schools in Hong Kong signed up for this program (Table 1). Due to the packed schedule of our curriculum, we only administered the instrument at the end of the program to gauge the effectiveness of the reading program. However, the focus of this paper is not on measuring the effectiveness of the reading program, but rather on exploring students' reading of socio-scientific image-texts created by GenAI. The data reported in this study comes from three different subsidized secondary schools in Hong Kong: 19 students (12 females and 7 males; average age: 13.16, <emph>SD</emph> = 0.83) from School A, 24 students (10 females, 13 males, and 1 unreported; average age: 14.54, <emph>SD</emph> = 0.59) from School B, and 25 students (14 females, 10 males, and 1 unreported; average age: 13.28, <emph>SD</emph> = 0.98) from School C. The majority (<emph>n</emph> = 62) identify themselves as English as a second language learners, while a small number of students (<emph>n</emph> = 6) report English as their first language.</p> <p>Table 1 Demographics of participants (<emph>n</emph> = 68)</p> <p> <ephtml> &lt;table rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" rowspan="2"&gt;&lt;p&gt;Schools (grade level)&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="3"&gt;&lt;p&gt;Gender&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Language status&lt;/p&gt;&lt;/th&gt;&lt;th align="left" rowspan="2"&gt;&lt;p&gt;Average age (Standard Deviation)&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Male&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Female&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Unreported&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;English as a first language&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;English as a second language&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;A (grade 8)&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;7&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;12&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;0&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;18&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;13.16 (0.83)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;B (grade 9)&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;13&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;10&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;21&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;14.54 (0.59)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;C (grade 8)&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;10&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;14&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;23&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;13.28 (0.98)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Total&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;30&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;36&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;6&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;62&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;13.69 (1.03)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0191207466-12">Developing a Conversation about Climate Science in GenAI</hd> <p>To elicit students' reading of socio-scientific image-texts in GenAI, we inputted two debatable claims regarding climate change into GPT-4 with DALL.E plugins and asked GPT-4 to generate linguistic text with a picture to verify the claims. We then took screenshots of the two conversations and designed a set of items. The setting of the instrument places students in a scenario where a simulated student, Kevin, was engaged in a human-GPT conversation to verify claims on climate science. After reviewing several research papers regarding common conceptions related to climate change, the research team selected two claims that focus on different categories of the nature of science. In one claim, air pollution is identified as the major cause of global warming (Groves &amp; Pugh, [<reflink idref="bib29" id="ref81">29</reflink>]), and GPT-4 was prompted to give further information on how scientists could verify this. In another claim, GPT-4 was asked about the authenticity of the Oregon petition (Idso et al., [<reflink idref="bib33" id="ref82">33</reflink>]), in which scientists collectively signed the statement that global warming is not the cause of climate change. For the first claim, GPT-4 responded by discussing the <emph>diversity of scientific methods</emph> used to study the link between air pollution and global warming. For the second claim, GPT-4 responded by exploring the role of <emph>social organizations and interactions</emph> among scientists in reaching a consensus on human-driven climate change. As argued by Erduran and Dagher ([<reflink idref="bib19" id="ref83">19</reflink>]), science is embedded in both cognitive-epistemic and social-institutional systems. The two claims were situated in different systems in relation to the epistemic aspects of science, resulting in a holistic articulation of how scientists work in researching climate change.</p> <p>In previous studies, the analysis of socio-scientific texts among upper secondary students ranged from 210 to 240 words (Stang Lund et al., [<reflink idref="bib61" id="ref84">61</reflink>]). Due to the heightened cognitive demands of the interactive exchanges between ChatGPT and the simulated student, our recent research, which explored students' linguistic comprehension of ChatGPT output, limited such dialogues to under 200 words (Authors, 2024). Owing to the presence of an image in conversation regarding climate science, the co-interpretation of both an image and written text presents further cognitive demand to students (Ainsworth, [<reflink idref="bib1" id="ref85">1</reflink>]). In this study, given these considerations, we opted to restrict GPT-4 to generating outputs under 150 words. Despite our instructions, GPT-4 frequently exceeded this limit in its responses. To address this, we consistently reminded GPT-4 to adhere to the word count constraint. We prompted GPT-4 "Could you please limit your response up to 150 words?" before its generation of socio-scientific image-texts. Hence, the conversation (Conversation A) on the <emph>diversity of scientific methods</emph> was 115 words, while the other conversation (Conversation B) on <emph>social organisations and interactions</emph> was 120 words. Both conversations, targeting junior form students, exhibit similar Flesch-Kincaid Readability (Flesch, [<reflink idref="bib25" id="ref86">25</reflink>]).</p> <hd id="AN0191207466-13">Designing Items According to the Construct Map</hd> <p>After capturing screenshots of the two conversations, we developed a set of items targeting the three domains: the Interpretation (IN), Exchange (EX), and Epistemic Evaluation (EV). In each conversation, two items were crafted for each domain. In the IN domain, items CA-IN1 and CB-IN1 focused on identifying and elaborating on visual and linguistic elements within the conversations, while CA-IN2 and CB-IN2 delved into recognizing aspects of the nature of science within the visual and linguistic elements. Within the EX domain, items CA-EX1 and CB-EX1 prompted students to input linguistic commands, while CA-EX2 and CB-EX2 required students to upload a file for follow-up. For the EV domain, items CA-EV1 and CB-EV1 encouraged students to evaluate linguistic output in GenAI, while CA-EV2 and CB-EV2 prompted evaluation of the images in GenAI. These items were then tested with 19 eighth-grade students. Most students successfully completed the items, confirming the appropriateness of the instrument design based on our prior experience (Authors, 2024). The responses from the pilot study, part of a sample of 68 students, were also included in our analysis.</p> <hd id="AN0191207466-14">Specifying the Outcome Spaces</hd> <p>Following a construct-driven approach, we systematically developed a hierarchy of constructs targeting various levels of competence across the three domains (Appendix Table S1), with higher student competence indicating success in image-text reading. Additionally, specific performance expectations were assigned to the two conversations in the two domains (Appendix Table S2). Using these performance expectations, we constructed an outcome space for each conversation and item (Appendix Table S3 and S4). The outcome space underwent multiple revisions to align with the construct map while also considering students' performance.</p> <p>The items underwent review by three experienced in-service teachers who facilitated communication with the research team regarding student participation. While two teachers found the instrument clear, one suggested that it might be too lengthy for students. Despite receiving this feedback, the research team chose to retain all items in order to comprehensively address multiple domains within the cognitive-epistemic and social-institutional systems.</p> <hd id="AN0191207466-15">Data Analysis</hd> <p>The second and third authors held multiple meetings to discuss the rating rubric. To familiarize themselves with it, both authors rated the responses of five students and deliberated on any ambiguities in the rubric. For instance, in CA-EX1, responses in Chinese were accepted since the study focuses on reading rather than writing in science. Similarly, for items CB-EX1/CB-EV2, responses mentioning "human activities" and "petition" without directly articulating climate change were considered as references to climate change. After resolving ambiguities in the rubric, both authors rated the responses of 25 students, which accounted for 37% of the total responses.</p> <p>Interrater reliability was assessed to ensure consistency between the two raters. As Cheung and Tai ([<reflink idref="bib12" id="ref87">12</reflink>]) suggested, Cohen's κ needs to be calculated for each item instead of a summation of items to inform the readers about the reliability of each item. For CA-IN1, a Cohen's κ of 0.875 (<emph>p</emph> &lt;.001) was obtained; for CA-IN2, a Cohen's κ of 0.843 (<emph>p</emph> &lt;.001) was obtained; for CA-EX1, a Cohen's κ of 0.851 (<emph>p</emph> &lt;.001) was obtained; for CA-EX2, a Cohen's κ of 0.853 (<emph>p</emph> &lt;.001) was obtained; for CA-EV1, a Cohen's κ of 0.838 (<emph>p</emph> &lt;.001) was obtained; for CA-EV2, a Cohen's κ of 0.938 (<emph>p</emph> &lt;.001) was obtained. Regarding conversation B, for CB-IN1, a Cohen's κ of 0.891 (<emph>p</emph> &lt;.001) was obtained; for CB-IN2, a Cohen's κ of 0.813 (<emph>p</emph> &lt;.001) was obtained; for CB-EX1, a Cohen's κ of 0.951 (<emph>p</emph> &lt;.001) was obtained; for CB-EX2, a Cohen's κ of 0.943 (<emph>p</emph> &lt;.001) was obtained; for CB-EV1, a Cohen's κ of 0.921 (<emph>p</emph> &lt;.001) was obtained; for CB-EV2, a Cohen's κ of 0.935 (<emph>p</emph> &lt;.001) was obtained. The range of Cohen's κ, from 0.813 to 0.938, indicated good interrater reliability.</p> <p>In addition to the quantitative analysis, qualitative content analysis (Elo &amp; Kyngäs, [<reflink idref="bib18" id="ref88">18</reflink>]) was conducted to explore students' epistemic understanding of image-text GenAI outputs in socio-scientific texts. In the Exchange domain, we are interested in what type of file students will upload to prompt GPT-4, as well as the reason for this. We inductively coded students' responses in CA-EX2 and CB-EX2, unfolding what file students would upload and their explanation (see Fig. 2 in the result section). A Sankey diagram was plotted to visualize the link between the file they uploaded and the corresponding reason. More importantly, we also inductively analysed students' responses regarding their understanding of nature of science and nature of GenAI among CA-EV1, CA-EV2, CB-EV1 and CB-EV2 (see Fig. 3 in the result section). We visualized the distribution of different kinds of epistemic understanding across items using two heat maps, with one heat map on nature of science and another heat map on nature of GenAI.</p> <p>Graph: Fig. 2 Type of file and function of the input in (a) item CA-EX2 and (b) item CB-EX2</p> <p>Graph: Fig. 3 Students' mentioning of (a) nature of science and (b) nature of GenAI across items</p> <hd id="AN0191207466-16">Results</hd> <p></p> <hd id="AN0191207466-17">Students' Performance in Reading Socio-Scientific Image-Texts</hd> <p>The mean score of each domain for each conversation was summed up and visualized in the bar plot (Fig. 4(a)). The IN domain has the highest mean (Conversation A: <emph>M</emph> = 3.81, <emph>SD</emph> = 1.83; Conversation B: <emph>M</emph> = 3.29, <emph>SD</emph> = 1.68). Next, students' performances in the IN domain in both conversations (Conversation A: <emph>M</emph> = 3.25, <emph>SD</emph> = 1.50; Conversation B: <emph>M</emph> = 3.16, <emph>SD</emph> = 1.60) lie between their performance in the EX and EV domains. The EV domain is the lowest performing domain (Conversation A: <emph>M</emph> = 3.14, <emph>SD</emph> = 1.08; Conversation B: <emph>M</emph> = 2.89, <emph>SD</emph> = 1.06) among all the three domains in the two conversations. In Conversation A, the EX domain exhibited a significantly higher mean than the IN domain (<emph>t</emph> (<reflink idref="bib67" id="ref89">67</reflink>) = 2.068, <emph>p</emph> &lt;.05), while the MI domain did not show a significantly higher mean than the EV domain (<emph>t</emph> (<reflink idref="bib67" id="ref90">67</reflink>) = 0.448, <emph>p</emph> =.655). No significant differences were found across domains within Conversation B.</p> <p>Graph: Fig. 4 (a) Bar plot of mean score of each domain for each conversation; (b) Bar plot of item mean score for each conversation (N = 68)</p> <p>A one-way repeated measures ANOVA with Greenhouse-Geisser correction was conducted to explore any significant differences between different domains in the two passages. Students' image-text reading significantly differed across the six domains (3 domains x 2 conversations) (<emph>F</emph>(4.168) = 31.049, <emph>p</emph> &lt;.05). Given this finding, the variances in students' performance between different conversations within the same domain of image-text reading were further examined. Three post-hoc paired sample t-tests were carried out to compare students' performance between the two passages in the identical domains of image-text reading. No significant differences were observed in performance in image-text reading between different conversations in the IN (<emph>t</emph> (<reflink idref="bib67" id="ref91">67</reflink>) = 0.322, <emph>p</emph> =.748) and EX (<emph>t</emph> (<reflink idref="bib67" id="ref92">67</reflink>) = 1.744, <emph>p</emph> =.086) domains. However, there was a statistically significant difference in the EV domain between conversations A and B (<emph>t</emph> (<reflink idref="bib67" id="ref93">67</reflink>) = 2.240, <emph>p</emph> &lt;.05). This indicates that students exhibited stronger performance in drawing upon their epistemic understanding of science and GenAI when reading a cognitive-epistemic image-text text on climate science than on a social-institutional image-texts on climate science.</p> <p>Students' performance on individual items were also analysed (Fig. 4(b)). In Conversation A, the best performing item is EX1 (<emph>M</emph> = 1.97, <emph>SD</emph> = 1.21) which requires students to input texts to follow the conversation on the diversity of scientific methods on the causes of global warming; in Conversation B, the best performing item is IN1 (<emph>M</emph> = 2.04, <emph>SD</emph> = 1.15) which requires students to interpret elements in linguistic and visual representations regarding the Oregon petition. In Conversation A, students found it easier to follow up with linguistic input in comparison with other items; in Conversation B, students found it easier to select elements from image-texts. For both conversations, IN2 had the lowest performance (Conversation A: <emph>M</emph> = 1.38, <emph>SD</emph> = 0.93; Conversation B: <emph>M</emph> = 1.11, <emph>SD</emph> = 0.76). This might be attributed to students finding it challenging to provide a more elaborated description of how more than two elements in the picture and words are related to global warming.</p> <p>The intra-domain differences in item performance within the same conversation were also examined (Table 2). Such an examination is important, particularly for the EX and EV domains, as one item in these two domains targets linguistic output while another item targets image output. There were no significant differences between items within the EX domain in both Conversation A (<emph>t</emph> (<reflink idref="bib67" id="ref94">67</reflink>) = 0.837, <emph>p</emph> =.405) and Conversation B (<emph>t</emph> (<reflink idref="bib67" id="ref95">67</reflink>) = 0.472, <emph>p</emph> =.639), as well as the EV domain in both Conversation A (<emph>t</emph> (<reflink idref="bib67" id="ref96">67</reflink>) = 0.191, <emph>p</emph> =.849) and Conversation B (<emph>t</emph> (<reflink idref="bib67" id="ref97">67</reflink>) = -1.628, <emph>p</emph> =.108). This indicates that students' performance in interacting with GPT-4 and drawing on their epistemic understanding to evaluate outputs is similar regardless of visual or linguistic output. In contrast, within the IN domain, IN2 showed significantly lower performance than IN1 in both Conversation A (<emph>t</emph> (<reflink idref="bib67" id="ref98">67</reflink>) = 3.260, <emph>p</emph> =.002) and Conversation B (<emph>t</emph> (<reflink idref="bib67" id="ref99">67</reflink>) = 6.798, <emph>p</emph> &lt;.001). This might be primarily due to the fact that students might need to involve more writing in the IN2 item compared to the IN1 item.</p> <p>Table 2 Paired sampled t-tests for items within the same domain of the same conversation</p> <p> <ephtml> &lt;table rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Paired differences&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Mean&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;SD&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;t&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;df&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Two-sided &lt;italic&gt;p&lt;/italic&gt; value&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CA-IN1 &amp;#8211; CA&amp;#8211;IN2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.49&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.23&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;3.260&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.002&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CB-IN1 &amp;#8211; CB-IN2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.93&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.12&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;6.798&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#60; 0.001&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CA-EX1 &amp;#8211; CA-EX2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.13&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.30&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.837&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.405&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CB-EX1 &amp;#8211; CB-EX2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.09&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.54&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.472&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.639&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CA-EV1 &amp;#8211; CA-EV2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.03&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.27&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.191&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.849&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;CB-EV1 &amp;#8211; CB-EV2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#8722; 0.22&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;1.12&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;-1.628&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;67&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.108&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0191207466-18">The Interpretation Domain</hd> <p>Content analysis was performed on individual items that draw our attention. As the IN2 item in each conversation had the lowest score among all items, we qualitatively examined students' responses to these items (Appendix 2 Table S1). In Conversation A, despite the emphasis on the diversity of scientific methods by both visual and linguistic elements in the GPT-4 conversation, six students (8.8%) expressed that there was only a single kind of scientific method. Interestingly, when students were asked to explain their reasons for deciding if there is a single kind of scientific method, most students (<reflink idref="bib42" id="ref100">42</reflink>, 61.8%) based their decisions on linguistic elements in Conversation A, despite the depiction of scientists using computer models and collecting samples from the ground in the image. In contrast, in Conversation B, more than one-fifth of the students (<reflink idref="bib16" id="ref101">16</reflink>, 23.5%) identified that scientists did not collaborate or communicate directly, potentially reflecting challenges in interpreting the question or the generated image-text texts, especially considering the majority of students identified as English as a second language learners. However, the interaction between scientists was clearly depicted both in the image and the text in Conversation B. The reasons for students selecting this choice were mainly grounded in their interpretation of the visual elements (<reflink idref="bib52" id="ref102">52</reflink>, 76.5%).</p> <p>Two salient findings emerged from this analysis. The first one was that certain students struggled to recognize the accurate attributes of science within the discussions, despite GPT-4 outputs presenting this information cohesively through both images and text. The other one was that students leaned towards relying on the textual content in GPT-4 outputs to comprehend the diversity of scientific methods, while they predominantly used the visual elements in GPT-4 outputs to interpret the interactions between scientists.</p> <hd id="AN0191207466-19">The Exchange Domain</hd> <p>Inductive content analysis was employed to analyze students' responses in the EX domain. In the item EX1 (Appendix 2 Table S2), students were immersed in the GPT-4 scenario concerning a human-AI conversation on methods to study the link between air pollution and global warming (Conversation A) and the Oregon petition (Conversation B). Regarding IN1, which involved students' follow-up actions using linguistic inputs, the majority of students in both Conversations A (<reflink idref="bib16" id="ref103">16</reflink>, 23.5%) and B (<reflink idref="bib10" id="ref104">10</reflink>, 4.7%) provided generic responses unrelated to climate science. In the second most common linguistic input scenario in the GPT-4 scenario, in Conversation A, nine students (13.2%) sought explanations for the link between air pollution and global warming, while in Conversation B, 12 students (17.6%) asked for more details about the petition.</p> <p>Interestingly, certain response categories emerged in students' reactions to Conversation B rather than Conversation A, such as "expressing stance" (<reflink idref="bib2" id="ref105">2</reflink>, 2.9%) and "asking for a solution" (<reflink idref="bib1" id="ref106">1</reflink>, 1.5%). These categories, specific to Conversation B and related to the climate scenario presented, did not align with the conversational context and did not aim to prompt a human-GenAI conversation. Additionally, irrelevant prompts were observed in both conversations, such as "Could people live in space" (CA-EX1, ID11) and "When can we move to the moon?" (CB-EX1, ID34). Despite students engaging in scientific reading, their prompts appeared detached from the GPT-4 scenario and only reflected personal preferences.</p> <p>In addition to the item EX1, students' responses to EX2 were analysed regarding the type of file they would upload and their justifications for such uploads (Fig. 2). In Conversation A, the majority of students (<reflink idref="bib18" id="ref107">18</reflink>, 26.4%) chose to upload an image without providing justifications; in Conversation B, most students (<reflink idref="bib21" id="ref108">21</reflink>, 30.9%) selected an image without a clear purpose such as training or prompting GPT-4. Some specific functions emerged from students' responses, tailored to each conversation. For instance, in Conversation A, one student opted to upload text for science communication, believing that "<emph>This file is about a list of scientists I attach this file because it can ask chatgpt to tell inform about the scientists</emph>". Conversely, in Conversation B, a student wished to upload "<emph>a news about moving to the Moon</emph>" to engage in two-way dialogue with ChatGPT.</p> <hd id="AN0191207466-20">The Epistemic Evaluation Domain</hd> <p>We also explored the ideas about the nature of GenAI and science that students drew upon when evaluating the linguistic output (CA-EV1 and CB-EV1) and image output (CA-EV1 and CB-EV2). The concepts related to the nature of science included uncertainty, claims based on textbooks, reliance on textual data, tentativeness, scientific consensus, reliability/unreliability, publication in papers, limitations, Internet searches, embracing the diversity of methods, and empirical or non-empirical nature. The predominant ideas about the nature of science across all four items (CA-EV1: 16.2%; CA-EV2: 13.2%; CB-EV1: 13.2%; CB-EV2: 11.8%) were that science is limited when students evaluated both linguistic and image outputs in the GPT-4 scenarios.</p> <p>In contrast to students' ideas about the nature of science, students drew upon a more diverse range of ideas about the nature of GenAI. These ideas included word processing tools, unrealism, uncertainty, reliance on textual data, tentativeness, image generation tools, machines/robots, Internet searches, intelligent systems, fact checkers, human-like entities, datasets, controllers, information sourced from climate news, and information based on climate models. The prevailing notion about the nature of GenAI was that GPT-4 with DALL.E plugins retrieves information from the Internet (CA-EV1: 22.1%; CA-EV2: 25%; CB-EV1: 27.9%; CB-EV2: 19.1%). Evidently, students did not possess knowledge about the diffusion models behind DALL.E, as indicated by the second prevailing idea of "image generation tools" in items CA-EV2 (14.7%) and CB-EV2 (19.1%). Interestingly, the third prevailing ideas in these items were "human" (CA-EV2: 7.4%; CB-EV2: 10.3%), such as photographers. Students attributed human characteristics to GenAI, influencing their assessment of the image output in the GPT-4 scenario.</p> <hd id="AN0191207466-21">Students' Image-Text Reading Across Grade Levels and Prior Exposure to ChatGPT</hd> <p>According to a series of independent sample <emph>t</emph>-tests, prior exposure did not significantly impact students' image-text performance in all domains and items (<emph>p</emph> &gt;.05). Concerning differences across grade levels, ninth graders achieved a significantly higher score in the MI domain in Conversation B compared to eighth graders (<emph>t</emph> (4.355, 66) = 2.815, <emph>p</emph> &lt;.01). This difference might be attributed to the fact that ninth graders had greater exposure to the socio-political dimensions of climate change through news media, enabling them to better interpret the visual and written components of Conversation B compared to Conversation A. However, for other domains and conversations, there was no significant difference between eighth and ninth graders.</p> <hd id="AN0191207466-22">Discussion and Conclusion</hd> <p>Based on our previous work regarding how students read linguistic mode of socio-scientific texts created by ChatGPT (Cheung et al., [<reflink idref="bib11" id="ref109">11</reflink>]), this present study furthers GenAI research in science education by theorising reading of socio-scientific image-texts in GenAI, as well as developing an instrument to explore students' performance. The theoretical and practical implications were discussed below.</p> <hd id="AN0191207466-23">Domains of Reading of GenAI-Created Socio-Scientific Image-Texts</hd> <p>The three domains of reading of socio-scientific image-texts created by GenAI did not exhibit a hierarchal structure. This claim is supported by the fact that students' performance in the domains within the conversations did not significantly differ from one another, except for the observation that the EX domain was significantly higher than the IN domain within Conversation A. Among the three domains, students' performance in one domain of reading is not necessarily higher than that on another domain of reading. Yet, this study indicates that students' performance in all three domains is not statistically significantly different from one another. Students' skills to interpret, exchange and evaluate might be independent of each other and can be taught separately. Also, there was an increasing number of opportunities for students to interact with ChatGPT in learning the cognitive-epistemic dimension of science, such as prompting image- texts for evidence on acceleration (Ng et al., [<reflink idref="bib49" id="ref110">49</reflink>]) and for engaging with scientific practices like tabulating data (Alasadi &amp; Baiz, [<reflink idref="bib2" id="ref111">2</reflink>]). This phenomenon might be attributed to the fact that within the cognitive-epistemic texts, students were already familiarised with the human-GenAI exchange.</p> <p>Another observation, consistent with findings from previous studies (Cheung et al., [<reflink idref="bib12" id="ref112">12</reflink>], [<reflink idref="bib11" id="ref113">11</reflink>]), was that students' performance in the EV domain in the GenAI socio-institutional image-texts was significantly lower than that in the cognitive-epistemic texts. This finding was also observed in students' reading of linguistic outputs on socio-scientific issues created by the ChatGPT scenario (Cheung et al., [<reflink idref="bib11" id="ref114">11</reflink>]). Two possible reasons could be attributed to this finding: firstly, the curriculum did not provide many opportunities for immersing students in broader epistemologies of science (Cheung, [<reflink idref="bib7" id="ref115">7</reflink>]; Caramaschi et al., [<reflink idref="bib6" id="ref116">6</reflink>]; Kaya &amp; Erduran, [<reflink idref="bib36" id="ref117">36</reflink>]), such as scientists signing petitions against claims of human-driven climate change. Although articulations of these social-institutional aspects of epistemologies of science were common on social media during different crisis such as Covid-19 (Cheung, Chan and Erduran, [<reflink idref="bib8" id="ref118">8</reflink>]), students lacked the competence to connect these epistemic aspects in their <emph>reading</emph> of these socio-scientific texts. Apart from their understanding of the epistemic aspects of science, students had misunderstandings about the epistemic aspects of multimodal GenAI, such as perceiving DALL.E as a Google search engine, human, or a picture creator. This lack of understanding was particularly worrying as students did not know how such images related to the works of scientists were generated.</p> <hd id="AN0191207466-24">Students' Performance Across Grade Levels and Prior Exposure to ChatGPT</hd> <p>Our findings also indicated that students' prior exposure to ChatGPT did not make a significant difference in their reading of socio-scientific image-texts created by GPT-4. Regarding the IN domain, it seems that many students lacked training in interpreting elements of image-text outputs. This was evidenced by the finding that 21 students (30.9%) (Appendix 2 Table S1) provided irrelevant reasons for their interpretations in CA-IN1, rather than describing and elaborating on the visual and linguistic elements in GenAI's image-text outputs. Additionally, in the EX domain, students struggled with giving specific prompts, with a higher percentage of students (Appendix 2 Table S2) offering generic and irrelevant prompts in response to the GenAI image-text outputs. As for the EV domain, this observation could be attributed to the notion that merely engaging students in scientific activities, including reading scientific texts, may not be sufficient to develop their epistemic understanding of science (Khishfe, [<reflink idref="bib37" id="ref119">37</reflink>]; Khishfe &amp; Abd-El-Khalick, [<reflink idref="bib38" id="ref120">38</reflink>]).</p> <hd id="AN0191207466-25">Practical Implications</hd> <p>Our findings underscore the need to develop students' image-text and epistemic reading skills when engaging with socio-scientific texts in GenAI outputs. For example, explicit literacy instruction could integrate the three domains stipulated in our construct map: the interpretation, exchange, and epistemic-evaluation domains. Science teachers, for instance, may guide students to identify the visual and linguistic components within GenAI outputs and prompt them to analyze the similarities and differences between these elements. Subsequently, teachers may engage students in discussions about the types of prompts—whether linguistic prompts or file uploads—that could be used to elicit more scientifically informed conversations with GPT. Teachers may then guide students in reflecting on how GenAI generates these scientific texts, drawing comparisons with how scientists create, validate, and revise scientific claims (Cheung et al., [<reflink idref="bib11" id="ref121">11</reflink>]). This instructional model can also be supported by genre-based teaching methods that target various forms of scientific texts (Tang, [<reflink idref="bib62" id="ref122">62</reflink>]), such as argumentation, informational reports, experimental accounts, and scientific explanations. The newly developed instrument in this study can be used to assess the effectiveness of such explicit literacy instruction within an experimental design.</p> <hd id="AN0191207466-26">Limitations and Future Research Directions</hd> <p>Two limitations of this study need to be acknowledged. Firstly, due to the small sample size, we did not conduct Rasch analysis to validate the differences across levels within the same domain. However, as an exploratory study, we combined quantitative and qualitative evidence to reveal diverse students' performances in reading of GenAI image-text outputs. Further research could be conducted with a larger sample of students using a wider variety of scientific texts. Additionally, to ensure consistency in measuring students' reading performance, we did not engage students in real-time interaction with GenAI. The screenshotted conversation did not capture a series of back-and-forth interaction between the user and GenAI, as we did not have such data. Future studies can first collect data of real-time interaction between GenAI and students. Then, the screenshotted interaction can be used to design such an instrument. Yet, this study presents one of the fruitful methodological approaches for assessing students' reading of socio-scientific image-texts in GenAI. Such a methodological approach provides an alternative to Likert-scale questionnaire (e.g., Wang et al., [<reflink idref="bib68" id="ref123">68</reflink>]). The combination of quantitative scoring and qualitative content analysis can unveil diverse student responses to the three domains of reading of socio-scientific image-texts, a depth of insight that may not be achievable through Likert-scale questionnaires. Importantly, the instrument provided in this study elicited students' <emph>epistemic adherence</emph> (Leung, [<reflink idref="bib43" id="ref124">43</reflink>], [<reflink idref="bib44" id="ref125">44</reflink>]) as they drew upon their epistemic understanding to interpret such scientific texts. Future research studies could compare students' performances using our instrument and their real-time interactions with GenAI.</p> <hd id="AN0191207466-27">Author Contributions</hd> <p>All authors contributed to the study conception and design. The first draft of the manuscript was written by JP and KC, and all authors commented on versions of the manuscript. All authors read and approved the final manuscript.</p> <hd id="AN0191207466-28">Funding</hd> <p>The study reported in this manuscript is based upon work supported by the Quality Education Fund, Hong Kong SAR Government (Grant Number: 9420033).</p> <hd id="AN0191207466-29">Data Availability</hd> <p>The data is not available publicly owing to restrictions of the ethical approval.</p> <hd id="AN0191207466-30">Declarations</hd> <p></p> <hd id="AN0191207466-31">Ethical Approval</hd> <p>The first author obtained ethical approval from his institution. This project was approved by Human and Artefacts Ethics Sub-Committee at City University of Hong Kong (Application Number: HU-STA-00000011).</p> <hd id="AN0191207466-32">Informed Consent</hd> <p>Informed consent was obtained from all individual participants included in the study.</p> <hd id="AN0191207466-33">Consent to Participate</hd> <p>Informed consent was obtained from all individual participants included in the study.</p> <hd id="AN0191207466-34">Consent for Publication</hd> <p>Informed consent was obtained from all individual participants included in the study.</p> <hd id="AN0191207466-35">Competing Interests</hd> <p>The authors declare no competing interests.</p> <hd id="AN0191207466-36">Electronic Supplementary Material</hd> <p>Below is the link to the electronic supplementary material.</p> <p>Graph: Supplementary Material 1</p> <hd id="AN0191207466-37">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0191207466-38"> <title> References </title> <blist> <bibl id="bib1" idref="ref16" type="bt">1</bibl> <bibtext> Ainsworth S. DeFT: A conceptual framework for considering learning with multiple representations. Learning and Instruction. 2006; 16; 3: 183-198. 10.1016/j.learninstruc.2006.03.001</bibtext> </blist> <blist> <bibl id="bib2" idref="ref3" type="bt">2</bibl> <bibtext> Alasadi, E. A, &amp; Baiz, C. R. (2024). Multimodal generative artificial intelligence tackles visual problems in chemistry. Journal of Chemical Education.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref18" type="bt">3</bibl> <bibtext> Avraamidou, L. (2024). Can we disrupt the momentum of the AI colonization of science education? Journal of Research in Science Teaching.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref60" type="bt">4</bibl> <bibtext> Berg, C, Omsén, L, Hansson, H, &amp; Mozelius, P. (2024). Students' AI-generated Images: Impact on Motivation, Learning and, Satisfaction. Paper presented at the ICAIR 2024.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref4" type="bt">5</bibl> <bibtext> Bewersdorff, A, Hartmann, C, Hornberger, M, Seßler, K, Bannert, M, Kasneci, E, &amp; Nerdel, C. (2024). Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education. arXiv preprint arXiv:2401.00832.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref116" type="bt">6</bibl> <bibtext> Caramaschi M, Cullinane A, Levrini O, Erduran S. Mapping the nature of science in the Italian physics curriculum: From missing links to opportunities for reform. International Journal of Science Education. 2022; 44; 1: 115-135. 10.1080/09500693.2021.2017061</bibtext> </blist> <blist> <bibl id="bib7" idref="ref115" type="bt">7</bibl> <bibtext> Cheung, K. K. C. (2020). Exploring the inclusion of nature of science in biology curriculum and high-stakes assessments in Hong Kong: Epistemic network analysis. Science &amp; Education, 29(3), 491–512.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref118" type="bt">8</bibl> <bibtext> Cheung, K. K. C, Chan, H. Y, &amp; Erduran, S. (2023). Communicating science in the COVID-19 news in the UK during Omicron waves: exploring representations of nature of science with epistemic network analysis. Humanities and Social Sciences Communications, 10(1), 1–14.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref34" type="bt">9</bibl> <bibtext> Cheung, K. K. C, Pun, J. K, &amp; Fu, X. (2024). Development and validation of a Reading in Science Holistic Assessment (RISHA): A Rasch measurement study. International Journal of Science and Mathematics Education, 22(7), 1537–1561.</bibtext> </blist> <blist> <bibtext> Cheung, K. K. C, Pun, J. K, &amp; Li, W. (2024). Students' holistic reading of socio-scientific texts on climate change in a ChatGPT scenario. Research in Science Education, 54(5), 957–976.</bibtext> </blist> <blist> <bibtext> Cheung, K. K. C, Long, Y, Liu, Q, &amp; Chan, H. Y. (2024). Unpacking epistemic insights of artificial intelligence (AI) in science education: A systematic review. Science &amp; Education, 1–31.</bibtext> </blist> <blist> <bibtext> Cheung, K. K. C, &amp; Tai, K. W. (2023). The use of intercoder reliability in qualitative interview data analysis in science education. Research in Science &amp; Technological Education, 41(3), 1155–1175.</bibtext> </blist> <blist> <bibtext> Clark, D, Ranney, M, &amp; Felipe, J. (2013). Knowledge helps: Mechanistic information and numeric evidence as cognitive levers to overcome stasis and build public consensus on climate change. Paper presented at the Proceedings of the Annual Meeting of the Cognitive Science Society.</bibtext> </blist> <blist> <bibtext> Cook J, Ellerton P, Kinkead D. Deconstructing climate misinformation to identify reasoning errors. Environmental Research Letters. 2018; 13; 2: 024018. 10.1088/1748-9326/aaa49f</bibtext> </blist> <blist> <bibtext> Cooper, G, &amp; Tang, K. S. (2024). Pixels and pedagogy: Examining science education imagery by generative artificial intelligence. Journal of Science Education and Technology, 1–13.</bibtext> </blist> <blist> <bibtext> Derevyanko N, Zalevska O. Comparative analysis of neural networks midjourney, stable diffusion, and DALL-E and ways of their implementation in the educational process of students of design specialities. Scientific Bulletin of Mukachevo State University Series Pedagogy and Psychology. 2023; 9; 3: 36-44</bibtext> </blist> <blist> <bibtext> Diakidoy IAN, Stylianou P, Karefillidou C, Papageorgiou P. The relationship between listening and reading comprehension of different types of text at increasing grade levels. Reading Psychology. 2005; 26; 1: 55-80. 10.1080/02702710590910584</bibtext> </blist> <blist> <bibtext> Elo S, Kyngäs H. The qualitative content analysis process. Journal of Advanced Nursing. 2008; 62; 1: 107-115. 10.1111/j.1365-2648.2007.04569.x</bibtext> </blist> <blist> <bibtext> Erduran, S, &amp; Dagher, Z. R. (2014). Reconceptualizing nature of science for science education. Springer.</bibtext> </blist> <blist> <bibtext> Farrell J, McConnell K, Brulle R. Evidence-based strategies to combat scientific misinformation. Nature Climate Change. 2019; 9; 3: 191-195. 10.1038/s41558-018-0368-6</bibtext> </blist> <blist> <bibtext> Fatima, N, Kinger, P, &amp; Kumar, A. (2024). Detection of real versus fake images on social media through generative. Generative AI: Current Trends and Applications, 87.</bibtext> </blist> <blist> <bibtext> Fazio X, Gallagher TL, DeKlerk C. Exploring adolescents' critical reading of socioscientific topics using multimodal texts. International Journal of Science and Mathematics Education. 2022; 20; Suppl 1: 93-116. 10.1007/s10763-022-10280-8</bibtext> </blist> <blist> <bibtext> Ferguson LE, Bråten I, Strømsø HI, Anmarkrud Ø. Epistemic beliefs and comprehension in the context of reading multiple documents: Examining the role of conflict. International Journal of Educational Research. 2013; 62: 100-114. 10.1016/j.ijer.2013.07.001</bibtext> </blist> <blist> <bibtext> Ferrara, E. (2024). GenAI against humanity: Nefarious applications of generative artificial intelligence and large Language models. Journal of Computational Social Science, 1–21.</bibtext> </blist> <blist> <bibtext> Flesch, R. (2007). Flesch-Kincaid readability test. Retrieved October, 26(3), 2007.</bibtext> </blist> <blist> <bibtext> Franzen, C. (2024). Science journal retracts peer-reviewed article containing AI generated 'nonsensical' images. Retrieved from https://venturebeat.com/ai/science-journal-retracts-peer-reviewed-article-containing-ai-generated-nonsensical-images/</bibtext> </blist> <blist> <bibtext> Gill SS, Xu M, Patros P, Wu H, Kaur R, Kaur K, Parlikad AK. Transformative effects of ChatGPT on modern education: Emerging era of AI chatbots. Internet of Things and Cyber-Physical Systems. 2024; 4: 19-23. 10.1016/j.iotcps.2023.06.002</bibtext> </blist> <blist> <bibtext> Glynn SM, Muth KD. Reading and writing to learn science: Achieving scientific literacy. Journal of Research in Science Teaching. 1994; 31; 9: 1057-1073. 10.1002/tea.3660310915</bibtext> </blist> <blist> <bibtext> Groves, F. H, &amp; Pugh, A. F. (1996). College Students' Misconceptions of Environmental Issues Related to Global Warming.</bibtext> </blist> <blist> <bibtext> Guo X, Dong L, Hao D. RETRACTED: Cellular functions of spermatogonial stem cells in relation to JAK/STAT signaling pathway. Frontiers in Cell and Developmental Biology. 2024; 11: 1339390. 10.3389/fcell.2023.1339390</bibtext> </blist> <blist> <bibtext> Harris KR. Liars and Trolls and bots online: The problem of fake persons. Philosophy &amp; Technology. 2023; 36; 2: 35. 10.1007/s13347-023-00640-9</bibtext> </blist> <blist> <bibtext> Hatakeyama-Sato K, Yamane N, Igarashi Y, Nabae Y, Hayakawa T. Prompt engineering of GPT-4 for chemical research: What can/cannot be done?. Science and Technology of Advanced Materials: Methods. 2023; 3; 1: 2260300</bibtext> </blist> <blist> <bibtext> Idso, C. D, Carter, R. M, &amp; Singer, S. F. (2015). Why scientists disagree about global warming. The Heartland Institute. Non Profit Research Organization.</bibtext> </blist> <blist> <bibtext> Jarman, R, &amp; McClune, B. (2007). Developing scientific literacy: Using news media in the classroom: Using news media in the classroom. McGraw-Hill Education (UK).</bibtext> </blist> <blist> <bibtext> Kasneci E, Seßler K, Küchemann S, Bannert M, Dementieva D, Fischer F, Hüllermeier E. ChatGPT for good? On opportunities and challenges of large Language models for education. Learning and Individual Differences. 2023; 103: 102274. 10.1016/j.lindif.2023.102274</bibtext> </blist> <blist> <bibtext> Kaya E, Erduran S. From FRA to RFN, or how the family resemblance approach can be transformed for science curriculum analysis on nature of science. Science &amp; Education. 2016; 25: 1115-1133. 10.1007/s11191-016-9861-3</bibtext> </blist> <blist> <bibtext> Khishfe R. Transfer of nature of science Understandings into similar contexts: Promises and possibilities of an explicit reflective approach. International Journal of Science Education. 2013; 35; 17: 2928-2953. 10.1080/09500693.2012.672774</bibtext> </blist> <blist> <bibtext> Khishfe R, Abd-El-Khalick F. Influence of explicit and reflective versus implicit inquiry-oriented instruction on sixth graders' views of nature of science. Journal of Research in Science Teaching. 2002; 39; 7: 551-578. 10.1002/tea.10036</bibtext> </blist> <blist> <bibtext> Krist C, Kubsch M. Bias, bias everywhere: A response to Li et al. and Zhai and Nehm. Journal of Research in Science Teaching. 2023; 60: 2395-2399. 10.1002/tea.21913</bibtext> </blist> <blist> <bibtext> Law, L. (2024). Application of generative artificial intelligence (GenAI) in Language teaching and learning: A scoping literature review. Computers and Education Open, 100174.</bibtext> </blist> <blist> <bibtext> LC, R, Liu, S, Hendra, L. B, &amp; Fu, K. (2024). TIME ENOUGH: Generative AI Visions of Climate Change as Cave Paintings of the Future. Paper presented at the Proceedings of the 16th Conference on Creativity &amp; Cognition.</bibtext> </blist> <blist> <bibtext> Lee, G, &amp; Zhai, X. (2025). Realizing visual question answering for education: GPT-4V as a multimodal AI. TechTrends, 1–17.</bibtext> </blist> <blist> <bibtext> Leung JSC. Promoting students' use of epistemic Understanding in the evaluation of socioscientific issues through a practice-based approach. Instructional Science. 2020; 48; 5: 591-622. 10.1007/s11251-020-09522-5</bibtext> </blist> <blist> <bibtext> Leung JSC. Students' adherences to epistemic Understanding in evaluating scientific claims. Science Education. 2020; 104; 2: 164-192. 10.1002/sce.21563</bibtext> </blist> <blist> <bibtext> Liu M, Zhang LJ, Biebricher C. Investigating students' cognitive processes in generative AI-assisted digital multimodal composing and traditional writing. Computers &amp; Education. 2024; 211: 104977. 10.1016/j.compedu.2023.104977</bibtext> </blist> <blist> <bibtext> Marchal, N, Xu, R, Elasmar, R, Gabriel, I, Goldberg, B, &amp; Isaac, W. (2024). Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data. arXiv preprint arXiv:2406.13843.</bibtext> </blist> <blist> <bibtext> Masataka N. Development of reading ability is facilitated by intensive exposure to a digital children's picture book. Frontiers in Psychology. 2014; 5: 396. 10.3389/fpsyg.2014.00396</bibtext> </blist> <blist> <bibtext> Meneses A, Escobar JP, Véliz S. The effects of multimodal Texts on science reading comprehension in Chilean fifth-graders: Text scaffolding and comprehension skills. International Journal of Science Education. 2018; 40; 18: 2226-2244. 10.1080/09500693.2018.1527472</bibtext> </blist> <blist> <bibtext> Ng, D. T. K, Tan, C. W, &amp; Leung, J. K. L. (2024). Empowering student self-regulated learning and science education through ChatGPT: A pioneering pilot study. British Journal of Educational Technology.</bibtext> </blist> <blist> <bibtext> Oliveras B, Márquez C, Sanmartí N. The use of newspaper articles as a tool to develop critical thinking in science classes. International Journal of Science Education. 2013; 35; 6: 885-905. 10.1080/09500693.2011.586736</bibtext> </blist> <blist> <bibtext> Oliveras B, Márquez C, Sanmartí N. Students' attitudes to information in the press: Critical reading of a newspaper Article with scientific content. Research in Science Education. 2014; 44: 603-626. 10.1007/s11165-013-9397-3</bibtext> </blist> <blist> <bibtext> OpenAI (2023). March 14, 2023). GPT-4.</bibtext> </blist> <blist> <bibtext> Rowsell, J, Kress, G, Pahl, K, &amp; Street, B. (2018). The social practice of multimodal reading: A new literacy studies—Multimodal perspective on reading. Theoretical models and processes of literacy (pp. 514–532). Routledge.</bibtext> </blist> <blist> <bibtext> Sadler TD. Situated learning in science education: socio-scientific issues as contexts for practice. Studies in Science Education. 2009; 45; 1: 1-42. 10.1080/03057260802681839</bibtext> </blist> <blist> <bibtext> Sadler, T. D, &amp; Dawson, V. (2012). Socio-scientific issues in science education: Contexts for the promotion of key learning outcomes. Second International Handbook of Science Education, 799–809.</bibtext> </blist> <blist> <bibtext> Samantray, A, &amp; Pin, P. (2019). Credibility of climate change denial in social media. Palgrave Communications, 5(1).</bibtext> </blist> <blist> <bibtext> Schäpers P, Freudenstein JP, Mussel P, Lievens F, Krumm S. Effects of situation descriptions on the construct-related validity of construct-driven situational judgment tests. Journal of Research in Personality. 2020; 87: 103963. 10.1016/j.jrp.2020.103963</bibtext> </blist> <blist> <bibtext> Singh, G, Deng, F, &amp; Ahn, S. (2021). Illiterate dall-e learns to compose. arXiv preprint arXiv:2110.11405.</bibtext> </blist> <blist> <bibtext> Sjöström, J. (2024). Vision III of scientific literacy And science education: An alternative vision for science education emphasising the ethico-socio-political And relational-existential. Studies in Science Education, 1–36.</bibtext> </blist> <blist> <bibtext> Sommer B, von Querfurth S. In the end, the story of climate change was one of hope and redemption: ChatGPT's narrative on global warming. Ambio. 2024; 53; 7: 951-959. 10.1007/s13280-024-01997-7</bibtext> </blist> <blist> <bibtext> Stang Lund E, Bråten I, Brandmo C, Brante EW, Strømsø HI. Direct and indirect effects of textual and individual factors on source-content integration when reading about a socio-scientific issue. Reading and Writing. 2019; 32: 335-356. 10.1007/s11145-018-9868-z</bibtext> </blist> <blist> <bibtext> Tang KS. Distribution of visual representations across scientific genres in secondary science textbooks: Analysing multimodal genre pattern of Verbal-Visual texts. Research in Science Education. 2022. 10.1007/s11165-022-10058-6</bibtext> </blist> <blist> <bibtext> Tang, K. S. (2024). Informing research on generative artificial intelligence from a Language and literacy perspective: A meta-synthesis of studies in science education. Science Education, n/a(n/a). https://doi.org/10.1002/sce.21875</bibtext> </blist> <blist> <bibtext> Tang KS, Lin SW, Kaur B. Mapping and extending the theoretical perspectives of reading in science and mathematics education research. International Journal of Science and Mathematics Education. 2022; 20; Suppl 1: 1-15. 10.1007/s10763-022-10322-1</bibtext> </blist> <blist> <bibtext> Tilak, S, Bagley, B, Cantu, J, Cosby, M, Engelbert, G, Hickman, G, &amp; Kennedy, S. (2024). Using text-to-image generative AI to create storyboards: Insights from a college psychology classroom. Journal of Sociocybernetics, 19(1).</bibtext> </blist> <blist> <bibtext> Treen KMI, Williams HT, O'Neill SJ. Online misinformation about climate change. Wiley Interdisciplinary Reviews: Climate Change. 2020; 11; 5: e665</bibtext> </blist> <blist> <bibtext> Vaghefi SA, Stammbach D, Muccione V, Bingler J, Ni J, Kraus M, Schimanski T. ChatClimate: Grounding conversational AI in climate science. Communications Earth &amp; Environment. 2023; 4; 1: 480. 10.1038/s43247-023-01084-x</bibtext> </blist> <blist> <bibtext> Wang B, Rau PLP, Yuan T. Measuring user competence in using artificial intelligence: Validity and reliability of artificial intelligence literacy scale. Behaviour &amp; Information Technology. 2023; 42; 9: 1324-1337. 10.1080/0144929X.2022.2072768</bibtext> </blist> <blist> <bibtext> Wilson, M. (2023). Constructing measures: An item response modeling approach. Routledge.</bibtext> </blist> <blist> <bibtext> Yang FY, Chang CC, Chen LL, Chen YC. Exploring learners' beliefs about science reading and scientific epistemic beliefs, and their relations with science text Understanding. International Journal of Science Education. 2016; 38; 10: 1591-1606. 10.1080/09500693.2016.1200763</bibtext> </blist> <blist> <bibtext> Yore LD, Tang KS. Foundations, insights, and future considerations of reading in science and mathematics education. International Journal of Science and Mathematics Education. 2022; 20; Suppl 1: 237-260. 10.1007/s10763-022-10321-2</bibtext> </blist> </ref> <aug> <p>By Jack Pun; Kason Ka Ching Cheung and Wangyin Kenneth-Li</p> <p>Reported by Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib24" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib52" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib15" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib63" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib26" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib31" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib46" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib35" firstref="ref10"></nolink> <nolink nlid="nl9" bibid="bib16" firstref="ref11"></nolink> <nolink nlid="nl10" bibid="bib60" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib12" firstref="ref14"></nolink> <nolink nlid="nl12" bibid="bib11" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib69" firstref="ref19"></nolink> <nolink nlid="nl14" bibid="bib36" firstref="ref20"></nolink> <nolink nlid="nl15" bibid="bib17" firstref="ref21"></nolink> <nolink nlid="nl16" bibid="bib47" firstref="ref22"></nolink> <nolink nlid="nl17" bibid="bib54" firstref="ref23"></nolink> <nolink nlid="nl18" bibid="bib55" firstref="ref25"></nolink> <nolink nlid="nl19" bibid="bib13" firstref="ref26"></nolink> <nolink nlid="nl20" bibid="bib56" firstref="ref28"></nolink> <nolink nlid="nl21" bibid="bib66" firstref="ref29"></nolink> <nolink nlid="nl22" bibid="bib14" firstref="ref30"></nolink> <nolink nlid="nl23" bibid="bib20" firstref="ref31"></nolink> <nolink nlid="nl24" bibid="bib64" firstref="ref32"></nolink> <nolink nlid="nl25" bibid="bib71" firstref="ref33"></nolink> <nolink nlid="nl26" bibid="bib23" firstref="ref35"></nolink> <nolink nlid="nl27" bibid="bib70" firstref="ref36"></nolink> <nolink nlid="nl28" bibid="bib28" firstref="ref37"></nolink> <nolink nlid="nl29" bibid="bib34" firstref="ref38"></nolink> <nolink nlid="nl30" bibid="bib59" firstref="ref39"></nolink> <nolink nlid="nl31" bibid="bib27" firstref="ref41"></nolink> <nolink nlid="nl32" bibid="bib45" firstref="ref43"></nolink> <nolink nlid="nl33" bibid="bib48" firstref="ref45"></nolink> <nolink nlid="nl34" bibid="bib22" firstref="ref47"></nolink> <nolink nlid="nl35" bibid="bib42" firstref="ref48"></nolink> <nolink nlid="nl36" bibid="bib32" firstref="ref50"></nolink> <nolink nlid="nl37" bibid="bib58" firstref="ref52"></nolink> <nolink nlid="nl38" bibid="bib39" firstref="ref55"></nolink> <nolink nlid="nl39" bibid="bib30" firstref="ref56"></nolink> <nolink nlid="nl40" bibid="bib50" firstref="ref57"></nolink> <nolink nlid="nl41" bibid="bib51" firstref="ref58"></nolink> <nolink nlid="nl42" bibid="bib10" firstref="ref59"></nolink> <nolink nlid="nl43" bibid="bib65" firstref="ref61"></nolink> <nolink nlid="nl44" bibid="bib40" firstref="ref62"></nolink> <nolink nlid="nl45" bibid="bib21" firstref="ref63"></nolink> <nolink nlid="nl46" bibid="bib53" firstref="ref64"></nolink> <nolink nlid="nl47" bibid="bib43" firstref="ref66"></nolink> <nolink nlid="nl48" bibid="bib44" firstref="ref67"></nolink> <nolink nlid="nl49" bibid="bib67" firstref="ref68"></nolink> <nolink nlid="nl50" bibid="bib41" firstref="ref70"></nolink> <nolink nlid="nl51" bibid="bib57" firstref="ref77"></nolink> <nolink nlid="nl52" bibid="bib29" firstref="ref81"></nolink> <nolink nlid="nl53" bibid="bib33" firstref="ref82"></nolink> <nolink nlid="nl54" bibid="bib19" firstref="ref83"></nolink> <nolink nlid="nl55" bibid="bib61" firstref="ref84"></nolink> <nolink nlid="nl56" bibid="bib25" firstref="ref86"></nolink> <nolink nlid="nl57" bibid="bib18" firstref="ref88"></nolink> <nolink nlid="nl58" bibid="bib49" firstref="ref110"></nolink> <nolink nlid="nl59" bibid="bib37" firstref="ref119"></nolink> <nolink nlid="nl60" bibid="bib38" firstref="ref120"></nolink> <nolink nlid="nl61" bibid="bib62" firstref="ref122"></nolink> <nolink nlid="nl62" bibid="bib68" firstref="ref123"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1504722 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Secondary Students' Reading of Socio-Scientific Image-Texts on Climate Change in a GPT-4 Scenario – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Jack+Pun%22">Jack Pun</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-8043-7645">0000-0002-8043-7645</externalLink>)<br /><searchLink fieldCode="AR" term="%22Kason+Ka+Ching+Cheung%22">Kason Ka Ching Cheung</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-6431-1129">0000-0002-6431-1129</externalLink>)<br /><searchLink fieldCode="AR" term="%22Wangyin+Kenneth-Li%22">Wangyin Kenneth-Li</searchLink> (ORCID <externalLink term="http://orcid.org/0009-0007-5273-2827">0009-0007-5273-2827</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Research+in+Science+Education%22"><i>Research in Science Education</i></searchLink>. 2026 56(1):183-202. – Name: Avail Label: Availability Group: Avail Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 20 – Name: DatePubCY Label: Publication Date Group: Date Data: 2026 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Science+and+Society%22">Science and Society</searchLink><br /><searchLink fieldCode="DE" term="%22Climate%22">Climate</searchLink><br /><searchLink fieldCode="DE" term="%22Reader+Text+Relationship%22">Reader Text Relationship</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Comprehension%22">Reading Comprehension</searchLink><br /><searchLink fieldCode="DE" term="%22Internet%22">Internet</searchLink><br /><searchLink fieldCode="DE" term="%22Search+Engines%22">Search Engines</searchLink><br /><searchLink fieldCode="DE" term="%22Imagery%22">Imagery</searchLink><br /><searchLink fieldCode="DE" term="%22Multimedia+Materials%22">Multimedia Materials</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1007/s11165-025-10258-w – Name: ISSN Label: ISSN Group: ISSN Data: 0157-244X<br />1573-1898 – Name: Abstract Label: Abstract Group: Ab Data: The prominence of multimodal generative artificial intelligence (GenAI) facilitates students' comprehension of scientific knowledge through linguistic and visual modes. However, there is a lack of research that investigates how students read image-text outputs created in GenAI. We conceptualize a model of image-text reading of GenAI scientific texts that comprises the interpretation, exchange, and evaluation domains. Based on this theoretical model, we explored how 68 junior secondary students read two image-text socio-scientific texts created by GPT-4 with DALL.E plugins, one focusing on cognitive-epistemic aspects and another focusing on social-institutional aspects of climate change. Our findings indicated that these domains did not exhibit a hierarchical structure, while students' performance in the evaluation domain in the cognitive-epistemic text was better than that in the social-institutional text. More importantly, students expressed a range of uninformed ideas regarding the nature of GenAI when they read the two texts, including equating GenAI to an Internet search engine, picture creators, and human. We discussed how teaching and learning can foster students' "image-text and epistemic" reading by targeting the three domains of our theoretical model. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2026 – Name: AN Label: Accession Number Group: ID Data: EJ1504722 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1504722 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11165-025-10258-w Languages: – Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 183 Subjects: – SubjectFull: Secondary School Students Type: general – SubjectFull: Science and Society Type: general – SubjectFull: Climate Type: general – SubjectFull: Reader Text Relationship Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Reading Comprehension Type: general – SubjectFull: Internet Type: general – SubjectFull: Search Engines Type: general – SubjectFull: Imagery Type: general – SubjectFull: Multimedia Materials Type: general Titles: – TitleFull: Secondary Students' Reading of Socio-Scientific Image-Texts on Climate Change in a GPT-4 Scenario Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Jack Pun – PersonEntity: Name: NameFull: Kason Ka Ching Cheung – PersonEntity: Name: NameFull: Wangyin Kenneth-Li IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 0157-244X – Type: issn-electronic Value: 1573-1898 Numbering: – Type: volume Value: 56 – Type: issue Value: 1 Titles: – TitleFull: Research in Science Education Type: main |
| ResultId | 1 |