Validation of a Peer Observation and Evaluation Tool for Online Teaching in the U.S.

Saved in:
Bibliographic Details
Title: Validation of a Peer Observation and Evaluation Tool for Online Teaching in the U.S.
Language: English
Authors: Yuane Jia (ORCID 0000-0002-1792-0431), Amy B. Spagnolo, Nora Barrett, Ann A. Murphy, Peter M. Basto, Pamela Rothpletz-Puglia, Stuart Luther
Source: Educational Technology Research and Development. 2025 73(1):615-639.
Availability: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/
Peer Reviewed: Y
Page Count: 25
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Tests/Questionnaires
Education Level: Higher Education
Postsecondary Education
Descriptors: Higher Education, Peer Evaluation, Lesson Observation Criteria, Test Construction, Test Reliability, Test Validity, Online Courses, Teacher Evaluation, Electronic Learning, Health Education, Teacher Effectiveness
DOI: 10.1007/s11423-024-10428-z
ISSN: 1042-1629
1556-6501
Abstract: The benefits of peer evaluation of teaching effectiveness and quality in higher education are well documented. While instruments exist for the review and evaluation of entire online courses, there is no standardized single-lesson, peer evaluation instrument available for online instruction. This pilot study focused on the validation of a peer observation and evaluation tool for use with single lessons in both synchronous and asynchronous online courses in an inter-professional school of health professions. The researchers modified a psychometrically validated instrument developed for in-person peer observation by adding items from a renowned online course rubric to create a peer observation tool, entitled the Peer Observation and Evaluation Tool-Online (POET-O). The resulting instrument demonstrated adequate construct validity and reliability by using the many-facet Rasch measurement (MFRM) technique. MFRM results also indicated potential places to revise and improve the instrument. Recommendations for implementing the peer evaluation process of teaching are provided.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1462806
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHLYF-Yeww6J8-AIynYPDETAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDBjt92tgnZKwxSqkJgIBEICBm33FAVS9RFZugGPLyRgXjpqJ5I4fkFRbmuL3JUpON3MMJLjDxQ4szc3EXPkJp-nMJKWTjTaw-E90S5z7UKM0PAassTbZGvKFC2YERw-9XF1zbHo5U15k_SgggQ5Sl6Gbqhilzki88XDU9mRBWCwkmUueI6doribS6ot525_p0yg_WhCoblGsugJq5fyrSQwmfcuUTebuz1XELsuP
Text:
  Availability: 1
  Value: <anid>AN0183751073;etr01feb.25;2025Mar19.03:45;v2.2.500</anid> <title id="AN0183751073-1">Validation of a peer observation and evaluation tool for online teaching in the U.S.: Validation of a peer observation and evaluation...: Y. Jia et al </title> <p>The benefits of peer evaluation of teaching effectiveness and quality in higher education are well documented. While instruments exist for the review and evaluation of entire online courses, there is no standardized single-lesson, peer evaluation instrument available for online instruction. This pilot study focused on the validation of a peer observation and evaluation tool for use with single lessons in both synchronous and asynchronous online courses in an inter-professional school of health professions. The researchers modified a psychometrically validated instrument developed for in-person peer observation by adding items from a renowned online course rubric to create a peer observation tool, entitled the Peer Observation and Evaluation Tool-Online (POET-O). The resulting instrument demonstrated adequate construct validity and reliability by using the many-facet Rasch measurement (MFRM) technique. MFRM results also indicated potential places to revise and improve the instrument. Recommendations for implementing the peer evaluation process of teaching are provided.</p> <p>Keywords: Peer observation; Teaching evaluation; Online instruction; Validation study; Many-facet Rasch measurement; Health professions education; Education Specialist Studies In Education</p> <hd id="AN0183751073-2">Introduction</hd> <p>Academics understand the concept of peer review in research and quality assurance, but traditionally, teaching has not undergone peer review to the same degree. Prior research has indicated that peer review of teaching has numerous benefits (Bell & Mladenovic, [<reflink idref="bib2" id="ref1">2</reflink>]; Yiend et al., [<reflink idref="bib56" id="ref2">56</reflink>]). For example, it promotes and/or spreads quality teaching practices (Chao et al., [<reflink idref="bib9" id="ref3">9</reflink>]; Little, [<reflink idref="bib36" id="ref4">36</reflink>]; McGahan et al., [<reflink idref="bib39" id="ref5">39</reflink>]; Ridge & Lavigne, [<reflink idref="bib45" id="ref6">45</reflink>]; Yiend et al., [<reflink idref="bib56" id="ref7">56</reflink>]), enhances professional development (Bell, [<reflink idref="bib3" id="ref8">3</reflink>]; Fletcher, [<reflink idref="bib20" id="ref9">20</reflink>]; Hammersley-Fletcher & Orsmond, [<reflink idref="bib23" id="ref10">23</reflink>]; Thomas et al., [<reflink idref="bib52" id="ref11">52</reflink>]), and fosters the growth of professional relationships over time (Shortland, [<reflink idref="bib47" id="ref12">47</reflink>]). Depending on who conducts the observation and its purpose, three models of teaching observation—evaluation, developmental, and collaborative—are commonly recognized in the literature (Yiend et al., [<reflink idref="bib56" id="ref13">56</reflink>]). These models reflect varying assumptions about the function of peer review and its impact on authority and power relationships among academics (Sachs & Parsell, [<reflink idref="bib46" id="ref14">46</reflink>]). The evaluation model is mainly managerial and judgmental, where managerial or academic staff oversee teaching quality to ensure compliance and promote best practices. The developmental model, which is less judgmental and more formative, involves an educational expert as the observer. The collaborative model, also formative, entails mutual observation between academic colleagues (Yiend et al., [<reflink idref="bib56" id="ref15">56</reflink>]). A recent study by Fletcher (2018) suggested that a collaborative model in nature has a greater chance of ongoing success.</p> <p>Historically in U.S. higher education, evaluation of teaching effectiveness and quality has relied heavily on student evaluations of teaching (SET). These evaluations often carry significant weight informing faculty merit increases and academic rank and tenure promotions. Many studies question the validity, reliability, and inherent bias of student evaluations as an indicator of teaching effectiveness (Boring et al., [<reflink idref="bib6" id="ref16">6</reflink>]; Esarey & Valdes, [<reflink idref="bib18" id="ref17">18</reflink>]; Hornstein, [<reflink idref="bib25" id="ref18">25</reflink>]; Spooren et al., [<reflink idref="bib50" id="ref19">50</reflink>]) and cast doubt on the use of student ratings to accurately evaluate teaching quality (Hornstein, [<reflink idref="bib25" id="ref20">25</reflink>]; Linse, [<reflink idref="bib35" id="ref21">35</reflink>]).</p> <p>More recently, some U.S. universities have initiated a multidimensional comprehensive teaching evaluation practice (TEval) to guide in assessing teaching quality and supporting systemic faculty growth in teaching (Weaver et al., [<reflink idref="bib54" id="ref22">54</reflink>]). They proposed the Teaching Quality Framework (TQF) that provided a standardized and consistent method for evaluating and improving teaching across higher education institutions (Finkelstein et al., [<reflink idref="bib19" id="ref23">19</reflink>]). TQF posits the following principles: (<reflink idref="bib1" id="ref24">1</reflink>) Adopting evidence-based teaching assessment practices is important to improve teaching; (<reflink idref="bib2" id="ref25">2</reflink>) Multi-perspective data (three key source of data from instructors themselves, their students, and their faculty peers) is essential in effective teaching assessment; (<reflink idref="bib3" id="ref26">3</reflink>) The process of evaluating teaching should lead to improvements by encompassing both formative and summative aspects. Peer review is a key component of evaluating faculty's effectiveness in teaching (Brown & Ward-Griffin, [<reflink idref="bib7" id="ref27">7</reflink>]). As noted for more than 30 years, objective methods, trained observers, faculty involvement, and constructive feedback are important factors that contribute to a successful peer evaluation. Despite many institutions of higher education recognizing the benefits of peer review (Chism, [<reflink idref="bib10" id="ref28">10</reflink>]; Sachs & Parsell, [<reflink idref="bib46" id="ref29">46</reflink>]), beyond other barriers of peer review such as fear of bias, lack of time and uncertainty about what and how should be reviewed, standardized and validated instruments/rubrics are underutilized (Cox et al., [<reflink idref="bib11" id="ref30">11</reflink>]; Ridge & Lavigne, [<reflink idref="bib45" id="ref31">45</reflink>]; Thomas et al., [<reflink idref="bib52" id="ref32">52</reflink>]; Le & Howard, [<reflink idref="bib26" id="ref33">26</reflink>]).</p> <p>There are variety of peer evaluation tools existing in specific disciplines for traditional face to face classroom, for example, the Practical Observation Rubric to Assess Active Learning (PORTAAL; Eddy et al., [<reflink idref="bib17" id="ref34">17</reflink>]), The Classroom Observation Protocol for Undergraduate STEM (COPUS; Smith et al., [<reflink idref="bib48" id="ref35">48</reflink>]), Peer Review of Teaching (PRoT) program for nursing faculty (Mager et al., [<reflink idref="bib11" id="ref36">11</reflink>]). As suggested by Dawson and Hocker ([<reflink idref="bib13" id="ref37">13</reflink>]), a systematic approach grounded in a shared definition of teaching excellence and appropriate infrastructure as guide to the peer review process is required.</p> <p>As an ongoing and steady increase in distance education and the accelerated implementation of online instruction, there is increasing need of standardized tool for evaluations of online instruction. Baldwin et al., ([<reflink idref="bib1" id="ref38">1</reflink>]) reviewed six national or statewide higher education online course evaluation instruments available for offering standards that can encourage quality and promote best practices in online course design, for example, Blackboard's Exemplary Course Program Rubric (2012), Quality Matters (QM) Higher Education Rubric (2016), the Open SUNY Course Quality Review Rubric (OSCQR) (2016). Each has their unique features and shared similarities, but these tools were not necessary designed for peer evaluation purpose. Moreover, all instruments for online instruction were designed for reviewing a whole course. Both course and single lesson evaluations assess teaching quality, evaluating a lesson vs a course focuses in on the nuances of teaching. The evaluation of a lesson focuses on evaluating the effectiveness of the teacher in delivering content, engaging students, and achieving specific learning objectives for that particular session. It involves analyzing individual performance and gathering data to assess instructional effectiveness in a specific instance. In contrast, the evaluation of a course involves a broader examination of the overall structure, organization of content, alignment with learning outcomes, appropriateness of assessments, pacing, and student engagement throughout the entire duration of the course. It utilizes assessment data gathered from various lessons and activities to provide a comprehensive assessment of the course's effectiveness. Given practical considerations, there is interest in tools designed specifically for evaluating single lessons, which can be more feasible for faculty to implement compared to assessing entire courses. This consideration acknowledges time constraints as a significant barrier in U.S. schools when implementing peer observation practices (Le & Howard, [<reflink idref="bib26" id="ref39">26</reflink>]; Ridge & Lavigne, [<reflink idref="bib45" id="ref40">45</reflink>]). After reviewing a number of options and focusing on tools that had been used effectively in post-secondary health education programs for single lesson evaluation, the Peer Observation Evaluation Tool (POET) was chosen (Trujillo et al., [<reflink idref="bib53" id="ref41">53</reflink>]).</p> <p>The POET, which was developed by the School of Pharmacy at Northeastern University, was designed to foster peer observation of an in-person, single classroom-based lesson and to provide formative feedback to faculty regarding their teaching effectiveness. It is divided into four parts (pre-observation, observation, climate of the class, and post-observation). The POET was tested with faculty and determined to have face validity and inter-rater reliability (Trujillo et al., [<reflink idref="bib53" id="ref42">53</reflink>]). The tool was further examined by Crabtree et al. ([<reflink idref="bib12" id="ref43">12</reflink>]) in an occupational therapy department where content validity and inter-rater reliability were further established. The strength of POET lies in advancing peer evaluation of teaching through the development of standardized and validated evaluation tools (Le & Howard, [<reflink idref="bib26" id="ref44">26</reflink>]). However, it was not designed for application in web-based learning environments. Given the lack of comparable online instruction evaluations that ensure equitable, objective, and valid peer observation and evaluation of both in-person and online instruction (Fox et al., [<reflink idref="bib21" id="ref45">21</reflink>]), this represents a significant gap in our institution's peer review evaluation practice. This need has become even more pronounced in the post-pandemic educational context. At the onset of this study, the authors could not identify a psychometrically sound, single-lesson, peer evaluation instrument for online instruction. This lack of equity across in-person and online lessons poses a challenge for the peer review process for faculty who teach only online or those who teach both online and in-person, particularly within programs that utilize both instructional formats. To address these challenges, the authors chose to modify the POET instrument by adding components from the online instruction rubric–Quality Matters (Maryland Online, [<reflink idref="bib37" id="ref46">37</reflink>]).</p> <p>This pilot study was initially focused on the validation of a modified peer observation and evaluation tool for use in multiple learning environments, as well as best practices for peer review in an inter-professional school of health professions. The researchers modified the psychometrically validated POET instrument (Trujillo et al., [<reflink idref="bib53" id="ref47">53</reflink>]), designed for in-person peer observation, by adding items from a renowned online course rubric, Quality Matters (Maryland Online, [<reflink idref="bib37" id="ref48">37</reflink>]), to create a peer observation tool for use with both in-person and online lessons. The resulting instrument, entitled the Peer Observation and Evaluation Tool-Online (POET-O), was initially intended to be used in its entirety for peer evaluation of single lessons offered synchronously in-person or asynchronously online. However, by the time the pilot study began and peer observations were conducted, it was the spring semester of 2020, and the spread of the COVID-19 pandemic was impacting colleges and universities across the U.S. and around the world. The school's in-person classes quickly shifted to being delivered online. Due to the lack of in-person lessons, the focus of this pilot study shifted to an analysis of the POET-O for peer review of online lessons only, comparing those lessons delivered synchronously via Zoom or asynchronously via the Canvas learning management system.</p> <p>Ultimately, the specific aims of this pilot study were to: (<reflink idref="bib1" id="ref49">1</reflink>) validate the POET-O for online synchronous and asynchronous lessons peer review of teaching; (<reflink idref="bib2" id="ref50">2</reflink>) establish the reliability and validity of the POET-O beyond inter-rater reliability and content validity (e.g., expert review); and (<reflink idref="bib3" id="ref51">3</reflink>) explore best practices of evaluating teaching performance (including rater recruitment and training, implementation considerations, and providing beneficial and quality feedback) that produce reliable and comparable teaching evaluation scores.</p> <hd id="AN0183751073-3">Methods</hd> <p></p> <hd id="AN0183751073-4">Context of the study</hd> <p>In the spring of 2017, the senior administration of a large northeastern university in the U.S. convened a task force charged with improving the teaching evaluation process. Faculty forums were held to discuss methods beyond student assessments of instruction that utilize reliable evaluation mechanisms such as peer review and teaching portfolios. The same year, an allied health school located within the university embarked on a five-year strategic plan that prioritized strategies to ensure excellence in teaching. Additionally, the school's reappointment and promotion guidelines had recently become more stringent, with teaching-track faculty progressing through the ranks needing mechanisms to demonstrate excellence in teaching that went beyond the SET. This convergence of events prompted further exploration of peer review options within the allied health school.</p> <p>The structure of the school presented some challenges to the adoption of a peer review process. The school offers a wide variety of health profession degree programs, as well as multiple mechanisms for delivering instruction. The development of online courses began more than twenty years ago, and eventually, led to the creation of fully online degree programs. Other degree programs offered hybrid learning environments that allowed students to take some online courses, particularly during semesters when they also were engaged in a complex schedule of clinical rotations or internship placements, in addition to traditional in-person classes. Thus, faculty members involved in the development and implementation of a robust system for evaluating teaching believed that it was important to choose a peer review tool that could be utilized with diverse degree programs and learning environments. It was also important to choose a standardized tool that had demonstrated reliability and validity to ensure consistency and comparability across programs, departments, and delivery methods.</p> <hd id="AN0183751073-5">Participants</hd> <p>A total of 33 faculty out of 140 at the school were recruited from across all eight departments in a school of health professions to participate in the pilot project, three (9%) out of 33 were male, more details see Table 1. Faculty were informed of the pilot study at a faculty meeting in Fall 2020, via email from the Dean's office, and by their department chairs who were informed by the co-principal investigators (co-PIs). One of the two co-PIs of the project did not review any lessons, another co-PI did participate the review lessons. Faculty volunteered to participate as reviewers (i.e., individuals who reviewed lessons using the POET-O), lesson submitters (i.e., individuals who submitted and/or provided access to asynchronous and synchronous lessons), or both (i.e., reviewers and lesson submitters). Fifteen out of 33 were both reviewers and lesson submitters. Six faculty had submitted both asynchronous and synchronous lessons. The lessons varied from undergraduate lessons to graduate level lessons, and from theoretical to practical skills. Twenty-six reviewers were randomly assigned to evaluate 14 asynchronous lessons, and 22 reviewers were randomly assigned to evaluate 12 synchronous lessons. The average time took for reviewers to review a single lesson were 137 min (SD = 68) for asynchronous lesson, and 170 min (SD = 84) for synchronous lesson, respectively.</p> <p>Table 1 Characteristics of participants</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left" rowspan="2"><p>Variables</p></th><th align="left" colspan="2"><p>Phase I: Asynchronous</p></th><th align="left" colspan="2"><p>Phase II: Synchronous</p></th></tr><tr><th align="left"><p>Reviewer (N = 26)</p></th><th align="left"><p>Instructor<sup>#</sup> (N = 14)</p></th><th align="left"><p>Reviewer (N = 22)</p></th><th align="left"><p>Instructor<sup>#</sup> (N = 12)</p></th></tr></thead><tbody><tr><td align="left"><p><italic>Gender</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Female</p></td><td align="left"><p>23</p></td><td align="left"><p>12</p></td><td align="left"><p>20</p></td><td align="left"><p>9</p></td></tr><tr><td align="left"><p>Male</p></td><td align="left"><p>3</p></td><td align="left"><p>2</p></td><td align="left"><p>2</p></td><td align="left"><p>3</p></td></tr><tr><td align="left"><p><italic>Department</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Clinical and preventive nutrition</p></td><td align="left"><p>4</p></td><td align="left"><p>3</p></td><td align="left"><p>4</p></td><td align="left"><p>1</p></td></tr><tr><td align="left"><p>Clinical Lab. and medical imaging sciences</p></td><td align="left"><p>2</p></td><td align="left"><p>2</p></td><td align="left"><p>1</p></td><td align="left"><p>0</p></td></tr><tr><td align="left"><p>Dean's office</p></td><td align="left"><p>1</p></td><td align="left"><p>0</p></td><td align="left"><p>0</p></td><td align="left"><p>0</p></td></tr><tr><td align="left"><p>Health informatics</p></td><td align="left"><p>3</p></td><td align="left"><p>2</p></td><td align="left"><p>3</p></td><td align="left"><p>1</p></td></tr><tr><td align="left"><p>Interdisciplinary studies</p></td><td align="left"><p>4</p></td><td align="left"><p>2</p></td><td align="left"><p>4</p></td><td align="left"><p>1</p></td></tr><tr><td align="left"><p>Physician Asst. studies</p></td><td align="left"><p>3</p></td><td align="left"><p>0</p></td><td align="left"><p>1</p></td><td align="left"><p>1</p></td></tr><tr><td align="left"><p>Psych Rehab and Counsel Pro</p></td><td align="left"><p>8</p></td><td align="left"><p>5</p></td><td align="left"><p>8</p></td><td align="left"><p>6</p></td></tr><tr><td align="left"><p>Rehab. and movement sciences</p></td><td align="left"><p>1</p></td><td align="left"><p>0</p></td><td align="left"><p>1</p></td><td align="left"><p>2</p></td></tr><tr><td align="left"><p><italic>Education level</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Doctorate</p></td><td align="left"><p>18</p></td><td align="left"><p>11</p></td><td align="left"><p>16</p></td><td align="left"><p>10</p></td></tr><tr><td align="left"><p>Masters</p></td><td align="left"><p>8</p></td><td align="left"><p>3</p></td><td align="left"><p>6</p></td><td align="left"><p>2</p></td></tr></tbody></table> </ephtml> </p> <p># indicates lesson submitter Fifteen out of 33 were both reviewers and submitters Each phase includes a two-step process</p> <hd id="AN0183751073-6">Instrument</hd> <p>The POET-O was developed for this pilot project by modifying and adapting the POET (Trujillo et al., [<reflink idref="bib53" id="ref52">53</reflink>]) by adding selected items from the Quality Matters rubric (Maryland Online, [<reflink idref="bib37" id="ref53">37</reflink>]). The first draft of POET-O addressed 7 criteria and had 56 items: Lesson Objectives (4 items), Instructional Materials (8 items), Activities (9 items), Use of Technology (5 items), Accessibility (4 items), Assessment (7 items), and Instructor (19 items). Each item was rated on a 4-point scale (from 1, "Not Present," to 4, "Accomplished Well"). An additional option, "Not Applicable," was provided in case the question is not applicable to the lesson being evaluated. Faculty (n = 8) from the Dean's Taskforce on Peer Review were enlisted to ensure the instrument covered the essential components of teaching performance and the quality of materials. The co-PIs collected the Taskforce's comments and suggestions regarding individual instrument items and collaborated with the peer review group until consensus was reached regarding further modification. The resulting instrument had 32 items and addressed 5 criteria. The Presentation and Activities section (7 items) addresses the design and development of materials and student engagement that are appropriate and effective. The Use of Technology section (5 items) assesses the effective use of technological tools to enhance learning. The ADA Compliance/Accessibility section (4 items) addresses the presentation of materials and use of technologies to facilitate learning for those with disabilities and those who benefit from the utilization of diverse formats. The Assessment section (3 items) assess the use of evaluative tools. The Instructor section (13 items) addresses the knowledge, preparation, and skill of the instructor. See Table 2 for sample items form the POET-O (see Appendix for the full POET-O instrument). Of note, the complete tool has five sections: pre-observation, observation, general reviewer comments, post-observation, and faculty reflection and action plan.</p> <p>Table 2 Sample items from the POET-O</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left"><p>POET-O section</p></th><th align="left"><p>Sample items</p></th></tr></thead><tbody><tr><td align="left"><p>Presentation and activities</p></td><td align="left"><p>The lesson and learning activities are well organized</p></td></tr><tr><td align="left" /><td align="left"><p>Planned student activities reflect appropriate lesson objectives</p></td></tr><tr><td align="left" /><td align="left"><p>Discussion forum and other learning tools promote critical thinking</p></td></tr><tr><td align="left"><p>Use of technology</p></td><td align="left"><p>The instructor effectively uses audio/visual learning tools and technologies to support the learning objectives or competencies</p></td></tr><tr><td align="left" /><td align="left"><p>Learning tools and technologies used promote learner engagement and active learning</p></td></tr><tr><td align="left" /><td align="left"><p>Access to links and external resources provided are functional</p></td></tr><tr><td align="left"><p>ADA compliance/accessibility</p></td><td align="left"><p>Information is provided regarding accessibility of all learning tools and technologies required in the lesson</p></td></tr><tr><td align="left" /><td align="left"><p>Alternative means of access to lesson materials in formats that meet the needs of diverse learners are provided</p></td></tr><tr><td align="left" /><td align="left"><p>All materials are designed to facilitate readability</p></td></tr><tr><td align="left"><p>Assessment</p></td><td align="left"><p>Planned assessment strategies measure the stated learning objectives or competencies</p></td></tr><tr><td align="left" /><td align="left"><p>Specific and descriptive criteria are provided for the evaluation of learners' work, i.e. rubrics etc</p></td></tr><tr><td align="left" /><td align="left"><p>Students are given ample time to complete the assignments and assessments for this lesson</p></td></tr><tr><td align="left"><p>Instructor</p></td><td align="left"><p>Instructor establishes the relevance of information</p></td></tr><tr><td align="left" /><td align="left"><p>Instructor makes connections with prior learning (from previous lessons and courses) when applicable</p></td></tr><tr><td align="left" /><td align="left"><p>Instructor encourages critical thinking and expression of divergent opinions or conflicting views when appropriate</p></td></tr></tbody></table> </ephtml> </p> <hd id="AN0183751073-7">Study design</hd> <p>The full POET-O tool includes multiple sections, as mentioned above. For this pilot study, we focused on the quantitative validity evidence of the measurement tool and item functions, that is, section 2 (peer observation) of the tool. The study design relied on random assignment of faculty reviewers to rate synchronous and asynchronous lessons. This limited the potential for reviewer bias. Specifically, various faculty submitted either asynchronous lessons, synchronous lessons, or both, and each lesson received ratings from two randomly selected reviewers. Ideally, if all lesson reviewers assess all lessons submitted on all tasks (i.e., complete design), it will lead to the highest precision of model estimation (Eckes, [<reflink idref="bib16" id="ref54">16</reflink>]). Unfortunately, practical considerations such as time constraints, reviewers' workloads, budget, et cetera, led us to narrow the design. One way to achieve both estimation precision and reduce reviewer burden is through an incomplete rating design. The crucial aspect of the incomplete design is to provide sufficient connection between facet elements to calibrate all elements of a particular facet (e.g., reviewers) on the same scale. In order to create the linkage among all the reviewers, one asynchronous and one synchronous lesson were rated by all the reviewers following the incomplete connected rating design (Eckes, [<reflink idref="bib16" id="ref55">16</reflink>]). Table 3 illustrates the basic structure of the rating design that provided sufficient links between facet elements, but also took into account the practical considerations mentioned above.</p> <p>Table 3 Schematic representation of the incomplete connected rating design</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left" rowspan="2" /><th align="left" /><th align="left" /><th align="left" /><th align="left" /><th align="left"><p>Reviewers</p></th><th align="left" /><th align="left" /><th align="left" /></tr><tr><th align="left"><p>1</p></th><th align="left"><p>2</p></th><th align="left"><p>3</p></th><th align="left"><p>4</p></th><th align="left"><p>5</p></th><th align="left"><p>6</p></th><th align="left"><p>7</p></th><th align="left"><p>...26</p></th></tr></thead><tbody><tr><td align="left"><p><italic>Instructor 1</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Asynchronous</p></td><td align="left"><p>x</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Synchronous</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left" /><td align="left"><p>x</p></td><td align="left" /></tr><tr><td align="left"><p><italic>Instructor 2</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Asynchronous</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td></tr><tr><td align="left"><p><italic>Instructor 3</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Synchronous</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td><td align="left"><p>x</p></td></tr><tr><td align="left"><p>......</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p><italic>Instructor 20</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Asynchronous</p></td><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left" /></tr><tr><td align="left"><p><italic>Instructor 21</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>Synchronous</p></td><td align="left" /><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left"><p>x</p></td><td align="left" /><td align="left" /><td align="left" /></tr></tbody></table> </ephtml> </p> <hd id="AN0183751073-8">Procedure</hd> <p>Faculty reviewers were trained via a video education module. The training provided an overview of the POET-O, existing research on the POET, the aims of the pilot study, how to use the POET-O, how to access lesson materials for review, and how to submit completed assessments. Additional materials, including strategies for approaching and conducting the review, were provided. A two-step process was used for both the asynchronous lessons phase and the synchronous lessons phase. The first step was the calibration, in which all reviewers assessed the same lesson and item scores from those reviews were compared for consistency. POET-O items with the most inconsistent scores were discussed with the faculty reviewers as a group during a debriefing session. Following the debriefing session, faculty reviewers were asked to review their ratings and adjust if necessary and as they deemed appropriate based on the clarifications discussed in the debriefing and resubmit. Rater consistency was assessed again with the revised scores. The debriefing meetings largely focused on reminding faculty reviewers of the meaning attached to each of the scores and the not applicable option. For example, many reviewers used not applicable to mean the item under review was not present, when it was intended to mean it does not apply and reasonably should not be present. Only one debriefing session for the asynchronous lessons and one for the synchronous lessons was needed to reach an acceptable level of rater consistency, defined as 90% agreement within one point of the mode.</p> <p>When an acceptable level of rater consistency was reached, the study moved into the second step, the independent review. During the independent review, two reviewers were randomly assigned to each lesson and evaluated it independently using the POET-O. This two-step process was utilized twice, first with asynchronous and then with synchronous lessons. Twenty-six reviewers were randomly assigned to evaluate 14 asynchronous lessons, and 22 reviewers were randomly assigned to evaluate 12 synchronous lessons. All evaluations were completed via a fillable pdf of the POET-O and uploaded to a shared folder in Canvas. The research team then compiled the rating data into Excel and readied it for data analysis.</p> <p>Faculty lesson submitters were asked to provide access to their asynchronous lessons (via their Canvas course) or submit their recorded synchronous lessons for review. Additionally, faculty lesson submitters provided their course syllabus, presentation materials (e.g., PowerPoint slide deck), additional materials that were made available to students, any assessment items that are associated with the lesson (e.g., exam/quiz items or paper/discussion topics), and specific aspects of the lesson on which they would like to receive reviewer feedback.</p> <hd id="AN0183751073-9">Data analyses</hd> <p>Despite its common use, rater agreement (i.e., a kappa statistic), which is used to quantify the quality of raters or rater training effectiveness, is not sufficient for informing an instrument's reliability (Hill et al., [<reflink idref="bib24" id="ref56">24</reflink>]; Wind & Jones, [<reflink idref="bib55" id="ref57">55</reflink>]). Reliability of a specific instrument using a kappa statistic is misleading, especially with the teacher evaluation systems based on classroom observations. There are multiple sources of variance in the observational rating score due to sampling of lessons, raters, and the instrument itself (Hill et al., [<reflink idref="bib24" id="ref58">24</reflink>]). Many-Facet Rasch Measurement (MFRM; Linacre, [<reflink idref="bib27" id="ref59">27</reflink>]), based on item response theory (IRT), can model the rater effects and quantify multiple sources of variance (i.e., instructors, items, and raters) in observational scores and the interaction among them, making it a good fit for subjectively rated performance evaluation (Eckes, [<reflink idref="bib16" id="ref60">16</reflink>]; Mulqueen et al., [<reflink idref="bib41" id="ref61">41</reflink>]). The MFRM technique can account for the rater severity differences for estimates of instructor's teaching performance (Wind & Jones, [<reflink idref="bib55" id="ref62">55</reflink>]). In addition to the multifaceted nature of reliability, the MFRM can be used to gather evidence (rater severity, model-data fit, rating category use) to inform the improvement of the instrument either through better rater training and/or item design, and therefore produce psychometrically sound evaluations of teaching performance (Wind & Jones, [<reflink idref="bib55" id="ref63">55</reflink>]). MFRM modeling is utilized to calibrate all elements of each facet on the same scale and account for each facet's impact on the performance evaluation. A four-facet Rasch model—instructors, raters, lesson format, and scoring criteria items were conducted in FACETS Version No. 3.83.6 (Linacre, [<reflink idref="bib32" id="ref64">32</reflink>]). In addition, a Principal Component Analysis (PCA) of the residuals to examine unidimensionality was conducted via Winsteps 3.72 (Linacre, [<reflink idref="bib31" id="ref65">31</reflink>]).</p> <hd id="AN0183751073-10">Results</hd> <p>A four-facet Rasch model was fitted in FACETS. The program used the rating scores to estimate individual lesson submitter's teaching proficiency, rater severities, item difficulties, and rating scale functioning by calibrating the specified facets onto a single linear scale, creating a single frame of reference for interpretating the results. The graphic representation of the output is shown in a variable map.</p> <hd id="AN0183751073-11">Calibration of facets: variable map</hd> <p>The variable map (Fig. 1) provided a graphical representation of the four facets: instructors, raters, items, lesson format (synchronous vs. asynchronous) and the 4-point rating scales used to score instructor teaching performance on a common "ruler" with an equal-interval scale (i.e., the logit scale, a measure of instructor's teaching abilities, see the first column of Fig. 1), creating a single frame of reference for interpreting the results of the analysis. The columns two to six represent distribution of instructors' abilities, rater severity, lesson format, item difficulty, and rating scale categories, respectively. It can be seen that the shape of the distribution of instructors approached a normal distribution and spanned from 3.25 to 0.74 logits (M = 1.82, SD = 0.59, N = 21). All participating instructors located above the average difficulty of items. There were no items to measure the instructors of certain ability ranges (e.g. above 1.5 logits). Thus, some more difficult items might need to be added to the instrument to differentiate those high teaching proficiency instructors. MFRM analysis reported two mean-square statistics (i.e., infit, and outfit statistics) indicating a data–model fit for each facet. Both statistics have an expected value of 1 and can range from 0 to infinity (Linacre, [<reflink idref="bib29" id="ref66">29</reflink>]; Myford & Wolfe, [<reflink idref="bib42" id="ref67">42</reflink>]). Linacre ([<reflink idref="bib29" id="ref68">29</reflink>]) suggested using a range of 0.50 to 1.50 as an indication of useful fit for infit and outfit mean-square statistics. Two instructors (<reflink idref="bib4" id="ref69">4</reflink>,<reflink idref="bib12" id="ref70">12</reflink>) showed evidence of misfit based on outfit and infit statistics.</p> <p>Graph: Fig. 1 Variable map from the many-facet Rasch measurement analysis</p> <p>Column three represented the calibration of the rater severity, ranging from 1.25 logits (Rater 24 most severe) to − 5.00 logits (Rater 8 most lenient) (SD = 1.24, N = 26). Three out of 26 raters (<reflink idref="bib19" id="ref71">19</reflink>, 22, 23) with infit values slightly greater than 1.5 showed more variation than expected in their ratings, suggesting possible haphazard ratings, or erratic use of the rating scale. MFRM models the raters to be "independent experts". The observed (58.5%) inter-rater agreement was found to be slightly higher than expected (57.4%), which is expected for trained raters. This demonstrates that our raters showed agreement with others but behaved like independent experts.</p> <p>The fourth column included the calibration of the format of the lessons, asynchronous and synchronous. Asynchronous lessons scored lower (M = − 0.19, SE = 0.06) than synchronous lessons (M = 0.19, SE = 0.04) in the logit scale, indicating that raters scored asynchronous lessons higher than synchronous lessons.</p> <p>Column five represented the item difficulties, ranging from 1.36 logits (Item 21 as easiest item) to − 2.11 logits (Item 14 as most difficult item) (SD = 0.77, N = 32). The variability across items in their level of difficulty was substantial. The item difficulty measures showed a 3.47-logit spread. Only one item 31 was showing misfit based on the infit and outfit statistics.</p> <p>Finally, the rightmost column represented the rating scale structure that raters actually used to evaluate instructor's teaching performance. Inspection of the rating scale structure based on the quantitative guidelines set forth by Linacre ([<reflink idref="bib28" id="ref72">28</reflink>], [<reflink idref="bib29" id="ref73">29</reflink>]), the locations of the thresholds (category boundaries) are not equally spaced, indicating the distances between categories of rating scales have not been used equally. Summary statistics for calibration of instructors, raters, lesson format, and items.</p> <p>Rasch summary statistics for the four facets are shown in Table 4. In order to define a frame of reference for interpreting the performance locations, all of the three facets except the object of measurement (i.e., instructor's teaching proficiency) are centered on the logit scale (mean set to zero). The Rasch model provides standard error (SE) estimates for each element within each facet (e.g., instructor, rater, item, and lesson format). Smaller values of SE indicate more precise estimates, such that the logit-scale locations would be expected to remain stable across repeated administrations of an assessment. As can be seen in Table 4, all SE values were close to zero, indicating precise and stable estimates.</p> <p>Table 4 Summary statistics for the many-facet Rasch measurement analysis</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left" rowspan="2"><p>Statistics</p></th><th align="left" colspan="4"><p>Facets</p></th></tr><tr><th align="left"><p>Instructor</p></th><th align="left"><p>Rater<sup>a</sup></p></th><th align="left"><p>Format</p></th><th align="left"><p>Item</p></th></tr></thead><tbody><tr><td align="left"><p>Mean measure</p></td><td align="left"><p>1.82</p></td><td align="left"><p>0</p></td><td align="left"><p>0</p></td><td align="left"><p>0</p></td></tr><tr><td align="left"><p>Mean SE</p></td><td align="left"><p>.19</p></td><td align="left"><p>.27</p></td><td align="left"><p>.04</p></td><td align="left"><p>.19</p></td></tr><tr><td align="left"><p>Chi-square</p></td><td align="left"><p>224.10*</p></td><td align="left"><p>291.80*</p></td><td align="left"><p>43.60*</p></td><td align="left"><p>421.40*</p></td></tr><tr><td align="left"><p>df</p></td><td align="left"><p>20</p></td><td align="left"><p>25</p></td><td align="left"><p>1</p></td><td align="left"><p>31</p></td></tr><tr><td align="left"><p>Separation index</p></td><td align="left"><p>2.9</p></td><td align="left"><p>2.46</p></td><td align="left"><p>6.53</p></td><td align="left"><p>3.69</p></td></tr><tr><td align="left"><p>Reliability of Separation</p></td><td align="left"><p>.89</p></td><td align="left"><p>.86</p></td><td align="left"><p>.98</p></td><td align="left"><p>.93</p></td></tr></tbody></table> </ephtml> </p> <p> <sups>a</sups>Raters with non-extreme scores only *<emph>p</emph> <.01</p> <p>After the logit-scale locations are estimated for each facet, the reliability of separation and a chi-square statistic can be used to describe the degree to which the elements within a facet can be reliably differentiated from one another. Reliability of separation is interpreted in a comparable fashion to Cronbach's alpha coefficient. For person and item facets, higher reliability is preferable. However, for the rater facet, lower reliability indicates more similarity among raters. The reliability and separation index for each facet is favorable according to the benchmark used in practice where person separation is greater than 2 and reliability is greater than.8 (Linacre, n.d.). Not surprisingly, with a large group of raters, the rater separation reliability was.86, indicating a heterogeneity of rater severity measures. The chi-square statistic explores whether the differences among logit-scale locations for elements of each facet are statistically significant. The chi-square statistics for all facets were significant (<emph>p</emph> <.01), indicating that the locations for elements of each facet (i.e., the sampled instructors, raters, items, lesson formats) on the logit scale are different.</p> <hd id="AN0183751073-12">Many-facet Rasch measurement model global model fit</hd> <p>Empirical observations never exactly fit the Rasch model. However, the model-data fit or misfit indices inform us of the quality of the measurement tool, and help us determine where any misfit comes from and how to improve it (Eckes, [<reflink idref="bib15" id="ref74">15</reflink>]). The overall data-model fit can be assessed by examining the unexpected responses given the assumptions of the model, that is, the (absolute) standardized residuals. A satisfactory overall model fit is indicated if about or less than 5% of standardized residuals are ≥ 2, and about or less than 1% of standardized residuals are ≥ 3 (Linacre, n.d.). There were 2,891 valid responses (i.e., responses used for estimation of model parameters) included in the analysis. Of these, 100 responses (or 3.5%) were associated with (absolute) standardized residuals ≥ 2, and 58 responses (or 2.0%) were associated with (absolute) standardized residuals ≥ 3. These findings indicated satisfactory model fit.</p> <p>Bubble charts (see Fig. 2) show item measures and fit values graphically. It is one way of visualizing the overall functioning of an instrument and the fit of its items. It is a graph of item difficulty vs item outfit and infit statistics. Ideally, items should be as close as possible to 0 for standardized outfit and infit. Most items included in the POET-O demonstrate good fit, except items 8, 24, and 32 where the t value of standardized fit is greater than 2. Overall, the graph shows a good fit for the majority items in the measurement instrument.</p> <p>Graph: Fig. 2 Bubble Chart of POET-O (32 items). Note Plotted from Winsteps; each bubble represents an item, whose size is proportional to the standard error of item difficulty calibration; well-fitting items are close to the central vertical line</p> <hd id="AN0183751073-13">Construct validity: unidimensionality</hd> <p>MFRM is utilized to measure a single latent trait, therefore the assumption of unidimensionality of the measurement tool needs to be justified before conducting the many-facet Rasch analysis. "Unidimensionality refers to the existence of a single trait or construct underlying a set of measures" (Gerbing & Anderson, [<reflink idref="bib22" id="ref75">22</reflink>], p. 186). Simply stated, MFRM tests if all items on an instrument are measuring the same thing. One of the approaches to detect unidimensionality is the PCA of the residuals. Figure 3 shows the loading scatterplot for the PCA of the residuals. The horizontal axis is the estimated item difficulty on the Rasch scale; the left hand vertical axis is the correlation coefficient between item scores and another potential construct after the primary construct (i.e., teaching proficiency) is controlled; and letters in the box represent items. In factor analysis literature, values of ± 0.4 or more are considered substantive (Stevens, [<reflink idref="bib51" id="ref76">51</reflink>]). PCA of the residuals showed that 27 items had a loading (i.e., correlation) within the − 0.4 to + 0.4 range, while five items (i.e., 14, 15, 16, 26, and 27 in the Appendix, also see Fig. 3 item A, B, C, a, b) were out of the range. Additionally, three items (i.e., 14, 15, and 16, see Fig. 3 item A, B, C) that assessed ADA Compliance/Accessibility were grouped as one strong construct measured by the items and not related to the main construct. Overall, no significant constructs seem to be underlying the residuals besides the main construct, suggesting a moderately strong construct underlying the measurement instrument. Rasch measures explained 39.3% of total variance. It was considered that data met the unidimensionality requirement, which provides evidence for the construct validity of the instrument. However, it is necessary to further investigate the measurement disturbance, caused by items 14, 15, and 16, to improve the unidimensionality of this instrument.</p> <p>Graph: Fig. 3 Loading scatterplot of PCA of the residuals. Note Each star in the third and fifth column represents one rater/item. The horizontal dashed lines in the rightmost column indicate the category threshold measures</p> <hd id="AN0183751073-14">Rating scale structure</hd> <p>To examine whether the four categories on the rating scale (i.e., 1 = not present, 2 = needs development, 3 = accomplished, and 4 = accomplished well) functioned as intended, various statistical indicators are available (Bond & Fox, [<reflink idref="bib4" id="ref77">4</reflink>]; Linacre, [<reflink idref="bib30" id="ref78">30</reflink>]). One of the indicators is the mean-square outfit statistics which captures the difference between the average and the expected measures for each category. It should not be greater than 2 (Eckes, [<reflink idref="bib15" id="ref79">15</reflink>]). Table 5 summarizes the findings regarding these indices. As the table shows, average measures of instructor's proficiency advanced as the rating categories increased, and values of the outfit mean-square statistic were equal, or close to the expected value of 1. So, we concluded that the rating scale functions properly with higher ratings associated with more proficiency of the variable being measured.</p> <p>Table 5 Category statistics for the POET-O rating scale</p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left"><p>Category</p></th><th align="left"><p>Frequency</p></th><th align="left"><p>%</p></th><th align="left"><p>Average measure</p></th><th align="left"><p>Expect measure</p></th><th align="left"><p>Outfit</p></th><th align="left"><p>Threshold</p></th><th align="left"><p>S.E</p></th><th align="left"><p>Expect measure</p></th></tr></thead><tbody><tr><td align="left"><p>1</p></td><td align="left"><p>128</p></td><td align="left"><p>5</p></td><td char="." align="char"><p>0.29</p></td><td char="." align="char"><p>0.19</p></td><td align="left"><p>1.7</p></td><td align="left"><p>–</p></td><td align="left"><p>–</p></td><td align="left"><p>(− 1.63)</p></td></tr><tr><td align="left"><p>2</p></td><td align="left"><p>183</p></td><td align="left"><p>7</p></td><td char="." align="char"><p>0.57</p></td><td char="." align="char"><p>0.59</p></td><td align="left"><p>0.9</p></td><td align="left"><p>0.03</p></td><td align="left"><p>0.1</p></td><td align="left"><p>− 0.49</p></td></tr><tr><td align="left"><p>3</p></td><td align="left"><p>603</p></td><td align="left"><p>22</p></td><td char="." align="char"><p>1.05</p></td><td char="." align="char"><p>1.1</p></td><td align="left"><p>0.8</p></td><td align="left"><p>− 0.36</p></td><td align="left"><p>0.07</p></td><td align="left"><p>0.41</p></td></tr><tr><td align="left"><p>4</p></td><td align="left"><p>1854</p></td><td align="left"><p>67</p></td><td char="." align="char"><p>1.88</p></td><td char="." align="char"><p>1.87</p></td><td align="left"><p>1</p></td><td align="left"><p>0.33</p></td><td align="left"><p>0.05</p></td><td align="left"><p>1.73</p></td></tr></tbody></table> </ephtml> </p> <p>Thresholds are Rasch-Andrich thresholds <emph>S.E.</emph> Standard error</p> <p>Figure 4 shows the category probability curves for the four-category scale that the raters used to rate instructors on the 32 items. The horizontal axis is the instructor's proficiency scale; the vertical axis is the probability of being rated in each category. The cross of two adjacent rating scale category's curves denotes the category threshold. Another indicator of rating scale quality is the ordering of category thresholds (Eckes, [<reflink idref="bib15" id="ref80">15</reflink>]). The threshold between category 2 and 3 was smaller than the threshold between category 1 and 2, which was disordered. Inspection of the item characteristic curves (Fig. 4) shows, the probability curve of category 2 is almost subsumed under the probability curve of category 1, indicating that category 2 is the single most probable response for very few instructors. Likewise, the probability curve of category 3 is almost subsumed under the probability curve of category 4, indicating that category 3 is the single most probable response for very few instructors. Taken together, the 4-category rating structure can be improved for the POET-O. A better rating scale structure (e.g., 3-category, combining categories with low frequencies) might be needed for further investigation.</p> <p>Graph: Fig. 4 Item characteristic curves for the 4-category rating scale</p> <hd id="AN0183751073-15">Discussion</hd> <p>Rasch analysis has been widely used to validate measurement tools, but rarely seen in previous literature for peer evaluation tools of teaching. Based on the Rasch model analyses, overall, the modified POET-O is a reliable and valid peer observation measure of both online synchronous and asynchronous lessons. Specifically, unidimensionality analyses demonstrated that the POET-O measures a single underlying construct of teaching proficiency, which provided evidence for the construct validity of the instrument. MFRM findings also indicated satisfactory overall data-model fit, established favorable multi-facet reliability (all > 0.85) and separation indices (all > 2) indicating good discrimination and separation ability in each facet. The majority of POET-O items performed very well, with moderate difficulty but substantial variability in difficulty measures, and were able to differentiate varying teaching proficiency levels. Additionally, all rating scale categories and average teaching performance measures on the POET-O advanced together, concluding that the rating scale functioned properly. However, equal intervals between thresholds among rating scale categories would be ideal. Therefore, a better rating scale structure (e.g., 3-categories, combining categories with low frequencies) may be needed. Admittedly, due to the relatively high proficiency of instructors in the pilot, more instructors with lower teaching proficiency are also needed to further validate the rating scale functionality.</p> <p>The variable map results showed teaching proficiency ratings were spread above the mean of the logit scale for all lessons/instructors evaluated. This is likely related to the disproportionately seasoned faculty with high proficiency that self-selected for this pilot study, creating a relatively homogenous sample. Additionally, the competency areas measured by the POET-O items were easily achieved by instructors who submitted lessons. As a pilot study, the MFRM results also indicated a need for improvement of a few items, primarily the ADA accessibility items. The ADA Accessibility items were the most difficult questions and impacted the construct validity (e.g., unidimensionality) of the instrument. Also, faculty in this pilot study reported lack of expertise in rating the ADA Accessibility items. While the school has done ADA accessibility training for faculty, there has been no follow-up to determine if faculty understand how to ensure that course design meets ADA accessibility criteria. Peer reviewers reported that they were not knowledgeable enough about ADA accessibility requirements to determine whether a lesson met the requirements as outlined in the POET-O. Therefore, more training in accessible formats for course materials and general strategies for compliance with ADA accessibility standards is needed. The increased trend of accommodation needs in higher education (Baldwin et al., [<reflink idref="bib1" id="ref81">1</reflink>]; Snyder et al., [<reflink idref="bib49" id="ref82">49</reflink>]) requires that institutions of higher education provide equal and accessible learning environments. Assistive technology and distance learning can advance the accessibility of higher education, yet, the technological aspects of access and related training for both instructors and students remain challenges (Parker Harris et al., [<reflink idref="bib44" id="ref83">44</reflink>]). Nevertheless, revising these items to provide more detailed guidance on how to evaluate ADA implementation will improve the tool. Accessibility items should be revised, followed by a reevaluation of the quality of the POET-O.</p> <p>Bubble chart indicated that a few items (i.e., 8, 24, 32) are right at the edge of the misfit, however, inspection of the item measure report in MFRM analysis, items 24 and 32 did not show the misfit in terms of mean-square (MNSQ) values and z-standardized (ZStd) fit statistics (< 2). The infit ZStd for item 8 was -2, indicating potential issue with this item. To improve the clarity, we suggest revising the item 8 to "The instructor effectively uses audio/visual learning tools to support the learning objectives." Similarly, item 31 showed misfit in the MFRM analysis. We suggest revising this item to "Instructor treats students respectfully, offering encouragement in public ways". Revalidate the quality of the POET-O with revised items is warranted with another sample. For more details on the original items, refer to the Appendix section 2.</p> <p>The MFRM analysis demonstrated a significant difference between asynchronous and synchronous lessons, where the asynchronous lessons were rated higher. There are several potential explanations for this finding. The researchers believe that reviewers may have found it easier to locate materials, review discussions among instructors and students, assess instructions for student activities, and evaluate methods of student assessment when viewed in the asynchronous course materials. Recorded live lessons required attention to detail and ratings were subject to what the faculty submitting lesson materials for live instruction provided. The average total amount of time it took for synchronous lesson review was 2.8 h and the average time it took to review asynchronous lessons was 2.3 h. This difference may be related to the findings showing asynchronous lessons were rated higher than synchronous lessons. Two practical implications arise from these findings: (<reflink idref="bib1" id="ref84">1</reflink>) given the ratings may be influenced by the quality of materials submitted by instructors, it is crucial to provide guidance and support to instructors to ensure the submission of high-quality materials for both synchronous and asynchronous lessons; (<reflink idref="bib2" id="ref85">2</reflink>) there is a need to encourage instructors to submit both synchronous and asynchronous lessons for evaluation, thereby promoting a comprehensive and equitable assessment process.</p> <p>Results of the current pilot study are consistent with previous research on the original POET finding inter-rater reliability, and that the tool can be used for interdisciplinary faculty peer reviews of teaching (Crabtree et al., [<reflink idref="bib12" id="ref86">12</reflink>]; DiVall et al., [<reflink idref="bib14" id="ref87">14</reflink>]; Trujillo et al., [<reflink idref="bib53" id="ref88">53</reflink>]). The current project confirmed the utility of the POET-O for peer reviews with faculty from various disciplines demonstrating that the tool may be adopted by universities and colleges for peer reviews of all faculty. Both Trujillo et al. ([<reflink idref="bib53" id="ref89">53</reflink>]) and DiVall et al. ([<reflink idref="bib14" id="ref90">14</reflink>]) utilized the POET with traditional in-person classes while Crabtree et al. ([<reflink idref="bib12" id="ref91">12</reflink>]) used a recorded video of a live class. The current project with POET-O used synchronous and Zoom-recorded synchronous lessons. Future research is needed to determine if the POET-O tool works similarly with traditional in-person class observations.</p> <p>Finally, the POET-O assumes the collaborative model with the purpose of improve teaching through constructive dialogue (Fletcher, 2018). Therefore, the peer review process meant to mutual benefit for both reviewer and reviewee. Aligned with an evidence-based framework for peer review of teaching (Dawson & Hocker, [<reflink idref="bib13" id="ref92">13</reflink>]), the goal of the current study was to implement a peer review system that promoted formative teaching reviews, engaged faculty in deeper conversations about evidence-based teaching practices, and produced a useful report for formal evaluations by the department chair and personnel committees for promotion, tenure, and merit. Although time consuming, if done well, it will lead to continuous improvement of teaching and better collegial relationship over time (Shortland, [<reflink idref="bib47" id="ref93">47</reflink>]).</p> <hd id="AN0183751073-16">Limitations</hd> <p>The voluntary sample of the current study posed a possible bias to the model estimation because the ideal situation for Rasch analysis is to have adequate representation of respondents for all possible response patterns across items (Cappelleri et al., [<reflink idref="bib8" id="ref94">8</reflink>]). If fewer respondents located at the ends of the construct being measured, items positioned at the ends of the construct will have higher standard errors. Thus, a more heterogeneous sample with greater variability in teaching proficiency will be needed to further test the tool and rating scale functionality. The majority of the faculty reviewers and lesson submitters who participated in the study were female, which represents the faculty population of the health professions school. However, literature has documented potential gender bias in teaching evaluations (e.g., Boring, [<reflink idref="bib5" id="ref95">5</reflink>]; Mengel et al., [<reflink idref="bib40" id="ref96">40</reflink>]; Özgümüs et al., [<reflink idref="bib43" id="ref97">43</reflink>]). For example, a recent study (Özgümüs et al., [<reflink idref="bib43" id="ref98">43</reflink>]) found that female instructors were rated lower than male instructors from male raters, while female instructors were rated higher from female raters. Therefore, the tool needs to be further validated with more gender-balanced sample in other professional schools. Finally, the original tool was validated in traditional in-person setting (lab, practice-based, simulation etc.). Our original intention was to modify the tool to be able to evaluate both online and in-person lessons. Due to the pandemic, when we conducted the study, we could not include the face-to-face lessons. That could be our future direction to expand the tool to those lab/practice-based lessons. Therefore, the POET-O should be further evaluated for these types of instruction and to test equivalency of the peer evaluation of online and in-person lessons. Finally, the POET-O tool does not include peer assessment of the use of learning theory, or the teaching philosophy employed in teaching, and this may be an opportunity for future research.</p> <hd id="AN0183751073-17">Implementation recommendations</hd> <p>The research team proposes the following recommendations for implementing the peer evaluation process using the POET-O tool: (a) All faculty reviewers should receive standardized training and/or debriefing to ensure reliable and consistent assessment regardless of reviewers. For example, in this pilot project, all faculty reviewers received standardized training to ensure inter-rater reliability for administering the POET-O tool. A calibration meeting was also held where discrepancies between reviewers were discussed to ensure everyone understood the meaning of each item of the tool for consistency in ratings; (b) A committee that oversees policy and procedures for peer review should be established. This committee would be responsible for developing the peer review process including the procedures to conduct the pre-observation, observation, and post observations reviews. It would also establish who qualifies for a peer review along with the scheduling and administration of the reviews; (c) Completion of the total POET-O process takes approximately 3.5 h. Faculty reviewers conducting POET-O reviews should receive release time and/or have the time spent completing reviews included in their annual evaluation under the category of Service; (d) Whether it is voluntary or required, peer review of teaching is a suitable and valuable commitment to ongoing faculty development. (e) All faculty should receive training on accessibility to ensure all course content is accessible according to the ADA section 508 guidelines and reviewers should receive clear instruction on rating ADA compliance.</p> <hd id="AN0183751073-18">Conclusion</hd> <p>The POET-O has been found to be a reliable and valid measure for evaluating both asynchronous and synchronous online lessons across various health disciplines. The tool is not specifically worded for health education settings, making it easily adaptable for peer evaluation of teaching in other higher education disciplines. The tool addresses the existing gap in standardized peer evaluation tools for individual online lessons with the consideration of time efficiency. Although additional research will be necessary to assess the impact of future modifications to the tool, the current findings suggest that the POET-O can be reasonably used for peer evaluation of teaching in higher education.</p> <hd id="AN0183751073-19">Acknowledgements</hd> <p>The project has received funding from the Rutgers University School of Health Professions Dean's Intramural Grants Programs--Teaching Innovation Awards. We thank the Dean's Taskforce on Peer Review and all the faculty members who participated in the project as lesson reviewers and/or lesson providers.</p> <hd id="AN0183751073-20">Data availability</hd> <p>The datasets generated or analyzed during the current study are available from the corresponding author on reasonable request.</p> <hd id="AN0183751073-21">Declarations</hd> <p></p> <hd id="AN0183751073-22">Conflict of interest</hd> <p>The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.</p> <hd id="AN0183751073-23">Ethical approval</hd> <p>This article does not contain any studies involving human participants performed by any of the authors.</p> <hd id="AN0183751073-24">Appendix</hd> <p></p> <hd id="AN0183751073-25">Peer Observation and Evaluation Tool–Online (POET-O)</hd> <p></p> <p> <ephtml> <table frame="hsides" rules="groups"><tbody><tr><td align="left"><p>Faculty (person being observed)</p></td></tr><tr><td align="left"><p>Reviewer</p></td></tr><tr><td align="left"><p>Course number/name/semester</p></td></tr><tr><td align="left"><p>Lesson title/topic</p></td></tr><tr><td align="left"><p>Date of review</p></td></tr></tbody></table> </ephtml> </p> <p>Time to Complete <bold><emph>Preparation (in minutes)</emph></bold>:</p> <p>Time to Complete <bold><emph>Viewing the Lesson (in minutes)</emph></bold>:</p> <p>Time to Complete <bold><emph>Assessment (in minutes)</emph></bold>:</p> <p>Peer observation is an opportunity for faculty to receive feedback from other faculty. This is a formative process designed to help faculty develop their skills as an educator. The process is three-parts: pre- observation, observation with completion of the POET-O, and post observation. The final product will be written feedback by the reviewer, discussion with the faculty member, and development of an action plan.</p> <hd id="AN0183751073-26">Section 1: Pre-observation</hd> <p>The purpose of the pre-observation is for the faculty to inform the reviewer of the desired goals or emphasis of the observation. The faculty and the reviewer should discuss areas in which the faculty member seeks feedback. A review of the course syllabus and learning outcomes is recommended. Prior to the observation, faculty should:</p> <p></p> <ulist> <item> Share syllabus, lesson materials (handouts, resources, etc.) and the assessment strategy (assignment, exam, paper) at least 1 week prior to the pre-observation to assist the reviewer in preparing for the observation.</item> <p></p> <item> Provide answers to the following questions:</item> <p></p> <item> What questions/concerns do you have? What would you particularly like feedback on?</item> <p></p> <item> How does this lesson's content fit within the entire course (e.g., is this lesson one part of several lessons on the same topic)?</item> <p></p> <item> What are the learning activities planned for this lesson? Which lesson objectives will these activities meet? How will these activities facilitate student learning?</item> </ulist> <hd id="AN0183751073-27">Section 2: Observation</hd> <p></p> <p> <ephtml> <table frame="hsides" rules="groups"><thead><tr><th align="left" /><th align="left"><p>NA</p></th><th align="left"><p>1</p></th><th align="left"><p>2</p></th><th align="left"><p>3</p></th><th align="left"><p>4</p></th><th align="left"><p>Comments</p></th></tr></thead><tbody><tr><td align="left" colspan="7"><p><italic>Presentation and Activities</italic></p></td></tr><tr><td align="left"><p>1. The lesson and learning activities are well organized</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>2. Planned student activities reflect appropriate lesson objectives</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>3. Depth and breadth of material presented appears appropriate to type of course, student level, and time dedicated to the topic</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>4. Learning activities provide opportunities for interaction that support active learning</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>5. Appropriate amount of content and course requirements are delivered</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>6. Discussion forum is interactive</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>7. Discussion forum and other learning tools promote critical thinking</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p><italic>Use of Technology</italic></p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>8. The instructor effectively uses audio/visual learning tools and technologies to support the learning objectives or competencies</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>9. Learning tools and technologies used promote learner engagement and active learning</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>10. Learning tools and technologies required in the module/unit are provided and/or readily obtainable</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>11. Access to links and external resources provided are functional</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>12. Learning tools and technologies used are current</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left" colspan="7"><p><italic>ADA Compliance/Accessibility</italic></p></td></tr><tr><td align="left"><p>13. Presentation and organization of materials facilitate ease of use</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>14. Information is provided regarding accessibility of all learning tools and technologies required in the lesson</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>15. Alternative means of access to lesson materials in formats that meet the needs of diverse learners are provided</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>16. All materials are designed to facilitate readability</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left" colspan="7"><p><italic>Assessment</italic></p></td></tr><tr><td align="left"><p>17. Planned assessment strategies measure the stated learning objectives or competencies</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>18. Specific and descriptive criteria are provided for the evaluation of learners' work, i.e. rubrics etc</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>19. Students are given ample time to complete the assignments</p><p>and assessments for this lesson</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left" colspan="7"><p><italic>Instructor</italic></p></td></tr><tr><td align="left"><p>20. Instructor provides an overview of what is planned for the lesson</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>21. Instructor appears well prepared for lesson</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>22. Instructor establishes the relevance of information</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>23. Lesson content is up to date and there is evidence that instructor is knowledgeable about topics presented</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>24. Content is explained clearly, providing examples when appropriate</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>25. Instructor makes connections with prior learning (from previous lessons and courses) when applicable</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>26. Instructor facilitates productive discussions related to content</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>27. Instructor encourages critical thinking and expression of divergent opinions or conflicting views when appropriate</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>28. Instructor contributes to discussion, providing encouragement, reflection, and corrective feedback</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>29. Instructor provides periodic summaries of the most important ideas and ties things together at the end of the lesson</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>30. Instructor uses active learning techniques</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>31. Instructor treats students respectfully, offering encouragement in public ways and does not criticize student publicly</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr><tr><td align="left"><p>32. Instructor utilizes diverse methods of instruction to accommodate all learners</p></td><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /><td align="left" /></tr></tbody></table> </ephtml> </p> <p>NA = Not Applicable</p> <p>1 = Not Present –Should consider adding</p> <ulist> <item>2 = Need Development–Major revisions are recommended</item> <item>3 = Accomplished–Minor revisions are recommended</item> <item>4 = Accomplished Well–No recommendations for improvement</item> </ulist> <hd id="AN0183751073-28">Section 3: General reviewer comments</hd> <p></p> <hd id="AN0183751073-29">Section 4: Post observation</hd> <p>The Post Observation feedback includes a discussion with the faculty member to review the checklist and any additional comments from the reviewer.</p> <p>The following are items to consider for that discussion:</p> <p>How do you think the lesson went? Provide an example(s) of something that went well and something that may need improvement.</p> <p>Is there anything you wanted to accomplish but were unable to do so? If yes, what was it and was it critical? What would you do differently next time to accomplish it?</p> <hd id="AN0183751073-30">Section 5: Faculty Reflection and Action Plan (to be completed by Faculty who was observed)....</hd> <p></p> <ulist> <item> Reflection (SWOT-Strengths, Weaknesses, Opportunities, Threats</item> <p></p> <item> Action Plan: (SMART-Specific, Measurable, Attainable, Relevant, Time)</item> </ulist> <hd id="AN0183751073-31">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0183751073-32"> <title> References </title> <blist> <bibl id="bib1" idref="ref24" type="bt">1</bibl> <bibtext> Baldwin S, Ching YH, Hsu YC. Online course design in higher education: A review of national and statewide evaluation instruments. TechTrends. 2018; 62: 46-57. 10.1007/s11528-017-0215-z</bibtext> </blist> <blist> <bibl id="bib2" idref="ref1" type="bt">2</bibl> <bibtext> Bell A, Mladenovic R. The benefits of peer observation of teaching for tutor development. Higher Education. 2008; 55; 6: 735-752. 10.1007/s10734-007-9093-1</bibtext> </blist> <blist> <bibl id="bib3" idref="ref8" type="bt">3</bibl> <bibtext> Bell M. Supported reflective practice: A programme of peer observation and feedback for academic teaching development. International Journal for Academic Development. 2001; 6; 1: 29-39. 10.1080/13601440110033643</bibtext> </blist> <blist> <bibl id="bib4" idref="ref69" type="bt">4</bibl> <bibtext> Bond, T. G, & Fox, C. M. (2007). Applying the Rasch model: Fundamental measurement in the human sciences. Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref95" type="bt">5</bibl> <bibtext> Boring A. Gender biases in student evaluations of teaching. Journal of Public Economics. 2017; 145: 27-41. 10.1016/j.jpubeco.2016.11.006</bibtext> </blist> <blist> <bibl id="bib6" idref="ref16" type="bt">6</bibl> <bibtext> Boring, A, Ottoboni, K, & Stark, P. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research.https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AETBZC.v1</bibtext> </blist> <blist> <bibl id="bib7" idref="ref27" type="bt">7</bibl> <bibtext> Brown B, Ward-Griffin C. The use of peer evaluation in promoting nursing faculty teaching effectiveness: A review of the literature. Nurse Education Today. 1994; 14; 4: 299-305. 10.1016/0260-6917(94)90141-4</bibtext> </blist> <blist> <bibl id="bib8" idref="ref94" type="bt">8</bibl> <bibtext> Cappelleri JC, Lundy JJ, Hays RD. Overview of classical test theory and item response theory for the quantitative assessment of items in developing patient-reported outcomes measures. Clinical Therapeutics. 2014; 36; 5: 648-662. 10.1016/j.clinthera.2014.04.006</bibtext> </blist> <blist> <bibl id="bib9" idref="ref3" type="bt">9</bibl> <bibtext> Chao T, Saj T, Tessier F. Establishing a quality review for online courses. Educause Quarterly. 2006; 29; 3: 32-39</bibtext> </blist> <blist> <bibtext> Chism, N. V. N. (2007). Peer review of teaching: A sourcebook (2nd ed.). Bolton, MA: Anker.</bibtext> </blist> <blist> <bibtext> Cox, C. D, Peeters, M. J, Stanford, B. L, & Seifert, C. F. (2013). Pilot of peer assessment within experiential teaching and learning. Currents in Pharmacy Teaching and Learning,5(4), 311–320. https://doi.org/10.1016/j.cptl.2013.02.003</bibtext> </blist> <blist> <bibtext> Crabtree, J. L, Scott, P. J, & Kuo, F. (2016). Peer observation and evaluation tool (POET): A formative peer review supporting teaching. The Open Journal of Occupational Therapy. https://doi.org/10.15453/2168-6408.1273</bibtext> </blist> <blist> <bibtext> Dawson SM, Hocker AD. An evidence-based framework for peer review of teaching. Advances in Physiology Education. 2020; 44; 1: 26-31. 10.1152/advan.00088.2019</bibtext> </blist> <blist> <bibtext> DiVall M, Barr J, Gonyeau M, Matthews JV, Amburgh J, Qualters D, Trujillo J. Follow-up assessment of a faculty peer observation and evaluation program. American Journal of Pharmaceutical Education. 2012. 10.5688/ajpe76461</bibtext> </blist> <blist> <bibtext> Eckes, T. (2009). On common ground? How raters perceive scoring criteria in oral proficiency testing. In A. Brown & K. Hill (Eds.), Tasks and criteria in performance assessment: Proceedings of the 28th language testing research colloquium (pp. 43–73). Frankfurt am Main: Peter Lang.</bibtext> </blist> <blist> <bibtext> Eckes, T. (2015). Introduction to many-facet Rasch measurement: Analyzing and evaluating rater-mediated assessments (2nd ed.). Peter Lang</bibtext> </blist> <blist> <bibtext> Eddy, S. L, Converse, M, & Wenderoth, M. P. (2015). PORTAAL: A classroom observation tool assessing evidence-based teaching practices for active learning in large science, technology, engineering, and mathematics classes. CBE—Life Sciences Education, 14(2), 1–16. https://doi.org/10.1187/cbe.14-06-0095</bibtext> </blist> <blist> <bibtext> Esarey J, Valdes N. Unbiased, reliable and valid evaluations can still be unfair. Assessment and Evaluation in Higher Education. 2020; 45; 8: 1106-1120. 10.1080/02602938.2020.1724875</bibtext> </blist> <blist> <bibtext> Finkelstein, N, Corbo, J. C, Reinholz, D. L, Gammon, M, & Keating, J. (2017). Evaluating teaching in a scholarly manner: A model and call for an evidence-based, departmentally-defined approach to enhance teaching evaluation for CU Boulder. Retrieved January, 5, 2019 from https://<ulink href="http://www.colorado.edu/teaching-qualityframework/sites/default/files/attached-files/2017-11%5ftqf-white-paper%5fnorecs.pdf">www.colorado.edu/teaching-qualityframework/sites/default/files/attached-files/2017-11%5ftqf-white-paper%5fnorecs.pdf</ulink></bibtext> </blist> <blist> <bibtext> Fletcher, J. A. (2018). Peer observation of teaching: A practical tool in higher education. The Journal of Faculty Development,32(1), 51–64.</bibtext> </blist> <blist> <bibtext> Fox, K, Srinivasan, N, Lin, N, Nguyen, A, & Gates, B. (2020). Time for class: COVID-19 edition part 2. https://tytonpartners.com/library/time-for-class-covid-19-edition-part-2/.</bibtext> </blist> <blist> <bibtext> Gerbing DW, Anderson JC. An updated paradigm for scale development incorporating unidimensionality and its assessment. Journal of Marketing Research. 1988; 25; 2: 186-192. 10.2307/3172650</bibtext> </blist> <blist> <bibtext> Hammersley-Fletcher L, Orsmond P. Reflecting on reflective practices within peer observation. Studies in Higher Education. 2005; 30; 2: 213-224. 10.1080/03075070500043358</bibtext> </blist> <blist> <bibtext> Hill HC, Charalambous CY, Kraft MA. When rater reliability is not enough: Teacher observation systems and a case for the generalizability study. Educational Researcher. 2012; 41; 2: 56-64. 10.3102/0013189X12437203</bibtext> </blist> <blist> <bibtext> Hornstein HA. Student evaluations of teaching are an inadequate assessment tool for evaluating faculty performance. Cogent Education. 2017; 4; 1: 1304016. 10.1080/2331186X.2017.1304016</bibtext> </blist> <blist> <bibtext> Le S, Howard ML. Peer evaluation of teaching programs within pharmacy education: A review of the literature. Currents in Pharmacy Teaching and Learning. 2023. 10.1016/j.cptl.2023.09.009</bibtext> </blist> <blist> <bibtext> Linacre JM. Many-facet Rasch measurement. 19942; MESA Press</bibtext> </blist> <blist> <bibtext> Linacre JM. Understanding Rasch measurement: Estimation methods for Rasch measures. Journal of Outcome Measurement. 1999; 3: 381-405</bibtext> </blist> <blist> <bibtext> Linacre JM. What do infit and outfit, mean-square and standardized mean?. Rasch Measurement Transaction. 2002; 7; 4: 878</bibtext> </blist> <blist> <bibtext> Linacre JM. Rasch model estimation: Further topics. Journal of Applied Measurement. 2004; 5; 1: 95-110</bibtext> </blist> <blist> <bibtext> Linacre, J. M. (2011). Winsteps Rasch measurement computer program (Version 3.72). Winsteps.com</bibtext> </blist> <blist> <bibtext> Linacre, J. M. (2021). Facets computer program for many-facet Rasch measurement, version 3.83.6. Winsteps.com</bibtext> </blist> <blist> <bibtext> Linacre, J. M. (n.d.). Unexpected (standardized residuals reported if not less than) = 3. Retrieved September 1, 2021 from https://<ulink href="http://www.winsteps.com/facetman/unexpected.htm">www.winsteps.com/facetman/unexpected.htm</ulink></bibtext> </blist> <blist> <bibtext> Linacre, J. M. (n.d.). Reliability and separation of measures. https://<ulink href="http://www.winsteps.com/winman/reliability.htm">www.winsteps.com/winman/reliability.htm</ulink></bibtext> </blist> <blist> <bibtext> Linse AR. Interpreting and using student ratings data: Guidance for faculty serving as administrators and on evaluation committees. Studies in Educational Evaluation. 2017; 54: 94-106. 10.1016/j.stueduc.2016.12.004</bibtext> </blist> <blist> <bibtext> Little BB. Quality assurance for online nursing courses. Journal of Nursing Education. 2009; 48; 7: 381-387. 10.3928/01484834-20090615-05</bibtext> </blist> <blist> <bibtext> Maryland Online, Inc. (2017). Quality Matters non-annotated standards, rubric, third edition. https://<ulink href="http://www.qualitymatters.org/qa-resources/rubric-standards">www.qualitymatters.org/qa-resources/rubric-standards</ulink></bibtext> </blist> <blist> <bibtext> Mager, D. R, Kazer, M. W, Conelius, J, Shea, J, Lippman, D. T, Torosyan, R, & Nantz, K. (2014). Development, implementation and evaluation of a peer review of teaching (PRoT) initiative in nursing education. International journal of nursing education scholarship,11(1), 113–120. https://doi.org/10.1515/ijnes-2013-0019</bibtext> </blist> <blist> <bibtext> McGahan SJ, Jackson CM, Premer K. Online course quality assurance: Development of a quality checklist. InSight: A Journal of Scholarly Teaching. 2015; 10: 126-140. 10.46504/10201510mc</bibtext> </blist> <blist> <bibtext> Mengel F, Sauermann J, Zölitz U. Gender bias in teaching evaluations. Journal of the European Economic Association. 2019; 17; 2: 535-566. 10.1093/jeea/jvx057</bibtext> </blist> <blist> <bibtext> Mulqueen, C, Baker, D, & Dismukes, R. K. (2000, April). Using multifaceted Rasch analysis to examine the effectiveness of rater training. In 15th annual conference for the society for industrial and organizational psychology (SIOP), New Orleans. <ulink href="http://www.air.org/files/multifacet%5fRasch.Pdf">http://www.air.org/files/multifacet%5fRasch.Pdf</ulink></bibtext> </blist> <blist> <bibtext> Myford C, Wolfe EW. Detecting and measuring rater effects using many-facet Rasch measurement: Part I. Journal of Applied Measurement. 2003; 4; 4: 386-422</bibtext> </blist> <blist> <bibtext> Özgümüs A, Rau HA, Trautmann ST, König-Kersting C. Gender bias in the evaluation of teaching materials. Frontiers in Psychology. 2020; 11: 1074. 10.3389/fpsyg.2020.01074</bibtext> </blist> <blist> <bibtext> Parker Harris, S, Gould, R, and Mullin, C. (2019). ADA research brief: Higher education and the ADA (pp. 1–6). ADA National Network Knowledge Translation Center. https://adata.org/research_brief/higher-education-and-ada</bibtext> </blist> <blist> <bibtext> Ridge, B. L, & Lavigne, A. L. (2020). Improving instructional practice through peer observation and feedback. Education Policy Analysis Archives. https://doi.org/10.14507/epaa.28.5023</bibtext> </blist> <blist> <bibtext> Sachs, J, & Parsell, M. (2014). Peer review of learning and teaching in higher education: International perspectives. Springer Science & Business Media.</bibtext> </blist> <blist> <bibtext> Shortland S. Feedback within peer observation: Continuing professional development and unexpected consequences. Innovations in Education and Teaching International. 2010; 47; 3: 295-304. 10.1080/14703297.2010.498181</bibtext> </blist> <blist> <bibtext> Smith, M. K, Jones, F. H, Gilbert, S. L, & Wieman, C. E. (2013). The classroom observation protocol for undergraduate STEM (COPUS): A new instrument to characterize university STEM classroom practices. CBE—Life Sciences Education, 12, 618–627. https://doi.org/10.1187/cbe.13-08-0154</bibtext> </blist> <blist> <bibtext> Snyder, T. D, de Brey, C, & Dillow, S. A. (2019). Digest of education statistics 2017 (NCES 2018–070). National Center for Education Statistics, Institute of Education Sciences, U.S. Department of Education</bibtext> </blist> <blist> <bibtext> Spooren P, Brockx B, Mortelmans D. On the validity of student evaluation of teaching: The state of the art. Review of Educational Research. 2013; 83; 4: 598-642. 10.3102/0034654313496870</bibtext> </blist> <blist> <bibtext> Stevens JP. Applied multivariate statistics for the social sciences. 19922; Erlbaum</bibtext> </blist> <blist> <bibtext> Thomas S, Chie QT, Abraham M, Jalarajan Raj S, Beh LS. A qualitative review of literature on peer review of teaching in higher education: An application of the SWOT framework. Review of Educational Research. 2014; 84; 1: 112-159. 10.3102/0034654313499617</bibtext> </blist> <blist> <bibtext> Trujillo JM, DiVall MV, Barr J, Gonyeau M, Van Amburgh JA, Matthews SJ, Qualters D. Development of a peer teaching-assessment program and a peer observation and evaluation tool. American Journal of Pharmaceutical Education. 2008. 10.5688/aj7206147</bibtext> </blist> <blist> <bibtext> Weaver GC, Austin AE, Greenhoot AF, Finkelstein ND. Establishing a better approach for evaluating teaching: The TEval Project. Change: The Magazine of Higher Learning. 2020; 52; 3: 25-31. 10.1080/00091383.2020.1745575</bibtext> </blist> <blist> <bibtext> Wind SA, Jones E. Not just generalizability: A case for multifaceted latent trait models in teacher observation systems. Educational Researcher. 2019; 48; 8: 521-533. 10.3102/0013189X19874084</bibtext> </blist> <blist> <bibtext> Yiend J, Weller S, Kinchin I. Peer observation of teaching: The interaction between peer review and developmental models of practice. Journal of Further and Higher Education. 2014; 38; 4: 465-484. 10.1080/0309877X.2012.726967</bibtext> </blist> </ref> <aug> <p>By Yuane Jia; Amy B. Spagnolo; Nora Barrett; Ann A. Murphy; Peter M. Basto; Pamela Rothpletz-Puglia and Stuart Luther</p> <p>Reported by Author; Author; Author; Author; Author; Author; Author</p> <p></p> <p>Yuane Jia Yuane Jia is an Assistant Professor in the Rutgers School of Health Professions' Department of Interdisciplinary Studies. Additionally, Dr. Jia is on the school's Methodology and Statistical Support Team and specializes in quantitative methods for advancing health science research. Her current research interests are on scale development and validation, advanced statistical modeling, and scholarship of teaching and learning in health science.</p> <p>Amy B. Spagnolo Amy B. Spagnolo is an Associate Professor in the Rutgers School of Health Professions' Department of Psychiatric Rehabilitation and Counseling Professions. Her research interests include best practice approaches for online curriculum development, instructor and course evaluation, peer support workforce development and wellness initiatives.</p> <p>Nora Barrett Nora Barrett is an Associate Professor in the Rutgers School of Health Professions' Department of Psychiatric Rehabilitation and Counseling Professions. Her research interests include evidence-based teaching and professional education strategies, models of faculty mentorship, and intervention strategies for older adults with serious mental health conditions.</p> <p>Ann A. Murphy Ann A. Murphy is an Associate Professor and Director in the Rutgers School of Health Professions' Department of Psychiatric Rehabilitation and Counseling Professions. Her research interests include developing the behavioral health workforce, implementation science, and interdisciplinary education to enhance the capacity of future health professionals to adequately address the needs of individuals with serious mental illnesses.</p> <p>Peter M. Basto Peter M. Basto is an Assistant Professor and an Undergraduate Program Director of Psychiatric Rehabilitation at Rutgers University. He has published on the psychiatric rehabilitation workforce, and on peer run services. His current research interests are on student learning outcomes and faculty peer review.</p> <p>Pamela Rothpletz-Puglia Pamela Rothpletz-Puglia is a Professor and a Director of a Ph.D. program at Rutgers School of Health Professions. In addition to experience with the scholarship of teaching and learning, Dr. Rothpletz-Puglia is on the school's Methodology and Statistical Support Team and specializes in qualitative and mixed methods for advancing health science research. For more information about her scholarship and expertise, please see the following website:  https://sites.google.com/scarletmail.rutgers.edu/bps/home</p> <p>Stuart Luther Stuart Luther is a senior training and consultation specialist for the Rutgers-Center for Comprehensive School Mental Health.</p> </aug> <nolink nlid="nl1" bibid="bib56" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib36" firstref="ref4"></nolink> <nolink nlid="nl3" bibid="bib39" firstref="ref5"></nolink> <nolink nlid="nl4" bibid="bib45" firstref="ref6"></nolink> <nolink nlid="nl5" bibid="bib20" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib23" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib52" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib47" firstref="ref12"></nolink> <nolink nlid="nl9" bibid="bib46" firstref="ref14"></nolink> <nolink nlid="nl10" bibid="bib18" firstref="ref17"></nolink> <nolink nlid="nl11" bibid="bib25" firstref="ref18"></nolink> <nolink nlid="nl12" bibid="bib50" firstref="ref19"></nolink> <nolink nlid="nl13" bibid="bib35" firstref="ref21"></nolink> <nolink nlid="nl14" bibid="bib54" firstref="ref22"></nolink> <nolink nlid="nl15" bibid="bib19" firstref="ref23"></nolink> <nolink nlid="nl16" bibid="bib10" firstref="ref28"></nolink> <nolink nlid="nl17" bibid="bib11" firstref="ref30"></nolink> <nolink nlid="nl18" bibid="bib26" firstref="ref33"></nolink> <nolink nlid="nl19" bibid="bib17" firstref="ref34"></nolink> <nolink nlid="nl20" bibid="bib48" firstref="ref35"></nolink> <nolink nlid="nl21" bibid="bib13" firstref="ref37"></nolink> <nolink nlid="nl22" bibid="bib53" firstref="ref41"></nolink> <nolink nlid="nl23" bibid="bib12" firstref="ref43"></nolink> <nolink nlid="nl24" bibid="bib21" firstref="ref45"></nolink> <nolink nlid="nl25" bibid="bib37" firstref="ref46"></nolink> <nolink nlid="nl26" bibid="bib16" firstref="ref54"></nolink> <nolink nlid="nl27" bibid="bib24" firstref="ref56"></nolink> <nolink nlid="nl28" bibid="bib55" firstref="ref57"></nolink> <nolink nlid="nl29" bibid="bib27" firstref="ref59"></nolink> <nolink nlid="nl30" bibid="bib41" firstref="ref61"></nolink> <nolink nlid="nl31" bibid="bib32" firstref="ref64"></nolink> <nolink nlid="nl32" bibid="bib31" firstref="ref65"></nolink> <nolink nlid="nl33" bibid="bib29" firstref="ref66"></nolink> <nolink nlid="nl34" bibid="bib42" firstref="ref67"></nolink> <nolink nlid="nl35" bibid="bib28" firstref="ref72"></nolink> <nolink nlid="nl36" bibid="bib15" firstref="ref74"></nolink> <nolink nlid="nl37" bibid="bib22" firstref="ref75"></nolink> <nolink nlid="nl38" bibid="bib51" firstref="ref76"></nolink> <nolink nlid="nl39" bibid="bib30" firstref="ref78"></nolink> <nolink nlid="nl40" bibid="bib49" firstref="ref82"></nolink> <nolink nlid="nl41" bibid="bib44" firstref="ref83"></nolink> <nolink nlid="nl42" bibid="bib14" firstref="ref87"></nolink> <nolink nlid="nl43" bibid="bib40" firstref="ref96"></nolink> <nolink nlid="nl44" bibid="bib43" firstref="ref97"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1462806
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Validation of a Peer Observation and Evaluation Tool for Online Teaching in the U.S.
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yuane+Jia%22">Yuane Jia</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-1792-0431">0000-0002-1792-0431</externalLink>)<br /><searchLink fieldCode="AR" term="%22Amy+B%2E+Spagnolo%22">Amy B. Spagnolo</searchLink><br /><searchLink fieldCode="AR" term="%22Nora+Barrett%22">Nora Barrett</searchLink><br /><searchLink fieldCode="AR" term="%22Ann+A%2E+Murphy%22">Ann A. Murphy</searchLink><br /><searchLink fieldCode="AR" term="%22Peter+M%2E+Basto%22">Peter M. Basto</searchLink><br /><searchLink fieldCode="AR" term="%22Pamela+Rothpletz-Puglia%22">Pamela Rothpletz-Puglia</searchLink><br /><searchLink fieldCode="AR" term="%22Stuart+Luther%22">Stuart Luther</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Educational+Technology+Research+and+Development%22"><i>Educational Technology Research and Development</i></searchLink>. 2025 73(1):615-639.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 25
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research<br />Tests/Questionnaires
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="DE" term="%22Peer+Evaluation%22">Peer Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Lesson+Observation+Criteria%22">Lesson Observation Criteria</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Reliability%22">Test Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Validity%22">Test Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Online+Courses%22">Online Courses</searchLink><br /><searchLink fieldCode="DE" term="%22Teacher+Evaluation%22">Teacher Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+Learning%22">Electronic Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Health+Education%22">Health Education</searchLink><br /><searchLink fieldCode="DE" term="%22Teacher+Effectiveness%22">Teacher Effectiveness</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1007/s11423-024-10428-z
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1042-1629<br />1556-6501
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The benefits of peer evaluation of teaching effectiveness and quality in higher education are well documented. While instruments exist for the review and evaluation of entire online courses, there is no standardized single-lesson, peer evaluation instrument available for online instruction. This pilot study focused on the validation of a peer observation and evaluation tool for use with single lessons in both synchronous and asynchronous online courses in an inter-professional school of health professions. The researchers modified a psychometrically validated instrument developed for in-person peer observation by adding items from a renowned online course rubric to create a peer observation tool, entitled the Peer Observation and Evaluation Tool-Online (POET-O). The resulting instrument demonstrated adequate construct validity and reliability by using the many-facet Rasch measurement (MFRM) technique. MFRM results also indicated potential places to revise and improve the instrument. Recommendations for implementing the peer evaluation process of teaching are provided.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1462806
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1462806
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11423-024-10428-z
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
        StartPage: 615
    Subjects:
      – SubjectFull: Higher Education
        Type: general
      – SubjectFull: Peer Evaluation
        Type: general
      – SubjectFull: Lesson Observation Criteria
        Type: general
      – SubjectFull: Test Construction
        Type: general
      – SubjectFull: Test Reliability
        Type: general
      – SubjectFull: Test Validity
        Type: general
      – SubjectFull: Online Courses
        Type: general
      – SubjectFull: Teacher Evaluation
        Type: general
      – SubjectFull: Electronic Learning
        Type: general
      – SubjectFull: Health Education
        Type: general
      – SubjectFull: Teacher Effectiveness
        Type: general
    Titles:
      – TitleFull: Validation of a Peer Observation and Evaluation Tool for Online Teaching in the U.S.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yuane Jia
      – PersonEntity:
          Name:
            NameFull: Amy B. Spagnolo
      – PersonEntity:
          Name:
            NameFull: Nora Barrett
      – PersonEntity:
          Name:
            NameFull: Ann A. Murphy
      – PersonEntity:
          Name:
            NameFull: Peter M. Basto
      – PersonEntity:
          Name:
            NameFull: Pamela Rothpletz-Puglia
      – PersonEntity:
          Name:
            NameFull: Stuart Luther
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1042-1629
            – Type: issn-electronic
              Value: 1556-6501
          Numbering:
            – Type: volume
              Value: 73
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Educational Technology Research and Development
              Type: main
ResultId 1