How Events Enter (or Not) Data Sets: The Pitfalls and Guidelines of Using Newspapers in the Study of Conflict

Saved in:
Bibliographic Details
Title: How Events Enter (or Not) Data Sets: The Pitfalls and Guidelines of Using Newspapers in the Study of Conflict
Language: English
Authors: Demarest, Leila (ORCID 0000-0001-6887-9937), Langer, Arnim
Source: Sociological Methods & Research. May 2022 51(2):632-666.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: http://sagepub.com
Peer Reviewed: Y
Page Count: 35
Publication Date: 2022
Document Type: Journal Articles
Reports - Descriptive
Descriptors: Guidelines, Research Methodology, Conflict, Social Science Research, Newspapers, Research Problems, Error of Measurement, Data Analysis
DOI: 10.1177/0049124119882453
ISSN: 0049-1241
Abstract: While conflict event data sets are increasingly used in contemporary conflict research, important concerns persist regarding the quality of the collected data. Such concerns are not necessarily new. Yet, because the methodological debate and evidence on potential errors remains scattered across different subdisciplines of social sciences, there is little consensus concerning proper reporting practices in codebooks, how best to deal with the different types of errors, and which types of errors should be prioritised. In this article, we introduce a new analytical framework--that is, the Total Event Error (TEE) framework--which aims to elucidate the methodological challenges and errors that may affect whether and how events are entered into conflict event data sets, drawing on different fields of study. Potential errors are diverse and may range from errors arising from the rationale of the media source (e.g., selection of certain types of events into the news) to errors occurring during the data collection process or the analysis phase. Based on the TEE framework, we propose a set of strategies to mitigate errors associated with the construction and use of conflict event data sets. We also identify a number of important avenues for future research concerning the methodology of creating conflict event data sets.
Abstractor: As Provided
Entry Date: 2022
Accession Number: EJ1337281
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHcI7ElDRJe_Jd_qLrhsHY4AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDLuuJSFjwPAlnVxdEwIBEICBmzspfr4UaxmIBNEkcQe8DhgX4PBNV6oyimmIzVXeULXrPL8VG7OJxhAe6D-zif6yn9e02lsu6oSjBtIlLzzWDqy8mn6F5xmGp_4FPGF6nvGRYbGiEldbtIe-xJ7IlT6vfsd0io-nKyfu6NGH7Km85TFwhXVv7afsU8fLQ0uepkOZlgcV70NUFuhR9PEsxz70yruvaksbTxg4kJfX
Text:
  Availability: 1
  Value: <anid>AN0156346852;som01may.22;2022Apr19.05:10;v2.2.500</anid> <title id="AN0156346852-1">How Events Enter (or Not) Data Sets: The Pitfalls and Guidelines of Using Newspapers in the Study of Conflict </title> <p>While conflict event data sets are increasingly used in contemporary conflict research, important concerns persist regarding the quality of the collected data. Such concerns are not necessarily new. Yet, because the methodological debate and evidence on potential errors remains scattered across different subdisciplines of social sciences, there is little consensus concerning proper reporting practices in codebooks, how best to deal with the different types of errors, and which types of errors should be prioritised. In this article, we introduce a new analytical framework—that is, the Total Event Error (TEE) framework—which aims to elucidate the methodological challenges and errors that may affect whether and how events are entered into conflict event data sets, drawing on different fields of study. Potential errors are diverse and may range from errors arising from the rationale of the media source (e.g., selection of certain types of events into the news) to errors occurring during the data collection process or the analysis phase. Based on the TEE framework, we propose a set of strategies to mitigate errors associated with the construction and use of conflict event data sets. We also identify a number of important avenues for future research concerning the methodology of creating conflict event data sets.</p> <p>Keywords: conflict events; data error; bias; unreliability; media data; total survey error</p> <hd id="AN0156346852-2">Introduction</hd> <p>With the quantitative turn in peace and conflict studies (e.g., [<reflink idref="bib11" id="ref1">11</reflink>]; [<reflink idref="bib25" id="ref2">25</reflink>]; [<reflink idref="bib39" id="ref3">39</reflink>]:3), the collection of detailed data on conflict events, actors, and casualty numbers has spurred many research projects in the field. The development of conflict event data sets has also accelerated in recent years. Major trends in new data projects include an increased focus on disaggregated conflict events—both in time and in space—as well as a focus on low-level forms of conflict such as protests, as opposed to civil war ([<reflink idref="bib3" id="ref4">3</reflink>]:375-78; Gleditsch et al. 2014:303-5, 308-9). News reports have been the most important source of data on conflict events, as they are widely available and often accessible at low cost. However, the widespread use of media reports as an empirical source raises important concerns about the quality of event data.</p> <p>The collection of political event data has a long history, both in the social movement (e.g., [<reflink idref="bib23" id="ref5">23</reflink>]) and in the international relations literature (e.g., [<reflink idref="bib1" id="ref6">1</reflink>]; [<reflink idref="bib52" id="ref7">52</reflink>]). Concerns about the validity and reliability of media data for political science research are hence not necessarily new (e.g., [<reflink idref="bib16" id="ref8">16</reflink>]; [<reflink idref="bib26" id="ref9">26</reflink>]). Nonetheless, the increased availability of (online) media sources and the development of new data sets has spurred new debates in this field. An important characteristic of many new data sets is their geographical focus on the developing world. Examples include the Armed Conflict Location and Event Dataset (ACLED; [<reflink idref="bib58" id="ref10">58</reflink>]), the Social Conflict in Africa Database (SCAD; [<reflink idref="bib62" id="ref11">62</reflink>]), the Urban Social Disturbance in Africa and Asia (USDAA) data set ([<reflink idref="bib72" id="ref12">72</reflink>]), the UCDP Georeferenced Event Dataset (UCDP GED; [<reflink idref="bib71" id="ref13">71</reflink>]),[<reflink idref="bib4" id="ref14">4</reflink>] the Global Terrorism Database ([<reflink idref="bib49" id="ref15">49</reflink>]), Political Instability Task Force (PITF) Worldwide Atrocities Dataset ([<reflink idref="bib67" id="ref16">67</reflink>]), the Mass Mobilization in Autocracies Database ([<reflink idref="bib76" id="ref17">76</reflink>]), and the Konstanz One-Sided Violence Event Dataset ([<reflink idref="bib69" id="ref18">69</reflink>]). By contrast, much of the methodological debate and evidence with respect to the use of media reports to construct event data is found in Western-focused social movement research (e.g., [<reflink idref="bib21" id="ref19">21</reflink>]; [<reflink idref="bib41" id="ref20">41</reflink>]) as well as in communications studies ([<reflink idref="bib27" id="ref21">27</reflink>]; [<reflink idref="bib33" id="ref22">33</reflink>]; [<reflink idref="bib47" id="ref23">47</reflink>]). Further, while in recent years, different methodological challenges associated with generating new conflict event data sets have been critically assessed (e.g., [<reflink idref="bib22" id="ref24">22</reflink>]; [<reflink idref="bib61" id="ref25">61</reflink>]; [<reflink idref="bib74" id="ref26">74</reflink>], [<reflink idref="bib75" id="ref27">75</reflink>]), this field of study stands to benefit from further systematization of research findings.</p> <p>To this aim, the current article introduces a new analytical framework that captures the methodological challenges of using news reports for generating conflict event data and recognizes a broad range of errors that may affect whether and how events are entered into data sets. The Total Event Error (TEE) framework draws on insights from the survey research literature and the Total Survey Error (TSE) framework. In analogy with [<reflink idref="bib29" id="ref28">29</reflink>]:41-63), we distinguish between measurement errors and errors of representation. The framework encompasses well-known forms of error mentioned in the literature, such as selection bias, which arises when newspapers deliberately select some events for publication, while leaving other events unreported ([<reflink idref="bib21" id="ref29">21</reflink>]:68-72; [<reflink idref="bib42" id="ref30">42</reflink>]; [<reflink idref="bib56" id="ref31">56</reflink>]; [<reflink idref="bib75" id="ref32">75</reflink>]). However, we also consider errors that are not necessarily caused by the rationale of media sources and have received much less attention in the literature. These errors arise during the data collection process, such as the coding of key variables, or in the analysis phase, when researchers make use of imputed values for missing data (e.g., the location of an event). Further, while bias, or a systematic difference between the measured value and the real value, is an important form of error, we also direct attention toward unreliability or random deviation from the real value, which undermines precision.</p> <p>The TEE framework offers a bridge between methodological insights from conflict event studies and Western-focused social movement and communications studies. This has the advantage that insights, methods, and procedures that are common in these latter literatures are introduced and discussed with regard to conflict events in developing countries. We devote particular attention to the implications of focusing on developing contexts as opposed to Western contexts. Indeed, while Western-centred studies have commonly focused on protest events, we focus on a wider range of events, including protests, but also violent armed conflict events. Finally, we discuss and compare human as well as automated forms of data collection and coding. Although optimism has often been expressed with regard to the potential opportunities and advantages of automated coding (e.g., [<reflink idref="bib6" id="ref33">6</reflink>]; [<reflink idref="bib44" id="ref34">44</reflink>]; [<reflink idref="bib48" id="ref35">48</reflink>]), so far, it is not yet widely used in conflict studies. Illustratively, human coding is used by all data projects cited above. Arguably, the main reason for why human coding has remained the common practice is that in recent years, conflict scholars have aimed to construct conflict event data sets on the basis of increasingly complex information drawn from media reports (e.g., [<reflink idref="bib31" id="ref36">31</reflink>]). Having said this, automated coding has important advantages compared to human coding and is therefore likely to gain more relevance in conflict studies in the coming years.</p> <p>As an analytical framework, TEE offers an important methodological basis for studies on conflict events and gives guidance to developers and users of old and new data sets. For developers, the TEE framework systematically sets out the different types of errors to be reflected upon in data codebooks, or articles introducing new data sets, and supports standardization of reporting practices in the field. Indeed, as the collection of event data and the use of media data have been taken up by different subdisciplines and areas of social science, the types of errors researchers are concerned with, or report on, appear to differ widely. As will also become clear from our discussion of the state of the art, relatively little is known concerning errors that may arise when collecting data on conflict events in the developing world. To fill this gap, new empirical research is necessary. On the basis of the TEE framework, we are able to identify a number of important avenues for future research. The TEE framework is also extremely useful for conflict event data users because it provides important insights concerning the range of errors one has to take into account when using a specific conflict event data set, and how these errors may potentially affect research findings.</p> <p>In the following section, we develop the TEE framework and discuss in depth the measurement and representation errors that can arise during each step in the research process. Our discussion is supported by (necessarily eclectic) empirical examples drawn from literature. In the third section, based on the TEE framework, we introduce guidelines and strategies for event data collection and future research. The fourth section concludes.</p> <hd id="AN0156346852-3">The TEE Framework</hd> <p>The TEE framework is inspired by the well-known TSE framework used in survey research ([<reflink idref="bib29" id="ref37">29</reflink>]:41-63). In the TSE framework, measurement errors occur when the measured value deviates from the real value. This can arise from unclear question wording and answer scales, the presence of an interviewer, which inhibits the respondent from answering truthfully (i.e., social desirability bias), or the incorrect processing of data. Errors of representation occur when not all existing observations are sampled in the survey. An essential characteristic of a survey is the sampling of only a subsection of the population, implying that this form of error always occurs. What is important is that observations are sampled randomly. This randomness can be jeopardized by a flawed sampling frame, nonresponse, or data adjustments based on a flawed external source (e.g., an outdated census). Clearly, two forms of error can occur: bias, which causes a systematic deviation from the real value, and unreliability, which arises from random errors, making the results less precise.</p> <p>The collection and use of event data resembles the survey process in important ways. An important similarity is that sampling is inherent to the process. By selecting media sources to capture conflict events, one is aware that not all events that have taken place are necessarily reported. The challenge arises from the nonrandom processes steering media event inclusion, a debate which can be related to concerns about nonresponse error in surveys. Like respondents, news sources and reports can present information in a biased way or they may simply not be able to provide the necessary information, leading to missing data. While the interviewer commonly plays a key role in the sampling of respondents for public opinion polls, the same is true for a coder in charge of sampling relevant events into a data set. Furthermore, unclear coding instructions or variable definitions and categories can lead to unreliable or biased data, as can unclear survey questions. For both types of data, researchers can attempt to validate data against an external source. This can be a census or medical records for surveys or police and nongovernmental organization (NGO) reports for event data. Lastly, in the analysis phase, researchers can choose to weight the data to compensate for nonresponse or biased selection or they can choose to impute missing data to preserve the number of cases in the analysis.</p> <p>Figure 1 visualizes the TEE framework. Central to the figure are the research steps taken in event data collection and analysis. These steps are not necessarily sequentially taken and may interact in important ways. For example, the development of the codebook is not necessarily finalized before the coding process, as a coding pilot test helps in refining the codebook. In case of automated coding, the coder has no role or at least a much more limited one. The development of the codebook (or dictionary in automated applications) becomes all the more important. In addition, comparisons to nonmedia sources are not often realized, simply because of a lack of such external data. In line with [<reflink idref="bib29" id="ref38">29</reflink>]:48) study, we associate each research step with both measurement and representation error. Several sources of error have been touched upon above, but they are addressed in greater depth in the following sections. We structure the discussion according to the research steps identified.</p> <p>Graph: Figure 1. Total event error.</p> <hd id="AN0156346852-4">News Source Sampling</hd> <p></p> <hd id="AN0156346852-5">News coverage</hd> <p>News source sampling can give rise to both measurement and representation error. We start the discussion with representation error caused by news coverage effects, as this relates to the relatively well-known problem of selection bias ([<reflink idref="bib21" id="ref39">21</reflink>]:68-72; [<reflink idref="bib42" id="ref40">42</reflink>]; [<reflink idref="bib56" id="ref41">56</reflink>]).[<reflink idref="bib5" id="ref42">5</reflink>] Nevertheless, while bias has been widely studied in the literature ([<reflink idref="bib21" id="ref43">21</reflink>]:68-72; [<reflink idref="bib42" id="ref44">42</reflink>]; [<reflink idref="bib56" id="ref45">56</reflink>]), coverage effects can also be associated with unreliability, much as is the case with sampling error.</p> <p>When deciding on collecting event data for specific types of conflict, time periods, and geographical settings, researchers first decide on the news source from which to extract data. This can be a newspaper, a news wire service, or even television and radio news. The choice of a news source implies that events included in the data set are dependent on media selection (or sampling) of events into the news. As a multitude of studies has shown, this selection is far from random. A dual problem is apparent: News source coverage can be determined by the characteristics of an event but also by the characteristics of the news source itself. The first is seen as a coverage effect common to different media outlets, but the second can be source-specific and underscores that the question from which news sources to extract data is an important one. We first discuss general, then source-specific selection effects.</p> <p>The seminal paper by [<reflink idref="bib27" id="ref46">27</reflink>] on the presentation of the Congo, Cuba, and Cyprus crises in Norwegian newspapers set the basis for news value theory in communications science, which investigates the characteristics of an event that are likely to make it newsworthy ([<reflink idref="bib32" id="ref47">32</reflink>]). Galtung and Ruge propose 12 news factors that determine whether a foreign crisis event will be reported, including the event's amplitude or importance and the involvement of elite actors. Following their work, other communications scientists have investigated the news values that determine selection into the news media and have increased or reduced the number of relevant factors (e.g., [<reflink idref="bib32" id="ref48">32</reflink>], [<reflink idref="bib33" id="ref49">33</reflink>]).</p> <p>Social movement scholars focus specifically on protest events ([<reflink idref="bib41" id="ref50">41</reflink>]; [<reflink idref="bib45" id="ref51">45</reflink>]). Summarizing the findings of previous studies, [<reflink idref="bib21" id="ref52">21</reflink>]:69), [<reflink idref="bib56" id="ref53">56</reflink>]:398-400), [<reflink idref="bib42" id="ref54">42</reflink>]:45-46), and [<reflink idref="bib41" id="ref55">41</reflink>]:350-51) note that large-scale protest events with many participants, events characterized by violence (property or physical damage, police repression, arrests, etc.), events organized by movements with professional (public relations) staff, and events involving high-profile actors are all more likely to be reported. These findings are supported by comparisons of event inclusion between media sources but also by comparisons of media reports with external sources such as police records.</p> <p>Representation error has been far less investigated with respect to conflict event data in the developing world. Recent interest in low-level conflicts, including protests in developing countries, can perhaps assume the same coverage preferences. For armed conflict events, we could assume that because of the level of violence, selection into the news is highly likely. However, the contexts in which armed conflicts arise are often different from the Western settings commonly investigated. Civil wars often erupt in rural areas away from the government's center of power (e.g., [<reflink idref="bib43" id="ref56">43</reflink>]:38-48), which has implications for the communications infrastructure present in the region. Furthermore, although armed conflict attracts journalistic attention, a climate of violence and infrastructure damage can obstruct event coverage. In this regard, a study by [<reflink idref="bib75" id="ref57">75</reflink>] is highly instructive. He compares a data set on armed conflict in Afghanistan collected by the U.S. military—and revealed by WikiLeaks—with the UCDP GED data set (which solely used media sources for this conflict) and finds that cell phone coverage significantly increases the likelihood of events being reported in the media, suggesting a systematic underrepresentation of events in remote rural areas.</p> <p>Important source-specific selection effects are related to the ideological or political orientation of a news source, as well as its geographical scope (e.g., [<reflink idref="bib17" id="ref58">17</reflink>]:107-26). For example, in the analysis of protest events, several studies find that conservative newspapers underreport violent demonstrations to limit copycat behavior (for an overview see [<reflink idref="bib56" id="ref59">56</reflink>]:401). The second factor, the geographical scope of the news source, relates to whether a local, national, or international target audience is reached. We devote more attention to this issue here, as many recent data sets make use of multisource inventories, such as Factiva, LexisNexis, or Keesing's Record of World Events,[<reflink idref="bib6" id="ref60">6</reflink>] which rely to an important extent on international news wire services to code events occurring in a wide range of developing countries, including violent armed conflict as well as protests.</p> <p>Local news sources can cover local conflicts more extensively than national sources, which implement an additional selection procedure. International sources have an even more stringent selection process. However, for some events, such as ongoing armed conflict, professional international news wire services could potentially be more valuable than (disrupted) local media services. Exactly how selection bias plays out when the scales of conflict and news source scope interact is a highly relevant and perhaps insufficiently addressed empirical question. Several studies do indicate its importance. [<reflink idref="bib37" id="ref61">37</reflink>], for example, find that international newspapers report substantially less protest events than national newspapers in Argentina, Mexico, and Paraguay and that these differences are related to, among others things, the use of violence but also to a general difference in international media attention toward these three countries.</p> <p>[<reflink idref="bib8" id="ref62">8</reflink>] developed a data set on political violence in Pakistan based on national newspapers and record a higher number of incidents than data sets relying on Factiva. [<reflink idref="bib20" id="ref63">20</reflink>] compared data sets based on international versus national news sources on conflict events in Nigeria. They find that international sources underrepresent conflict events, in particular protest events. Both studies also find that relative underreporting affects the subnational distribution of events, an increasingly important research line in conflict studies (e.g., [<reflink idref="bib30" id="ref64">30</reflink>]:303-05). Lastly, [<reflink idref="bib2" id="ref65">2</reflink>] used district-level news sources to capture violent events in Indonesia and show that these record more incidents than provincial newspapers and hence provide greater insights into local causes of conflict.</p> <p>In general, multisource inventories are argued to be more reliable than single sources ([<reflink idref="bib42" id="ref66">42</reflink>]:47-49), but it is important to keep in mind that multisource inventories do not include the "universe of media reports" ([<reflink idref="bib56" id="ref67">56</reflink>]:402). Especially for conflict in the developing world, it is important to consider that the international, English sources included in these inventories might not cover these settings sufficiently ([<reflink idref="bib64" id="ref68">64</reflink>]:552-53). Automated coding procedures in principle are not sensitive to representation error, but they do require machine readable text and predominantly draw on international news wire reporting services such as Reuters and LexisNexis (e.g., Integrated Data for Events Analysis data set, [<reflink idref="bib5" id="ref69">5</reflink>]; Global Data on Events, Location and Tone [GDELT] data set, [<reflink idref="bib48" id="ref70">48</reflink>]; Kansas Event Data System [KEDS] data set, [<reflink idref="bib63" id="ref71">63</reflink>]), which is an important characteristic to consider. The use of local newspapers to investigate conflict events in developing countries is likely to emerge as an important research line in conflict studies, not in the least because local newspapers are increasingly available online (e.g., AllAfrica repository).[<reflink idref="bib7" id="ref72">7</reflink>] However, the use of local sources brings with it new challenges, for example, related to media ownership and state control of media sources. So far, little systematic research appears to have been conducted to assess the impact of these issues on the quality of conflict event data.</p> <hd id="AN0156346852-6">News reporting</hd> <p>We now turn to the problem of measurement error arising from the news source, which concerns the information news sources report with regard to an event. This form of error can be linked to the concept of description bias ([<reflink idref="bib21" id="ref73">21</reflink>]:72-73). When discussing description bias problems, several scholars make a distinction between "hard news" and "soft news" and argue that the former is less subject to bias than the latter ([<reflink idref="bib21" id="ref74">21</reflink>]:72; [<reflink idref="bib26" id="ref75">26</reflink>]:7; [<reflink idref="bib58" id="ref76">58</reflink>]:656).[<reflink idref="bib8" id="ref77">8</reflink>] Hard news is suggested to include the "who, what, when, where, and why of the event" ([<reflink idref="bib21" id="ref78">21</reflink>]:72), whereas soft news is said to include interpretations of causes and consequences, portrayals of the actors, and so on. We first discuss research on soft news dimensions and then focus on hard news. As argued below, the distinction made in the literature between hard and soft news is, however, not straightforward. Furthermore, not all reporting inaccuracies are necessarily signs of bias but can also indicate unreliability due to challenges for media sources to acquire certain types of information.</p> <p>Soft news effects can be related to the concept of framing. Several definitions of framing exist; as an example, we cite [<reflink idref="bib24" id="ref79">24</reflink>]:52): "To frame is to select some aspects of a perceived reality and make them more salient in a communicating context, in such a way as to promote a particular problem definition, causal interpretation, moral evaluation, and/or treatment recommendation." A substantial amount of research has focused on the way in which media represent protest actions, albeit predominantly focused on Western settings (e.g., [<reflink idref="bib15" id="ref80">15</reflink>]; [<reflink idref="bib53" id="ref81">53</reflink>]). A major research line focuses on differences in framing according to the ideological profile (conservative or liberal) of the news source and whether conservative newspapers are more likely to depict protesters negatively (e.g., [<reflink idref="bib9" id="ref82">9</reflink>]; [<reflink idref="bib50" id="ref83">50</reflink>]; [<reflink idref="bib73" id="ref84">73</reflink>]).</p> <p>Although literature on the framing of conflict events provides interesting insights into the orientations of different news sources, it is less clear to what extent different representations of conflict can affect event data sets that focus on dates, locations, and actors. Nonetheless, the line between soft news and hard news is not necessarily clear-cut. A well-known example of a commonly contested "hard fact" is the number of participants at a protest, which can be exaggerated by activists or understated by police authorities ([<reflink idref="bib19" id="ref85">19</reflink>]:130). Furthermore, fatality estimates are often regarded as hard facts that are difficult to establish ([<reflink idref="bib58" id="ref86">58</reflink>]:656; [<reflink idref="bib71" id="ref87">71</reflink>]:527). The source—police versus protesters or government versus rebels—that is preferred by a sampled newspaper can bias event data statistics.</p> <p>Relatively few studies have compared external data with the reporting of hard facts in the media. [<reflink idref="bib51" id="ref88">51</reflink>]:117-26) compared data from police records with print (and electronic) media reports for protest events in Washington, DC. They found good correspondence for protest dates and purpose but weaker agreement concerning protest size. The latter could be related to bias or unreliability, however, as the analysis does not describe how the variables relate to each other. [<reflink idref="bib74" id="ref89">74</reflink>] investigated differences in the reporting of hard facts between the U.S. military data set on armed conflict events in Afghanistan and UCDP GED. He finds that for most events, the casualty numbers of the military data set fall within the low–high casualty estimate of the UCDP GED data set. There are more events for which UCDP GED gives a higher estimate though, which could indicate a slight bias toward reporting higher casualty numbers in news reports. There are also differences between UCDP GED and the military data set in the reported location of an event. Based on his analyses, [<reflink idref="bib74" id="ref90">74</reflink>]:1143) argues that researchers should not use data for analyses below a range of 50 km. His research indicates that even a hard fact such as "location" is also not always reported reliably, in particular when considering armed conflict events. Although fatalities are considered difficult to establish reliably, Weidmann's research suggests that their reporting appears relatively free of error.</p> <p>News coverage and reporting relate to some of the best-known errors described in the literature. Nonetheless, many studies focus on Western contexts and protest events and only to a lesser extent on developing contexts and events of violent (armed) conflict. While important principles and lessons can be drawn from social movement and communications studies, there is a clear need for more empirical research on these forms of errors in conflict studies. In the following sections, we turn to errors that are arguably less widely discussed in current scholarship. These errors do not necessarily arise from the workings of media sources but are more related to data collection procedures.</p> <hd id="AN0156346852-7">News Report Sampling</hd> <p></p> <hd id="AN0156346852-8">Issue and page sampling</hd> <p>While some researchers draw on all reports available from a specific source, others rely on the additional sampling of specific newspaper issues or pages ([<reflink idref="bib21" id="ref91">21</reflink>]:68; [<reflink idref="bib47" id="ref92">47</reflink>]:112-25). When it comes to recent conflict data sets (see Introduction), this additional sampling stage is not included, as they commonly rely on reports drawn from multisource inventories, using key terms and date and country specifications. For studies relying on national or local newspapers, especially when a relatively extensive period is being studied, this additional sampling may be necessary to reduce coding costs. For example, for their seminal study on protest events in four Western-European countries, [<reflink idref="bib46" id="ref93">46</reflink>]:253-63) used one national newspaper per country but only the Monday edition. They covered the period from 1975 to 1989. Even if no systematic biases are associated with specific newspaper editions, this additional sampling engenders further unreliability. It is also possible to select only the first page of a newspaper issue, which could for instance reinforce bias toward the inclusion of high-profile events characterized by violence.</p> <hd id="AN0156346852-9">Report content</hd> <p>Journalistic or editorial preferences can also be an important source of error at the level of the news report. For example, [<reflink idref="bib10" id="ref94">10</reflink>]:390-92) draw attention to the fact that news reports often quote sources that have incentives to provide biased information. They use the example of a report in which a rebel leader claimed to have killed 30 government soldiers. In their data set (Event Data on Conflict and Security [EDACS]), they created an additional variable, indicating that the information might be biased if doubtful sources are used.</p> <p>In addition to the biases that can arise from reporting preferences, news reports themselves can be important sources of unreliability. First, reports on events can be detailed or vague. Some reports might provide information on the size of a group of protesters, whereas another report on the same event might only mention that the protest occurred. Similarly, the capture of territory by a rebel group can be reported but not necessarily whether there were any casualties. In some cases, multiple reports can provide valuable additional information, yet for others, vague reports might be the sole source of information and a substantial degree of missing data can result. The newsworthiness of an event can also affect the depth of reporting and the length of the news article devoted to it. For example, a large-scale protest can attract more news attention than an event of limited size, and hence, more information on the event might also be reported. Nonetheless, while some events can gain strong news attention, such as grave human rights abuses in armed conflict, the "fog of war" can also prevent the collection of reliable information.</p> <p>Second, reports can also explicitly cast doubts on whether and how an event occurred, on the identity of the actors, or on the validity of a casualty estimate. Reports can, for example, state that the identity of attackers or suspected rebels is uncertain. These forms of measurement error can only be captured if such indicator variables are included in the codebook.</p> <p>Third, coding challenges can arise from conflicting reports. While the incompleteness of news reports leads many researchers to draw from multiple reports to construct event data variables, this can also raise additional questions concerning the way in which reports are combined ([<reflink idref="bib76" id="ref95">76</reflink>]:125-26). A crucial problem arises when information is inconsistent. Some data sets provide instructions to coders to aggregate the information in particular ways. For example, SCAD states that in the case of multiple casualty estimates, the mean is taken (Codebook version 3.1.), whereas ACLED states that the lowest number should be used ([<reflink idref="bib57" id="ref96">57</reflink>]:20). Other solutions to conflicting reports suggest coding each report individually. Based on their work on protest events for the Nonviolent and Violent Campaigns and Outcomes data set, [<reflink idref="bib19" id="ref97">19</reflink>]:130-31) recommend the coding of different reports, together with including a metric ambiguity range variable in the final event data set.</p> <p>Similarly, [<reflink idref="bib76" id="ref98">76</reflink>] propose the creation of an intermediate data set, which includes the event coding by news report, and an event data set, which aggregates the information across reports. As all reporting information is provided, aggregation rules (mean, minimum, etc.) can be altered. Coding news reports separately can increase transparency and replicability, as opposed to allowing coders to aggregate news reports themselves. This coding choice can also have important implications for the monitoring of the coding process and intercoder reliability scores (see below). Coding news reports separately can, however, increase research costs.</p> <hd id="AN0156346852-10">Codebook Development</hd> <p></p> <hd id="AN0156346852-11">Sampling instructions</hd> <p>Codebook instructions are crucial to avoid coder confusion and to support consistent sampling as well as coding of relevant events. When using machine coding, the dictionary and coding program determine selection and coding of cases into the data set based on the identification of relevant actors, wordings, and so on, rather than a coder.[<reflink idref="bib9" id="ref99">9</reflink>] Generally, codebooks and dictionaries are revised after an initial coding test phase, in which potential sources of error are revealed. Sampling instructions are an important concern: Which events should be included in the data set and which should be excluded?</p> <p>When developing instructions for human coders, researchers can either adopt a definition or provide a list of eligible events (e.g., [<reflink idref="bib46" id="ref100">46</reflink>]:263-69). Many conflict event data sets mainly rely on event definitions, but a potential caveat is that the stricter the definition, the more difficult it becomes to consistently code vague reports of events. Reports do not always give details on actors, which actor used violence, or the number of participants at an event, for example, which can create confusion and sampling inconsistencies when categorization requires this information. It can also be important to include instructions on how to handle cases for which a report casts doubts on its occurrence or eligibility for inclusion.</p> <p>When sampling events from online repositories, the same concerns apply. In databases such as LexisNexis, one can develop a search string of relevant key words and apply these to extract news stories about a specific topic or event. Afterward, a subsample can be manually verified by coders to select the usability and efficiency of the search string and the amount of "noise." Nevertheless, a coder's decision to include or exclude events still requires consistency and replicability and consideration of the aforementioned issues. A news report that includes relevant key words such as "violence" might report more than one event, for example, all of which need to be sampled consistently. Furthermore, the use of search strings does not assure that all events sampled from a news source are also sampled by using specific terms. Although search strings often include many key terms, some events can still be overlooked.</p> <p>The issues of noise and the overlooking of events are also a major concern when using automated coding procedures. It is useful, however, to first point out the benefits of machine coding. While the development of dictionaries is time-consuming, including as many key verbs and phrases, variations, names of actors (e.g., United States, US, USA, President Trump) as possible, once developed, they offer the potential to go through large volumes of data in seconds ([<reflink idref="bib6" id="ref101">6</reflink>]; [<reflink idref="bib65" id="ref102">65</reflink>]). Further, a revision of the dictionary does not result in a time-consuming recoding process. Instead, the program can just rerun on the same data with the revised dictionary. Finally, dictionaries can be shared between researchers and be used for new projects. A major point of discussion is however whether machine coding is able to identify the "right" events and whether these events are coded correctly (see below), with human coding often taken as the standard.</p> <p>The ability of machine coding procedures to include a sufficient high number of relevant events ("recall"), while at the same excluding irrelevant events ("precision")—events related to sports competitions are common false positives—is an important sampling challenge.[<reflink idref="bib10" id="ref103">10</reflink>] Several researchers have empirically investigated recall and/or precision for machine coding applications compared to a training set developed by human coders. For instance, [<reflink idref="bib6" id="ref104">6</reflink>] find that the original KEDS's sparse parsing program[<reflink idref="bib11" id="ref105">11</reflink>] performs as least as well as (new) human coders in identifying relevant events (around 80 percent). [<reflink idref="bib44" id="ref106">44</reflink>] test the VRA reader and find that it performs as well as human coders for recall (93 percent correct) but less for precision (23 percent correct). Overall, they are positive about the potential of machine coding, however.</p> <p>Besides comparisons with human coders, there have also been comparisons between programs which are continually developing. [<reflink idref="bib7" id="ref107">7</reflink>] compare the TABARI program developed by Schrodt as a follow-up to the original KEDS program and find that with regard to recall and precision, its sparse parsing procedure is significantly outperformed by the BBN SERIF program that relies on natural language processing.[<reflink idref="bib12" id="ref108">12</reflink>] Most recently, [<reflink idref="bib14" id="ref109">14</reflink>] developed a machine learning classifier system that shows recall and precision percentages of around 90 and 50, respectively, again as compared to human coders. [<reflink idref="bib34" id="ref110">34</reflink>] propose a joint human/machine process for the selection of relevant text by supervised machine learning to improve recall and precision. Besides natural language processing and machine learning, another area of progress in automated coding lies with conditional random fields ([<reflink idref="bib48" id="ref111">48</reflink>]:38; [<reflink idref="bib70" id="ref112">70</reflink>]).</p> <p>It appears that automated coding has important and increasing benefits for event sampling. There continue to be a number of challenges to consider, however. The first crucial challenge concerns duplication or the inclusion of the same event into the data set multiple times ([<reflink idref="bib5" id="ref113">5</reflink>]:737-38; [<reflink idref="bib48" id="ref114">48</reflink>]). There is no real automatic procedure yet to filter out duplicates, except to discard events with the same time, location, actors, and so on. Human review of the data set can be required to exclude further duplicates and can still be a costly exercise when considering large volumes of data. Another challenge concerns language, as most dictionaries and applications predominately focus on the English language ([<reflink idref="bib48" id="ref115">48</reflink>]:45), while extensions to other languages can lead to the inclusion of more diverse and non-Western sources. Nonetheless, the use of English is also not uniform, and specific word choices and sentence structures can also vary across regions or countries and can be more pronounced for domestic than international events ([<reflink idref="bib66" id="ref116">66</reflink>]). Even the news source itself can vary in language use ([<reflink idref="bib7" id="ref117">7</reflink>]).</p> <p>Automated coding has primarily been used for the collection of political event data in the field of international relations (e.g., [<reflink idref="bib63" id="ref118">63</reflink>]). Increasingly, the use of automated coding is also used to investigate domestic conflicts, including in developing contexts (e.g., [<reflink idref="bib48" id="ref119">48</reflink>]). This implies that the challenges with regard to dictionary construction described above are becoming increasingly pertinent to deal with, both when concerning the selection of events and the coding of events, as will be discussed below.</p> <hd id="AN0156346852-12">Coding instructions</hd> <p>Unclear coding instructions can create representation errors as well as measurement errors. Again, the problem of defining events arises. For example, the USDAA codebook includes 12 event definitions, but it is argued that these conflict types "are by no means mutually exclusive categories. [...] While we have tried to be consistent in the coding of such events, one should be careful in treating the categories as clearly distinguishable phenomena" ([<reflink idref="bib72" id="ref120">72</reflink>]:11). This problem stems from missing or conflicting information in event reports. In some data sets, for example, the mentioning of an association behind the protest can make the difference between categorization as a spontaneous or as an organized protest (e.g., SCAD).[<reflink idref="bib13" id="ref121">13</reflink>] Yet, this can also be influenced by the depth of reporting.</p> <p>When developing the codebook, researchers potentially have to choose between very generic categories of events, actors, and so on, which can be coded reliably, or very specific categories, for which coding is more unreliable. This is an important trade-off to be made. While broad or generic categories might create more consistency, they might not provide the level of information precision that researchers strive for. A generic actor category such as "attackers" might be coded very reliably, for example, but one would also want to know, where possible, whether the attackers were particular rebel groups or ethnic militias, political parties, and so on. Unfortunately, the need for detailed event information to pursue particular research questions is not always accommodated by the information provided in media sources.</p> <p>For automated coding, the complexity of event coding is not only challenged by the information available in news reports but also by the dictionary and the nuances predefined sentence structures can capture. [<reflink idref="bib6" id="ref122">6</reflink>] also analyzed event categorization besides event sampling and find again that machine coding performs similar to human coding. [<reflink idref="bib44" id="ref123">44</reflink>] have similar findings but also show that more general event classifications are coded more reliably than detailed ones. The fact that detailed event definitions are not always workable is also discussed by [<reflink idref="bib48" id="ref124">48</reflink>]:33).</p> <p>In general, automated coding is deemed to work better when the variables that need to be extracted are not too complex. One challenge here is that the field of peace and conflict studies is increasingly moving toward more complex event definitions and characteristics, as well as detailed collection of time and location information. As discussed above, subnational location information for events is increasingly sought after in empirical research, yet automated coding is argued to work better on the country level ([<reflink idref="bib5" id="ref125">5</reflink>]:739; [<reflink idref="bib48" id="ref126">48</reflink>]:46). [<reflink idref="bib31" id="ref127">31</reflink>], for instance, argue that the GDELT data set should be used with caution for subnational analyses as it differs substantially from human coding and seems to show a bias toward country capitals. [<reflink idref="bib38" id="ref128">38</reflink>] are more optimistic when comparing spatial information for human and machine coded data in the framework of the EDACS data set, yet concerns are still raised.</p> <p>When human coding is used, the development of the codebook is an important start, yet how it is implemented is to a large extent the responsibility of the coders. Machine coding rules out coders or gives them a more limited (supervising) role (e.g., [<reflink idref="bib34" id="ref129">34</reflink>]). In the following section, we will focus on errors arising from the coder in a typical human coding project. Interestingly, even though machine coding is commonly compared to a human coding benchmark, human coding itself is also subjected to substantial errors. This is indeed the core argument of [<reflink idref="bib6" id="ref130">6</reflink>]:555) who early on lamented the poor quality of human coding.</p> <hd id="AN0156346852-13">Coding Process</hd> <p></p> <hd id="AN0156346852-14">Coder sampling</hd> <p>Following codebook instructions, coders sample events into the data set and extract information on key variables. Thus, the coder can also be a source of representation error and measurement error, and both unreliability and bias can arise. When sampling, coders can overlook events completely at random due to, for example, inattentiveness. Bias occurs when coders routinely overlook certain events or regularly misinterpret instructions on what constitutes a relevant event. Unfortunately, it is likely that smaller, low-scale events more often go unnoticed than high-profile events announced in headlines (e.g., [<reflink idref="bib46" id="ref131">46</reflink>]:270), which is why coder sampling error can potentially reinforce selection bias. Coder sampling error is not often measured (or reported), but some researchers have attempted to quantify it. In their work on social movements in four Western-European countries, [<reflink idref="bib46" id="ref132">46</reflink>]:270) report that in paired comparisons, around 60 percent of protest events were registered by both coders. A follow-up project reached about 70 percent identification agreement between coders ([<reflink idref="bib41" id="ref133">41</reflink>]:355). Although they used a different data source than news reports—reports from the United Nations Secretary General on peacekeeping operations—[<reflink idref="bib60" id="ref134">60</reflink>]:348-51) also note severe coder sampling error. They find that independent coders only double-identified 18–41 percent of relevant events. This necessitated the research team switching strategies and having the team leaders identify and highlight relevant events, which were then coded by the assistants.</p> <hd id="AN0156346852-15">Coder reliability</hd> <p>For research that makes use of media content analysis, the calculation of intercoder reliability to indicate measurement error is regarded as a methodological imperative in communications science ([<reflink idref="bib47" id="ref135">47</reflink>]:272-73).[<reflink idref="bib14" id="ref136">14</reflink>] This imperative has also made its way into protest event analyses in (Western) social movement studies ([<reflink idref="bib41" id="ref137">41</reflink>]:354-55). However, many conflict event data sets focusing on the developing world do not report such measurements ([<reflink idref="bib60" id="ref138">60</reflink>]:356-59; [<reflink idref="bib61" id="ref139">61</reflink>]:107-08). By conducting intercoder reliability tests, however, one can check whether the same measurement instrument (the codebook) leads independent coders to reach similar results ([<reflink idref="bib47" id="ref140">47</reflink>]:273-75). Common measures are Krippendorff's α and Cohen's κ, which both correct for chance agreement by weighing inconsistency in less frequent response categories more heavily in the final coefficient. Intercoder reliability checks can be used to refine the codebook or select the "better" coders after a pilot stage. It is recommended to conduct tests regularly throughout the coding process as only conducting postdata collection tests can reveal the need to discard or recode a substantial amount of data. The tests can be conducted on a small subset of the data (5–10 percent).</p> <p>When interpreting intercoder reliability, it is also important to be aware that low intercoder reliability can arise if each coder makes random errors (coder unreliability) or if each coder routinely interprets rules in a different way (coder bias). However, if all the coders routinely misinterpret a coding rule, this bias will not be captured by the reliability statistic. The calculation of intercoder reliability statistics can be particularly important for research into causal interpretations or framing in media reports. Nevertheless, it is not necessarily safe to assume that hard facts are coded relatively free from errors (e.g., [<reflink idref="bib22" id="ref141">22</reflink>]:130-35).</p> <p>Lastly, it is worth noting that intercoder reliability is generally calculated at the level of the news report in communications studies. Indeed, this level allows for the closest monitoring of coder work. However, when coders are instructed to aggregate event reports and information, this monitoring process can become more complicated. Key challenges can arise when attempting to retrace coder decisions: For example, did coders notice all reports of an event, are all reports indeed about the same event, have all reports been processed consistently, and so on. Hence, aggregation by coders, without the separate coding of news reports, can make it difficult to establish the source of low intercoder agreement in event inclusion and coding.</p> <hd id="AN0156346852-16">Nonmedia Data Comparison</hd> <p>To investigate errors arising from media preferences, several researchers have compared event data with nonmedia data sources. Although such data and comparisons are rare, they can give important indications of media errors. However, the external data themselves may have significant (and unknown) errors, which can jeopardize the validity of findings from media comparisons.</p> <p>Police records are most frequently used to investigate the media coverage of protest events in Western contexts. [<reflink idref="bib42" id="ref142">42</reflink>]:44) note that studies generally find a single newspaper covers no more (and often less) than 20–40 percent of events identified in police records. While many studies have used police records to investigate coverage error (confirming the selection effects discussed in News Coverage section), we noted that they have also been used to study reporting error. We refer in particular to the study of [<reflink idref="bib51" id="ref143">51</reflink>]:117-26) with regard to "hard facts" about demonstrations in Washington, DC (see News Reporting section).</p> <p>Caution is nevertheless needed to avoid overly relying on the quality of police records, as they are not necessarily collected systematically and can lack important details of events ([<reflink idref="bib55" id="ref144">55</reflink>]:48). For events in developing contexts—the geographical focus of many conflict event data sets—police records might be subject to more serious errors than in Western contexts as well as having access problems (e.g., [<reflink idref="bib4" id="ref145">4</reflink>]:332).</p> <p>For studies focusing on armed conflict or violence against civilians, NGO reports are another external source and are commonly used for the construction of conflict event data sets (often in addition to media data). As [<reflink idref="bib18" id="ref146">18</reflink>] show for state violence in Guatemala, NGO reports document more state violations and different trends in state violence over time than newspaper accounts, although whether this is due to measurement or representation error cannot be established. Interestingly, interview data show yet another picture. Further, although the purpose of many NGOs in the field is to provide independent, reliable information, reporting can be dependent on donor attention to "hot topic" events or deliberately created to draw international media attention. In turn, NGO reports often rely on media reports. Hence, NGO reports could potentially reinforce media bias toward particular countries or conflicts in event data sets. While a military data set can reveal important insights into the coverage of armed conflict ([<reflink idref="bib74" id="ref147">74</reflink>], [<reflink idref="bib75" id="ref148">75</reflink>]), it can also serve particular organizational goals and does not necessarily provide a true reflection of reality.</p> <hd id="AN0156346852-17">Data Adjustments</hd> <p></p> <hd id="AN0156346852-18">Data weighting</hd> <p>The last step in the event research process is the analysis stage, during which researchers can apply corrections to the event data set to compensate for sampling or measurement error. A first type of correction involves weighting the data to correct for underrepresentation of specific events. Although the intention is to reduce error, this type of correction can also create it. Indeed, there is often no external data that match the media-based event data set. Corrections are then made based on different studies, and these findings are assumed to hold over space and time. [<reflink idref="bib40" id="ref149">40</reflink>], for example, propose statistical corrections for selection bias (e.g., weighting) based on a comparative study of police records and local newspapers from four Swiss cities. They argue that corrections for coverage preferences of violent events and events with more participants might also be useful in other contexts. [<reflink idref="bib56" id="ref150">56</reflink>]:408-11) argue to the contrary that this can be a bias-increasing procedure if the relevance of selection factors as well as their magnitude does not translate to other contexts.</p> <p>Other recently proposed corrections do not rely on comparisons with external sources but rather with other media data. [<reflink idref="bib35" id="ref151">35</reflink>] use a mark and recapture method to estimate the true number of events based on information from multiple media sources. SCAD draws on Associated Press (AP) and Agence France-Presse (AFP) reports. The coding scheme, starting from 2012, records whether an event was reported in AFP, AP, or both sources. By estimating the correspondence between the sources, it is possible to make corrections to the data for events not covered in both data sets. A similar approach is used proposed by [<reflink idref="bib12" id="ref152">12</reflink>]. Importantly, the method requires data sets to consistently report all sources that have reported on an event, which is not common practice. Indeed, while data sets often cite a particular source, this does not imply that the event was not included in other sources. SCAD is a notable exception. However, it does rely on the same types of media sources, international news wire services, while local newspapers could capture a substantial number of additional events (see Data Coverage section). To correct for differential attention toward particular countries by international news media (e.g., [<reflink idref="bib37" id="ref153">37</reflink>]), it has also been proposed to include a variable for the total number of nonconflict related news reports devoted to a particular country in a given year as a control variable in substantive analyses ([<reflink idref="bib36" id="ref154">36</reflink>]:1664-65).</p> <hd id="AN0156346852-19">Missing data imputation</hd> <p>A second type of correction that can be made to conflict event data is the imputation of missing data to compensate for measurement error. These adjustments can in turn lead to erroneous statistics. Missing data corrections are often performed for dates, geolocations, and fatality estimates. For example, UCDP GED gives a date and a time to each event, but for some events, uncertainty arises about the precision of these variables ([<reflink idref="bib13" id="ref155">13</reflink>]:5-6). Sometimes only the week, month, or year of an event is known. In these cases, UCDP GED accords the earliest possible date to the event. It is also common to give the geographical coordinates of the center of the administrative unit or country when exact locations are unknown. Imputation of time and location data is often accompanied by variables indicating a level of uncertainty in the coding. Similar approaches are taken by ACLED ([<reflink idref="bib57" id="ref156">57</reflink>]). Importantly, it is not clear to what extent precision indicators are actually used in empirical applications of conflict event data, for example, by excluding uncertain events as a robustness check.</p> <p>Lastly, imputations for fatality estimates also exist. One example is the splitting of the casualty count when an event occurred at multiple locations or over the course of multiple dates (e.g., ACLED but not SCAD). Another example relates to words being used to describe casualty numbers, which is relatively common (e.g., several, some, dozens). [<reflink idref="bib10" id="ref157">10</reflink>]:391-92) choose to write the word down in the data set but not to quantify it. SCAD (Codebook 3.1., updated November 20, 2017:5) makes use of a distinction for missing but more ("probably large") or less ("probably small") than 10. ACLED ([<reflink idref="bib57" id="ref158">57</reflink>]) chooses to quantify the description: Several, many, plural, or unknown is set to 10, dozens is set to 12, hundreds is set to 100. Such quantification could potentially risk jeopardizing data quality.</p> <hd id="AN0156346852-20">Event Data: A Way Forward</hd> <p>The TEE framework outlined in the previous section has allowed for a comprehensive discussion of the sources of error that can affect the quality of conflict event data, cutting across subdisciplines of specialization. We have also discussed potential strategies to mitigate these errors proposed in the literature as well as their limits. Table 1 offers an overview of these error sources, available estimates of their size, and mitigation strategies. The estimates of the degree of error are based on the studies reviewed here and hence on different geographical contexts, time periods, (automated) coding procedures, and so on. Moreover, the estimates show the extent to which information can diverge but not necessarily how this impacts substantive research results. While this should be taken into account, they do provide researchers indications on how to assess data quality. Finally, the available (and unavailable) estimates also indicate where more empirical research is needed. In this section, we focus mostly on the methodological questions which have so far been insufficiently addressed in the literature. The last column of Table 1 contains an extensive list of questions which require further research and which together constitute a research agenda concerning the methodology of creating and using conflict event data sets.</p> <p>Graph</p> <p>Table 1. Total Event Error: Errors, Guidelines, and Future Research.</p> <p> <ephtml> <table><thead><tr><th>Research Step</th><th>Error Risks</th><th>Available Estimates<sup>a</sup></th><th>Solutions</th><th>Future research Directions</th></tr></thead><tbody><tr><td>News source sampling</td><td><p>Representation and Measurement:</p><p>News coverage and reporting are dependent on:</p><list list-type="Bullet"><list-item><p>– Characteristics of event: more probable selection of violent events, events with a higher number of participants/information of protests easier to establish than armed conflict events (e.g., location)</p></list-item><list-item><p>– Characteristics news source: ideology, geographical scale of target audience (local, national, international), source preference (e.g., government sources)</p></list-item><list-item><p>– Characteristics contexts: for example, poor infrastructure for access to (or verification of) information, government control on information</p></list-item></list></td><td>Representation:<list list-type="Bullet"><list-item><p>– National news reports versus police recordsb: 20–60 percent of protest events</p></list-item><list-item><p>– National news reports record around 10 times less lethal violence than nongovernmental organization (NGO) documentary sources and around 2 times less than interviewsc</p></list-item><list-item><p>– International news reports versus military datad: 28.5 percent of lethal events</p></list-item><list-item><p>– International versus national news reportse: 1.5–5 percent of protest/riots; around 25 percent of lethal events; around 0.7 times the number of terrorist events</p></list-item><list-item><p>– National versus provincial news reportsf: around 0.3 times the number of deaths</p></list-item></list><p>Measurement:</p><list list-type="Bullet"><list-item><p>– Protest report correspondence with police recordsg: >.98 (date), >.65 (purpose), >.61 (size)</p></list-item><list-item><p>– Armed conflict report correspondence with military datah: - 80 percent within 50 km of real location; 50 percent correspondence casualties, small differences</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Draw data from multiple news sources, including media sources (with different political orientations) as well as external sources (e.g., NGO reports)</p></list-item><list-item><p>– Use national or local news sources for single-country studies. International sources can be necessary for large-scale cross-national studies but hold representational risks</p></list-item><list-item><p>– Adapt the scale of conflict events studied to the scale of the news source. Local sources can be better suited to study low-level events (e.g., protests versus terrorist attacks) and to conduct subnational analyses</p></list-item><list-item><p>– Apply corrective weights to compensate for coverage effects based on comparisons between media sources or with external data (e.g., police records)</p></list-item><list-item><p>– Control for the amount of non-conflict-related reports to account for differences in media attention toward countries or regions</p></list-item><list-item><p>– Explicitly code differences between sources and take these into account for substantive analyses to account for reporting error</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Which conflict events are more likely to be selected into the news? How do biases differ according to the type of event (e.g., protest versus violent conflict)?</p></list-item><list-item><p>– How do coverage and reporting differ between local, national, and international news sources?</p></list-item><list-item><p>– How can we use automated coding procedures on local news sources taking into account access to digital data, language differences etc.?</p></list-item><list-item><p>– How does political orientation of the news source affect event coverage and reporting?</p></list-item><list-item><p>– How does press freedom affect event coverage and reporting?</p></list-item></list></td></tr><tr><td>News report sampling</td><td>Representation:<list list-type="Bullet"><list-item><p>– Issue sampling (e.g., Monday issues) and page sampling (e.g., first page) increase unreliability and can reinforce coverage bias</p></list-item></list><p>Measurement:</p><list list-type="Bullet"><list-item><p>– Reports can rely on biased sources (e.g., government, rebels, political parties)</p></list-item><list-item><p>– Reports have missing information on date, location, actors, casualties</p></list-item><list-item><p>– Reports explicitly cast doubt on event occurrence, actor identity, fatality estimates</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– 11–18 percent of events without precise time, 23–49 percent without precise locationi</p></list-item><list-item><p>– Uncertainty indicators make little differencej</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Code all issues and pages for a (random) subset of the data and compare with reduced sample. Potentially apply corrective weights</p></list-item><list-item><p>– Include indicators for unreliability and bias in the codebook for event occurrence, actor identities, fatality estimates etc.</p></list-item><list-item><p>– Code reports on the same event separately and include an ambiguity range in the final data set or allow users to test different specifications (e.g., minimum and maximum fatalities)</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– What are the effects of different report sampling mechanisms on event statistics (e.g., front page sampling, full issue sampling, search string in online repository, automated procedure)?</p></list-item><list-item><p>– How does the inclusion of unreliability and bias indicators affect substantive findings in conflict analyses?</p></list-item></list></td></tr><tr><td>Codebook development</td><td><p>Representation/Measurement</p><list list-type="Bullet"><list-item><p>– Unclear and/or ambiguous selection/coding instructions can create inconsistencies in the coding process and unsystematic as well as systematic differences between coders</p></list-item><list-item><p>– Dictionaries for automated coding can leave too many relevant events out and too many irrelevant ones in and code variables incorrectly</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Sampling of a relevant eventk: 80–97 percent recall, 23–58 percent precision</p></list-item><list-item><p>– Correct categorization of eventl: 7–96 percent</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Calculate intercoder selection/reliability measures between human coders in a pilot phase and throughout the coding process to adapt instructions where needed</p></list-item><list-item><p>– Refine the dictionary for automated procedures after data checks and rerun the coding program</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– How do codebook instructions affect intercoder selection/coding reliability?</p></list-item><list-item><p>– Which types of instructions work best given the reality of the data (e.g., vague reports)?</p></list-item><list-item><p>– What level of complexity can we reach with human and automated coding?</p></list-item><list-item><p>– What are the possibilities to improve and extend automated coding to other languages, contexts etc.?</p></list-item></list></td></tr><tr><td>Coding process</td><td><p>Representation:</p><list list-type="Bullet"><list-item><p>– Coders miss relevant events or include irrelevant events</p></list-item><list-item><p>– Causes unreliability but also potential bias when coders tend to miss specific events more often (e.g., low-profile protests)</p></list-item></list><p>Measurement:</p><list list-type="Bullet"><list-item><p>– Coders wrongly categorize events or actors, make mistakes in time and location data etc.</p></list-item><list-item><p>– Can cause unreliability or bias when coders systematically misinterpret instructions</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– 60–70 percent overlap between independent coders in the selection of protest eventsm</p></list-item></list><list list-type="Bullet"><list-item><p>– Cohen's κs show substantial agreement for event category, actors, fatalities, causen</p></list-item><list-item><p>– Krippendorff αs for event category, actors, fatalities <0.77o</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Monitoring of the coding process</p></list-item><list-item><p>– Calculate intercoder selection/reliability measures in a pilot phase and throughout the coding process</p></list-item><list-item><p>– Retain coders who spot the most relevant events/have higher reliability scores after an initial test phase</p></list-item><list-item><p>– Have a separately trained research team preselect relevant events to be coded by a different team</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– What are intercoder selection/reliability measures for currently well-known data sets?</p></list-item><list-item><p>– Are selection reliabilities higher when using search strings or semiautomated procedures?</p></list-item><list-item><p>– How does selection interact with the characteristics of the event (e.g., violent) and the news source (e.g., common usage of visualizations)?</p></list-item><list-item><p>– How reliable is the coding of hard facts (e.g., date, location, actor) versus soft facts (e.g., ascribed cause of the event)?</p></list-item><list-item><p>– What are the characteristics of "good coders" (e.g., education level, length of contract)?</p></list-item></list></td></tr><tr><td>Nonmedia data comparison</td><td><p>Representation and Measurement:</p><list list-type="Bullet"><list-item><p>– External data (e.g., police records, military data, NGO reports) can also suffer from bias and unreliability</p></list-item><list-item><p>– When used to correct media-based data, errors in external data can create additional uncertainty or reinforce biases</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– NGO reports contain about 4 times more deaths than interview datap</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Compare different nonmedia sources (e.g., NGO reports, interviews, surveys) to investigate source-specific coverage and reporting effects</p></list-item><list-item><p>– Drawn on media as well as external sources, code reports separately and include indicators of potential uncertainty and bias in the final data set</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– What choices do writers of NGO reports make and how can this affect event data (e.g., are they likely to exaggerate atrocities to advocate for more support)?</p></list-item><list-item><p>– To what extent can we rely on police records/military data in developing contexts?</p></list-item></list></td></tr><tr><td>Data adjustments</td><td><p>Representation:</p><list list-type="Bullet"><list-item><p>– Corrective weights to compensate for selection effects can induce bias when selection bias is not constant over time and space.</p></list-item></list><p>Measurement:</p><list list-type="Bullet"><list-item><p>– Imputation of missing data on dates, locations, fatalities, etc., can lead to a false sense of reliability</p></list-item><list-item><p>– Time and data imputation can affect in particular analyses requiring fine-grained time and location data</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Bias in significance and direction of regression coefficientsq</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– Create weights in the data set, ensure transparency on their creation, and allow users to incorporate weights or not</p></list-item><list-item><p>– Include indicators for imputation of variables</p></list-item></list></td><td><list list-type="Bullet"><list-item><p>– How do corrective measures affect substantive findings?</p></list-item><list-item><p>– How does selection bias differ over time, space, and news source?</p></list-item><list-item><p>– How does data imputation affect substantive findings?</p></list-item></list></td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>: ACLED = Armed Conflict Location and Event Dataset.</p> <p>2 <sups>a</sups>Percentages are used when the events/fatalities were matched, otherwise we calculate the difference in number of events/fatalities registered (X times less or more). <sups>b</sups>[<reflink idref="bib42" id="ref159">42</reflink>]; [<reflink idref="bib54" id="ref160">54</reflink>]. <sups>c</sups>[<reflink idref="bib18" id="ref161">18</reflink>]. <sups>d</sups>[<reflink idref="bib75" id="ref162">75</reflink>]. <sups>e</sups>[<reflink idref="bib20" id="ref163">20</reflink>], [<reflink idref="bib37" id="ref164">37</reflink>], Bueno de [<reflink idref="bib8" id="ref165">8</reflink>]. <sups>f</sups>[<reflink idref="bib2" id="ref166">2</reflink>]. <sups>g</sups>[<reflink idref="bib51" id="ref167">51</reflink>], print media estimates. <sups>h</sups>[<reflink idref="bib74" id="ref168">74</reflink>]. <sups>i</sups>ACLED and UCDP data, respectively. <sups>j</sups>[<reflink idref="bib67" id="ref169">67</reflink>]. <sups>k</sups>[<reflink idref="bib6" id="ref170">6</reflink>], Croicu and [<reflink idref="bib75" id="ref171">75</reflink>], [<reflink idref="bib44" id="ref172">44</reflink>]. <sups>l</sups>[<reflink idref="bib6" id="ref173">6</reflink>], [<reflink idref="bib7" id="ref174">7</reflink>], [<reflink idref="bib44" id="ref175">44</reflink>], [<reflink idref="bib70" id="ref176">70</reflink>]. <sups>m</sups>[<reflink idref="bib41" id="ref177">41</reflink>], [<reflink idref="bib46" id="ref178">46</reflink>]. <sups>n</sups>[<reflink idref="bib62" id="ref179">62</reflink>]. <sups>o</sups>[<reflink idref="bib20" id="ref180">20</reflink>]. <sups>p</sups>[<reflink idref="bib18" id="ref181">18</reflink>]. <sups>q</sups>[<reflink idref="bib56" id="ref182">56</reflink>]</p> <p>Most attention in the literature has been directed to coverage error and for important reasons. Indeed, the available estimates on event selection reveal that the distorting effects of coverage error on research findings may be substantial. While most evidence of such bias has been established in the context of protest movements in Western contexts (e.g., [<reflink idref="bib42" id="ref183">42</reflink>]), there is indication that violent conflict as well is underreported ([<reflink idref="bib18" id="ref184">18</reflink>]; [<reflink idref="bib75" id="ref185">75</reflink>]). Besides the form of conflict, another important challenge concerns the widespread use of international news wire reports to investigate (violent) conflict in the developing world. Evidence suggests that this may be problematic ([<reflink idref="bib8" id="ref186">8</reflink>]; [<reflink idref="bib20" id="ref187">20</reflink>]; [<reflink idref="bib37" id="ref188">37</reflink>]). This problem could be mitigated by the increased availability of online local sources and, potentially, new evolutions in automated coding. Interestingly, the available estimates on reporting error appear to indicate that the facts of protests ([<reflink idref="bib51" id="ref189">51</reflink>]) and violent conflict ([<reflink idref="bib74" id="ref190">74</reflink>]) may be reported relatively error free. Nevertheless, reporting error can also be dependent on the context. This is especially important to take into account when considering local media sources subjected to government control. No estimates appear to be available for these types of contexts, however.</p> <p>Errors arising from the logic of the media source require careful consideration and further research. Other features of the data collection protocol require attention too, however. This is also revealed by the estimates of errors related to the coding process. In this regard, it is important to point out that while different indicators can be used to assess particular methodological choices (selection agreement, intercoder reliability, recall, and precision), there is for now no real consensus in the literature concerning the use of such indicators and, consequently, their reporting. Many new data sets in peace and conflict studies for instance rarely provide information on coder selection and reliability or the general degree of imputed data in the data set. By contrast, automated coding developers appear to show more agreement on the need to report recall and precision rates.</p> <p>The measurement of such errors is important to establish where most data collection efforts should be directed in order to achieve the largest gains in terms of data quality. For example, [<reflink idref="bib67" id="ref191">67</reflink>]:29) argue that including indicators of uncertainty about events, actors, and so on (see Table 1, "codebook development"), did not add much value to PITF's atrocities data, and they left out these indicators in later versions. The same questions can be raised with regard to the coding of reports separately to account for differences between them ([<reflink idref="bib19" id="ref192">19</reflink>]; [<reflink idref="bib76" id="ref193">76</reflink>]). More research is needed in order to determine the merits of such procedures to be able to assess their use for new data sets.</p> <p>Users as well should direct sufficient attention to event error sources. This applies to the selection of particular data sets to address substantive research questions but also in reporting and robustness checks. Event error sources are necessary to understand the limits of particular studies, for instance, both in the academic and policy domains. Furthermore, when data developers provide indicators of data quality, we argue that researchers focusing on substantive questions should not only report these indicators in their studies but should also reflect upon the implications of these indicators for the validity of their findings and conclusions. This includes in particular missing data imputation indicators that are not commonly used in quantitative conflict studies even though the field is increasingly focusing on fine-grained details on events both in time and in place ([<reflink idref="bib30" id="ref194">30</reflink>]). Although event data weighting is still not commonly used, again the effect of weighting should be carefully compared with results based on nonweighted data, and a preference for some results over others should be explicitly motivated.</p> <hd id="AN0156346852-21">Conclusion</hd> <p>The quality of conflict event data can be affected by a wide range of errors. The discussion in this article was guided by the current state of the art concerning conflict event studies and also drew on social movement and communications studies. The major advantages of the TEE framework is that it offers a holistic perspective on the sources of error affecting conflict event data and, by consequence, analytical clarity into an arguably broad field of study. Indeed, while many error sources have been discussed in the literature, these debates have not always allowed for further systematization. By doing just this, the TEE framework offers a baseline tool for new and established data developers and users, as well as guidance for future research. Furthermore, while TEE has focused on human and automated event data collection practices, it can be extended into new areas. The emergence of "citizen reporting" via social media, for example, is becoming an important new source for event data collection but similar concerns with regard to coverage and reporting effects, as well as data collection procedures apply.</p> <p>Finally, it is worth noting that while errors can and should be minimized, they can hardly be ruled out completely. Hence, event data will never be a true reflection of reality. However, this is not unlike other empirical data sources in the social sciences, including public opinion surveys. Going back to our initial analogy, it is worth considering that the sources of error are widely recognized in survey research but also that the real exercise lies in minimizing errors by taking into account limited resources. As with "survey errors and survey costs" ([<reflink idref="bib28" id="ref195">28</reflink>]), the balance between event errors and costs constrains event data set developers. In order to improve guidelines and standards for data collection, however, conflict event data methodology needs to be considered as a research agenda in its own right. The TEE framework and the research questions laid out in Table 1 offer important directions to do so.</p> <hd id="AN0156346852-22">Notes</hd> <ref id="AN0156346852-23"> <title> Notes </title> <blist> <bibl id="bib1" idref="ref6" type="bt">1</bibl> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref65" type="bt">2</bibl> <bibtext> The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study received funding from the Research Foundation Flanders (FWO).</bibtext> </blist> <blist> <bibl id="bib3" idref="ref4" type="bt">3</bibl> <bibtext> Leila Demarest https://orcid.org/0000-0001-6887-9937</bibtext> </blist> <blist> <bibl id="bib4" idref="ref14" type="bt">4</bibl> <bibtext> Earlier versions of UCDP Georeferenced Event Dataset included only conflict events in Africa; a new global data set is currently available (version 5.0). Note that especially for violent conflict event data sets, developing countries are predominant, even if the data set has a global focus.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref42" type="bt">5</bibl> <bibtext> An extensive range of studies has been written on the topic of media selection bias alone. We provide an overview of key ideas here and direct readers to the references cited in this section, and Data Weighting section for corrections on selection bias, for further information.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref33" type="bt">6</bibl> <bibtext> All the conflict event data sets mentioned in the introduction predominantly use these inventories, except for Armed Conflict Location and Event Dataset, which also draws from local newspapers.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref72" type="bt">7</bibl> <bibtext> The use of local news sources has, however, been more frequently the case to investigate Hindu–Muslim riots in India (e.g., [77]).</bibtext> </blist> <blist> <bibl id="bib8" idref="ref62" type="bt">8</bibl> <bibtext> Note that in literature, hard news is also conceptualized as having a high degree of newsworthiness such as news regarding politics, economics, and social matters, whereas soft news has less substantive informational value, for example, gossip, human interest stories, and so on (e.g., [59]). This distinction is related to, but differs from, the one used in this article.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref82" type="bt">9</bibl> <bibtext> While the dictionary and the coding program are in principle separate entities in automated coding procedures ([48]:24), we do not explicitly separate the two in our discussion here. We also do not go into programming errors or characteristics (e.g., speed).</bibtext> </blist> <blist> <bibtext> Recall is equal to the number of true positives divided by the sum of the number of true positives and the number of false negatives. Precision is equal to the number of true positives divided by the sum of the number of true positives and the number of false positives. The F1 statistic captures the harmonic mean of recall and precision ([34]).</bibtext> </blist> <blist> <bibtext> The sparse parsing procedure breaks down sentences in relevant text based on actors, targets, and transient verbs.</bibtext> </blist> <blist> <bibtext> It is, however, important to mention that the use of TABARI in their work has been criticized by Schrodt ([48]:27).</bibtext> </blist> <blist> <bibtext> See Codebook 3.1, updated November 20, 2014, pp. 3-4.</bibtext> </blist> <blist> <bibtext> In addition to intercoder reliability one can pay attention to intracoder reliability or stability, that is, does a coder code previous reports in the same way at a later point in time ([47]:270-71).</bibtext> </blist> </ref> <ref id="AN0156346852-24"> <title> References </title> <blist> <bibtext> Azar Edward E.1980. " The Conflict and Peace Data Bank (COPDAB) Project." Journal of Conflict Resolution24:143–52.</bibtext> </blist> <blist> <bibtext> Barron Patrick, Sharpe Joanne. 2008. " Local Conflict in Post-Suharto Indonesia: Understanding Variations in Violence Levels and Forms through Local Newspapers." Journal of East Asian Studies8:395–423.</bibtext> </blist> <blist> <bibtext> Bernauer Thomas, Gleditsch Nils P.. 2012. " New Event Data in Conflict Research." International Interactions38:375–81.</bibtext> </blist> <blist> <bibtext> Bocquier Philippe, Maupeu Hervé. 2005. " Analysing Low Intensity Conflict in Africa Using Press Reports." European Journal of Population21:321–45.</bibtext> </blist> <blist> <bibtext> Bond Doug, Bond Joe, Oh Churl, Craig Jenkins J., Taylor Charles Lewis. 2003. " Integrated Data for Events Analysis (IDEA): An Event Typology for Automated Events Data Development." Journal of Peace Research40:733–45.</bibtext> </blist> <blist> <bibtext> Bond Doug, Craig Jenkins J., Taylor Charles L., Schock Kurt. 1997. " Mapping Mass Political Conflict and Civil Society: Issues and Prospects for the Automated Development of Event Data." Journal of Conflict Resolution41:553–79.</bibtext> </blist> <blist> <bibtext> Boschee Elizabeth, Natarajan Premkumar, Weischedel Ralph. 2013. "Automatic Extraction of Events from Open Source Text for Predictive Forecasting." Pp. 51–67 in Handbook of Computational Approaches to Counterterrorism, edited bySubrahmanian V. S.. New York: Springer Science + Business Media.</bibtext> </blist> <blist> <bibtext> Bueno de Mesquita Ethan, Christine Fair C., Jordan Jenna, Rais Rasul B., Shapiro Jacob N.. 2015. " Measuring Political Violence in Pakistan: Insights from the BFRS Dataset." Conflict Management and Peace Science32:536–58.</bibtext> </blist> <blist> <bibtext> Chan Joseph M., Lee Chi-Chuan. 1984. "Journalistic 'Paradigms' of Civil Protests: A Case Study in Hong Kong." Pp. 249–76 in The News Media in National and International Conflict, edited byArno Andrew, Dissanayake Wimal. Boulder, CO: Westview.</bibtext> </blist> <blist> <bibtext> Chojnacki Sven, Ickler Christian, Spies Michael, Wiesel John. 2012. " Event Data on Armed Conflict and Security: New Perspectives, Old Challenges, and Some Solutions." International Interactions38:382–401.</bibtext> </blist> <blist> <bibtext> Collier Paul, Hoeffler Anke. 2002. " Greed and Grievance in Civil War." Centre for the Study of African Economies Working Paper Series 2002-01, Oxford, England.</bibtext> </blist> <blist> <bibtext> Cook Scott J., Blas Betsabe, Carroll Raymond J., Sinha Samiran. 2017. " Two Wrongs Make a Right: Addressing Underreporting in Binary Data from Multiple Sources." Political Analysis25:223–40.</bibtext> </blist> <blist> <bibtext> Croicu Mihai, Sundberg Ralph. 2016. UCDP GED Codebook Version 5.0. Uppsala, Sweden: Department of Peace and Conflict Research, Uppsala University.</bibtext> </blist> <blist> <bibtext> Croicu Mihai, Weidmann Nils B.. 2015. " Improving the Selection of News Reports for Event Coding Using Ensemble Classification." Research & Politics2:1–8.</bibtext> </blist> <blist> <bibtext> Dardis Frank E.2006. " Marginalization Devices in U.S. Press Coverage of Iraq War Protest: A Content Analysis." Mass Communication and Society9:117–35.</bibtext> </blist> <blist> <bibtext> Danzger M. Herbert. 1975. " Validating Conflict Data." American Sociological Review40:570–84.</bibtext> </blist> <blist> <bibtext> Davenport Christian. 2010. Media Bias, Perspective, and State Repression: The Black Panther Party. Cambridge, MA: Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Davenport Christian, Ball Patrick. 2002. " Views to a Kill: Exploring the Implications of Source Selection in the Case of Guatemalan State Terror, 1977-1995." Journal of Conflict Resolution46:427–50.</bibtext> </blist> <blist> <bibtext> Day Joel, Pinckney Jonathan, Chenoweth Erica. 2015. " Collecting Data on Nonviolent Action: Lessons Learned and Ways Forward." Journal of Peace Research52:129–33.</bibtext> </blist> <blist> <bibtext> Demarest Leila, Langer Arnim. 2018. " The Study of Violence and Social Unrest in Africa: A Comparative Analysis of Three Conflict Event Datasets." African Affairs117:310–25.</bibtext> </blist> <blist> <bibtext> Earl Jennifer, Martin Andrew, McCarthy John D., Soule Sarah A.. 2004. " The Use of Newspaper Data in the Study of Collective Action." Annual Review of Sociology30:65–80.</bibtext> </blist> <blist> <bibtext> Eck Kristine. 2012. " In Data We Trust? A Comparison of UCDP GED and ACLED Conflict Event Datasets." Cooperation and Conflict47:124–41.</bibtext> </blist> <blist> <bibtext> Eisinger Peter K.1973. " The Conditions of Protest Behavior in American Cities." The American Political Science Review67:11–28.</bibtext> </blist> <blist> <bibtext> Entman Robert M.1993. " Framing: Toward Clarification of a Fractured Paradigm." Journal of Communication43:51–58.</bibtext> </blist> <blist> <bibtext> Fearon James D., Laitin David D.. 2003. " Ethnicity, Insurgency, and Civil War." American Political Science Review97:75–90.</bibtext> </blist> <blist> <bibtext> Franzosi Roberto. 1987. " The Press as a Source of Socio-historical Data: Issues in the Methodology of Data Collection from Newspapers." Historical Methods: A Journal of Quantitative and Interdisciplinary History20:5–16.</bibtext> </blist> <blist> <bibtext> Galtung Johan, Ruge Mari H.. 1965. " The Structure of Foreign News: The Presentation of the Congo, Cuba and Cyprus Crises in Four Norwegian Newspapers." Journal of Peace Research2:64–90.</bibtext> </blist> <blist> <bibtext> Groves Robert M.1989. Survey Errors and Survey Costs. New York: Wiley.</bibtext> </blist> <blist> <bibtext> Groves Robert M., Fowler Floyd J., Couper Mick P., Lepkowski James M., Singer Eleanor, Tourangeau Roger. 2004. Survey Methodology. Hoboken, NJ: Wiley.</bibtext> </blist> <blist> <bibtext> Gleditsch Kristian S., Metternich Nils W., Ruggeri Andrea. 2014. " Data and Progress in Peace and Conflict Research." Journal of Peace Research51:301–14.</bibtext> </blist> <blist> <bibtext> Hammond Jesse, Weidmann Nils B.. 2014. " Using Machine-coded Event Data for the Micro-level Study of Political Violence." Research & Politics1:1–8.</bibtext> </blist> <blist> <bibtext> Harcup Tony, O'Neill Deirdre. 2001. " What Is News? Galtung and Ruge Revisited." Journalism Studies2:261–80.</bibtext> </blist> <blist> <bibtext> Harcup Tony, O'Neill Deirdre. 2016. " What Is News? News Values Revisited (Again)." Journalism Studies18:1–19. doi: 10.1080/1461670X.2016.1150193.</bibtext> </blist> <blist> <bibtext> Heap Bradford, Krzywicki Alfred, Schmeidl Susanne, Wobcke Wayne, Bain Michael. 2017. " A Joint Human/Machine Process for Coding Events and Conflict Drivers." Conference Paper, International Conference on Advanced Data Mining and Applications, ADMA 2017, Singapore. October.</bibtext> </blist> <blist> <bibtext> Hendrix Cullen S., Salehyan Idean. 2015. " No News Is Good News: Mark and Recapture for Event Data When Reporting Probabilities Are Less Than One." International Interactions41:392–406.</bibtext> </blist> <blist> <bibtext> Hendrix Cullen S., Salehyan Idean. 2017. " A House Divided: Threat Perception, Military Factionalism, and Repression in Africa." Journal of Conflict Resolution61:1653–81.</bibtext> </blist> <blist> <bibtext> Herkenrath Mark, Knoll Alex. 2011. " Protest Events in International Press Coverage: An Empirical Critique of Cross-national Conflict Databases." International Journal of Comparative Sociology52:163–80.</bibtext> </blist> <blist> <bibtext> Hickler Christian, Wiesel John. 2012. " New Method, Different War? Evaluating Supervised Machine Learning by Coding Armed Conflict." SFB-Governance Working Paper Series, No.39. Collaborative Research Center (SFB) 700, Berlin. 2012. Retrieved October 12, 2019 (https://refubium.fu-berlin.de/bitstream/handle/fub188/18738/WP39.pdf?sequence=1&isAllowed=y).</bibtext> </blist> <blist> <bibtext> Hirshleifer Jack. 1994. " The Dark Side of the Force." Economic Inquiry32:1–10.</bibtext> </blist> <blist> <bibtext> Hug Simon, Wisler Dominique. 1998. " Correcting for Selection Bias in Social Movement Research." Mobilization: An International Quarterly3:141–61.</bibtext> </blist> <blist> <bibtext> Hutter Sven. 2014. "Protest Event Analysis and Its Offspring." Pp. 335–67 in Methodological Practices in Social Movement Research, edited byPorta Donatella della. Oxford, UK: Oxford University Press.</bibtext> </blist> <blist> <bibtext> Jenkins J. Craig, Maher Thomas V.. 2016. " What Should We Do about Source Selection in Event Data? Challenges, Progress, and Possible Solutions." International Journal of Sociology46:42–57.</bibtext> </blist> <blist> <bibtext> Kalyvas Stathis N.2006. The Logic of Violence in Civil War. Cambridge, MA: Cambridge University Press.</bibtext> </blist> <blist> <bibtext> King Gary, Lowe Will. 2003. " An Automated Information Extraction Tool for International Conflict Data with Performance as Good as Human Coders: A Rare Events Evaluation Design." International Organization57:617–42.</bibtext> </blist> <blist> <bibtext> Koopmans Ruud, Rucht Dieter. 2002. "Protest Event Analysis." Pp. 231–59 in Methods of Social Movement Research, edited byKlandermans B., Staggenborg S.. Minneapolis: University of Minnesota.</bibtext> </blist> <blist> <bibtext> Kriesi Hanspeter, Koopmans Ruud, Duyvendak Jan W., Giugni Marco G.. 1998. New Social Movements in Western Europe: A Comparative Analysis. Minneapolis: University of Minnesota.</bibtext> </blist> <blist> <bibtext> Krippendorff Klaus. 2013. Content Analysis: An Introduction to Its Methodology. Thousand Oaks, CA: Sage.</bibtext> </blist> <blist> <bibtext> Leetaru Kalev, Schrodt Philip A.. 2013. " GDELT: Global Data on Events, Location and Tone, 1979-2012." Paper presented at the International Studies Association Meetings, San Francisco, CA, April.</bibtext> </blist> <blist> <bibtext> LaFree Gary, Dugan Laura. 2007. " Introducing the Global Terrorism Database." Terrorism and Political Violence19:181–204.</bibtext> </blist> <blist> <bibtext> Lee Francis L. F.2014. " Triggering the Protest Paradigm: Examining Factors Affecting News Coverage of Protests." International Journal of Communication8:2725–46.</bibtext> </blist> <blist> <bibtext> McCarthy John D., McPhail Clark, Smith Jackie, Crishock Louis J.. 1999. "Electronic and Print Media Representations of Washington D.C. Demonstrations, 1982 and 1991: A Demography of Description Bias." Pp. 113–30 in Acts of Dissent: New Developments in the Study of Protest, edited byRucht Dieter, Koopmans Ruud, Neidhart Friedhelm. Oxford, England: Rowman & Littlefield.</bibtext> </blist> <blist> <bibtext> McClelland Charles A.1976. World Event/Interaction Survey Codebook (ICPSR 5211). Ann Arbor, MI: Inter-University Consortium for Political and Social Research.</bibtext> </blist> <blist> <bibtext> McLeod Douglas, Hertog James. 1992. " The Manufacture of 'Public Opinion' by Reporters: Informal Cues for Public Perceptions of Protest Groups." Discourse & Society3:259–75.</bibtext> </blist> <blist> <bibtext> Myers Daniel J., Caniglia Beth S.. 2004. " All the Rioting That's Fit to Print: Selection Effects in National Newspaper Coverage of Civil Disorders, 1968-1969." American Sociological Review69:519–43.</bibtext> </blist> <blist> <bibtext> Oliver Pamela E., Myers Daniel J.. 1999. " How Events Enter the Public Sphere: Conflict, Location, and Sponsorship in Local Newspaper Coverage of Public Events." American Journal of Sociology105:38–87.</bibtext> </blist> <blist> <bibtext> Ortiz David G., Myers Daniel J., Eugene Walls N., Diaz Maria-Elena D.. 2005. " Where Do We Stand with Newspaper Data? " Mobilization: An International Journal10:397–419.</bibtext> </blist> <blist> <bibtext> Raleigh Clionadh, Dowd Caitriona. 2017. " Armed Conflict and Event Location (ACLED) Codebook." Retrieved 12 October 2019 (https://<ulink href="http://www.acleddata.com/wp-content/uploads/2017/01/ACLED%5fCodebook%5f2017.pdf">www.acleddata.com/wp-content/uploads/2017/01/ACLED%5fCodebook%5f2017.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Raleigh Clionadh, Linke Andrew, Hegre Håvard, Karlsen Joakim. 2010. " Introducing ACLED-Armed Conflict Location and Event Data." Journal of Peace Research47:1–10.</bibtext> </blist> <blist> <bibtext> Reinemann Carsten, Stanyer James, Scherr Sebastian, Legnante Guido. 2011. " Hard and Soft News: A Review of Concepts, Operationalizations and Key Findings." Journalism13:221–39.</bibtext> </blist> <blist> <bibtext> Ruggeri Andrea, Gizelis Theodora-Ismene, Dorussen Han. 2011. " Events Data as Bismarck's Sausages? Intercoder Reliability, Coders' Selection, and Data Quality." International Interactions37:340–61.</bibtext> </blist> <blist> <bibtext> Salehyan Idean. 2015. " Best Practices in the Collection of Conflict Data." Journal of Peace Research52:105–9.</bibtext> </blist> <blist> <bibtext> Salehyan Idean, Hendrix Cullen S., Hamner Jesse, Case Christina, Linebarger Christopher, Stull Emily, Williams Jennifer. 2012. " Social Conflict in Africa: A New Database." International Interactions38:503–11.</bibtext> </blist> <blist> <bibtext> Schrodt Philip A.2006. " Twenty Years of the Kansas Event Data System Project." The Political Methodologist14:2–8.</bibtext> </blist> <blist> <bibtext> Schrodt Philip A.2012. " Precedents, Progress, and Prospects in Political Event Data." International Interactions38:546–69.</bibtext> </blist> <blist> <bibtext> Schrodt Philip A., Brackle David Van. 2013. "Automated Coding of Political Event Data." Pp. 23–49 in Handbook of Computational Approaches to Counterterrorism, edited bySubrahmanian V. S.. New York: Springer Science + Business Media.</bibtext> </blist> <blist> <bibtext> Schrodt Philip A., Simpson Erin M., Gerner Deborah J.. 2001. " Monitoring Conflict Using Automated Coding of Newswire Reports: A Comparison of Five Geographical Regions." Conference paper, June 8-9, Uppsala, Sweden.</bibtext> </blist> <blist> <bibtext> Schrodt Philip A., Ulfelder Jay. 2016. " Political Instability Task Force Atrocities Event Data Collection Codebook Version 1.1b1." Retrieved October 12, 2019 (<ulink href="http://eventdata.parusanalytics.com/data.dir/PITF%5fAtrocities.codebook.1.1B1.pdf">http://eventdata.parusanalytics.com/data.dir/PITF%5fAtrocities.codebook.1.1B1.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Smith Jackie, McCarthy John D., McPhail Clark, Boguslaw Augustyn. 2001. " From Protest to Agenda Building: Description Bias in Media Coverage of Protest Events in Washington, D.C." Social Forces79:1397–423.</bibtext> </blist> <blist> <bibtext> Schneider Gerald, Bussmann Margit. 2013. " Accounting for the dynamics of one-sided violence: Introducing KOSVED." Journal of Peace Research50:635–44.</bibtext> </blist> <blist> <bibtext> Stepinski Adam, Stoll Richard, Subramanian Devika. 2006. " Automated Event Coding Using Conditional Random Fields." Retrieved June13, 2018 (https://<ulink href="http://www.cs.rice.edu/∼devika/conflict/papers/draft1.pdf">www.cs.rice.edu/∼devika/conflict/papers/draft1.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Sundberg Ralph, Melander Erik. 2013. " Introducing the UCDP Georeferenced Event Dataset." Journal of Peace Research50:523–32.</bibtext> </blist> <blist> <bibtext> Urdal Henrik. 2008. " Urban Social Disturbance in Africa and Asia: report on a New Dataset." PRIO Papers, Oslo, Norway.</bibtext> </blist> <blist> <bibtext> Weaver David A., Scacco Joshua M.. 2012. " Revisiting the Protest Paradigm." The International Journal of Press/Politics18:61–84.</bibtext> </blist> <blist> <bibtext> Weidmann Nils B.2015. " On the Accuracy of Media-based Conflict Event Data." The Journal of Conflict Resolution59:1129–49.</bibtext> </blist> <blist> <bibtext> Weidmann Nils B.2016. " A Closer Look at Reporting Bias in Conflict Event Data." American Journal of Political Science60:206–18.</bibtext> </blist> <blist> <bibtext> Weidmann Nils B., Rød Espen G.. 2015. " Making Uncertainty Explicit: Separating Reports and Events in the Coding of Violence and Contention." Journal of Peace Research52:125–28.</bibtext> </blist> <blist> <bibtext> Wilkinson Steven I.2004. Votes and Violence: Ethnic Competition and Ethnic Riots in India. New York: Cambridge University Press.</bibtext> </blist> </ref> <aug> <p>By Leila Demarest and Arnim Langer</p> <p>Reported by Author; Author</p> <p></p> <p>Leila Demarest is an assistant professor of African politics at the Department of Political Science, Leiden University, the Netherlands. Her research interests include social movements and political mobilization in Africa, political communication, and quantitative and qualitative social science research methodology. Her most recent publications (with Arnim Langer) are "The Study of Violence and Social Unrest in Africa: A Comparative Analysis of Three Conflict Event Datasets" in African Affairs (April 2018) and "Peace Journalism on a Shoestring? Conflict Reporting in Nigeria's National News Media" in Journalism (first online, August 2018).</p> <p>Arnim Langer is a professor in international relations at KU Leuven and director of the Centre for Research on Peace and Development (CRPD) at the Faculty of Social Sciences. Currently, he is also Humboldt Research Fellow at the University of Heidelberg, Germany. He has published extensively on the causes of violent conflict in heterogeneous societies and the challenges to sustainable peacebuilding. Some of his recent publications include "Conceptualising and Measuring Social Cohesion in Africa: Towards a Perceptions-based Index" (published in Social Indicators Research) and "A General Class of Social Distance Measures" (published in Political Analysis).</p> </aug> <nolink nlid="nl1" bibid="bib11" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib25" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib39" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib23" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib52" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib16" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib26" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib58" firstref="ref10"></nolink> <nolink nlid="nl9" bibid="bib62" firstref="ref11"></nolink> <nolink nlid="nl10" bibid="bib72" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib71" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib49" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib67" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib76" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib69" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib21" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib41" firstref="ref20"></nolink> <nolink nlid="nl18" bibid="bib27" firstref="ref21"></nolink> <nolink nlid="nl19" bibid="bib33" firstref="ref22"></nolink> <nolink nlid="nl20" bibid="bib47" firstref="ref23"></nolink> <nolink nlid="nl21" bibid="bib22" firstref="ref24"></nolink> <nolink nlid="nl22" bibid="bib61" firstref="ref25"></nolink> <nolink nlid="nl23" bibid="bib74" firstref="ref26"></nolink> <nolink nlid="nl24" bibid="bib75" firstref="ref27"></nolink> <nolink nlid="nl25" bibid="bib29" firstref="ref28"></nolink> <nolink nlid="nl26" bibid="bib42" firstref="ref30"></nolink> <nolink nlid="nl27" bibid="bib56" firstref="ref31"></nolink> <nolink nlid="nl28" bibid="bib44" firstref="ref34"></nolink> <nolink nlid="nl29" bibid="bib48" firstref="ref35"></nolink> <nolink nlid="nl30" bibid="bib31" firstref="ref36"></nolink> <nolink nlid="nl31" bibid="bib32" firstref="ref47"></nolink> <nolink nlid="nl32" bibid="bib45" firstref="ref51"></nolink> <nolink nlid="nl33" bibid="bib43" firstref="ref56"></nolink> <nolink nlid="nl34" bibid="bib17" firstref="ref58"></nolink> <nolink nlid="nl35" bibid="bib37" firstref="ref61"></nolink> <nolink nlid="nl36" bibid="bib20" firstref="ref63"></nolink> <nolink nlid="nl37" bibid="bib30" firstref="ref64"></nolink> <nolink nlid="nl38" bibid="bib64" firstref="ref68"></nolink> <nolink nlid="nl39" bibid="bib63" firstref="ref71"></nolink> <nolink nlid="nl40" bibid="bib24" firstref="ref79"></nolink> <nolink nlid="nl41" bibid="bib15" firstref="ref80"></nolink> <nolink nlid="nl42" bibid="bib53" firstref="ref81"></nolink> <nolink nlid="nl43" bibid="bib50" firstref="ref83"></nolink> <nolink nlid="nl44" bibid="bib73" firstref="ref84"></nolink> <nolink nlid="nl45" bibid="bib19" firstref="ref85"></nolink> <nolink nlid="nl46" bibid="bib51" firstref="ref88"></nolink> <nolink nlid="nl47" bibid="bib46" firstref="ref93"></nolink> <nolink nlid="nl48" bibid="bib10" firstref="ref94"></nolink> <nolink nlid="nl49" bibid="bib57" firstref="ref96"></nolink> <nolink nlid="nl50" bibid="bib65" firstref="ref102"></nolink> <nolink nlid="nl51" bibid="bib12" firstref="ref108"></nolink> <nolink nlid="nl52" bibid="bib14" firstref="ref109"></nolink> <nolink nlid="nl53" bibid="bib34" firstref="ref110"></nolink> <nolink nlid="nl54" bibid="bib70" firstref="ref112"></nolink> <nolink nlid="nl55" bibid="bib66" firstref="ref116"></nolink> <nolink nlid="nl56" bibid="bib13" firstref="ref121"></nolink> <nolink nlid="nl57" bibid="bib38" firstref="ref128"></nolink> <nolink nlid="nl58" bibid="bib60" firstref="ref134"></nolink> <nolink nlid="nl59" bibid="bib55" firstref="ref144"></nolink> <nolink nlid="nl60" bibid="bib18" firstref="ref146"></nolink> <nolink nlid="nl61" bibid="bib40" firstref="ref149"></nolink> <nolink nlid="nl62" bibid="bib35" firstref="ref151"></nolink> <nolink nlid="nl63" bibid="bib36" firstref="ref154"></nolink> <nolink nlid="nl64" bibid="bib54" firstref="ref160"></nolink> <nolink nlid="nl65" bibid="bib28" firstref="ref195"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1337281
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: How Events Enter (or Not) Data Sets: The Pitfalls and Guidelines of Using Newspapers in the Study of Conflict
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Demarest%2C+Leila%22">Demarest, Leila</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6887-9937">0000-0001-6887-9937</externalLink>)<br /><searchLink fieldCode="AR" term="%22Langer%2C+Arnim%22">Langer, Arnim</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. May 2022 51(2):632-666.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: http://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 35
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Descriptive
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Guidelines%22">Guidelines</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Conflict%22">Conflict</searchLink><br /><searchLink fieldCode="DE" term="%22Social+Science+Research%22">Social Science Research</searchLink><br /><searchLink fieldCode="DE" term="%22Newspapers%22">Newspapers</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Problems%22">Research Problems</searchLink><br /><searchLink fieldCode="DE" term="%22Error+of+Measurement%22">Error of Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/0049124119882453
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: While conflict event data sets are increasingly used in contemporary conflict research, important concerns persist regarding the quality of the collected data. Such concerns are not necessarily new. Yet, because the methodological debate and evidence on potential errors remains scattered across different subdisciplines of social sciences, there is little consensus concerning proper reporting practices in codebooks, how best to deal with the different types of errors, and which types of errors should be prioritised. In this article, we introduce a new analytical framework--that is, the Total Event Error (TEE) framework--which aims to elucidate the methodological challenges and errors that may affect whether and how events are entered into conflict event data sets, drawing on different fields of study. Potential errors are diverse and may range from errors arising from the rationale of the media source (e.g., selection of certain types of events into the news) to errors occurring during the data collection process or the analysis phase. Based on the TEE framework, we propose a set of strategies to mitigate errors associated with the construction and use of conflict event data sets. We also identify a number of important avenues for future research concerning the methodology of creating conflict event data sets.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2022
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1337281
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1337281
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/0049124119882453
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 35
        StartPage: 632
    Subjects:
      – SubjectFull: Guidelines
        Type: general
      – SubjectFull: Research Methodology
        Type: general
      – SubjectFull: Conflict
        Type: general
      – SubjectFull: Social Science Research
        Type: general
      – SubjectFull: Newspapers
        Type: general
      – SubjectFull: Research Problems
        Type: general
      – SubjectFull: Error of Measurement
        Type: general
      – SubjectFull: Data Analysis
        Type: general
    Titles:
      – TitleFull: How Events Enter (or Not) Data Sets: The Pitfalls and Guidelines of Using Newspapers in the Study of Conflict
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Demarest, Leila
      – PersonEntity:
          Name:
            NameFull: Langer, Arnim
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
          Numbering:
            – Type: volume
              Value: 51
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1