Comparing Methods for Estimating Demographics in Racially Polarized Voting Analyses

Saved in:
Bibliographic Details
Title: Comparing Methods for Estimating Demographics in Racially Polarized Voting Analyses
Language: English
Authors: Ari Decter-Frain (ORCID 0000-0001-9635-3334), Pratik Sachdeva (ORCID 0000-0002-6809-2437), Loren Collingwood (ORCID 0000-0002-4447-8204), Hikari Murayama (ORCID 0000-0002-4067-4734), Juandalyn Burke (ORCID 0000-0002-6345-7505), Matt Barreto, Scott Henderson (ORCID 0000-0003-0624-4965), Spencer Wood (ORCID 0000-0002-5794-2619), Joshua Zingher (ORCID 0000-0003-3928-4269)
Source: Sociological Methods & Research. 2025 54(2):706-738.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 33
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Voting, Computation, Racial Composition, Bayesian Statistics, Statistical Inference, Elections, Statistical Analysis
Geographic Terms: Georgia, New York
DOI: 10.1177/00491241231192383
ISSN: 0049-1241
1552-8294
Abstract: We consider the cascading effects of researcher decisions throughout the process of quantifying racially polarized voting (RPV). We contrast three methods of estimating precinct racial composition, Bayesian Improved Surname Geocoding (BISG), fully Bayesian BISG, and Citizen Voting Age Population (CVAP), and two algorithms for performing ecological inference (EI), King's EI and EI:RxC using eiCompare. Using data from two different elections we identify circumstances in which different combinations of methods produce divergent results, comparing against ground-truth data where available. We first find that BISG outperforms CVAP at estimating racial composition, though fully Bayesian BISG does not yield further improvements. Next, in a statewide election, we find that all combinations of methods yield similarly reliable estimates of RPV. However, county-level analyses and results from a non-partisan school board election reveal that BISG and CVAP produce divergent estimates of Black preferences in elections with low turnout and few precincts. Our results suggest that methodological choices can meaningfully alter conclusions about RPV, particularly in smaller, low-turnout elections.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1473565
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHq-XuBoKahFIsgNqTgw913AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNpLZdKYecju6ZNI7wIBEICBm-7-JAQgJZthEVgtEoZ4IgcV0sN29m1kSBLd-IlrFR_0xIfIJtXId_s_f77E2RgBpLWCJQCRS1drjzSVtgoZOOs0C67uWwyDiszpQYOJseQ7iHvmL6bjOFdaHBwRKbuAkE2r4CgOVKBB5dDjm72BLHnVOffP2G206YA35IVehE9Scwj1mrmVLiW8-8hHaClq_aSTDuFmdKTaebiq
Text:
  Availability: 1
  Value: <anid>AN0184233955;som01may.25;2025Apr07.05:58;v2.2.500</anid> <title id="AN0184233955-1">Comparing Methods for Estimating Demographics in Racially Polarized Voting Analyses </title> <p>We consider the cascading effects of researcher decisions throughout the process of quantifying racially polarized voting (RPV). We contrast three methods of estimating precinct racial composition, Bayesian Improved Surname Geocoding (BISG), fully Bayesian BISG, and Citizen Voting Age Population (CVAP), and two algorithms for performing ecological inference (EI), King's EI and EI:RxC using eiCompare. Using data from two different elections we identify circumstances in which different combinations of methods produce divergent results, comparing against ground-truth data where available. We first find that BISG outperforms CVAP at estimating racial composition, though fully Bayesian BISG does not yield further improvements. Next, in a statewide election, we find that all combinations of methods yield similarly reliable estimates of RPV. However, county-level analyses and results from a non-partisan school board election reveal that BISG and CVAP produce divergent estimates of Black preferences in elections with low turnout and few precincts. Our results suggest that methodological choices can meaningfully alter conclusions about RPV, particularly in smaller, low-turnout elections.</p> <p>Keywords: ecological inference; Bayesian statistics; voting rights; race imputation; elections</p> <hd id="AN0184233955-2">Introduction</hd> <p>Since the passage of the Voting Rights Act (VRA), scholars and legal practitioners have studied the presence (or absence) of racially polarized voting (RPV) in jurisdictions around the country (e.g., school boards, legislative districts, etc.). To do so, scholars typically gather publicly available precinct voting results (i.e., voter turnout and candidate totals) and racial demographics from the U.S. Decennial Census or American Community Survey (ACS). One of the most relied upon sources of demographic data is ACS' Citizen Voting Age Population (CVAP), a tabulation by geographic unit of the eligible voting pool. Analysts then use racial demographics as the independent variable to estimate candidate choice as the dependent variable.</p> <p>RPV is said to have occurred when, for example, Black voters consistently support one candidate or slate of candidates, and White voters consistently support another candidate or slate of candidates. Because individual-level public opinion data remain unavailable in many election contests, correctly estimating the level of RPV using aggregate data is fundamental to determining whether a jurisdiction has violated the VRA and whether race/ethnicity continues to play an overriding role in political behavior and representation. In this paper, we ask whether different methods for estimating racial composition, which serve as the inputs to ecological inference (EI) statistical models, impact subsequent conclusions about whether RPV characterizes an election. Our main goal is to examine how newly developed methods compare to the more established methods using CVAP as a data source.</p> <p>Numerous commonly used datasets across the social sciences lack a self-reported measure of race/ethnicity, which presents an obstacle to the study of race and racial disparities. A method of combining surname race probabilities with aggregate Census data—called Bayesian Improved Surname Geocoding (BISG)—can more accurately predict individual level and aggregate race classification ([<reflink idref="bib14" id="ref1">14</reflink>], [<reflink idref="bib15" id="ref2">15</reflink>]; [<reflink idref="bib31" id="ref3">31</reflink>]). The U.S. Census Bureau enables this method by providing a dataset containing the racial distribution of surnames recorded during the Decennial Census. For example, as of the 2010 census, of all people named Smith in the U.S., 70.9% are White, 23.1% are Black, and 2.4% are Hispanic. This methodological innovation potentially has significant consequences for how scholars and practitioners conduct RPV analyses. For instance, combining Florida's voter file with precinct data, [<reflink idref="bib28" id="ref4">28</reflink>]) show that aggregated BISG estimates of turnout by race outperform estimates using CVAP data and EI methods. Yet, to date, scholars have not examined how BISG and CVAP compare in predicting vote choice by race.</p> <p>The inclusion of racial probabilities obtained via BISG as the source of demographic input data is a potentially consequential advancement in RPV methodology because heretofore scholars have generally used CVAP (or VAP) data. One reason to expect that BISG might improve EI estimates is that all data inputs (race and candidate vote) emerge from the same geographic unit (the precinct), thus reducing spatial misclassification inherent in joining CVAP data with voting data. Second, we can be confident that all inputs reflect the votes of actual voters. More specifically, using data from voters alone will narrow the information bounds at the precinct level on who did and did not vote—thereby potentially enhancing EI performance ([<reflink idref="bib34" id="ref5">34</reflink>]; [<reflink idref="bib4" id="ref6">4</reflink>]). On the other hand, BISG entails geocoding voters based oon supplied addresses, then estimating individuals' race using surname and census-block racial information before aggregating such data to the precinct prior to implementing EI. Errors can occur at each step: (a) a poor geocoder or data-entry mistake can misplace voters into incorrect geographic units, (b) individual race classification errors—potentially in racially diverse precincts—can occur, leading to (c) racial aggregation bias. This latter error can mask or attenuate the presence of RPV.</p> <p>The potential pitfalls of BISG notwithstanding, in some contexts CVAP data may disproportionately produce RPV false negatives, for two reasons. First, particularly in low-turnout elections with a small number of precincts, White voters typically vote at higher rates ([<reflink idref="bib20" id="ref7">20</reflink>]; [<reflink idref="bib27" id="ref8">27</reflink>]; [<reflink idref="bib26" id="ref9">26</reflink>]). Therefore, CVAP input may over-estimate the number of non-White constituents that actually voted, even after attempting to account for turnout in the modeling strategy. Minority vote share can also fluctuate substantially between elections ([<reflink idref="bib41" id="ref10">41</reflink>]; [<reflink idref="bib48" id="ref11">48</reflink>]; [<reflink idref="bib55" id="ref12">55</reflink>]), and CVAP will not capture these fluctuations. Second, because CVAP data exists at the block-group level, spatial joining to the precinct—which has different spatial boundaries—can lead to measurement error. In a large jurisdiction like the state of Georgia, outliers may balance out, but in small jurisdictions, a few outstanding outliers could lead to noisy EI estimation of candidate selection by race.</p> <p>We bring data to bear on this emerging debate to begin to answer two research questions. First, do BISG and CVAP perform similarly in the estimation of RPV? We hypothesize that in jurisdictions with a large number of precincts such as statewide contests, BISG and CVAP will estimate RPV at similar levels when compared to a known benchmark. We test this by analyzing RPV in Georgia's 2018 gubernatorial election. Our datasets include BISG estimates of all Georgia voters aggregated by precinct, CVAP racial estimates aggregated by precinct, self-reported race aggregated by precinct, and publicly available exit poll data. Relative to an exit poll benchmark, we find that BISG, CVAP, and self-reported race perform similarly: All three methods demonstrate overwhelming White support for Republican Brian Kemp, and overwhelming Black support for Democrat Stacey Abrams.</p> <p>Second, relative to CVAP demographic data, does BISG demographic data more accurately detect the presence of RPV in low-turnout elections with a small number of precincts? We hypothesize that in these contests, BISG may more accurately capture the racial composition of the electorate thereby revealing true voting patterns, whereas CVAP may produce outliers that can wildly over-estimate minority vote share—even though we know that minority voters historically turn out at significantly lower rates than do White voters ([<reflink idref="bib19" id="ref13">19</reflink>], [<reflink idref="bib20" id="ref14">20</reflink>]). We test this question in two contexts. We first conduct EI within each county in Georgia for the 2018 Gubernatorial election and use heterogeneity across counties to explore the conditions under which CVAP and BISG yield divergent conclusions about RPV. Next, we examine a nonpartisan, low-turnout, local school board election in East Ramapo, NY. For this analysis, our data include voter files, BISG estimates aggregated by precinct, precinct election results, and CVAP estimates aggregated to the precinct. We find that CVAP, when paired with the EI:rows by columns (RxC) method, consistently reports less polarization in candidate preference among Blacks and Whites when compared to all other East Ramapo results, which are inconsistent with campaign dynamics. Our findings help contribute to and advance the methodologies practitioners and election scientists are called upon to employ in court challenges around the country, and to better understand the contours of RPV when testing theories of voting in an increasingly diverse populace.</p> <hd id="AN0184233955-3">Background</hd> <p> <emph>Thornburg v. Gingles, 478 U.S. 30, 1986</emph> established a legal framework to guide voting rights challenges in cases where a plaintiff suspects a jurisdiction of diluting minority votes. A common challenge occurs when a jurisdiction uses an at-large system to elect council members, which, scholarship reveals, can dilute minority votes and representation ([<reflink idref="bib9" id="ref15">9</reflink>]; [<reflink idref="bib57" id="ref16">57</reflink>]; [<reflink idref="bib58" id="ref17">58</reflink>]; [<reflink idref="bib1" id="ref18">1</reflink>]).</p> <p>According to <emph>Gingles</emph>, plaintiffs must establish three prongs to prove a voting rights violation: (a) The minority group is geographically compact and large enough to win a single-member district (assuming RPV), (b) the minority group tends to vote cohesively, and (c) the majority group consistently prefers different candidates who consistently win ([<reflink idref="bib51" id="ref19">51</reflink>]). Based on this framework and the court's prescribed methods, social scientists employ voting analyses by relying on a combination of precinct voting data and demographic data from multiple elections to examine whether a jurisdiction is in violation of the VRA.</p> <p>Due to statistical uncertainty inherent in making claims about individual voting behavior from aggregate data, much of the academic literature on EI examines the precision of statistical models and algorithms with the goal of avoiding ecological fallacy. Early evidence presented at trial relied on Goodman's regression that regressed vote choice on race, controlling for unit population size ([<reflink idref="bib22" id="ref20">22</reflink>], [<reflink idref="bib23" id="ref21">23</reflink>]), and homogeneous precinct analysis ([<reflink idref="bib35" id="ref22">35</reflink>]; [<reflink idref="bib39" id="ref23">39</reflink>]). However, over the decades, data accessibility, computing capacity, and statistical tools evolved considerably. For instance, [<reflink idref="bib33" id="ref24">33</reflink>]) advocated for a more precise measurement of racial voting patterns using EI methods that became known as King's EI. No longer facing a strictly Black-Anglo hyper-segregated environment, others, notably [<reflink idref="bib49" id="ref25">49</reflink>]), developed algorithms to simultaneously estimate candidate preference by race for scenarios with multiple candidates and racial groups, referred to as EI - RxC.</p> <p>As it stands, social scientists—and the courts—tend to rely on iterative EI and the newer and computationally intensive approach, RxC. Scholarship investigating which method produces more reliable estimates is mixed, with [<reflink idref="bib18" id="ref26">18</reflink>]) showing that RxC is preferred in certain contexts, whereas [<reflink idref="bib24" id="ref27">24</reflink>]); [<reflink idref="bib11" id="ref28">11</reflink>]); [<reflink idref="bib5" id="ref29">5</reflink>]) show that multiple EI estimation methods tend to produce similar results. No research, however, has examined whether the data comprising racial demographic estimates that serve as inputs to EI affects the subsequent results.</p> <p>Our main goal is to examine how BISG methods compare to the more established CVAP data source in EI estimation. Therefore, we provide a brief review of BISG's development. BISG as applied in political science research and assessing RPV has its origins in the health sciences. [<reflink idref="bib14" id="ref30">14</reflink>]) presented Bayesian Surname and Geocoding, a method to improve racial predictions of individual enrollees in a national health plan. Specifically, they found that this new approach vastly outperformed earlier approaches based only on surname or geographic location. Using a new Census surname list, [<reflink idref="bib15" id="ref31">15</reflink>]) showed even greater improvement over previous analyses that used a less comprehensive surname list. [<reflink idref="bib2" id="ref32">2</reflink>]) implemented BISG and examined ideal race probability cutoffs for racial classification, and found variation by race/ethnicity. Thus, when it comes to estimating individual-level race, the literature clearly indicates that BISG is an improvement over previous individual-level race predictions such as surname matching—and that the improvements come specifically for estimates of White and Black voters.</p> <p>Although political scientists have not yet compared BISG and CVAP methods to estimate candidate preference, [<reflink idref="bib28" id="ref33">28</reflink>])'s Who Are You (WRU) software implements the BISG algorithm and incorporates Census data into an R statistical package. The package includes access to 2010 Census block, block-group, tract, and county data. They test WRU on the Florida voter file—which contains a column for self-reported race/ethnicity—confirming earlier studies that BISG improves individual-level race estimates relative to surname only, or geographic-only estimates. Then, they compare BISG to CVAP methods to estimate voter turnout by race. They show that BISG-estimated individual race probabilities aggregated to the precinct do a better job estimating turnout by race than the more traditional method of estimating turnout using EI with CVAP data.</p> <p>Political scientists have begun to incorporate BISG—or related racial prediction algorithms—into various analyses. For example, in a paper showing that the demolition of a largely Black housing project in Chicago changed Whites' voting preferences, [<reflink idref="bib16" id="ref34">16</reflink>]) uses BISG before estimating a series of King's EI models. In a study of candidates or districts, [<reflink idref="bib19" id="ref35">19</reflink>]) implements individual-level race prediction developed by Catalyst—a well-known political database company—to estimate racial turnout across the entire United States. [<reflink idref="bib17" id="ref36">17</reflink>]) also combine BISG and EI to estimate the effect of protests on left-leaning ballot support. [<reflink idref="bib56" id="ref37">56</reflink>]) compare a Bayesian classifier against a crowd-sourced estimate of candidates' race. And, in their study showing that minority donors prefer to donate to minority candidates, [<reflink idref="bib25" id="ref38">25</reflink>]) apply BISG to donor lists. To begin to assess how BISG performs in RPV analyses using EI, we apply BISG to voter files in two very different contexts: (a) a highly competitive statewide contest with relatively high voter turnout (Georgia's 2018 governor's race) and (b) a local contest with low voter turnout and a small number of precincts (East Ramapo Central School Board [ERCSB], NY).</p> <hd id="AN0184233955-4">Data and Methods</hd> <p>This section discusses our data and methods. We begin with a discussion of how we construct the main dataset, which involves combining various sources to analyze the 2018 Georgia gubernatorial election between Republican Brian Kemp and Democrat Stacey Abrams. We then discuss the datasets used in our East Ramapo analysis.</p> <p>All code used to conduct the analyses and create the figures in this paper will be made available on GitHub before publication.</p> <hd id="AN0184233955-5">Georgia Statewide</hd> <p>From the Georgia Secretary of State, we gathered the Georgia voter file in June 2020. This file contains voters' names (first, middle, and surname), addresses, and self-reported race/ethnicity. Additionally, we obtained the voter history file for the November 6, 2018 general election. The voter history file contains a list of the registration IDs for all voters who participated in the election. We subset the voter file to people who voted in the 2018 general election by matching registration IDs between the two files. This produces a voter pool of 3,396,471 voters, lower than the 3,939,328 that voted in the election according to the voter history file. The 542,857 missing voters may correspond to residents who were removed from voter rolls by Georgia's Secretary of State ([<reflink idref="bib13" id="ref39">13</reflink>]).</p> <p>One of the primary reasons we utilize the Georgia voter file is because Georgians record their race when they register to vote: White, Black, Asian-Pacific Islander, Latino/Hispanic, Native-American, other, and unknown for voters who decline to mark a race. To match the categories used by WRU, we recode the relevant data into the following five categories: White, Black, Hispanic, Asian-Pacific Islander, and other. In all relevant BISG analyses we treat this variable as our benchmark "truth." We also drop the seven percent of voters who mark race unknown (though including them in the analyses resulted in no discernible differences in the subsequent results: see Appendix A, supplemental material)[<reflink idref="bib6" id="ref40">6</reflink>].</p> <p>We geocoded each voter's residential address into latitude-longitude coordinates using a combination of the censusxy R package ([<reflink idref="bib45" id="ref41">45</reflink>]) and the commercial vendor opencage ([<reflink idref="bib54" id="ref42">54</reflink>]). Consult Appendix B in the supplemental material for a detailed discussion of our geocoding strategy.</p> <p>Next, we gathered the Georgia block-level shape file, which denotes spatial block boundaries on a mapping plane. We then reverse-geocoded each voter to the Census block (which includes columns for block-group, tract, county, and any other relevant geo-boundary information). [<reflink idref="bib8" id="ref43">8</reflink>]) showed that ZIP-level Census data can yield racial composition estimates with only a minor deterioration of performance. Nonetheless, we focus on block-level geographic data because it provides the most granular, accurate data available. We used each geographic unit's Federal Information Processing Standard (FIPS) identifier to merge in geographic/race estimates implemented in eiCompare, which borrows from WRU's BISG race prediction model. WRU contains five "geography given race" values for each census block in the United States. We connected the individual voter with the appropriate WRU vector of probabilities via the Census block identifier. WRU also contains the 2010 surname file that lists all surnames appearing more than 100 times in the United States and their racial distribution according to the Census.[<reflink idref="bib7" id="ref44">7</reflink>] The package automatically matches each voter's surname against the surname probability and pulls that probability into the BISG equation. This is all of the necessary data to conduct BISG for every voter in the Georgia database. For each voter, BISG estimates five racial probabilities.[<reflink idref="bib8" id="ref45">8</reflink>]</p> <p>To evaluate the accuracy of BISG's racial classification in Georgia, we plotted receiver operating curves (ROC) and computed the area under the ROC curve (AUC) for classifying White and Black voters, and compared it with estimates reported in recent papers (e.g., [<reflink idref="bib28" id="ref46">28</reflink>])). We present these results in Appendix C in the supplemental material. We obtained an AUC of 0.91 for White voters and 0.93 for Black voters. This is consistent with [<reflink idref="bib28" id="ref47">28</reflink>])'s findings in Florida (0.90 and 0.94, respectively). Also consistent with Imai and Khanna (2016), BISG obtains a false negative rate among White and Black voters to 0.13 and 0.10, respectively, while at the same time keeping a true positive rate over 0.80. These results are also consistent with recent work validating BISG in different settings ([<reflink idref="bib31" id="ref48">31</reflink>]; [<reflink idref="bib8" id="ref49">8</reflink>]).</p> <p>After we estimated each voter's set of race probabilities, we aggregated race predictions up to the precinct level. Thus, for each precinct, we produced a BISG aggregate estimate for the number of White, Black, Hispanic, Asian, and other voters in that precinct. This method has been shown to produce more accurate estimates relative to the highest probability binary method ([<reflink idref="bib28" id="ref50">28</reflink>]). For each precinct, we then estimated the share of each racial group by dividing the estimated racial group number by the estimated total number of voters according to the BISG predictions. For each EI analysis, we inputted race proportion and control for total votes taken from the aggregate precinct candidate returns issued by the state.</p> <p>To develop our CVAP by race precinct estimates, we first gathered the Georgia precinct shape file from the website OpenPrecincts.[<reflink idref="bib9" id="ref51">9</reflink>] There are 2,626 number of precincts in the data.[<reflink idref="bib10" id="ref52">10</reflink>] Next, because CVAP is taken from the 5-year ACS, we used block-group as the geographic unit of analysis from which we gather our racial data. There are 5,533 block groups in the state. Because voting districts (precincts) and Census blocks do not perfectly overlap, we use areal weighted interpolation to estimate each precinct's racial/ethnic distribution using the block-group data ([<reflink idref="bib42" id="ref53">42</reflink>]).</p> <p>Finally, we aggregated self-reported race from the voter file to the precinct level by tabulating the share of each racial group by precinct. This served as the "truth" against which we compare all other precinct racial estimates. With our three racial estimates, we joined our new precinct data with 2018 candidate vote and total vote data accessed from the Secretary of State. We converted the raw numbers to percents and weighted each precinct by total vote.</p> <p>To assess RPV using BISG, CVAP, and known truth, we utilized both iterated King's EI and EI-RxC. [<reflink idref="bib33" id="ref54">33</reflink>]); [<reflink idref="bib49" id="ref55">49</reflink>]) and [<reflink idref="bib36" id="ref56">36</reflink>]) delineate these methods. We focused only on Abrams and Kemp, omitting minor third-party candidates. We compared our overall statewide racial voting estimates to the Georgia exit poll, which benchmarks vote preference by race ([<reflink idref="bib44" id="ref57">44</reflink>]). As a second point of reference, we also gathered survey data from the 2018 Cooperative Congressional Election Study (CCES), subsetting the data to just Georgia voters. See Appendix G in the supplemental material for our alternative measure of RPV based on this survey data, which yields results consistent with the exit polls.</p> <hd id="AN0184233955-6">East Ramapo, NY</hd> <p>Our second dataset comes from the East Ramapo School Board (East Ramapo School District (ERCSD)), a small school district outside New York City containing about 60,000 eligible voters. ERCSD is a useful case for our purposes for two reasons. First, these school board elections are small and typically have low turnout, providing an informative counterpoint to our statewide analyses of Georgia.</p> <p>Second, a recent lawsuit in ERCSD demonstrated through ample qualitative evidence that these school board elections were characterized by vote dilution and high levels of RPV. For years, the majority White voting population completely controlled the at-large school board and elected "pro-private school" candidates; while the sizable minority population's preferred candidates failed to win any competitive election (see Appendix H in the supplemental material for more details on the case). Given that small elections almost never have exit poll data, the qualitative evidence of RPV from this legal case makes ERCSD uniquely well-suited to evaluating EI methods in these types of elections.</p> <p>We gathered voter files from the May 2016 ERCSD school board elections. The voter file included each voter's first name, middle name, surname, and residential address. We gathered only names of voters who cast a ballot in each election we examined. Historically, East Ramapo elections are low turnout, with turnout ranging from 9,000 to 13,000 voters.</p> <p>Our methods for analyzing ERCSD elections proceeded analogously to Georgia. First, we geocoded all voters using the ggmap geocode function ([<reflink idref="bib30" id="ref58">30</reflink>]). Next, due to the relatively small pool of voters, we reverse geocoded each voter's latitude/longitude coordinates by sending addresses to the Federal Communications Commission application programming interface, which converted these coordinates to Census blocks.[<reflink idref="bib11" id="ref59">11</reflink>] From there we had each voter's requisite Census block, block-group, tract, and county information. As with Georgia, we used BISG to obtain estimates for all voters and then aggregated the probabilities by precinct.</p> <p>To develop our CVAP race by precinct estimates, first, we gathered the ERCSB precinct shape file, in which there are 10 voting districts. Next, we gathered the Census block-group shape file for Rockland County and subsetted the Census blocks to only those overlapping the ERCSB area. We then gathered CVAP data from 2012–2016 5-year ACS and appended it to the Census block shape file. However, because precincts and Census blocks do not perfectly overlap, we conducted an areal-weighted interpolation to estimate each precinct's racial/ethnic distribution.</p> <p>We gathered precinct candidate preference data from the ERCSB elections division.[<reflink idref="bib12" id="ref60">12</reflink>] Our unit of analysis is the precinct, of which there are only ten. We then joined the BISG-aggregated precinct data with the CVAP precinct data and the candidate preference data. Like for Georgia, we converted all raw numbers to percents and controlled for total voters by precinct in each analysis.</p> <p>Unfortunately, New York voter files do not contain a column for self-reported race. Furthermore, exit polls are unavailable in such a small jurisdiction. Thus, our analysis cannot validate each methodological combination against a known truth. Still, we argue a comparison of BISG and CVAP in a small jurisdiction with low turnout can begin to inform whether the methods yield divergent results in different contexts, and enable us to begin exploring the mechanisms behind these divergences.</p> <hd id="AN0184233955-7">Results</hd> <p>We sought to determine whether usage of CVAP or BISG in EI results in a similar estimation of RPV and whether recent improvements to BISG had an effect. We did so by examining two elections: The 2018 Georgia gubernatorial election, and the 2016 East Ramapo, NY School Board election. First, we quantified the accuracy of CVAP and BISG in estimating precinct-level racial composition. Specifically, we evaluated how these racial composition estimates compared to the self-reported data provided by the Georgia voter file. Next, we conducted EI to detect RPV using different combinations of EI method and racial composition input for the 2018 Georgia gubernatorial election. Subsequently, we explored whether CVAP and BISG yield divergent conclusions about RPV under different turnout conditions by (a) conducting EI separately within each county in Georgia (some of which have higher turnout than others) and (b) measuring RPV in the non-partisan, low-turnout 2016 election of the ERCSB.</p> <p>Throughout our results, we present the performance of three different implementations of BISG. First, we apply BISG as has been widely used in recent years ([<reflink idref="bib28" id="ref61">28</reflink>]), including new surname data supplements ([<reflink idref="bib50" id="ref62">50</reflink>]). Next, we present results using the "fully-Bayesian version of BISG (FBISG), where priors are specified on the probability distribution of race/ethnicities to account for measurement error within geographies ([<reflink idref="bib29" id="ref63">29</reflink>]). Finally, we include a third version of BISG which adds voters' first names. In Appendix C in the supplemental material we replicated the result that for individual classification, FBISG attains a higher AUC than BISG, and adding first names further improves AUC.</p> <hd id="AN0184233955-8">Evaluating Racial Composition Estimates</hd> <p>We examined the accuracy of BISG and CVAP in estimating precinct-level racial composition. Specifically, we applied BISG and CVAP to the 2018 Georgia voter file to obtain racial composition estimates at the precinct level, and compared these to the "known truth" as determined by the self-reported race (see the "Methods section). Figure 1 compares BISG and CVAP estimates ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi></math> </ephtml> -axes) against the known truth ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math> </ephtml> -axes) at the precinct level. This comparison is depicted for the five racial groups we used: White, Black, Hispanic/Latino, Asian/Pacific Islander (API), and other. Each point represents a precinct, with points lying close to the identity lines (Figure 1: red dashed lines) denoting precincts whose racial composition are precisely estimated by the estimation approach. Meanwhile, points below the line indicate precincts where racial composition is under-estimated, and points above the line indicate precincts where racial composition is over-estimated, relative to the self-reported race.</p> <p>Graph: Figure 1. Comparison of BISG and CVAP estimates of racial composition (y-axes) to self-reported racial composition (x-axes), per racial group. BISG was calculated using 2010 decennial Census data at the block level. CVAP estimates were calculated using extensive aerial-weighted interpolation. Each point denotes a precinct in Georgia (point radius scales with precinct size). Dash red lines denote the identity lines, where the prediction matches the observation fraction. BISG and CVAP approaches are compared in pairwise subplots, for each racial group. BISG = Bayesian Improved Surname Geocoding; CVAP = Citizen Voting Age Population.</p> <p>All three BISG variations appear more tightly coupled to the 45-degree line than CVAP, whose estimates are notably more noisy. At the same time, BISG appears to slightly underestimate the share of White voters in precincts that are majority White (second panel from the left, top row), and overestimate the share of black voters in neighborhoods that have few White voters (second panel from the left, second row). In other words, BISG appears to understate the level of homogeneity in highly White precincts. This underestimation in homogeneous precincts is particularly evident in the estimates provided by the fully Bayesian implementations of BISG (third and fourth columns). By contrast, CVAP and the simpler version of BISG have less obvious patterns of bias.</p> <p>Meanwhile, the CVAP and BISG methods tend to slightly over-estimate the share of Hispanic, Asian, and Other voters in a precinct relative to observed fraction (bottom three rows). In particular, for Hispanic voters, the estimates tend to fall above the identity line. We hypothesize this is due more to the tendency for Georgia Hispanics to select "Other" or not report their race as opposed to BISG over-estimating these groups' population shares. As newer immigrant-based groups to the state with a large White-Black population, their racial identity may be less socially codified than it is for White and Black voters, and so these groups are disproportionately likely to select "Other," "White," or not report their race relative to their Census classification. Appendix E in the supplemental material examines the pattern in greater detail, showing that BISG estimates are almost surely more representative than self-reported race.</p> <p>To formally measure error rates across the two comparisons, we computed a series of error metrics. First, we computed the overall Brier score for each method. The Brier score is a metric for evaluating overall performance at predicting proportional compositions ([<reflink idref="bib40" id="ref64">40</reflink>]); [<reflink idref="bib47" id="ref65">47</reflink>])). For each precinct <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi></math> </ephtml> , we computed the Brier score <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>b</mi><mi>j</mi></msub></math> </ephtml> as <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>b</mi><mi>j</mi></msub><mo>=</mo><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mn>5</mn></munderover><mo stretchy="false">(</mo><msub><mrow><mover><mrow><mspace width=".1em" /><mi>p</mi></mrow><mo stretchy="false">^</mo></mover></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>−</mo><msub><mi>p</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><msup><mo stretchy="false">)</mo><mn>2</mn></msup></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi></math> </ephtml> indexes the five race categories, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>p</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> denotes the self-reported composition of racial group <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi></math> </ephtml> in precinct <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mrow><mspace width=".1em" /><mi>p</mi></mrow><mo stretchy="false">^</mo></mover></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub></math> </ephtml> is the estimated racial composition. We then calculated an <emph>overall</emph> Brier score <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>b</mi></math> </ephtml> for each method by averaging the individual Brier scores across precincts, weighted by the log-voter population of the precincts: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>b</mi><mo>=</mo><mrow><mfrac><mrow><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>N</mi></munderover><msub><mi>b</mi><mi>j</mi></msub><mi>log</mi><mspace width="0.2em" /><msub><mi>n</mi><mi>j</mi></msub></mrow><mrow><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>N</mi></munderover><mi>log</mi><mspace width="0.2em" /><msub><mi>n</mi><mi>j</mi></msub></mrow></mfrac></mrow></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>n</mi><mi>j</mi></msub></math> </ephtml> is the population of the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi></math> </ephtml> th precinct, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>N</mi></math> </ephtml> is the total number of precincts. We weighted the contribution of each precinct to the overall Brier score by the log of the number of voters in each precinct, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>log</mi><mspace width="0.2em" /><msub><mi>n</mi><mi>j</mi></msub></math> </ephtml> , to account for how precinct turnout spans multiple orders of magnitude. Alongside the overall Brier score, we also computed the root-mean-squared error (RMSE) and bias for each race group across precincts, again weighting by <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>log</mi><mspace width="0.2em" /><msub><mi>n</mi><mi>j</mi></msub></math> </ephtml> .</p> <p>Table 1 presents the overall Brier score, RMSE, and bias for each approach. The basic BISG implementation has a Brier score 80% lower than using CVAP and has lower RMSE for estimating the electorate share of all five race/ethnic groups. BISG also attains lower bias for all but the "Other' racial/ethnic category. These results demonstrate that BISG is universally preferable to using CVAP for estimating precinct racial composition.</p> <p>Table 1. Error Statistics Comparing Racial Composition Estimates.</p> <p>Graph</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /><col align="center" /></colgroup><thead><tr><th align="left" /><th align="center">Brier</th><th align="center" colspan="2">White</th><th align="center" colspan="2">Black</th><th align="center" colspan="2">Hispanic</th><th align="center" colspan="2">Asian</th><th align="center" colspan="2">Other</th></tr><tr><th align="center" /><th align="center">score</th><th align="center">RMSE</th><th align="center">Bias</th><th align="center">RMSE</th><th align="center">Bias</th><th align="center">RMSE</th><th align="center">Bias</th><th align="center">RMSE</th><th align="center">Bias</th><th align="center">RMSE</th><th align="center">Bias</th></tr></thead><tbody><tr><td>CVAP</td><td>0.020</td><td>0.097</td><td>-0.030</td><td>0.092</td><td>0.003</td><td>0.032</td><td>0.017</td><td>0.024</td><td>0.008</td><td>0.016</td><td><bold>0.001</bold></td></tr><tr><td>BISG</td><td><bold>0.004</bold></td><td><bold>0.047</bold></td><td><bold>-0.022</bold></td><td>0.043</td><td><bold>0.011</bold></td><td>0.016</td><td>0.005</td><td>0.009</td><td><bold>0.001</bold></td><td><bold>0.010</bold></td><td>0.004</td></tr><tr><td>FBISG</td><td>0.013</td><td>0.090</td><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo xmlns="">−</mo><mn xmlns="">0.067</mn></math></p></td><td>0.059</td><td>0.037</td><td>0.013</td><td>0.006</td><td>0.011</td><td>0.005</td><td>0.027</td><td>0.020</td></tr><tr><td>FN-FBISG</td><td>0.005</td><td>0.060</td><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo xmlns="">−</mo><mn xmlns="">0.040</mn></math></p></td><td><bold>0.041</bold></td><td>0.020</td><td><bold>0.009</bold></td><td><bold>0.005</bold></td><td><bold>0.008</bold></td><td>0.003</td><td>0.021</td><td>0.015</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>: RMSE refers to root-mean-square error, and bias refers to the average difference between true and predicted race proportions. Both metrics are averaged across precincts, weighted by the log population. CVAP estimates are based on the 2014–2018 ACS (See Appendix D in the supplemental material for details). Bold values are the smallest in each column. RMSE = root-mean-squared error; ACS = American Community Survey; CVAP = Citizen Voting Age Population; BISG = Bayesian Improved Surname Geocoding; FBISG = fully Bayesian version of BISG.</p> <p>It is less clear, however, which implementation of BISG to use in which context. FBISG has a Brier score three times larger than BISG and exhibits relatively large biases for estimating White and Black electorate share. These biases are partially rectified by incorporating first-name information. Still, the results raise the question of whether a fully Bayesian approach is always preferable for estimating racial composition. This result is consistent with (Decter-Frain, 2022), who found that machine learning methods can always outperform the standard BISG at individual prediction, but not always at predicting group composition. We return to this issue in the discussion.</p> <hd id="AN0184233955-9">EI Analysis</hd> <p>Next, we turn to the downstream impact of using different measures of racial composition when quantifying RPV with EI. We first present a Georgia state-wide analysis that compares EI results to a "ground-truth" measure of voter preferences from exit polls. Subsequently, we conduct sub-state analyses in Georgia and East Ramapo, which illustrate some circumstances under which different race estimation methods may yield different EI results.</p> <p>We used the eiCompare package to conduct iterative EI and RxC EI ([<reflink idref="bib10" id="ref66">10</reflink>]). This package wraps around eiPack ([<reflink idref="bib36" id="ref67">36</reflink>]) and WRU ([<reflink idref="bib32" id="ref68">32</reflink>]), and facilitates easy comparison of different EI methods ([<reflink idref="bib5" id="ref69">5</reflink>]; [<reflink idref="bib37" id="ref70">37</reflink>]; [<reflink idref="bib53" id="ref71">53</reflink>]; [<reflink idref="bib52" id="ref72">52</reflink>]).</p> <hd id="AN0184233955-10">Georgia Statewide</hd> <p>We begin with a state-wide analysis in Georgia. We conducted EI analyses using all racial estimates available to us: BISG, CVAP, and the voter file data (self-reported race). We computed estimates using both iterated King's EI and EI:RxC estimates. Since we lacked the known truth of voter preference by race, we compared our results against exit poll data collected by the Washington Post ([<reflink idref="bib44" id="ref73">44</reflink>]), which is based on the responses of 3,984 Georgia voters. Since exit polls are known to have their own inaccuracies, particularly among minority voters that comprise a small subset of the population ([<reflink idref="bib6" id="ref74">6</reflink>]; [<reflink idref="bib43" id="ref75">43</reflink>]), we used the exit poll data as a benchmark rather than a known truth.</p> <p>Figure 2 presents our EI results estimating support for Stacey Abrams by race, method, and data source. The broad set of results are clear: for both White and Black voters, BISG, CVAP, and the self-reported race produce consistent results, indicating the presence of RPV. Specifically, according to the EI models, Black voters back Abrams somewhere between 82% and 98% while White voters back Abrams between 19% and 30%. Not only do all three demographic data sources reveal that Black voters support Abrams and White voters do not, the results are consistent with the exit poll (Black voters support Abrams 93%, whereas 25% of White voters support her). Furthermore, both King's EI and EI:RxC produce very similar results. Together, these results demonstrate that BISG and CVAP will result in similar EI estimates in a setting with large turnout and a large number of precincts. Furthermore, despite the differences in performance across racial composition methods shown above, all four methods yield mostly similar estimates of RPV. This result is consistent with previous work ([<reflink idref="bib5" id="ref76">5</reflink>]).</p> <p>Graph: Figure 2. Estimated support for Democratic candidate Stacey Abrams in the 2018 Georgia gubernatorial elections by race. Estimates generated using different combinations of EI method (Iterative and RxC), and racial composition prediction (Known, BISG, and CVAP). Error bars are 95% credible intervals. Exit poll results based on a 3,984 voter sample from the Washington Post, with confidence intervals computed as p∓1.96p(1−p)N. RXC = rows by columns; BISG = Bayesian Improved Surname Geocoding; CVAP = Citizen Voting Age Population.</p> <p>Figure 2 also reveals that both EI methods substantially overestimated support for Stacy Abrams among Hispanic and "Other" voters, relative to the exit poll data. Notably, the Hispanic and "Other" categories of voters make up 3% and 4% of the electorate, respectively. These discrepancies indicate that current EI methods struggle to capture the voting preferences of very small minority groups, although they may also result from inaccuracies in exit polling for small minority group voters ([<reflink idref="bib6" id="ref77">6</reflink>]). In Appendix J in the supplemental material, we conduct the same analysis combining together all non-White race/ethnic groups. When combining minority groups, we find that all methods perform well. In general, we recommend collapsing together small minority groups when performing EI.</p> <p>Practitioners often seek to account for differential turnout when using CVAP for EI by adding a third "no-vote category as a way to correct for differential turnout. We redid the analysis in this section making this correction and present the results in Appendix K in the supplemental material. We find that this adjustment has little effect on the differences between BISG and CVAP.</p> <hd id="AN0184233955-11">Intra-County Ecological Inference</hd> <p>The prior results suggest that we may expect to see similar voting preferences estimated by EI regardless of method and/or whether the inputs come from BISG and CVAP. However, Georgia's 2018 gubernatorial election was unique in three important ways. First, it was an unusually high-turnout election. Second, it was a highly partisan election, with candidates aligning closely to national political parties. Third, the state of Georgia is large and composed of many precincts, unlike many voting rights cases which can focus on single counties, school districts, and other small jurisdictions. The following two sections explore whether EI remains robust to the researcher's choice of racial composition measure across heterogeneity in these circumstances.</p> <p>First, to evaluate the robustness of EI results across population size, number of precincts, and turnout, we ran EI <emph>within</emph> each county in Georgia. Of the 159 counties in Georgia, we removed those with fewer than 5 precincts. With few precincts, the sampling procedures underlying EI become unstable and can sometimes fail ([<reflink idref="bib10" id="ref78">10</reflink>]). This left 102 counties.</p> <p>We do not have exit polls against which to compare at the county level. However, we can still explore the conditions associated with divergent results across methodologies. For each county and race, we calculated the difference in predicted vote share for Stacey Abrams as estimated using BISG and CVAP to generate race inputs to EI.</p> <p>In Figure 3 we show how this difference varies as a function of county size and county turnout. Each point in these figures represents a county, and the size of the point represents the number of precincts it contains. On the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi></math> </ephtml> -axis of each panel, we plot the difference between estimated vote share for Stacy Abrams as measured using BISG versus CVAP. For example, a point with a <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi></math> </ephtml> -axis value of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mn>0.25</mn></math> </ephtml> is a county where, when conducting EI, usage of BISG to measure racial composition yields an estimated support for Abrams that is 25 percentage points higher than when CVAP is used. On the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math> </ephtml> -axis we plot turnout as a share of CVAP.</p> <p>Graph: Figure 3. Difference in estimated support for Stacey Abrams between BISG and CVAP inputs as a function of turnout. Each point represents a county, with size corresponding to the number of precincts in each county. Counties with fewer than five precincts are excluded from this analysis. Turnout was measured as the number of votes in each county divided by the CVAP of that county. Black lines are ordinary least squares regression lines, and shaded regions denote 95% confidence intervals.</p> <p>Regarding the role of county size, most of the counties with large numbers of precincts lie quite close to zero on the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi></math> </ephtml> -axis. This suggests that EI results are more robust to the choice between CVAP and BISG when applied in regions with many subdivisions.[<reflink idref="bib13" id="ref79">13</reflink>] In other words, highly populated counties with more precincts produce similar results whether demographic inputs to the EI model are developed via CVAP or BISG.</p> <p>Figure 3 further indicates the importance of turnout. We find a positive relationship between BISG-CVAP divergence and turnout for White voters, and a negative relationship for Black voters. In the lowest-turnout counties ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo>≈</mo><mn>20</mn><mi mathvariant="normal">%</mi></math> </ephtml> turnout), estimated support for Abrams among Black voters was projected to be approximately <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mn>20</mn></math> </ephtml> to <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mn>25</mn></math> </ephtml> percentage points higher when using BISG to estimate racial composition than when using CVAP.</p> <p>Meanwhile, in higher-turnout counties, BISG and CVAP yield similar results for Black voters. For White voters, BISG estimates less support for Abrams in low-turnout counties, and greater support in high-turnout counties. Taken together, these turnout-related patterns suggest the BISG identifies greater levels of RPV and CVAP in low-turnout contexts. The reasoning is simple: BISG only includes those who turned out to vote in its racial composition estimates, whereas CVAP incorporates voters and non-voters alike. Using CVAP to predict racial composition essentially assumes that the share of turned out voters from each race/ethnic group matches their CVAP population share. Typically, in lower turnout contexts, White turnout will be higher than minority turnout, which could lead EI estimates based on CVAP to incorrectly attribute support for the White-preferred candidate to black voters. Consider the example of a precinct where Black voters strongly prefer Abrams and White voters strongly prefer Kemp, but Black voters turn out at a rate lower than their share of the CVAP population. In this case, using CVAP as a race input will lead to EI estimates that incorrectly attribute some of the precinct's support for Kemp to Black voters who did not actually turn out, thereby underestimating the RPV.</p> <p>In Appendix I in the supplemental material, we tested this differential turnout mechanism graphically by plotting the same <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi></math> </ephtml> -axis variable against the difference between each race/ethnicity group's turnout share and their population share. We found that indeed, in Georgia, CVAP yields lower estimates of Black support for Abrams in counties where Black voters turnout out at rates below their share of the CVAP population. This result provides evidence that through the mechanism of differential turnout, CVAP may underestimate RPV.</p> <p>Finally, we observe no meaningful relationships between turnout and BISG-CVAP difference for the small minority groups, Hispanic and Other. Nonetheless, we suspect that in jurisdictions and states with larger shares of Hispanic voters and voters and voters of other minority populations, the results we observe here for Black voters may hold for these other groups as well.</p> <hd id="AN0184233955-12">East Ramapo Central School District: Morales v. Weissmandl</hd> <p>Thus far, our analysis has focused on the Georgia 2018 Gubernatorial election at the statewide and county levels. To supplement these analyses, we considered a separate election characterized by low turnout and a small number of precincts. Specifically, we studied the May 2016 election for school board in ERCSD. We focus on an at-large seat contested by Natasha Morales (who reputedly represented the public school bloc) and Yehuda Weissmandl (who reputedly represented the private school bloc). Given the racial demographics of East Ramapo, these blocs may also have corresponded to voters' racial identity. ERCSD, therefore, represents a realistic scenario in which plaintiffs might bring VRA section-2 cases. Since New York State does not provide race on the voter file, this election serves as a useful test case for elucidating differences in EI output given a CVAP or BISG input.</p> <p>At the time of the election, ERCSD consisted of ten precincts that exhibit a large amount of segregation between White and Black constituents. To quantify the degree of segregation, we calculated the district's Black–White dissimilarity index. We obtained a dissimilarity index of 70.59, which indicates a high degree of segregation. To contextualize this index, if ERCSD were a city, it would be the fourth most segregated in New York state, after New York City, Buffalo, and Long Beach ([<reflink idref="bib7" id="ref80">7</reflink>]). Thus, given our Georgia analysis and results demonstrating that BISG classification error correlates with precinct diversity, we can hypothesize that BISG will accurately estimate the Black-White population at the aggregate (precinct) level.</p> <p>We computed the racial composition of the 10 precincts in ERCSD using CVAP and the different BISG implementations. Due to the small number of precincts, we restricted our analysis to Black and White voters. Figure 4 highlights precincts that exhibit large differences in the racial composition as estimated by the different methods. With the exception of precincts six and ten, CVAP estimates a lower White share of the electorate than any of the BISG methods. This discrepancy is most apparent in precinct five, where CVAP estimate of the White electorate share is 37 percentage points below BISG. Given the low turnout of the election, it is likely that CVAP data under-estimates the share of White voters and overestimates the share of Black voters. These errors could have adverse effects in a downstream EI analysis depending on these inputs.</p> <p>Graph: Figure 4. Predicted White share of voters using different measures of precinct racial composition across the ten precincts in the East Ramapo School District.</p> <p>As in Georgia, we performed EI using both iterative EI and RxC EI, with CVAP and the three BISG inputs. We report the estimated support for Morales across methods and inputs in Figure 5 (we omit the estimated support for Weissmandl given that he was the only other candidate). We found that both BISG and CVAP inputs result in EI estimates indicating the presence of RPV, with White voters largely backing Weissmandl and Black voters largely backing Morales. Specifically, both EI approaches, with either BISG or CVAP inputs, indicated that roughly 25% of White voters backed Morales (and thus 75% backed Weissmandl) (Figure 5: left). Meanwhile, both EI approaches, with either input, indicated that more than 50% of Black voters supported Morales. However, the choice of racial composition input influences the degree to which RPV is estimated. Usage of any BISG method as an input to the EI approach results in estimates of 75%–95% of Black voters backing Morales (depending on the EI procedure: Figure 5, orange points). On the other hand, usage of CVAP brings the estimate closer to 50%–60% of Black voters backing Morales. Importantly, the error bars on the CVAP estimates span the majority mark (50%), indicating that one could reasonably conclude a lack of sufficient evidence for RPV when using CVAP in this particular contest. Thus, in an election with low turnout, usage of CVAP versus BISG results in manifestly distinct estimates, informing the confidence in which one could assert the presence of RPV.</p> <p>Graph: Figure 5. Estimated support for Natasha Morales for the at-large seat in the May 2016 ERCSD School Board election. Subplots denote the voting preference by racial group, with White voters on the left and Black voters on the right. The y-axis denotes the proportion of each racial group supporting Morales (with Yehuda Weissmandl, the other candidate for the seat, implicitly receiving the other share). Estimates with each subplot are separated by racial estimation input, and are redundantly coded by color. The choice of EI method (iterative EI vs. RxC EI) is denoted by the symbol shape. Error bars denote 95% confidence intervals.</p> <p>Since we lacked ground truth, or an approximation to it via an exit poll, we cannot definitively claim which of BISG or CVAP results in superior EI estimates. However, we note that the observations in Figure 5 agree with the earlier results presented in Figure 3. Specifically, BISG provided higher estimates of the Black voting bloc backing Stacey Abrams relative to CVAP when turnout is low. Similarly, we observed that usage of BISG results in a higher estimate of the Black voting bloc backing Morales relative to CVAP. Together, these results concretely demonstrate that BISG and CVAP exhibit markedly different results when turnout is low. Given that low-turnout elections are the exact scenario in which we expect CVAP to exhibit errors, BISG may be preferable in such settings.</p> <hd id="AN0184233955-13">Discussion</hd> <p>This paper examined whether the type of racial demographic inputs into EI models influence conclusions drawn about candidate preference by racial group. To our knowledge, this is the first paper to compare BISG and CVAP methods as racial/ethnic data inputs into EI vote-choice models. We compiled the requisite data to examine this question in Georgia because the state has self-reported race on the voter file and an exit poll indicating candidate preference by race in 2018. In addition, we examined candidate preference trends by race in a small New York jurisdiction that did not have self-reported race on the voter file or an exit poll, but did represent a jurisdiction where our two data sources might diverge in candidate preference by race.</p> <p>In Georgia, we conducted four analyses. First, we compared BISG and CVAP race-share at the precinct level against self-reported race, focusing primarily on Blacks and Whites. We made two main discoveries. First, for measuring precinct racial composition, BISG always outperforms CVAP on traditional performance metrics. While both CVAP and BISG appear well-calibrated and unbiased, BISG is much less noisy. Second, the new, more complex fully Bayesian BISG model does not appear to improve estimates of precinct racial composition. In this case, the more standard BISG performs better in terms of RMSE and bias.</p> <p>It remains unclear why BISG outperforms FBISG in this context, since FBISG is meant to correct for measurement error that should be present in the inputs to both models ([<reflink idref="bib29" id="ref81">29</reflink>]). This result is consistent with other work showing performance at individual race/ethnicity classification does not always translate to predicting racial composition of groups or neighborhoods ([<reflink idref="bib12" id="ref82">12</reflink>]). We believe this is ultimately a question for future research, but offer some hypotheses here.</p> <p>The task of estimating aggregate composition has been called "quantification' or "prevalence estimation in machine learning literature (see [<reflink idref="bib21" id="ref83">21</reflink>]) for a review). According to this literature, prevalence estimation by simple aggregation may fail when training and test data are drawn from different distributions. This does directly explain why pooling across neighborhoods, as in FBISG, would harm the method's performance. Another explanation could be that the zero-counts over which FBISG smoothed do not occur due to measurement error, but are instead true zeros. Smoothing over true zeros may bias precinct aggregate estimates away from real extreme values, which does seem to match the pattern we observe in Figure 1.</p> <p>Although we observed that BISG outperformed CVAP in Georgia, we believe practitioners should still proceed with caution when deciding which method to use. We maintain this position because the Georgia voter file was used to generate the surname, first name, and middle name measures that supplement the Census surname lists in the latest implementation of BISG ([<reflink idref="bib29" id="ref84">29</reflink>]). Therefore, this evaluation of BISG takes place <emph>within-sample</emph> when an <emph>out-of-sample</emph> evaluation would provide additional validity. When applying BISG in states that were not used to supplement the data (the majority of all states), BISG may not perform differently than it did here. At the same time, [<reflink idref="bib12" id="ref85">12</reflink>]) performed an "out-of-state' validation and found BISG still performs relatively well.</p> <p>For our second analysis, we used the two mainstay EI algorithms (King's EI and EI:RxC), to estimate candidate preference for Stacey Abrams (Democrat) and Brian Kemp (Republican) among each race/ethnic group. Overall, both data inputs and methods produced the same outcome: Black voters preferred Stacey Abrams and White voters preferred Brian Kemp, which are consistent with an externally available exit poll. Thus, in large, high-turnout elections, conclusions about RPV appear robust to researcher decisions about which EI method and which race/ethnicity measurement method to use.</p> <p>Our third and fourth analyses explored the robustness of different approaches across jurisdictions of varying size and electoral engagement. We conducted EI within each county in Georgia for the 2018 Gubernatorial election and in a separate, small, school board election. In both cases we found that in low turnout elections, BISG yields more precise estimates of RPV than does CVAP. We hypothesize that this result is driven by differential turnout across race/ethnicities, and found suggestive evidence in support of this mechanism.</p> <p>Overall, this paper demonstrates that at the statewide level with reasonably high turnout, BISG and CVAP will likely perform similarly: either method will accurately identify RPV to the extent that it is present in the jurisdiction. However, our results cast some doubt on the robustness of CVAP-based RPV analysis for smaller, low turnout settings. While we lack exit poll data required to make definitive conclusions, we expect that in these types of elections BISG is preferable over CVAP for measuring racial composition, especially in jurisdictions characterized by high precinct homogeneity and differential turnout by race.</p> <p>We highlight three areas of future research. First, future research should consider the problem of estimating precinct racial composition within the framework of the prevalence estimation literature ([<reflink idref="bib21" id="ref86">21</reflink>]). Prevalence estimation is a distinct task from individual classification, and we observed here that simply aggregating individual probabilities does not always yield performance in proportion to the individual level. Accurately measuring minority vote share across precincts and districts is also important for identifying instances of partisan gerrymandering ([<reflink idref="bib3" id="ref87">3</reflink>]).</p> <p>Second, using CVAP to estimate voter preference in low turnout elections with a small number of precincts may produce results seemingly inconsistent with campaign dynamics. Future work could definitively test the phenomenon we uncovered by fielding exit polls in small jurisdictions, then conducting separate EI analyses with both BISG and CVAP data.</p> <p>Third, in states where voters do not self-report race, the racial data that serve as inputs to EI models are always estimates, either via BISG or via a spatial join from some type of Census data. In the latter, the source data of the spatial join come from CVAP data available at the Census block-group level. These data are themselves estimates. No work to date has developed error propagation methods to incorporate uncertainty emerging from the racial data input side of the EI equation. Understanding the standard error for EI by carrying errors forward from data collection can help researchers better understand the certainty of EI derived estimates—especially for datasets with large numbers of precincts. Moreover, this added sophistication will inform the validity this analysis in other regions/states where race is not self-reported in the voter file.</p> <hd id="AN0184233955-14">Acknowledgments</hd> <p>The authors wish to thank Perry Grossman and Jesse Barber from NYCLU, Chad Dunn from the UCLA VRP, members of the ACLU data science team, and the organizers and participants of the 2020 Data Science for Social Good program hosted by the University of Washington eScience Institute, and participants at the Fall 2020 Politics of Race, Immigration, and Ethnicity Consortium meeting for helpful feedback throughout. Portions of this data were used in a Section 2 voting rights lawsuit against East Ramapo Central School District. All data from the lawsuit is publicly available and the case is now resolved in favor of the NAACP plaintiffs. In particular, we are indebted to Russell Mangas, of Latham & Watkins who served as attorney for NAACP and handled portions of the case related to RPV and BISG. Both Mr. Mangas and Mr. Grossman were critical components of the NAACP plaintiffs legal team and commended for their commitment to justice and social science.</p> <ref id="AN0184233955-15"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref18" type="bt">1</bibl> <bibtext> The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref32" type="bt">2</bibl> <bibtext> The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref87" type="bt">3</bibl> <bibtext> Ari Decter-Frain https://orcid.org/0000-0001-9635-3334 Pratik Sachdeva https://orcid.org/0000-0002-6809-2437 Loren Collingwood https://orcid.org/0000-0002-4447-8204 Hikari Murayama https://orcid.org/0000-0002-4067-4734 Juandalyn Burke https://orcid.org/0000-0002-6345-7505 Scott Henderson https://orcid.org/0000-0003-0624-4965 Spencer Wood https://orcid.org/0000-0002-5794-2619 Joshua Zingher https://orcid.org/0000-0003-3928-4269</bibtext> </blist> <blist> <bibl id="bib4" idref="ref6" type="bt">4</bibl> <bibtext> The code and data to replicate the analyses in this paper are available online. To access these data, please request access at https://osf.io/e6awg. We have kept the replication repository private to protect the personal information of voters. Furthermore, within the repository we do not include the voter registration file for the East Ramapo School District because we received it as part of discovery during the legal proceedings for the case, and are bound not to share the personal information of voters. The repository contains the Georgia voter registration data, aggregate data for East Ramapo, and all code for the paper.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref29" type="bt">5</bibl> <bibtext> Supplemental material and Appendix for this article are available online.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref40" type="bt">6</bibl> <bibtext> For all EI analyses, we omit voters whose self-reported race is "unknown". This "unknown" classification does not fall neatly into the BISG approach, since it is unclear how surname or geographic location might be predictive of a voter not knowing or choosing not to self-report their race. However, these voters can still be given a probabilistic rating of race, since we have their surnames and residential locations. Thus, when estimating racial composition by geographic unit using the BISG estimates, we could include or omit these voters. Ultimately, this depends on the distribution of voters who do not self-report race. If we expect that distribution to be the same as the general population, then the inclusion or omission of these voters in the BISG aggregation will have no impact on the end results. We show this in Appendix A in the supplemental material.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref44" type="bt">7</bibl> <bibtext> BISG requires both surname and geographic databases to estimate the posterior probability of an individual's race. As of this writing, the 2010 Census database is the most recent information available concerning surname probabilities in the U.S. The geographic database, however, can come from multiple sources. If we used the 2010 decennial Census, we have access to nearly complete data available at the level of Census blocks. However, this data is 10 years old, and thus may not reflect demographic shifts in the following decade. An alternative data source is the ACS, which is reported every year. The downside to the ACS data, however, is that it is only a sample of the population (roughly a 2% sampling rate) and is only reported at the block group level. Thus, what these estimates gain in temporal accuracy may be affected by a lack of spatial resolution and coverage.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref43" type="bt">8</bibl> <bibtext> All analyses in this paper incorporate surname, geography, and first name to predict individual-level race.[28]) show that including party identification/registration can marginally improve individual race prediction. However, we do not include party for comparison purposes since not all of our datasets include party registration and many local elections in which section II cases are litigated do not have aparty on the voter file.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref15" type="bt">9</bibl> <bibtext> https://openprecincts.org/</bibtext> </blist> <blist> <bibtext> However, there are 8 precincts with 0 votes, which are dropped from the analysis.</bibtext> </blist> <blist> <bibtext> https://geo.fcc.gov/api/census/</bibtext> </blist> <blist> <bibtext> https://<ulink href="http://www.ercsd.org/Page/292">www.ercsd.org/Page/292</ulink></bibtext> </blist> <blist> <bibtext> The number of precincts in a county is highly correlated with the size of its electorate ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>r</mi><mo>=</mo>0.935</math> </ephtml> ).</bibtext> </blist> <blist> <bibtext> https://<ulink href="http://www.census.gov/programs-surveys/decennial-census/about/voting-rights/cvap.2018.html">www.census.gov/programs-surveys/decennial-census/about/voting-rights/cvap.2018.html</ulink></bibtext> </blist> <blist> <bibtext> According to data compiled by the University of Michigan's Population Studies Center, available at https://<ulink href="http://www.psc.isr.umich.edu/dis/census/segregation2010.html">www.psc.isr.umich.edu/dis/census/segregation2010.html</ulink></bibtext> </blist> <blist> <bibtext> https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi{\special{t4ht@@}\%\special{t4ht@@}}3A10.7910/DVN/ZSBZ7K</bibtext> </blist> <blist> <bibtext> [4]</bibtext> </blist> </ref> <ref id="AN0184233955-16"> <title> References </title> <blist> <bibtext> Abott C., Magazinnik A. 2020. " At-large Elections and Minority Representation in Local Government." American Journal of Political Science. 64(3):717-33.</bibtext> </blist> <blist> <bibtext> Adjaye-Gbewonyo D., Bednarczyk R. A., Davis R. L., Omer S. B. 2014. " Using the Bayesian Improved Surname Geocoding Method (bisg) to Create a Working Classification of Race and Ethnicity in a Diverse Managed Care Population: A Validation Study." Health Services Research. 49(1):268-83.</bibtext> </blist> <blist> <bibtext> Ansolabehere S., Palmer M. 2016. " A Two-Hundred Year Statistical History of the Gerrymander." Ohio State Law Journal. 4:741-762</bibtext> </blist> <blist> <bibtext> Barreto M., Cohen M., Collingwood L., Dunn C.W., Waknin S. 2022. " A Novel Method for Showing Racially Polarized Voting: Bayesian Improved Surname Geocoding." New York University Review of Law & Social Change. 46:1.</bibtext> </blist> <blist> <bibtext> Barreto M., Collingwood L., Garcia-Rios S. and K. A. Oskooii. 2019. " Estimating Candidate Support in Voting Rights Act Cases: Comparing Iterative EI and EI-Rx C Methods." Sociological Methods & Research0049124119852394.</bibtext> </blist> <blist> <bibtext> Barreto M. A., Guerra F., Marks M., Nuño S. A., Woods N. D. 2006. " Controversies in Exit Polling: Implementing a Racially Stratified Homogenous Precinct Approach." PS: Political Science & Politics. 39(3):477-83.</bibtext> </blist> <blist> <bibtext> CensusScope. 2011. Segregation: Dissimilarity indices.</bibtext> </blist> <blist> <bibtext> Clark J. T., Curiel J. A., Steelman T. S. 2021, November. " Minmaxing of Bayesian Improved Surname Geocoding and Geography Level Ups in Predicting Race." Political Analysis. 57(3):731-62.</bibtext> </blist> <blist> <bibtext> Collingwood L., Long S.2019. " Can States Promote Minority Representation? Assessing the Effects of the California Voting Rights Act." Urban Affairs Review: 92-101.</bibtext> </blist> <blist> <bibtext> Collingwood L., Oskooii K., Garcia-Rios S., Barreto M. 2016. " eicompare: Comparing Ecological Inference Estimates Across Ei and Ei: Rc." The R journal. 8(2):92.</bibtext> </blist> <blist> <bibtext> de Benedictis-Kessner J. 2015. " Evidence in Voting Rights Act Litigation: Producing Accurate Estimates of Racial Voting Patterns." Election Law Journal. 14(4):361-81.</bibtext> </blist> <blist> <bibtext> Decter-Frain A. 2022. "How Should we Proxy for Race/Ethnicity? Comparing Bayesian Improved Surname Ggeocoding to Machine Learning Methods." Publisher: arXiv Version Number: 2.</bibtext> </blist> <blist> <bibtext> Durkin E. 2018, Oct. Gop candidate improperly purged 340,000 from georgia voter rolls, investigation claims.</bibtext> </blist> <blist> <bibtext> Elliott M. N., Fremont A., Morrison P. A., Pantoja P., Lurie N. 2008. " A New Method for Estimating Race/ethnicity and Associated Disparities Where Administrative Records Lack Self-reported Race/ethnicity." Health Services Research. 43(5p1):1722-36.</bibtext> </blist> <blist> <bibtext> Elliott M. N., Morrison P. A., Fremont A., McCaffrey D. F., Pantoja P. and N. Lurie.2009. " Using the Census Bureau's Surname List to Improve Estimates of Race/ethnicity and Associated Disparities." Health Services and Outcomes Research Methodology. 9(2):69-83.</bibtext> </blist> <blist> <bibtext> Enos R. D. 2016. " What the Demolition of Public Housing Teaches Us About the Impact of Racial Threat on Political Behavior." American Journal of Political Science. 60(1):123-42.</bibtext> </blist> <blist> <bibtext> Enos R. D., Kaufman A. R., Sands M. L. 2019. " Can Violent Protest Change Local Policy Support? Evidence From the Aftermath of the 1992 Los Angeles Riot." American Political Science Review. 113(4):1012-28.</bibtext> </blist> <blist> <bibtext> Ferree K. E. 2004. " Iterative Approaches to R × C Ecological Inference Problems: Where they Can Go Wrong and One Quick Fix." Political Analysis. 12(2):143-59.</bibtext> </blist> <blist> <bibtext> Fraga B. L. 2016. " Candidates Or Districts? Reevaluating the Role of Race in Voter Turnout." American Journal of Political Science. 60(1):97-122.</bibtext> </blist> <blist> <bibtext> Fraga B. L. 2018. The Turnout Gap: Race, Ethnicity, and Political Inequality in a Diversifying AmericaCambridge University Press.</bibtext> </blist> <blist> <bibtext> González P., Castaño A., Chawla N. V., Coz J. J. D.2017, sep. " A Review on Quantification Learning." ACM Computing Surveys. 50(5):1-40.</bibtext> </blist> <blist> <bibtext> Goodman L. A. 1953. " Ecological Regressions and Behavior of Individuals." American Sociological Review. 18:663-664. https://doi.org/10.2307/2088121</bibtext> </blist> <blist> <bibtext> Goodman L. A. 1959. " Some Alternatives to Ecological Correlation." American Journal of Sociology. 64(6):610-25.</bibtext> </blist> <blist> <bibtext> Grofman B., Barreto M. 2009. " A Reply to Zax's (2002) Critique of Grofman and Migalski (1988) Double-equation Approaches to Ecological Inference when the Independent Variable is Misspecified." Sociological Methods & Research. 37(4):599-617.</bibtext> </blist> <blist> <bibtext> Grumbach J. M., Sahn A. 2020. " Race and Representation in Campaign Finance." American Political Science Review. 114(1):206-21.</bibtext> </blist> <blist> <bibtext> Hajnal Z., Trounstine J. 2005. " Where Turnout Matters: The Consequences of Uneven Turnout in City Politics." The Journal of Politics. 67(2):515-35.</bibtext> </blist> <blist> <bibtext> Hajnal Z. L., Lewis P. G. 2003. " Municipal Institutions and Voter Turnout in Local Elections." Urban Affairs Review. 38(5):645-68.</bibtext> </blist> <blist> <bibtext> Imai K., Khanna K. 2016. " Improving Ecological Inference by Predicting Individual Ethnicity From Voter Registration Records." Political Analysis. 24:263-72.</bibtext> </blist> <blist> <bibtext> Imai K., Olivella S., Rosenman E. T. R. 2022. "Addressing Census Data Problems in Race Imputation via Fully Bayesian Improved Surname Geocoding and Name Supplements." Publisher: arXiv Version Number: 3.</bibtext> </blist> <blist> <bibtext> Kahle D., Wickham H. 2013. " ggmap: Spatial Visualization with Ggplot2." The R Journal. 5(1):144-61.</bibtext> </blist> <blist> <bibtext> Kenny C. T., Kuriwaki S., McCartan C., Rosenman E. T. R., Simko T., Imai K. 2021, October. " The Use of Differential Privacy for Census Data and Its Impact on Redistricting: The Case of the 2020 U.S. Census." Science Advances. 7(41):eabk3283.</bibtext> </blist> <blist> <bibtext> Khanna K., Imai K. 2020. wru: Who are You? Bayesian Prediction of Racial Category Using Surname and Geolocation. R package version 0.1-10.</bibtext> </blist> <blist> <bibtext> King G. 1997. A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data. Princeton University Press.</bibtext> </blist> <blist> <bibtext> King G., Rosen O., Tanner M. A. 2004. "Information in Ecological Inference: An Introduction." Pp. 1-12. in Ecological Inference: New Methodological Strategies, Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Kousser J. M. 2001. " Ecological Inference From Goodman to King." Historical Methods: A Journal of Quantitative and Interdisciplinary History. 34(3):101-26.</bibtext> </blist> <blist> <bibtext> Lau O., Moore R. T., Kellermann M. 2007. " eipack: Rx C Ecological Inference and Higher-dimension Data Management." New Functions for Multivariate Analysis. 7(1):43.</bibtext> </blist> <blist> <bibtext> Leung V. 2022. " Asian American Candidate Preferences: Evidence from California." Political Behavior 44:1759-1788. https://doi.org/10.1007/s11109-020-09673-8</bibtext> </blist> <blist> <bibtext> Massey D. S., Denton N. 1988. " The Dimensions of Residential Segregation." Social Forces. 67(2):281-315.</bibtext> </blist> <blist> <bibtext> McCrary P. 1990. " Racially Polarized Voting in the South: Quantitative Evidence from the Courtroom." Social Science History. 14(04):507-31.</bibtext> </blist> <blist> <bibtext> Murphy A. H. 1973. " A New Vector Partition of the Probability Score." Journal of Applied Meteorology and Climatology. 12(4):595-600.</bibtext> </blist> <blist> <bibtext> Oskooii K. A. R. 2020, July. " Perceived Discrimination and Political Behavior." British Journal of Political Science. 50(3):867-92. Publisher: Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Pebesma E. 2018. " Simple Features for R: Standardized Support for Spatial Vector Data." The R Journal. 10(1):439-46.</bibtext> </blist> <blist> <bibtext> Pedraza F., Barreto M.2000. 6.3. 1 exit polls and ethnic diversity: How to improve estimates and reduce bias among minority voters.</bibtext> </blist> <blist> <bibtext> Post T. W. 2018, Nov. Georgia 2018 voter poll results.</bibtext> </blist> <blist> <bibtext> Prener C., Fox B. 2020. censusxy: Access the U.S. Census Bureau's Geocoding A.P.I. System. R package version 1.0.0.</bibtext> </blist> <blist> <bibtext> Prener C., Revord C.2019. " areal: An R Package for Areal Weighted Interpolation." Journal of Open Source Software. 4(37):1221.</bibtext> </blist> <blist> <bibtext> Redelmeier D. A., Bloch D. A., Hickam D. H. 1991. " Assessing Predictive Accuracy: How to Compare Brier Scores." Journal of Clinical Epidemiology. 44(11):1141-6.</bibtext> </blist> <blist> <bibtext> Reny T., Wilcox-Archuleta B., Nichols V. C. 2018, December. " Threat, Mobilization, and Latino Voting in the 2018 Election." The Forum. 16(4):573-99. Publisher: De Gruyter.</bibtext> </blist> <blist> <bibtext> Rosen O., Jiang W., King G., Tanner M. A. 2001. " Bayesian and Frequentist Inference for Ecological Inference: The Rx C Case." Statistica Neerlandica. 55(2):134-56.</bibtext> </blist> <blist> <bibtext> Rosenman E. T. R., Olivella S., Imai K. 2022. Race and Ethnicity Data for First, Middle, and Last Names. Publisher: arXiv Version Number: 1.</bibtext> </blist> <blist> <bibtext> Ross M. M. 1993. " The Voting Rights Act." The Urban Lawyer925-34.</bibtext> </blist> <blist> <bibtext> Sadhwani S. 2021. " The Influence of Candidate Race and Ethnicity: The Case of Asian Americans." Politics, Groups, and Identities. 10(4):1-37.</bibtext> </blist> <blist> <bibtext> Sadhwani S. and M. Mendez. 2018. " Candidate Ethnicity and Latino Voting in Co-partisan Elections." California Journal of Politics and Policy. 10(2):1–17.</bibtext> </blist> <blist> <bibtext> Salmon M. 2020. opencage: Interface to the OpenCage API.</bibtext> </blist> <blist> <bibtext> Shah P. 2014, June. " It Takes a Black Candidate: A Supply-Side Theory of Minority Representation." Political Research Quarterly. 67(2):266-79. Publisher: SAGE Publications Inc.</bibtext> </blist> <blist> <bibtext> Shah P. R., Davis N. R. 2017. " Comparing Three Methods of Measuring Race/ethnicity." The Journal of Race, Ethnicity, and Politics. 2(1):124-39.</bibtext> </blist> <blist> <bibtext> Trounstine J.M. E. Valdini.2008. " The Context Matters: The Effects of Single-member Versus At-large Districts on City Council Diversity." American Journal of Political Science. 52(3):554-69.</bibtext> </blist> <blist> <bibtext> Welch S. 1990. " The Impact of At-large Elections on the Representation of Blacks and Hispanics." The Journal of Politics. 52(4):1050-76.</bibtext> </blist> </ref> <aug> <p>By Ari Decter-Frain; Pratik Sachdeva; Loren Collingwood; Hikari Murayama; Juandalyn Burke; Matt Barreto; Scott Henderson; Spencer Wood and Joshua Zingher</p> <p>Reported by Author; Author; Author; Author; Author; Author; Author; Author; Author</p> <p></p> <p>Ari Decter-Frain is a PhD candidate in the Cornell University Brooks School of Public Policy. His research uses nontraditional data and computational methods for demographic estimation, particularly to study migration and politics in the United States. His research is funded by the Social Sciences and Humanities Research Council of Canada. In 2020, he was a Data Science for Social Good fellow at the University of Washington eScience Institute, working on the application of ecological inference statistical methodology to real world issues in the domain of voting rights research and the law.</p> <p>Pratik Sachdeva is a Senior Data Scientist in Social Sciences D-Lab at University of California, Berkeley. He leads D-Lab's special projects, provides mentorship to graduate students in the D-Lab community, and supports D-Lab's instructional and consulting programs. He has expertise in data science and machine learning, and has conducted data-driven, interdisciplinary research in a variety of fields, including computational social science, computational health sciences, hate speech, political science, physics, neuroscience, and software engineering.</p> <p>Loren Collingwood is associate professor of Political Science at University of New Mexico, where his research and teaching focuses on racial and ethnic politics, research methodology and applied statistics. He is co-author of the ecological inference software, eiCompare and has testified in Federal court on the voting rights act numerous times. He is the co-author of Sanctuary Cities: The Politics of Refuge (2019) and author of Campaigning in a Racially Diversifying America: When and How Cross-Racial Electoral Mobilization Works (2020) both with Oxford University Press. In 2020, Collingwood and Barreto led a Data Science for Social Good project in the University of Washington eScience Institute, on the application of ecological inference statistical methodology to real world issues in the domain of voting rights research and the law.</p> <p>Hikari Murayama is a PhD student in the Energy and Resources Group and University of California, Berkeley. Her research largely focuses on combining remote sensing and machine learning methods to study human impacts on the environment. She was selected to be a Fellow at the University of Washington eScience Institute's Data Science for Social Good Program in 2020 and has been a Data Science Fellow at the D-Lab also since 2020. She is currently a Research Fellow in the Global Policy Lab advised by Professor Solomon Hsiang.</p> <p>Juandalyn Burke , PhD, MPH is a recent postdoctoral fellow at Johns Hopkins University. Her work focuses on providing statistical approaches such as bayesian statistics to predictive modeling, benefit-harm models, qualitative design and database development. She joined the Data Science for Social Good program in 2020. She was selected to present her collaborative DSSG work at the Michigan Institute for Data Science. She also earned the position of Conference Director of the Learning & Doing Data for Good (LDDG) Conference which brought current and past students of Data for Social Good programs across the United States, Canada, and the United Kingdom to present their work at the eScience Institute in Seattle, WA. Her current interest and work involves using statistical approaches, modeling and analytics for biomedical education, and the healthcare industry.</p> <p>Matt Barreto is professor of Political Science and Chicana/o Studies at University of California, Los Angeles, and the faculty director of the UCLA Voting Rights Project. He is co-author of the ecological inference software, eiCompare and has testified in Federal court on the voting rights act dozens of times. His research and teaching focuses on research methodology, voting behavior, and voting rights and he teaches classes in the division of Social Sciences, School of Public Policy, and School of Law. He has published 4 books and more than 80 articles and book chapters on racial/ethnic politics, voting and research methodology. In 2020, he and Collingwood led a Data Science for Social Good project in the University of Washington eScience Institute, on the application of ecological inference statistical methodology to real world issues in the domain of voting rights research and the law.</p> <p>Scott Henderson is research scientist in the University of Washington (UW) Department of Earth and Space Sciences and Data Science Fellow at the eScience Institute. He specializes in geospatial analysis of satellite remote sensing datasets and cloud computing infrastructure for data intensive research. In 2020, he was a data science mentor for a Data Science for Social Good project in the University of Washington eScience Institute, on the application of ecological inference statistical methodology to real world issues in the domain of voting rights research and the law.</p> <p>Spencer Wood is a Research Scientist and Data Science Fellow in the University of Washington eScience Institute, and Affiliate Faculty in the School for Environmental and Forest Sciences. He is an interdisciplinary scientist with broad interests and experience in ecology, sustainability, computer science, statistics, and economics. His expertise lies in using data, technology, and quantitative methods to improve policy and management decisions. In 2020, he was a data science mentor for a Data Science for Social Good project in the University of Washington eScience Institute, on the application of ecological inference statistical methodology to real world issues in the domain of voting rights research and the law.</p> <p>Joshua Zingher is associate professor of political science and geography at Old Dominion University. His research focuses on mass political behavior, representation, and public opinion. His research has appeared in journals like the British Journal of Political Science, Journal of Politics, and Political Research Quarterly. His first book, Political Choice in a Polarized America is available from Oxford University Press.</p> </aug> <nolink nlid="nl1" bibid="bib14" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib15" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib31" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib28" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib34" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib20" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib27" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib26" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib41" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib48" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib55" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib19" firstref="ref13"></nolink> <nolink nlid="nl13" bibid="bib57" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib58" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib51" firstref="ref19"></nolink> <nolink nlid="nl16" bibid="bib22" firstref="ref20"></nolink> <nolink nlid="nl17" bibid="bib23" firstref="ref21"></nolink> <nolink nlid="nl18" bibid="bib35" firstref="ref22"></nolink> <nolink nlid="nl19" bibid="bib39" firstref="ref23"></nolink> <nolink nlid="nl20" bibid="bib33" firstref="ref24"></nolink> <nolink nlid="nl21" bibid="bib49" firstref="ref25"></nolink> <nolink nlid="nl22" bibid="bib18" firstref="ref26"></nolink> <nolink nlid="nl23" bibid="bib24" firstref="ref27"></nolink> <nolink nlid="nl24" bibid="bib11" firstref="ref28"></nolink> <nolink nlid="nl25" bibid="bib16" firstref="ref34"></nolink> <nolink nlid="nl26" bibid="bib17" firstref="ref36"></nolink> <nolink nlid="nl27" bibid="bib56" firstref="ref37"></nolink> <nolink nlid="nl28" bibid="bib25" firstref="ref38"></nolink> <nolink nlid="nl29" bibid="bib13" firstref="ref39"></nolink> <nolink nlid="nl30" bibid="bib45" firstref="ref41"></nolink> <nolink nlid="nl31" bibid="bib54" firstref="ref42"></nolink> <nolink nlid="nl32" bibid="bib10" firstref="ref52"></nolink> <nolink nlid="nl33" bibid="bib42" firstref="ref53"></nolink> <nolink nlid="nl34" bibid="bib36" firstref="ref56"></nolink> <nolink nlid="nl35" bibid="bib44" firstref="ref57"></nolink> <nolink nlid="nl36" bibid="bib30" firstref="ref58"></nolink> <nolink nlid="nl37" bibid="bib12" firstref="ref60"></nolink> <nolink nlid="nl38" bibid="bib50" firstref="ref62"></nolink> <nolink nlid="nl39" bibid="bib29" firstref="ref63"></nolink> <nolink nlid="nl40" bibid="bib40" firstref="ref64"></nolink> <nolink nlid="nl41" bibid="bib47" firstref="ref65"></nolink> <nolink nlid="nl42" bibid="bib32" firstref="ref68"></nolink> <nolink nlid="nl43" bibid="bib37" firstref="ref70"></nolink> <nolink nlid="nl44" bibid="bib53" firstref="ref71"></nolink> <nolink nlid="nl45" bibid="bib52" firstref="ref72"></nolink> <nolink nlid="nl46" bibid="bib43" firstref="ref75"></nolink> <nolink nlid="nl47" bibid="bib21" firstref="ref83"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1473565
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Comparing Methods for Estimating Demographics in Racially Polarized Voting Analyses
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Ari+Decter-Frain%22">Ari Decter-Frain</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9635-3334">0000-0001-9635-3334</externalLink>)<br /><searchLink fieldCode="AR" term="%22Pratik+Sachdeva%22">Pratik Sachdeva</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6809-2437">0000-0002-6809-2437</externalLink>)<br /><searchLink fieldCode="AR" term="%22Loren+Collingwood%22">Loren Collingwood</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4447-8204">0000-0002-4447-8204</externalLink>)<br /><searchLink fieldCode="AR" term="%22Hikari+Murayama%22">Hikari Murayama</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4067-4734">0000-0002-4067-4734</externalLink>)<br /><searchLink fieldCode="AR" term="%22Juandalyn+Burke%22">Juandalyn Burke</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6345-7505">0000-0002-6345-7505</externalLink>)<br /><searchLink fieldCode="AR" term="%22Matt+Barreto%22">Matt Barreto</searchLink><br /><searchLink fieldCode="AR" term="%22Scott+Henderson%22">Scott Henderson</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0624-4965">0000-0003-0624-4965</externalLink>)<br /><searchLink fieldCode="AR" term="%22Spencer+Wood%22">Spencer Wood</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-5794-2619">0000-0002-5794-2619</externalLink>)<br /><searchLink fieldCode="AR" term="%22Joshua+Zingher%22">Joshua Zingher</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3928-4269">0000-0003-3928-4269</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. 2025 54(2):706-738.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 33
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Voting%22">Voting</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Racial+Composition%22">Racial Composition</searchLink><br /><searchLink fieldCode="DE" term="%22Bayesian+Statistics%22">Bayesian Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Inference%22">Statistical Inference</searchLink><br /><searchLink fieldCode="DE" term="%22Elections%22">Elections</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+Analysis%22">Statistical Analysis</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Georgia%22">Georgia</searchLink><br /><searchLink fieldCode="DE" term="%22New+York%22">New York</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/00491241231192383
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241<br />1552-8294
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We consider the cascading effects of researcher decisions throughout the process of quantifying racially polarized voting (RPV). We contrast three methods of estimating precinct racial composition, Bayesian Improved Surname Geocoding (BISG), fully Bayesian BISG, and Citizen Voting Age Population (CVAP), and two algorithms for performing ecological inference (EI), King's EI and EI:RxC using eiCompare. Using data from two different elections we identify circumstances in which different combinations of methods produce divergent results, comparing against ground-truth data where available. We first find that BISG outperforms CVAP at estimating racial composition, though fully Bayesian BISG does not yield further improvements. Next, in a statewide election, we find that all combinations of methods yield similarly reliable estimates of RPV. However, county-level analyses and results from a non-partisan school board election reveal that BISG and CVAP produce divergent estimates of Black preferences in elections with low turnout and few precincts. Our results suggest that methodological choices can meaningfully alter conclusions about RPV, particularly in smaller, low-turnout elections.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1473565
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1473565
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/00491241231192383
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 33
        StartPage: 706
    Subjects:
      – SubjectFull: Voting
        Type: general
      – SubjectFull: Computation
        Type: general
      – SubjectFull: Racial Composition
        Type: general
      – SubjectFull: Bayesian Statistics
        Type: general
      – SubjectFull: Statistical Inference
        Type: general
      – SubjectFull: Elections
        Type: general
      – SubjectFull: Statistical Analysis
        Type: general
      – SubjectFull: Georgia
        Type: general
      – SubjectFull: New York
        Type: general
    Titles:
      – TitleFull: Comparing Methods for Estimating Demographics in Racially Polarized Voting Analyses
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Ari Decter-Frain
      – PersonEntity:
          Name:
            NameFull: Pratik Sachdeva
      – PersonEntity:
          Name:
            NameFull: Loren Collingwood
      – PersonEntity:
          Name:
            NameFull: Hikari Murayama
      – PersonEntity:
          Name:
            NameFull: Juandalyn Burke
      – PersonEntity:
          Name:
            NameFull: Matt Barreto
      – PersonEntity:
          Name:
            NameFull: Scott Henderson
      – PersonEntity:
          Name:
            NameFull: Spencer Wood
      – PersonEntity:
          Name:
            NameFull: Joshua Zingher
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
            – Type: issn-electronic
              Value: 1552-8294
          Numbering:
            – Type: volume
              Value: 54
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1