Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach

Saved in:
Bibliographic Details
Title: Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach
Language: English
Authors: Moretti, Angelo (ORCID 0000-0001-6543-9418), Whitworth, Adam
Source: Sociological Methods & Research. 2023 52(4):1785-1815.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 31
Publication Date: 2023
Document Type: Journal Articles
Reports - Evaluative
Descriptors: Simulation, Geometric Concepts, Computation, Measurement, Error of Measurement, Bias, Income, Municipalities, Foreign Countries
Geographic Terms: Italy
DOI: 10.1177/0049124120986199
ISSN: 0049-1241
1552-8294
Abstract: Spatial microsimulation encompasses a range of alternative methodological approaches for the small area estimation (SAE) of target population parameters from sample survey data down to target small areas in contexts where such data are desired but not otherwise available. Although widely used, an enduring limitation of spatial microsimulation SAE approaches is their current inability to deliver reliable measures of uncertainty--and hence confidence intervals--around the small area estimates produced. In this article, we overcome this key limitation via the development of a measure of uncertainty that takes into account both variance and bias, that is, the mean squared error. This new approach is evaluated via a simulation study and demonstrated in a practical application using European Union Statistics on Income and Living Conditions data to explore income levels across Italian municipalities. Evaluations show that the approach proposed delivers accurate estimates of uncertainty and is robust to nonnormal distributions. The approach provides a significant development to widely used spatial microsimulation SAE techniques.
Abstractor: As Provided
Entry Date: 2023
Accession Number: EJ1397535
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFY4YMt2SZikhizFlVLdoMjAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDPfbH22Awmr1aIlZtwIBEICBmzBNk9BdYZ5I7uczAlBAHJPjJYgMN-VCGNnrOU0dqjJb90lxo3dOkLAamePeX2AXA1jMly78qmHFoZ9U8GqeucjTAutRQIjwFan2XwgXUunJAoGQrHiv0gym0_mUeko8O1818_ZXm-kd-KrQn9aKLJ7DyrGUMAMhypBGp4uvmItqTGMahfp0f5aG7b3izUHKS6CJIXwrSRw13eJq
Text:
  Availability: 1
  Value: <anid>AN0173122298;som01nov.23;2023Oct25.06:20;v2.2.500</anid> <title id="AN0173122298-1">Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach </title> <p>Spatial microsimulation encompasses a range of alternative methodological approaches for the small area estimation (SAE) of target population parameters from sample survey data down to target small areas in contexts where such data are desired but not otherwise available. Although widely used, an enduring limitation of spatial microsimulation SAE approaches is their current inability to deliver reliable measures of uncertainty—and hence confidence intervals—around the small area estimates produced. In this article, we overcome this key limitation via the development of a measure of uncertainty that takes into account both variance and bias, that is, the mean squared error. This new approach is evaluated via a simulation study and demonstrated in a practical application using European Union Statistics on Income and Living Conditions data to explore income levels across Italian municipalities. Evaluations show that the approach proposed delivers accurate estimates of uncertainty and is robust to nonnormal distributions. The approach provides a significant development to widely used spatial microsimulation SAE techniques.</p> <p>Keywords: calibration; weighting; synthetic; indirect estimator; raking; resampling</p> <p>Large-scale surveys are designed to obtain reliable estimates at national level or, in some instances, for large subnational scales such as regions. These can be considered to be the planned domains of the survey sampling design ([<reflink idref="bib6" id="ref1">6</reflink>]). However, there is a growing demand from both research and policy communities for various local estimates at more detailed spatial resolutions such as municipalities or neighborhoods due to the absence of data at such small area scales from existing census or administrative sources. However, this small area desire frequently encounters a problem of unplanned domains, given that for cost reasons such small areas typically have small or zero sample sizes in the survey sampling design. In these circumstances, commonly used direct estimators such as the Horvitz–Thompson estimator ([<reflink idref="bib29" id="ref2">29</reflink>]) that only use sample survey information either cannot be used (in the case of zero sample size domains) or provide unacceptably large variability in the estimates to be practically useful (in the case of small sample size domains).</p> <p>In such scenarios, indirect small area estimation (SAE) of target population parameters has become a relatively widely used and increasingly demanded methodological technique via a range of SAE approaches. We refer to [<reflink idref="bib49" id="ref3">49</reflink>], [<reflink idref="bib60" id="ref4">60</reflink>], [<reflink idref="bib47" id="ref5">47</reflink>], and [<reflink idref="bib38" id="ref6">38</reflink>] for useful methodological reviews on both regression-based and microsimulation-based SAE methods. Spatial microsimulation approaches, sometimes referred to as survey calibration approaches ([<reflink idref="bib20" id="ref7">20</reflink>]), represent a family of reweighting approaches to SAE in which the challenge is to reweight the survey units such that they optimally fit the demographic and socioeconomic profile of each small area according to a selected set of benchmark constraints. Part of the appeal of spatial microsimulation approaches to SAE for both researchers and policy users is their intuitive and accessible appeal without much of the complex statistical expertise required within many regression-based SAE methods, particularly as assumptions fail or more complex outcomes are desired. In those circumstances, one advantage of spatial microsimulation approaches over regression-based SAE estimators is that they tend to be more robust to failures in model assumptions, given that as model-assisted estimators it is only necessary that the population be reasonably well described by an assumed model for that model to be valid ([<reflink idref="bib51" id="ref8">51</reflink>]).</p> <p>Spatial microsimulation SAE approaches have been used to produce small area estimates across a range of policy areas including child malnutrition ([<reflink idref="bib31" id="ref9">31</reflink>]), obesity ([<reflink idref="bib18" id="ref10">18</reflink>]), fuel poverty ([<reflink idref="bib45" id="ref11">45</reflink>]), income and poverty ([<reflink idref="bib5" id="ref12">5</reflink>]; [<reflink idref="bib46" id="ref13">46</reflink>]; [<reflink idref="bib64" id="ref14">64</reflink>]), regional planning ([<reflink idref="bib11" id="ref15">11</reflink>]), participation in sport ([<reflink idref="bib30" id="ref16">30</reflink>]), and transport ([<reflink idref="bib36" id="ref17">36</reflink>]; [<reflink idref="bib50" id="ref18">50</reflink>]; [<reflink idref="bib58" id="ref19">58</reflink>]). Spatial microsimulation SAE approaches have been well validated against known external data and against alternative regression-based SAE techniques ([<reflink idref="bib41" id="ref20">41</reflink>]; [<reflink idref="bib57" id="ref21">57</reflink>]; [<reflink idref="bib61" id="ref22">61</reflink>]). Spatial microsimulation SAE has also been used effectively to assess the spatial impacts of differing "what if" policy scenarios ([<reflink idref="bib10" id="ref23">10</reflink>]; [<reflink idref="bib13" id="ref24">13</reflink>]; [<reflink idref="bib55" id="ref25">55</reflink>]; [<reflink idref="bib63" id="ref26">63</reflink>]). For example, [<reflink idref="bib7" id="ref27">7</reflink>] introduce the SimAlba spatial microsimulation model Scotland in order to estimate the simulated impact of various policy scenarios on individual's health outcomes. In an Australian context, [<reflink idref="bib55" id="ref28">55</reflink>] use spatial microsimulation SAE to assess the geographical impacts on expected elderly poverty levels from changes to the state pension ([<reflink idref="bib56" id="ref29">56</reflink>]).</p> <p>However, despite these widespread applications and contributions, an important long-standing limitation of spatial microsimulation SAE approaches is their continued inability to deliver estimates of uncertainty around central point estimates at small area level. More specifically, one faces something of a trade-off with any SAE approach in seeking via the indirect SAE estimates to reduce variance compared to the direct estimates while acknowledging that those direct estimates are unbiased compared to the indirect SAE estimates that are inherently biased. As such, it is essential when calculating the uncertainty of any SAE estimator that the mean squared error (MSE) is used, given that this takes into account both variance and bias. However, calculating MSE is often challenging in an SAE content and particularly so in spatial microsimulation approaches. In the case of design-based estimators, the design weights are known before the sample selection and are therefore nonrandom meaning that analytical approximations of MSE are available ([<reflink idref="bib51" id="ref30">51</reflink>]). However, this is no longer the case in spatial microsimulation approaches to SAE when reweighing algorithms are used as the weights become random variables themselves. In these scenarios, analytical approximation of MSE becomes highly challenging, given that bias and variance cannot be computed in closed form such that empirically based resampling techniques are instead required to estimate their uncertainty ([<reflink idref="bib9" id="ref31">9</reflink>]). Some analytical approximations have been suggested in the literature ([<reflink idref="bib14" id="ref32">14</reflink>]; [<reflink idref="bib16" id="ref33">16</reflink>]), but [<reflink idref="bib9" id="ref34">9</reflink>] highlight important practical challenges around the requirement for joint selection probabilities that are rarely computed or known in practice.</p> <p>Empirical attempts to estimate the uncertainty around spatial microsimulation SAE estimates have been proposed in recent years ([<reflink idref="bib9" id="ref35">9</reflink>]; [<reflink idref="bib44" id="ref36">44</reflink>]; [<reflink idref="bib62" id="ref37">62</reflink>]), though none entirely successful. In response, this article develops a novel modified parametric bootstrap technique in order to estimate the uncertainty of spatial microsimulation small area estimates based on the MSE such that it captures both bias and variance components in the uncertainty estimate, unlike previous attempts. Our approach benefits from clear statistical properties under a linear model and can be used flexibly across alternative spatial microsimulation SAE techniques.</p> <p>The remainder of this article is structured as follows. In the second section, the general problem of SAE of the population mean using iterative proportional fitting (IPF) is outlined. In the third section, the challenge of uncertainty estimation in SAE contexts is further described, and our approach to MSE estimation via bootstrap is detailed. In the fourth section, the bootstrap approach is evaluated via a simulation study, and in the fifth section, a practical data application focusing on Italian Statistics on Income and Living Conditions (SILC) data is presented to illustrate the approach. The sixth section provides a concluding discussion focusing on wider implications for spatial microsimulation SAE and potential next steps for research.</p> <hd id="AN0173122298-2">SAE With IPF</hd> <p>This section sets out the SAE problem of the population mean and formally introducing the IPF spatial microsimulation approach used in the later development of our proposed uncertainty estimator.</p> <hd id="AN0173122298-3">The General SAE Problem for a Small Area Mean</hd> <p>Let us consider a sample <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>s</mi><mo>⊂</mo><mi>Ω</mi></mrow></math> </ephtml> of size <emph>n</emph> drawn from the target finite population <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>Ω</mi></math> </ephtml> of size <emph>N</emph>. Let <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> denote the small areas for which we want to compute the small area estimates. <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>N</mi><mo>−</mo><mi>n</mi></mrow></math> </ephtml> are the nonsampled units and these are denoted by <emph>r</emph>, hence <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>s</mi><mi>d</mi></msub><mo>=</mo><mi>s</mi><mo>∩</mo><msub><mi>Ω</mi><mi>d</mi></msub></mrow></math> </ephtml> is the subsample from the small area <emph>d</emph> of size <emph>n<subs>d</subs></emph>, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>n</mi><mo>=</mo><mstyle displaystyle="true"><msubsup><mo>∑</mo><mrow><mi>d</mi><mo>=</mo><mn>1</mn></mrow><mi>D</mi></msubsup><mrow><msub><mi>n</mi><mi>d</mi></msub></mrow></mstyle></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>s</mi><mo>=</mo><msub><mo>∪</mo><mi>d</mi></msub><msub><mi>s</mi><mi>d</mi></msub></mrow></math> </ephtml> . <emph>r<subs>d</subs></emph> denotes the nonsampled units in small area <emph>d</emph> with <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>N</mi><mi>d</mi></msub><mo>−</mo><msub><mi>n</mi><mi>d</mi></msub></mrow></math> </ephtml> dimension. Here, the target parameter is the population mean <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub><mo>=</mo><msubsup><mi>N</mi><mi>d</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mstyle displaystyle="true"><msubsup><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mrow><msub><mi>N</mi><mi>d</mi></msub></mrow></msubsup><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></mstyle></mrow></math> </ephtml> of the variable <emph>Y</emph> for area <emph>d</emph>, with <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> denoting the value of variable <emph>Y</emph> for <emph>i</emph>th unit from <emph>d</emph>th area.</p> <p>Due to the unplanned domain problem, <emph>n<subs>d</subs></emph> may be too small (even zero) for many small areas in the survey data to compute reliable direct estimates of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub></mrow></math> </ephtml> from the survey data based on <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup><mo>=</mo><msup><mrow><mfenced><mrow><mstyle displaystyle="true"><msub><mo>∑</mo><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></msub><mrow><msub><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></mstyle></mrow></mfenced></mrow><mrow><mo>−</mo><mn>1</mn></mrow></msup><mstyle displaystyle="true"><msub><mo>∑</mo><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></msub><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><msub><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></mstyle></mrow></math> </ephtml> , where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> denotes the design weight for <emph>i</emph>th unit from <emph>d</emph>th area in <emph>s<subs>d</subs></emph>. The direct survey estimate comes from a standard direct estimator and is based on sample survey information only. They are weighted averages where the weights are the design weight based on the complex survey design. As a consequence, there is a need in such circumstances to consider indirect SAE estimation techniques using auxiliary information if one wishes small area estimates either at all (in the case of zero survey sample sizes) or with reduced uncertainty (in the case of low small area survey sample sizes). There is a bias-variance trade-off in operation when doing so such that any reductions in variability must naturally be balanced with acknowledgment of increased bias in the indirect SAE estimates compared to the unbiased direct estimates ([<reflink idref="bib47" id="ref38">47</reflink>]; [<reflink idref="bib49" id="ref39">49</reflink>]).</p> <hd id="AN0173122298-4">Reweighing Using the IPF Algorithm</hd> <p>IPF is one of three main spatial microsimulation approaches to SAE—IPF, generalized regression reweighting and combinatorial optimization. To demonstrate our approach to the estimation of uncertainty in spatial microsimulation SAE, we focus on the IPF algorithm, given that this is both widely used and a constructively challenging test of our MSE estimator, given that its iterative nature renders the survey weights random variables themselves ([<reflink idref="bib9" id="ref40">9</reflink>]; [<reflink idref="bib48" id="ref41">48</reflink>]; [<reflink idref="bib53" id="ref42">53</reflink>]).</p> <p>Like all spatial microsimulation SAE approaches, IPF can be understood as a reweighting optimization problem where the aim is to reweight the survey units (e.g., individuals or households) such that they optimally fit the demographic and socioeconomic profile of each small area according to a selected set of benchmark constraints (e.g., age-sex, employment status, ethnicity, health, education). [<reflink idref="bib16" id="ref43">16</reflink>] provide a statistical theory of these reweighting techniques and alternative approaches. For each local area, the result is a tailored set of reweighted survey cases that fit to the benchmark characteristics of that small area in terms both of total population and the profile of that population across the benchmark constraints. A key data set created during IPF is a weights matrix giving new weights for each survey unit in each separate small area. For each small area, those final IPF weights show how representative each survey unit is of each area given their respective characteristics across the benchmarks. Across all survey individuals, these reweighted units sum to the small area population total and map onto its population profile across the benchmarks. As such, these reweighted data provide a valuable synthetic micropopulation for each small area that can be employed in further local analyses (e.g., "what if" policy simulations) as desired ([<reflink idref="bib2" id="ref44">2</reflink>]; [<reflink idref="bib36" id="ref45">36</reflink>]).</p> <p>Formally, IPF can be understood as follows. Let <emph>w<subs>i</subs></emph> be the initial weight (usually the survey design weight) for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>i</mi><mo>∈</mo><mi>s</mi></mrow></math> </ephtml> . The calibration problem is area-specific and therefore generates new weights denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mo>*</mo></msubsup></mrow></math> </ephtml> for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></math> </ephtml> for area <emph>d</emph> that satisfy the calibration equation given by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mstyle displaystyle="true"><msub><mo>∑</mo><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></msub><mrow><msubsup><mi>w</mi><mi>i</mi><mo>*</mo></msubsup><msub><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mi>i</mi></msub></mrow></mstyle><mo>=</mo><mstyle displaystyle="true"><msub><mo>∑</mo><mrow><mi>i</mi><mo>∈</mo><msub><mi>Ω</mi><mi>d</mi></msub></mrow></msub><mrow><msub><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mi>i</mi></msub><mo>=</mo><msub><mstyle mathsize="normal" mathvariant="bold"><mi>X</mi></mstyle><mi>d</mi></msub></mrow></mstyle></mrow></math> </ephtml> , where <emph>x<subs>i</subs></emph> is a vector of auxiliary variables. Here, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mo>*</mo></msubsup></mrow></math> </ephtml> minimizes a given distance function between <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfenced close="}" open="{"><mrow><msubsup><mi>w</mi><mi>i</mi><mo>*</mo></msubsup><mo>;</mo><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></mfenced></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfenced close="}" open="{"><mrow><msub><mi>w</mi><mi>i</mi></msub><mo>;</mo><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></mfenced></mrow></math> </ephtml> . IPF is the exponential case within a wider family of synthetic reweighting algorithms ([<reflink idref="bib16" id="ref46">16</reflink>]). Its constrained optimization problem is given as follows, where <emph>a<subs>i</subs></emph> denotes the initial weights, usually the design weights ([<reflink idref="bib9" id="ref47">9</reflink>]; [<reflink idref="bib16" id="ref48">16</reflink>]):</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><mtext>min</mtext><mo>:</mo><mtext /><munder><mrow><mstyle displaystyle="true"><mo>∑</mo></mstyle></mrow><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></munder><mfenced close="]" open="["><mrow><msub><mi>w</mi><mi>i</mi></msub><mtext>ln</mtext><mfenced><mrow><mfrac><mrow><msub><mi>w</mi><mi>i</mi></msub></mrow><mrow><msub><mi>a</mi><mi>i</mi></msub></mrow></mfrac></mrow></mfenced><mo>−</mo><msub><mi>w</mi><mi>i</mi></msub><mo>+</mo><msub><mi>a</mi><mi>i</mi></msub></mrow></mfenced><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><mtext>such that </mtext><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></munder><msub><mi>w</mi><mi>i</mi></msub><msub><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mi>i</mi></msub><mo>=</mo><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>∈</mo><msub><mi>Ω</mi><mi>d</mi></msub></mrow></munder><msub><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mi>i</mi></msub><mo>=</mo><msub><mstyle mathsize="normal" mathvariant="bold"><mi>X</mi></mstyle><mi>d</mi></msub><mo>.</mo></mrow></mtd></mtr></mtable></mrow></math> </ephtml> </p> <p>Graph</p> <p>As noted above, (<reflink idref="bib1" id="ref49">1</reflink>) unfortunately does not have a closed-form solution such that solution via analytical approximation is required. This is however highly challenging. The IPF algorithm is therefore employed iteratively across the benchmark constraints in order to estimate the final weights for each survey unit in order to derive a solution empirically. To describe the IPF method more fully, the terminology and notation used by [<reflink idref="bib34" id="ref50">34</reflink>] is followed:</p> <p></p> <ulist> <item> Initialize the iteration counter <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>←</mo><mn>0</mn></mrow></math> </ephtml> and the weights as <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mn>0</mn><mo>,</mo><mi>p</mi></mrow></msubsup><mo>←</mo><msub><mi>w</mi><mi>i</mi></msub></mrow></math> </ephtml> .</item> <p></p> <item> Increment the iteration counter <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>←</mo><mi>t</mi><mo>+</mo><mn>1</mn></mrow></math> </ephtml> , thus updating the weights as <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mn>0</mn></mrow></msubsup><mo>←</mo><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn><mo>,</mo><mi>p</mi></mrow></msubsup></mrow></math> </ephtml> .</item> <p></p> <item> Update the weights through each of the benchmark constraint variables in turn, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>v</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>p</mi></mrow></math> </ephtml> : <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>v</mi></mrow></msubsup><mo>=</mo><mfenced close="" open="{"><mrow><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>v</mi><mo>−</mo><mn>1</mn></mrow></msubsup><mfrac><mrow><mi>T</mi><mo stretchy="false">(</mo><msub><mstyle mathsize="normal" mathvariant="bold"><mi>X</mi></mstyle><mi>v</mi></msub><mo stretchy="false">)</mo></mrow><mrow><msub><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>l</mi><mo>∈</mo><mi>s</mi></mrow></msub><msubsup><mi>w</mi><mi>l</mi><mrow><mi>t</mi><mo>,</mo><mi>v</mi><mo>−</mo><mn>1</mn></mrow></msubsup><msub><mi>x</mi><mrow><mi>v</mi><mi>l</mi></mrow></msub></mrow></mfrac><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>x</mi><mrow><mi>v</mi><mi>i</mi></mrow></msub><mo>≠</mo><mn>0</mn></mrow></mtd></mtr><mtr><mtd><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>v</mi><mo>−</mo><mn>1</mn></mrow></msubsup><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>x</mi><mrow><mi>v</mi><mi>i</mi></mrow></msub><mo>=</mo><mn>0</mn></mrow></mtd></mtr></mtable></mrow></mfenced><mo>.</mo></mrow></math> </ephtml></item> </ulist> <p>Graph</p> <p></p> <ulist> <item> If the discrepancies between <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mstyle displaystyle="true"><msub><mo>∑</mo><mrow><mi>i</mi><mo>∈</mo><mi>s</mi></mrow></msub><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>p</mi></mrow></msubsup><msub><mi>x</mi><mi>v</mi></msub></mrow></mstyle></mrow></math> </ephtml> (i.e., the sample totals with the new weights) and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>T</mi><mfenced><mrow><msub><mstyle mathsize="normal" mathvariant="bold"><mi>X</mi></mstyle><mi>v</mi></msub></mrow></mfenced></mrow></math> </ephtml> are within a priori defined tolerance for all <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>v</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>p</mi></mrow></math> </ephtml> , then declare convergence and go to step 5, otherwise return to step 2.</item> <p></p> <item> The weights <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>p</mi></mrow></msubsup></mrow></math> </ephtml> are the final calibrated weights and are denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mi>i</mi><mrow><mi>t</mi><mo>,</mo><mi>p</mi></mrow></msubsup><mo>=</mo><msubsup><mi>w</mi><mi>i</mi><mo>*</mo></msubsup></mrow></math> </ephtml> .</item> </ulist> <p>The benchmark constraints used for the survey reweighting are usually categorical variables in real applications. Therefore,</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mi>i</mi><mtext>′</mtext></msubsup><mo>=</mo><mfenced><mrow><msubsup><mtext>δ</mtext><mrow><mn>1</mn><mi>i</mi></mrow><mrow><mfenced><mn>1</mn></mfenced></mrow></msubsup><mo>,</mo><mo>...</mo><mo>,</mo><msubsup><mtext>δ</mtext><mrow><msub><mi>F</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub></mrow><mrow><mfenced><mn>1</mn></mfenced></mrow></msubsup><mo>,</mo><msubsup><mtext>δ</mtext><mrow><mn>1</mn><mi>i</mi></mrow><mrow><mfenced><mn>2</mn></mfenced></mrow></msubsup><mo>,</mo><mo>...</mo><msubsup><mtext>δ</mtext><mrow><mn>1</mn><mi>i</mi></mrow><mrow><mfenced><mi>p</mi></mfenced></mrow></msubsup><mo>,</mo><mo>...</mo><mo>,</mo><msubsup><mtext>δ</mtext><mrow><msub><mi>F</mi><mi>p</mi></msub><mi>i</mi></mrow><mrow><mfenced><mi>p</mi></mfenced></mrow></msubsup></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>l</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>p</mi></mrow></math> </ephtml> denotes the <emph>l</emph>th benchmark constraint and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mtext>δ</mtext><mrow><mi>k</mi><mi>i</mi></mrow><mrow><mfenced><mi>l</mi></mfenced></mrow></msubsup><mo>=</mo><mn>1</mn></mrow></math> </ephtml> if <emph>I</emph> is in the category <emph>k</emph> of <emph>l</emph>th control variable. <emph>F<subs>l</subs></emph> is the number of categories of the <emph>l</emph>th benchmark constraint. [<reflink idref="bib1" id="ref51">1</reflink>] suggests that <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mo>=</mo><mn>20</mn><mo /></mrow></math> </ephtml> is sufficient as a conservative guide to the number of loops through the benchmark constraints in order to optimize the calibration to the set of benchmarks. The IPF algorithm is area-specific, and the IPF reweighting therefore needs to be iterated for each small area <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>d</mi><mo>=</mo><mo>,</mo><mn>1</mn><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> .</p> <p>The IPF estimator can therefore be defined as follows:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup><mo>=</mo><mfrac><mrow><msubsup><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup><msubsup><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow><mtext>*</mtext></msubsup><msub><mi>y</mi><mi>i</mi></msub></mrow><mrow><msubsup><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup><msubsup><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow><mtext>*</mtext></msubsup></mrow></mfrac><mo>,</mo><mtext /><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mtext /><mi>D</mi><mo>,</mo><mtext /><mi>i</mi><mo>=</mo><mo>,</mo><mn>1</mn><mo>...</mo><mo>,</mo><mi>n</mi><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>w</mi><mrow><mi>d</mi><mi>i</mi></mrow><mo>*</mo></msubsup></mrow></math> </ephtml> denotes the IPF-calibrated survey weight for unit <emph>i</emph>th from area <emph>d</emph>th. It can be noted that <emph>y<subs>i</subs></emph> appears for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>n</mi></mrow></math> </ephtml> , which means that <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> belongs to the class of small area synthetic estimators ([<reflink idref="bib49" id="ref52">49</reflink>]). Of course, in order for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> to be more efficient than <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> in terms of its combination of variance and bias as measured by the MSE, the auxiliary variables used in the calibration problem need to be related sufficiently to the target variable <emph>Y</emph>, as with all such model-based or model-assisted small area estimators ([<reflink idref="bib22" id="ref53">22</reflink>]).</p> <hd id="AN0173122298-5">Measuring the Uncertainty: the MSE Estimator of Y ¯ ^ d I P F</hd> <p>The quality of an estimate is assessed with reference both to its accuracy (bias) and to its precision (variability). It is therefore important to capture both aspects in any measure of uncertainty ([<reflink idref="bib17" id="ref54">17</reflink>]; [<reflink idref="bib54" id="ref55">54</reflink>]). The bias of an estimate can be defined as its degree to which it describes the measured phenomena correctly, in other words its difference from the true (though often unobserved) population value. In contrast, the variability of an estimate relates to how closely repeated observations confirm themselves (e.g., under random sampling). Figure 1 reproduces a visual summary of these two considerations from [<reflink idref="bib21" id="ref56">21</reflink>].</p> <p>Graph: Figure 1. Precision and accuracy in estimates.</p> <p>The MSE is the second moment about the origin of the error and thus takes into account both bias and variance. It is therefore an appropriate measure for our proposed estimation of uncertainty around spatial microsimulation SAE estimator. For the design unbiased direct survey estimates, the MSE is equal to the variance, while for the indirect SAE estimator, the MSE is equal to the bias squared plus the variance. As such, the attractiveness of any indirect SAE estimator compared to the unbiased direct estimator is dependent upon the reductions in variance in any SAE approach exceeding its increases to bias such that the MSE of the indirect SAE estimator is smaller than the MSE of the direct estimator. This is possible in SAE contexts where small areas have low or no survey sample sizes such that direct small area estimates are either nonviable or come with large variance.</p> <p>Capturing the MSE via resampling techniques such as the bootstrap is common within regression-based SAE approaches ([<reflink idref="bib27" id="ref57">27</reflink>]; [<reflink idref="bib37" id="ref58">37</reflink>]; [<reflink idref="bib40" id="ref59">40</reflink>]). Indeed, [<reflink idref="bib27" id="ref60">27</reflink>] point out that even when analytical approximations are available, bootstrap resampling might provide more accurate estimates due to its second-order accuracy, a property discussed further in [<reflink idref="bib19" id="ref61">19</reflink>]. However, no similar MSE measures have yet been considered in the spatial microsimulation SAE context where MSE expressions are not available in closed form and where empirical approaches are therefore necessary to explore ([<reflink idref="bib9" id="ref62">9</reflink>]; [<reflink idref="bib14" id="ref63">14</reflink>]). This is particularly relevant since the estimation of bias in particular has proven elusive in previous attempts ([<reflink idref="bib9" id="ref64">9</reflink>]; [<reflink idref="bib44" id="ref65">44</reflink>]; [<reflink idref="bib62" id="ref66">62</reflink>]).</p> <p>This article responds to this gap through its modification of bootstrap ideas for regression-based small area estimators set out in [<reflink idref="bib27" id="ref67">27</reflink>], so that they become suited to the differing technical processes and requirements of spatial microsimulation approaches. The use of models in the estimation of MSE of model-assisted estimators can be found in the regression estimator context where a model unbiased conditional MSE estimator is proposed ([<reflink idref="bib35" id="ref68">35</reflink>]). This provides initial motivation for this article to develop and adapt such an approach in the context of spatial microsimulation in order to estimate the MSE of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> .</p> <p>In order to provide an estimator of the MSE of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> , denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> , we assume that the observations <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> for unit <emph>i</emph> in area <emph>d</emph> are related to <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi mathvariant="bold-italic">x</mi><mrow><mi mathvariant="bold-italic">d</mi><mi mathvariant="bold-italic">i</mi></mrow></msub><mo>=</mo><msup><mrow><mfenced><mrow><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mn>1</mn></mrow></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mi>p</mi></mrow></msub></mrow></mfenced></mrow><mi>T</mi></msup></mrow></math> </ephtml> denoting a vector of <emph>p</emph> auxiliary variables, via the [<reflink idref="bib4" id="ref69">4</reflink>] linear nested-error (i.e., multilevel) regression model:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mo>=</mo><msubsup><mi mathvariant="bold-italic">x</mi><mrow><mi>d</mi><mi>i</mi></mrow><mi>T</mi></msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>+</mo><msub><mi>u</mi><mi>d</mi></msub><mo>+</mo><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mo>,</mo><mtext /><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>N</mi><mi>D</mi></msub><mo>,</mo><mtext /><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mi>d</mi></msub><mtext /><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mi>N</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mtext>σ</mtext><mi>u</mi><mn>2</mn></msubsup></mrow></mfenced><mo>,</mo><mtext /><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mi>N</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mtext>σ</mtext><mi>e</mi><mn>2</mn></msubsup></mrow></mfenced><mo>,</mo><mtext /><mi>i</mi><mi>n</mi><mi>d</mi><mi>e</mi><mi>p</mi><mi>e</mi><mi>n</mi><mi>d</mi><mi>e</mi><mi>n</mi><mi>t</mi><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <emph>u<subs>d</subs></emph> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> are the area random effect and the residual error term, respectively. As with all SAE methods, these are assumed to be independent ([<reflink idref="bib49" id="ref70">49</reflink>]). The model assumes that the population has a two-level structure where units are nested in areas. This is reasonable in the SAE context given the aim to estimate target parameters of small domains in the population. By doing so, the approach recognizes explicitly that the MSE will be based on a multilevel (two-level) structure and that the intraclass correlation (ICC) will therefore have relevance. The ICC describes the extent to which units (e.g., individuals) within the same higher level unit (e.g. areas) are similar to one another ([<reflink idref="bib33" id="ref71">33</reflink>]). To sensitivity test this issue, the simulation study below explicitly tests the effect of differing ICCs on the MSE estimator.</p> <p>Under model (<reflink idref="bib3" id="ref72">3</reflink>), our proposed estimator of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> can be derived via a parametric bootstrap by adapting the principles in [<reflink idref="bib27" id="ref73">27</reflink>] for the different challenges of the spatial microsimulation context. The algorithm steps for the bootstrap MSE for IPF are listed below for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>b</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>B</mi></mrow></math> </ephtml> bootstrap replications where the symbol * is used to denote the bootstrap quantities and for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> small areas:</p> <p></p> <ulist> <item> Fit model (<reflink idref="bib3" id="ref74">3</reflink>) to the observed sample data, denoted by <emph>s</emph>, and estimate the model parameters. The estimates are denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>^</mo></mover></math> </ephtml> , <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mtext>σ</mtext><mo>^</mo></mover><mi>u</mi><mn>2</mn></msubsup></mrow></math> </ephtml> , <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mtext>σ</mtext><mo>^</mo></mover><mi>e</mi><mn>2</mn></msubsup></mrow></math> </ephtml> ;</item> <p></p> <item> Generate the bootstrap area effects <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>u</mi><mi>d</mi><mrow><mo>*</mo><mfenced><mi>b</mi></mfenced></mrow></msubsup><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mi>N</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mover accent="true"><mtext>σ</mtext><mo>^</mo></mover><mi>u</mi><mn>2</mn></msubsup></mrow></mfenced></mrow></math> </ephtml> ;</item> <p></p> <item> Generate the bootstrap residual error term <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow><mrow><mo>*</mo><mfenced><mi>b</mi></mfenced></mrow></msubsup><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mo /><mi>N</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mover accent="true"><mtext>σ</mtext><mo>^</mo></mover><mi>e</mi><mn>2</mn></msubsup></mrow></mfenced></mrow></math> </ephtml> , independently of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>u</mi><mi>d</mi><mrow><mo>*</mo><mfenced><mi>b</mi></mfenced></mrow></msubsup></mrow></math> </ephtml> , for every unit <emph>i</emph> in the sample in area <emph>d</emph>, for the sample units, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></math> </ephtml> ;</item> <p></p> <item> Calculate the true population means for each small area of the bootstrap population as follows:</item> </ulist> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup><mo>=</mo><mtext /><msubsup><mover accent="true"><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mo>¯</mo></mover><mrow><mi>d</mi><mo>,</mo><mo /><mi>p</mi><mi>o</mi><mi>p</mi></mrow><mi>T</mi></msubsup><mover accent="true"><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>^</mo></mover><mo>+</mo><msubsup><mstyle mathsize="normal"><mi>u</mi></mstyle><mi>d</mi><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>x</mi><mo>¯</mo></mover><mrow><mi>d</mi><mo>,</mo><mo /><mi>p</mi><mi>o</mi><mi>p</mi></mrow></msub></mrow></math> </ephtml> denotes the means of the known population auxiliary variables for each area <emph>d</emph>. These may be taken, for instance, from the census or administrative data.</p> <p></p> <ulist> <item> 5. Generate the bootstrap data as follows <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>,</mo><mo /><mi>i</mi><mo>∈</mo><msub><mi>s</mi><mi>d</mi></msub></mrow></math> </ephtml> :</item> </ulist> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup><mo>=</mo><mtext /><msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mrow><mi>d</mi><mi>i</mi></mrow><mi>T</mi></msubsup><mover accent="true"><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>^</mo></mover><mo>+</mo><msubsup><mstyle mathsize="normal"><mi>u</mi></mstyle><mi>d</mi><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup><mo>+</mo><mtext /><msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>e</mi></mstyle><mrow><mi>d</mi><mi>i</mi></mrow><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>noting that (<reflink idref="bib5" id="ref75">5</reflink>) follows model (<reflink idref="bib3" id="ref76">3</reflink>).</p> <p></p> <ulist> <item> 6. Compute the IPF estimator defined in (<reflink idref="bib2" id="ref77">2</reflink>) on <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow><mrow><mo>*</mo><mfenced><mi>b</mi></mfenced></mrow></msubsup></mrow></math> </ephtml> and obtain the IPF estimates on the bootstrap data <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi><mo>*</mo><mfenced><mi>b</mi></mfenced></mrow></msubsup></mrow></math> </ephtml> ;</item> <p></p> <item> 7. Repeat steps (<reflink idref="bib2" id="ref78">2</reflink>) through (<reflink idref="bib6" id="ref79">6</reflink>) for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>b</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>B</mi></mrow></math> </ephtml> for each area <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> .</item> </ulist> <p>An estimator of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> is given by the following Monte Carlo approximation:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mover accent="true"><mrow><mi>M</mi><mi>S</mi><mi>E</mi></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msup><mi>B</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munderover><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>b</mi><mo>=</mo><mn>1</mn></mrow><mi>B</mi></munderover><msup><mrow><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup><mo>−</mo><msubsup><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi><mrow><mtext>*</mtext><mfenced><mi>b</mi></mfenced></mrow></msubsup></mrow></mfenced></mrow><mn>2</mn></msup><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <hd id="AN0173122298-6">Simulation Study</hd> <p>This section presents the findings from a simulation study to examine the performance of our proposed MSE bootstrap estimator for spatial microsimulation SAE. For a classificatory work on simulation studies in SAE and further theoretical details, we refer to [<reflink idref="bib43" id="ref80">43</reflink>]. In this model-based simulation, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>S</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>000</mn></mrow></math> </ephtml> populations are generated from model (<reflink idref="bib3" id="ref81">3</reflink>), given that estimators such as the MSE estimator depend on model assumptions and hence that it is important to evaluate the statistical properties of our proposed approach under the model. There are no significant computational barriers to the approach, and this is an important practical consideration. Using a standard modern machine, the simulation study took around 10 hours to perform and the application of municipality income in Tuscany in the fifth section took around 20 minutes to perform.</p> <p>The parameters for the simulation are selected from the LANDSAT data that are widely used in SAE simulation settings. These are survey and satellite data for corn and soybeans in 12 Iowa counties obtained from the 1978 June survey of the U.S. Department of Agriculture and from land observatory satellites (see [<reflink idref="bib4" id="ref82">4</reflink>]; [<reflink idref="bib15" id="ref83">15</reflink>]; [<reflink idref="bib40" id="ref84">40</reflink>]). The small area problem arises in these data since small area sample sizes are small. The simulations are computationally intensive in large population dimensions and are therefore controlled for the purposes of this simulation. All analyses are conducted in R, and details on the code and functions used for the bootstrap are provided in the Online Appendix (which can be found at <ulink href="http://smr.sagepub.com/supplemental/">http://smr.sagepub.com/supplemental/</ulink>).</p> <hd id="AN0173122298-7">Generating the Population</hd> <p>The population is generated using the following parameters: <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>N</mi><mo>=</mo><mn>20</mn><mo>,</mo><mn>000</mn></mrow></math> </ephtml> , <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>D</mi><mo>=</mo><mn>80</mn></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>130</mn><mo>≤</mo><msub><mi>N</mi><mi>d</mi></msub><mo>≤</mo><mn>420</mn></mrow></math> </ephtml> . <emph>N<subs>d</subs></emph>, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi></mrow></math> </ephtml> is generated from the discrete uniform distribution, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>N</mi><mi>d</mi></msub><mo>∼</mo><mi>d</mi><mi>U</mi><mi>n</mi><mi>i</mi><mi>f</mi><mfenced><mrow><mn>130</mn><mo>,</mo><mo /><mn>420</mn></mrow></mfenced></mrow></math> </ephtml> , with <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mstyle displaystyle="true"><msubsup><mo>∑</mo><mrow><mi>d</mi><mo>=</mo><mn>1</mn></mrow><mi>D</mi></msubsup><mrow><msub><mi>N</mi><mi>d</mi></msub></mrow></mstyle><mo>=</mo><mn>20</mn><mo>,</mo><mn>000</mn></mrow></math> </ephtml> . <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> observations are generated according to the following model:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mo>=</mo><msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mrow><mi>d</mi><mi>i</mi></mrow><mi>T</mi></msubsup><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>+</mo><msub><mi>u</mi><mi>d</mi></msub><mo>+</mo><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mo>,</mo><mtext /><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>N</mi><mi>D</mi></msub><mo>,</mo><mtext /><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mi>d</mi></msub><mtext /><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mi>N</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mtext>σ</mtext><mi>u</mi><mn>2</mn></msubsup></mrow></mfenced><mo>,</mo><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub><mover><mo>∼</mo><mrow><mi>i</mi><mi>i</mi><mi>d</mi></mrow></mover><mi>F</mi><mfenced><mrow><mn>0</mn><mo>,</mo><msubsup><mtext>σ</mtext><mi>e</mi><mn>2</mn></msubsup></mrow></mfenced><mo>,</mo><mtext /><mi>i</mi><mi>n</mi><mi>d</mi><mi>e</mi><mi>p</mi><mi>e</mi><mi>n</mi><mi>d</mi><mi>e</mi><mi>n</mi><mi>t</mi><mtext>,</mtext></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>F</mi><mo>∈</mo><mfenced close="}" open="{"><mrow><mi>N</mi><mi>o</mi><mi>r</mi><mi>m</mi><mi>a</mi><mi>l</mi><mo>,</mo><mo /><mi>G</mi><mi>u</mi><mi>m</mi><mi>b</mi><mi>e</mi><mi>l</mi><mo>,</mo><mi>L</mi><mi>o</mi><mi>g</mi><mi>i</mi><mi>s</mi><mi>t</mi><mi>i</mi><mi>c</mi></mrow></mfenced></mrow></math> </ephtml> . The rationale for sensitivity testing the performance of our bootstrap estimator across these three distribution types is that the MSE is based on a normality assumption. In line with good practice in previous SAE literature ([<reflink idref="bib27" id="ref85">27</reflink>]), distributions are chosen deliberately in order to sensitivity test how the estimators perform when the error term <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mrow><mi>d</mi><mi>i</mi></mrow></msub></mrow></math> </ephtml> is not normal but is instead skewed (Gumbel) or symmetric with heavy tails (Logistic). All three distribution types are common in real data applications.</p> <p>The auxiliary variables are defined as follows:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mstyle mathsize="normal" mathvariant="bold"><mi>x</mi></mstyle><mrow><mi>d</mi><mi>i</mi></mrow></msub><mo>=</mo><msup><mrow><mfenced><mrow><mn>1</mn><mtext /><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mn>1</mn></mrow></msub><mtext /><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mn>2</mn></mrow></msub></mrow></mfenced></mrow><mi>T</mi></msup></mrow><mo>,</mo><mtext> with </mtext><mrow><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mn>1</mn></mrow></msub><mo>∼</mo><mi>d</mi><mi>U</mi><mi>n</mi><mi>i</mi><mi>f</mi><mfenced><mrow><mn>145</mn><mo>,</mo><mtext /><mn>459</mn></mrow></mfenced></mrow><mtext> and </mtext><mrow><msub><mi>x</mi><mrow><mi>d</mi><mi>i</mi><mn>2</mn></mrow></msub><mo>∼</mo><mi>d</mi><mi>U</mi><mi>n</mi><mi>i</mi><mi>f</mi><mfenced><mrow><mn>55</mn><mo>,</mo><mtext /><mn>345</mn><mo>.</mo></mrow></mfenced></mrow></math> </ephtml> </p> <p>Graph</p> <p>and the regression coefficients are given by the following vector:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mstyle mathsize="normal" mathvariant="bold"><mi>β</mi></mstyle><mo>=</mo><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><msup><mrow><mfenced><mrow><mn>17.97</mn><mtext /><mn>0.36</mn><mtext /><mo>−</mo><mn>0.03</mn></mrow></mfenced></mrow><mi>T</mi></msup><mo>.</mo></mrow></mtd></mtr></mtable></mrow></math> </ephtml> </p> <p>Graph</p> <p>As noted above, since the data are assumed to take a multilevel structure (units inside areas), the ICC will have relevance and requires sensitivity testing. The ICC varies across applications dependent upon the extent to which the variability in the data observes a hierarchal structure with, for example, values ranging across 0.005, 0.05, and 0.2 with respect to mortality ([<reflink idref="bib3" id="ref86">3</reflink>]), fear of crime ([<reflink idref="bib59" id="ref87">59</reflink>]), and well-being ([<reflink idref="bib41" id="ref88">41</reflink>]), respectively. Given that the ICC plays a role in the performance of the model-based MSE estimator, it is important that sensitivity tests are performed around its value within the simulation ([<reflink idref="bib39" id="ref89">39</reflink>]; [<reflink idref="bib40" id="ref90">40</reflink>]). We use the following relationships to explore the role of the ICC in the case of the Normal distribution:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>ρ</mtext><mo>=</mo><mfrac><mrow><msubsup><mtext>σ</mtext><mi>u</mi><mn>2</mn></msubsup></mrow><mrow><msubsup><mtext>σ</mtext><mi>u</mi><mn>2</mn></msubsup><mo>+</mo><msubsup><mtext>σ</mtext><mi>e</mi><mn>2</mn></msubsup></mrow></mfrac></mrow><mtext>, </mtext><mrow><msubsup><mtext>σ</mtext><mi>u</mi><mn>2</mn></msubsup><mo>=</mo><mo>−</mo><mfrac><mtext>ρ</mtext><mrow><mtext>ρ</mtext><mo>−</mo><mn>1</mn></mrow></mfrac><msubsup><mtext>σ</mtext><mi>e</mi><mn>2</mn></msubsup></mrow><mtext> with </mtext><mrow><msubsup><mtext>σ</mtext><mi>e</mi><mn>2</mn></msubsup><mo>=</mo><mn>297.71</mn></mrow><mtext> and </mtext><mrow><mtext>ρ</mtext><mo>∈</mo><mfenced close="}" open="{"><mrow><mn>0.01</mn><mo>,</mo><mn>0.03</mn><mo>,</mo><mtext /><mn>0.05</mn><mo>,</mo><mn>0.08</mn><mo>,</mo><mtext /><mn>0.10</mn><mo>,</mo><mtext /><mn>0.15</mn><mo>,</mo><mtext /><mn>0.20</mn><mo>,</mo><mtext /><mn>0.50</mn></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Due to space constraints, in the cases of the Gumbel and Logistic distributions, the ICC is set at a realistic value of 0.05 only (see, e.g., [<reflink idref="bib41" id="ref91">41</reflink>]).</p> <p>In order to produce the small area IPF estimates, we create the following classes identifying the benchmark constraints related to the covariates <emph>x</emph><subs>1</subs> and <emph>x</emph><subs>2</subs>:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>145</mn><mo>≤</mo><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>224.20</mn><mo>,</mo><mtext /><mn>224.20</mn><mo><</mo><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>380.70</mn><mo>,</mo><mtext /><mn>380.70</mn><mo><</mo><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>459</mn><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>55</mn><mo>≤</mo><msub><mi>x</mi><mrow><mn>2</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>126.30</mn><mo>,</mo><mtext /><mn>126.30</mn><mo><</mo><msub><mi>x</mi><mrow><mn>2</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>272.10</mn><mo>,</mo><mtext /><mn>272.10</mn><mo><</mo><msub><mi>x</mi><mrow><mn>2</mn><mi>i</mi></mrow></msub><mo>≤</mo><mn>345</mn><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <hd id="AN0173122298-8">Simulation Steps</hd> <p>The simulation consists of the following steps:</p> <p></p> <ulist> <item> Population generation: Generate the responses <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi><mi>s</mi></mrow></msub></mrow></math> </ephtml> according to model (<reflink idref="bib7" id="ref92">7</reflink>) for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>s</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>S</mi></mrow></math> </ephtml> , ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>S</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>000</mn></mrow></math> </ephtml> ) with parameters presented above;</item> <p></p> <item> Draw a stratified random sample with simple random sample without replacement selection in each area <emph>d</emph> from each simulated population, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>n</mi><mi>d</mi></msub><mo>∼</mo><mi>d</mi><mi>U</mi><mi>n</mi><mi>i</mi><mi>f</mi><mfenced><mrow><mn>7</mn><mo>,</mo><mo /><mn>21</mn></mrow></mfenced></mrow></math> </ephtml> , with <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>n</mi><mo>=</mo><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mi>d</mi></munder><msub><mi>n</mi><mi>d</mi></msub><mo>=</mo><mn>1</mn><mo>,</mo><mn>129</mn></mrow></math> </ephtml> . The overall sampling fraction is given by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>f</mi><mo>=</mo><mfrac><mi>n</mi><mi>N</mi></mfrac><mo>=</mo><mn>5.6</mn><mi>%</mi></mrow></math> </ephtml> . Across the areas, these takes values between 1.91 percent and 15.33 percent;</item> <p></p> <item> Estimate <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub></mrow></math> </ephtml> via the IPF estimator given in (<reflink idref="bib2" id="ref93">2</reflink>) in each sample and obtain <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> ;</item> <p></p> <item> Estimate <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub></mrow></math> </ephtml> via the direct estimator given by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> in each sample and obtain <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> ;</item> <p></p> <item> Estimate the MSE of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> via parametric bootstrap described in the fourth section (with <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>B</mi><mo>=</mo><mn>500</mn></mrow></math> </ephtml> ), the estimate is denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mover accent="true"><mrow><mi>M</mi><mi>S</mi><mi>E</mi></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> .</item> </ulist> <p>In order to evaluate the performance of the proposed bootstrap MSE estimator, the following quality measures are calculated:</p> <p>Empirical MSE of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> (the true MSE):</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msup><mi>S</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mi>s</mi></munder><msup><mrow><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup><mo>−</mo><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mfenced></mrow><mn>2</mn></msup></mrow></mtd></mtr></mtable><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Empirical MSE of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> :</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msup><mi>S</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mi>s</mi></munder><msup><mrow><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup><mo>−</mo><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mfenced></mrow><mn>2</mn></msup><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Relative bias of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mover accent="true"><mrow><mtext>MSE</mtext></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> :</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><mi>R</mi><mi>B</mi><mfenced><mrow><mtext /><msub><mrow><mover accent="true"><mrow><mtext>MSE</mtext></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></mfenced><mo>=</mo><msup><mi>S</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mi>s</mi></munder><mfrac><mrow><mi>m</mi><mi>s</mi><msub><mi>e</mi><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>−</mo><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow><mrow><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></mfrac></mrow></mtd></mtr></mtable><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Relative bias of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> :</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mfrac><mrow><mi>R</mi><mi>B</mi><mtext /><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msup><mi>S</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munder><mstyle displaystyle="true"><mo>∑</mo></mstyle><mi>s</mi></munder><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup><mo>−</mo><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mrow><mrow><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mfrac><mo>.</mo></math> </ephtml> </p> <p>Graph</p> <p>Relative bias of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> :</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfrac><mrow><mi>R</mi><mi>B</mi><mtext /><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msup><mi>S</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><munder><mstyle displaystyle="true" mathsize="140%"><mo>∑</mo></mstyle><mi>s</mi></munder><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup><mo>−</mo><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mrow><mrow><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub></mrow></mfrac><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mrow><mi>d</mi><mi>s</mi></mrow></msub><mo>=</mo><munderover><mstyle displaystyle="true"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mrow><msub><mi>N</mi><mi>d</mi></msub></mrow></munderover><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi><mi>s</mi></mrow></msub><mo>/</mo><msub><mi>N</mi><mi>d</mi></msub></mrow></math> </ephtml> . The true small area means are denoted by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub><mo>=</mo><mstyle displaystyle="true"><msubsup><mo>∑</mo><mrow><mi>s</mi><mo>=</mo><mn>1</mn></mrow><mi>S</mi></msubsup><mrow><mstyle displaystyle="true"><msubsup><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mrow><msub><mi>N</mi><mi>d</mi></msub></mrow></msubsup><mrow><msub><mi>y</mi><mrow><mi>d</mi><mi>i</mi><mi>s</mi></mrow></msub></mrow></mstyle></mrow></mstyle><mo>/</mo><msub><mi>N</mi><mi>d</mi></msub></mrow></math> </ephtml></p> <p>In the following section, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mi>R</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><msqrt><mrow><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></msqrt></mrow></math> </ephtml> is used to denote the empirical root MSE (and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mi>R</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mtext /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> ). These quality measures are evaluated and compared across the areas using the median as a robust central tendency measure ([<reflink idref="bib8" id="ref94">8</reflink>]; [<reflink idref="bib24" id="ref95">24</reflink>]).</p> <hd id="AN0173122298-9">Results</hd> <p></p> <hd id="AN0173122298-10">Performance of the bootstrap MSE estimator under different distributional assumptions of e d...</hd> <p>Figure 2 presents the empirical root MSE (ERMSE) of the IPF estimator and the direct estimator. Since the performance is very similar across all three distributions, Figure 2 presents the findings for the Normal case only, ordered by increasing small area sample size.</p> <p>Graph: Figure 2. Empirical root mean squared error comparisons: Direct versus iterative proportional fitting estimator for the Normal case.</p> <p>It can be seen that the IPF synthetic estimator provides estimates with lower MSE than the direct estimator due to the use of related auxiliary variables in reducing variance. This is particularly true for small areas with smaller survey sample sizes where the performance gains of IPF are large relative to direct estimator. This occurs because the ERMSE of the direct estimator naturally depends on the survey sample size: When the survey sample size in the small area <emph>d</emph> is smaller, the ERMSE tends to increase due to the larger variance around such estimates. As the small area survey sample size increases, the performance gains of the synthetic IPF estimator decline relative to the direct estimator until a point where its performance converges with that of the direct estimator. For reference, Figure 2 looks identical when ordered by sampling fraction rather than sample size.</p> <p>For each distribution, Figure 3 shows the IPF point estimates across the small areas plotted against the true values observed in the population.</p> <p>Graph: Figure 3. Comparisons of iterative proportional fitting estimates versus true means.</p> <p>Table 1 shows the median estimate comparisons obtained using <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> under the different distributional scenarios. The true value <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub></mrow></math> </ephtml> is shown in the first row, while the direct and IPF central estimates are shown in rows 2 and 3, respectively. The relative bias of those direct and IPF estimates across all the small areas are then shown in the penultimate two rows. It can be seen that the IPF small area estimator returns only negligible biases across the small area even in cases of Gumbel and Logistic distributions of the error term.</p> <p>Graph</p> <p>Table 1. Point Estimates Comparisons and Relative Biases Across the Small Area, Median Values Where Intraclass Correlation = 0.05.</p> <p> <ephtml> <table><thead><tr><th rowspan="2">Performance Measure</th><th colspan="3">Scenario</th></tr><tr><th>Normal</th><th>Gumbel</th><th>Logistic</th></tr></thead><tbody><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mi>d</mi></msub></mrow></math></p></td><td>120.696</td><td>120.595</td><td>120.714</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math></p></td><td>120.720</td><td>120.756</td><td>120.964</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math></p></td><td>120.662</td><td>120.677</td><td>120.687</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>R</mi><mi>B</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>0.000</td><td>0.000</td><td>0.000</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>R</mi><mi>B</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>0.006</td><td>0.004</td><td>0.004</td></tr></tbody></table> </ephtml> </p> <p>Figure 4 moves on to focus on the performance of the bootstrap MSE estimator to calculate the uncertainty around those central small area point estimates in the three Normal, Gumbel, and Logistic distributions, respectively. It can be seen that our proposed bootstrap approach provides nearly unbiased estimates of the true MSE with relative bias centered on and close to zero across the small areas. No association is found between estimate bias and sampling fraction across Figure 4.</p> <p>Graph: Figure 4. RBMSE^bootY¯^dIPF, Normal, Gumbel, and Logistic cases.</p> <p>Table 2 provides further details of the performance of the bootstrap MSE estimator across the three distributions. In particular, the true MSE (i.e., empirical MSE) in each distribution is compared to our bootstrap MSE estimate and coverage rates (of 95 percent confidence intervals) are also presented. It can be seen that the relative bias values are close to zero and that the MSE bootstrap estimator is nearly unbiased across the small areas in each of the three distributions.</p> <p>Graph</p> <p>Table 2. Performance Measures of the Bootstrap MSE Estimates.</p> <p> <ephtml> <table><thead><tr><th rowspan="2">Performance Measure</th><th colspan="3">Scenario</th></tr><tr><th>Normal</th><th>Gumbel</th><th>Logistic</th></tr></thead><tbody><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>16.114</td><td>17.159</td><td>15.990</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mover accent="true"><mrow><mi>M</mi><mi>S</mi><mi>E</mi></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>14.592</td><td>15.865</td><td>14.001</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>R</mi><mi>B</mi><mfenced><mrow><msub><mrow><mover accent="true"><mrow><mi>M</mi><mi>S</mi><mi>E</mi></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></mfenced></mrow></math></p></td><td>−0.076</td><td>−0.080</td><td>−0.079</td></tr><tr><td>Coverage rates</td><td>0.918</td><td>0.915</td><td>0.914</td></tr></tbody></table> </ephtml> </p> <p>1 Note: MSE = mean squared error; EMSE = empirical mean squared error.</p> <hd id="AN0173122298-11">On the role of the ICC in the case of the Normal distribution</hd> <p>On the basis of these analyses, the performance of our proposed MSE estimator is encouraging in terms both of the small relative bias of the MSE estimates and the quality of the coverage. However, as noted above, it may be that the magnitude of the ICC affects the performance of the MSE estimator.</p> <p>Table 3 shows the findings of ICC sensitivity analyses with a focus on the Normal distribution where the ICC is varied at several points from 0.01 to 0.50. The results for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>ρ</mtext><mo>=</mo><mn>0.05</mn></mrow></math> </ephtml> shown above are repeated to aid comparison. It can be seen that the relative bias of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></math> </ephtml> increases slightly when the ICC increases beyond around 0.15, ranging from −0.001 when ρ = 0.01–0.006 when <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>ρ</mtext><mo>=</mo><mn>0.2</mn><mo /></mrow></math> </ephtml> and 0.021 when <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><mi>ρ</mi><mo>=</mo><mn>0.50</mn></mrow></math> </ephtml> . In terms of the bootstrap MSE estimator, the penultimate row shows that this estimator delivers consistently small relative bias in the MSE though with somewhat weaker performance in coverage at very low levels of ICC as displayed in the final row. When the ICC is small, the MSE is slightly underestimated as seen by looking the relative bias and by comparing the empirical MSE (line two) with our bootstrap MSE (line 3). However, these relative bias estimates remain acceptable.</p> <p>Graph</p> <p>Table 3. Performance Measures of Our Bootstrap MSE Estimator at Varying Levels of Intraclass Correlation in the Normal Distribution.</p> <p> <ephtml> <table><thead><tr><th rowspan="2">Performance Measure</th><th colspan="8">Scenario</th></tr><tr><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.01</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.03</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.05</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.08</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.10</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.15</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.20</mn></mrow></math></p></th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>ρ</mtext><mo>=</mo><mn>0.50</mn></mrow></math></p></th></tr></thead><tbody><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>R</mi><mi>B</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>−0.001</td><td>0.002</td><td>−0.001</td><td>0.002</td><td>0.001</td><td>0.002</td><td>0.006</td><td>0.021</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>E</mi><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>5.300</td><td>11.001</td><td>16.114</td><td>16.115</td><td>16.116</td><td>53.426</td><td>64.551</td><td>293.397</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mover accent="true"><mrow><mtext>MSE</mtext></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mtext>boot</mtext></mrow></msub><mfenced><mrow><msubsup><mrow><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo stretchy="true">^</mo></mover></mrow><mi>d</mi><mrow><mtext>IPF</mtext></mrow></msubsup></mrow></mfenced></mrow></math></p></td><td>4.734</td><td>10.197</td><td>14.592</td><td>14.650</td><td>14.600</td><td>52.247</td><td>63.220</td><td>293.835</td></tr><tr><td><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mi>R</mi><mi>B</mi><mfenced><mrow><msub><mrow><mover accent="true"><mrow><mi>M</mi><mi>S</mi><mi>E</mi></mrow><mo stretchy="true">^</mo></mover></mrow><mrow><mi>b</mi><mi>o</mi><mi>o</mi><mi>t</mi></mrow></msub><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced></mrow></mfenced></mrow></math></p></td><td>−0.080</td><td>−0.079</td><td>−0.076</td><td>−0.075</td><td>−0.075</td><td>−0.019</td><td>−0.018</td><td>0.006</td></tr><tr><td>Coverage rates</td><td>0.917</td><td>0.917</td><td>0.918</td><td>0.924</td><td>0.939</td><td>0.942</td><td>0.950</td><td>0.950</td></tr></tbody></table> </ephtml> </p> <p>2 Note: MSE = mean squared error; EMSE = empirical mean squared error.</p> <hd id="AN0173122298-12">Application to Small Area Income Estimation in Italian Municipalities</hd> <p>This section provides a real-world application of an IPF small area estimator and, more centrally for this article, of our proposed bootstrap estimator of its MSE. The application used is the estimation of mean equivalized annual household disposable income (in Euros) for the municipalities of Tuscany region ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>D</mi><mo>=</mo><mn>287</mn></mrow></math> </ephtml> ). The survey data used are provided from the 2009 European Union SILC (EU-SILC). These EU-SILC data contain a sample of 1,448 households for Tuscany. EU-SILC is designed to deliver estimates at the national and also regional (NUTS-2) level ([<reflink idref="bib23" id="ref96">23</reflink>]). Therefore, this situation is typical of most survey situations in that EU-SILC cannot be used to derive usable income estimates at smaller subregional geographies such as municipalities due to low or zero survey sample sizes. The household income variable of interest is given in the EU-SILC data and equivalized using Eurostat's official modified Organization for Economic Cooperation and Development equivalence scale ([<reflink idref="bib28" id="ref97">28</reflink>]; [<reflink idref="bib37" id="ref98">37</reflink>]). The auxiliary variables for the Tuscan municipalities come from the Population Census of Italy.</p> <hd id="AN0173122298-13">Model Fitting and Internal Validation</hd> <p>The explanatory variables used in this application are working status, years of education, gender, and age of the survey identified head of household. These have been informed by preliminary model investigations and findings from previous studies ([<reflink idref="bib24" id="ref99">24</reflink>]). Model diagnostics identified some skewness and outliers in the distribution of the income outcome variable, as is common with such distributions, and this variable was therefore log transformed. No evidence of leverage was found. Table 4 presents the results from the log-linear linear model in EU-SILC based on (<reflink idref="bib3" id="ref100">3</reflink>).</p> <p>Graph</p> <p>Table 4. Model Results.</p> <p> <ephtml> <table><thead><tr><th>Coefficient</th><th>Estimates</th><th><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><mtext>exp</mtext><mfenced><mtext>β</mtext></mfenced></mrow></math></p></th><th>Standard Error</th><th><italic>p</italic> Value</th></tr></thead><tbody><tr><td>Intercept</td><td>8.460</td><td>4,754.748</td><td>.105</td><td>.000</td></tr><tr><td>Gender</td><td>0.215</td><td>1.236</td><td>.031</td><td>.000</td></tr><tr><td>Working status</td><td>0.352</td><td>1.422</td><td>.041</td><td>.000</td></tr><tr><td>Age</td><td>0.010</td><td>1.012</td><td>.001</td><td>.000</td></tr><tr><td>Years of education</td><td>0.034</td><td>1.035</td><td>.003</td><td>.000</td></tr></tbody></table> </ephtml> </p> <p>Validation is an important step in any SAE study. SAE models can be validated internally in terms of the underlying model and externally against some known other external data of the target outcome variable. In terms of the internal validation, Figure 5 shows the fitted values versus the residuals as well as the <emph>Q</emph>–<emph>Q</emph> plots of the residuals from the log-linear model used to produce the MSE of the IPF estimates. These show good behavior with respect to the normality assumption. External validation is discussed below.</p> <p>Graph: Figure 5. Fitted values versus residuals (left) and Q – Q plot of the residuals from the model used to produce the mean squared error of the iterative proportional fitting estimates.</p> <hd id="AN0173122298-14">Estimating Municipality Income in Tuscany</hd> <p>This section discusses the results of the IPF SAE of the mean equivalized annual household disposable income across Tuscan municipalities along with their uncertainty estimates. Figure 6 maps the mean IPF estimates across the 287 Tuscan municipalities. The map displays municipalities in four quartiles and shows a range in 2009 municipal income estimates from a low of just over 16,000 Euros per annum to a high of just under 20,000 Euros per annum. Municipalities located in the provinces of Massa Carrara (North West), Grosseto (South), and Prato and Pistoia (North) show the lowest estimated municipal income levels. On the contrary, municipalities around Florence, Arezzo, Pisa, and Livorno show the highest estimated municipal income levels.</p> <p>Graph: Figure 6. Iterative proportional fitting income estimates for Tuscan municipalities.</p> <p>In terms of external validation of these IPF estimates, a frequent inherent challenge, as here, is the typical lack of any such existing small area data against which to validate (hence the motivation for the SAE). External validation of these IPF estimates is provided in two ways. Firstly, the spatial patterns in Figure 6 are in line with known geographical patterns of similar indicators across Tuscany seen in previously published research ([<reflink idref="bib23" id="ref101">23</reflink>]; [<reflink idref="bib42" id="ref102">42</reflink>]). Secondly, no identical income indicators exist at this municipality scale, and no direct survey estimates to municipality level are viable from the EU-SILC survey data. However, it is viable to produce direct survey estimates from the EU-SILC survey data to Tuscany's 10 larger provinces and to compare these with indirect IPF estimates also to province level. The Spearman's rank correlation between these two sets of estimates is.93, and this is statistically significant at below the 1 percent level, although acknowledging the limited sample size involved.</p> <p>Table 5 presents summary statistics of the uncertainty of the direct survey estimates compared with the IPF estimates. In particular, it shows the root mean squared error (column 1) and, expressed as a percentage of the estimates, the relative root mean squared error (RRMSE%; column three) of the small area estimates. Given that the direct estimates are unbiased, the standard deviation (<emph>SD</emph>; column 2) and coefficient of variation (CV%; column 4) of the direct estimates enable a comparison of uncertainty of the direct estimates with the bootstrap MSE estimates of the IPF estimator. The coefficient of variation (CV) is a standardized measure of the dispersion of a distribution and is calculated as a ratio of its <emph>SD</emph> to its mean. In the present analyses, it is obtained as the ratio between the <emph>SD</emph> of the direct survey estimate and the direct survey estimate for every area. Since direct estimates are unbiased, their CVs represent measures of uncertainty ([<reflink idref="bib49" id="ref103">49</reflink>]). RRMSE and CV are standard measures of uncertainty that are required in many official statistics institutes (see [<reflink idref="bib52" id="ref104">52</reflink>]; [<reflink idref="bib54" id="ref105">54</reflink>]).</p> <p>Graph</p> <p>Table 5. Summary Statistics of the Performance Gains from the Synthetic IPF Estimator Compared to the Direct Estimator for Small Area Income Estimates Across Tuscan Municipalities.</p> <p> <ephtml> <table><thead><tr><th>Areas with</th><th>Summary Statistic</th><th>RMSE IPF</th><th><italic>SD</italic> Direct</th><th>RRMSE IPF %</th><th>CV Direct %</th><th>Gains %</th></tr></thead><tbody><tr><td rowspan="4"><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mi>n</mi><mi>d</mi></msub><mo>></mo><mn>0</mn></mrow></math></p></td><td> Min.</td><td>2,197.50</td><td>2,444.02</td><td>10.63</td><td>11.21</td><td>26.46</td></tr><tr><td> Mean</td><td>3,289.50</td><td>6,663.51</td><td>17.50</td><td>32.10</td><td>32.71</td></tr><tr><td> Median</td><td>3,356.00</td><td>4,802.52</td><td>18.48</td><td>26.20</td><td>32.80</td></tr><tr><td> Max.</td><td>4,431.00</td><td>51,750.48</td><td>23.51</td><td>99.05</td><td>94.19</td></tr><tr><td rowspan="4"><p><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mi>n</mi><mi>d</mi></msub><mo>=</mo><mn>0</mn></mrow></math></p></td><td> Min.</td><td>2,196.80</td><td>—</td><td>13.11</td><td>—</td><td>—</td></tr><tr><td> Mean</td><td>3,308.03</td><td>—-</td><td>19.27</td><td>—</td><td>—</td></tr><tr><td> Median</td><td>3,346.46</td><td>—-</td><td>18.98</td><td>—</td><td>—</td></tr><tr><td> Max.</td><td>4,756.51</td><td>—-</td><td>24.51</td><td>—</td><td>—</td></tr></tbody></table> </ephtml> </p> <p>3 Note: MSE = mean squared error; CV = coefficient of variation; RRMSE = relative root mean squared error; IPF = iterative proportional fitting.</p> <p>Table 5 shows that the IPF estimates are more reliable than the direct estimates across all points of the income distribution, as depicted by the lower values of the RRMSE IPF (column 1) and RRMSE IPF (column 3) compared to <emph>SD</emph> direct (column 2) and CV direct (column 4), respectively. The IPF small area estimates can also be considered reliable in absolute terms. Values of RRMSE below a threshold of 20 percent are often taken by statistical agencies as acceptable ([<reflink idref="bib12" id="ref106">12</reflink>]), and almost all of this municipality distribution of small area estimates is below this level. The final column of Table 5 summarizes the gains in efficiency of the IPF estimates over the direct survey estimates by comparing the MSE for the IPF estimator with the variance of the unbiased direct estimator. Technically, these are calculated by</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtable equalcolumns="true" equalrows="true"><mtr><mtd><mrow><mi>G</mi><mi>a</mi><mi>i</mi><mi>n</mi><mo /><mfenced><mrow><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>=</mo><mfrac><mrow><mi>M</mi><mi>S</mi><mi>E</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>I</mi><mi>P</mi><mi>F</mi></mrow></msubsup></mrow></mfenced><mo>−</mo><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced></mrow><mrow><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced></mrow></mfrac></mrow></mtd></mtr></mtable><mo>×</mo><mn>100</mn><mo>,</mo><mtext /><mi>d</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>D</mi><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></mfenced></mrow></math> </ephtml> denotes the variance of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><msubsup><mover accent="true"><mover accent="true"><mi>Y</mi><mo>¯</mo></mover><mo>^</mo></mover><mi>d</mi><mrow><mi>D</mi><mi>i</mi><mi>r</mi><mi>e</mi><mi>c</mi><mi>t</mi></mrow></msubsup></mrow></math> </ephtml> . Equation (<reflink idref="bib13" id="ref107">13</reflink>) denotes a measure of gain in efficiency of using an estimator with higher precision compared to the direct estimator. We refer to [<reflink idref="bib51" id="ref108">51</reflink>] for measures of efficiency in survey statistics and to [<reflink idref="bib40" id="ref109">40</reflink>] and [<reflink idref="bib26" id="ref110">26</reflink>] for some examples of their use. Results for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>n</mi><mi>d</mi></msub><mo>=</mo><mn>0</mn></mrow></math> </ephtml> (municipalities with zero sample size) and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>n</mi><mi>d</mi></msub><mo>></mo><mn>0</mn></mrow></math> </ephtml> (municipalities with some households in the survey) have been separated because it is not possible to compute direct estimates (and as consequently the gains) for areas with zero sample size. Table 5 shows that the small area estimates from the indirect IPF estimator provide significant performance gains compared to the direct estimator at all points of the municipality income distribution.</p> <p>Figure 7 drills down to focus on the extent to which these performance gains vary according to the size of the municipality sample size in the EU-SILC survey, a key driver of the variance of the direct estimator and key limiter of the viability of using the direct estimator to produce reliable survey estimates for small areas. To aid comparison, Figure 7 is ordered from left to right by municipality survey sample size in the EU-SILC. RRMSE is shown for the IPF estimates, while CV is shown for the direct estimates. Figure 7 illustrates that the RRMSE of the indirect IPF estimates does not depend on the sample size in contrast to the direct estimator. As such, Figure 7 highlights that while performance gains from the IPF estimator are seen across the whole distribution, they increase as the municipality sample size in the survey decreases. Among those municipalities with the smallest survey sample sizes, there is a marked increase in the performance gains available from the IPF estimator compared to the direct survey estimator. Figure 7 is naturally only able to display comparative results for municipalities with nonzero sample sizes, given that direct estimates cannot be produced for small areas with zero sample size. Estimates for these municipalities do of course become viable with synthetic IPF estimator. For reference, Figure 7 looks identical when ordered by sampling fraction rather than sample size.</p> <p>Graph: Figure 7. Performance comparison sorted by small area sample size.</p> <hd id="AN0173122298-15">Discussion</hd> <p>The combination of high costs of survey data collection and increasing demands for ever more spatially detailed data from policy makers and scholars alike mean growing demands for SAE techniques. Spatial microsimulation approaches to SAE continue to be widely used across diverse domains including transport, health, physical activity, and income. However, their continuing inability to produce reliable estimates of uncertainty alongside their central point estimates remains a pressing limitation to their utility for practitioners and scholars alike. This is understandable in part given that the estimation of MSE is difficult in a spatial microsimulation context since it cannot be estimated in a closed form and analytical approximations are highly challenging.</p> <p>Widely discussed in the SAE literature is the importance of providing measures of uncertainty such as MSE or confidence intervals alongside the central point estimates in order to assess the reliability of the small area estimates ([<reflink idref="bib46" id="ref111">46</reflink>]). This is particularly important where policy decisions are taken on the basis of the small area estimates since the consequences of real-world decision making without a clear sense of the uncertainty around the point estimates can be misleading and potentially harmful ([<reflink idref="bib25" id="ref112">25</reflink>]). [<reflink idref="bib32" id="ref113">32</reflink>]:264) argues that "[T]he credibility of [microsimulation models] with the research community as well as with users will in the long run depend on the application of sound principles of inference in the estimation, testing and validation of these models."</p> <p>This article provides a significant development in this context by presenting a novel parametric bootstrap approach for the estimation of uncertainty in spatial microsimulation SAE techniques. Importantly, the measure of uncertainty estimated is the MSE that contains both the variance and bias of the estimate. Simulation results demonstrate that under model assumptions, our proposed MSE estimator is relatively unbiased and displays good coverage properties against known true population values. The approach delivers substantial performance gains compared to the direct estimator across all portions of the distribution. In doing so, our approach enables researchers and policy makers alike to quantify both the performance gains potentially available through the use of spatial microsimulation approaches to SAE compared to direct survey estimates and to quantify the extent of uncertainty around those small area estimates. The simulation results show that those performance gains exist irrespective of the target small area sample size but are especially large at low sample sizes (below 10 in this simulation) and, naturally, when small areas have zero sample size such that direct estimates are nonviable but synthetic small area estimates are possible. Sensitivity tests confirm that these performance gains are maintained both across nonnormal Gumbel and Logistic distributions as well as across differing values of ICC, though coverage performance falls slightly at very low levels of the ICC in our simulation. A practical data application using EU-SILC data to municipality income across Tuscany is presented and validated in order to demonstrate the applicability and similar performance of the MSE bootstrap estimator in a real-world setting.</p> <p>While the performance and sensitivity analyses confirm that our proposed approach evaluates well and marks a significant contribution to the field, it serves also to open up opportunities for further more advanced enquiry as a result. First, we focus here only on a linear model. Future work will need to explore the performance of other nonlinear models in this context. However, the bootstrap approach in these scenarios will follow the same steps as those proposed in our approach. Second, a clear next step is to extend the framework to different types of outcome variables beyond the scalar target variable assessed here. Third, in this study, the common case of nonnormal error terms is assessed, and our simulation study shows good performance of the MSE estimator in distributions with mild skew and heavy tails as well as in the normal case. However, future work could take into account more fully the implications of and potential responses to failures in model assumptions and of model failure. Our hope therefore is that this article's contributions will not only make a significant advance to research and policy practice in spatial microsimulation SAE, but that it will also stimulate further scholarly attention to these and other areas.</p> <hd id="AN0173122298-16">Supplemental Material</hd> <p>Graph: Supplemental Material, sj-docx-1-smr-10.1177_0049124120986199 for Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach by Angelo Moretti and Adam Whitworth in Sociological Methods & Research</p> <ref id="AN0173122298-17"> <title> References </title> <blist> <bibl id="bib1" idref="ref49" type="bt">1</bibl> <bibtext> Anderson B.2007. Creating Small Area Income Estimates for England: Spatial Microsimulation Modelling. London, England: Department of Communities and Local Government.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref44" type="bt">2</bibl> <bibtext> Anderson B.2013. " Estimating Small-area Income Deprivation: An Iterative Proportional Fitting Approach " Pp.19 in Spatial Microsimulation: A Reference Guide for Users, edited by Tanton R., Edwards K. L. the Netherlands: Springer.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref72" type="bt">3</bibl> <bibtext> Ambugo E. A., Hagen T. P. 2015. A multilevel analysis of mortality following acute myocardial infarction in Norway: do municipal health services make a difference? BMJ Open. 5:e008764.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref69" type="bt">4</bibl> <bibtext> Battese G. E., Harter R. M., Fuller W. A. 1988. " An Error-components Model for Prediction of County Crop Areas Using Survey and Satellite data." Journal of the American Statistical Association. 83(401):28–36.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref12" type="bt">5</bibl> <bibtext> Bell W. R., Basel W. W., Maples J. J. 2016. "An Overview of the U.S. Census Bureau's Small Area Income and Poverty Estimates Program" Pp. 349 in Analysis of Poverty Data by Small Area Estimation, edited by Pratesi M. London, England: Wiley.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref1" type="bt">6</bibl> <bibtext> Benavent R., Morales D. 2016. " Multivariate Fay–Herriot Models for Small Area Estimation." Computational Statistics & Data Analysis. 94:372–90.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref27" type="bt">7</bibl> <bibtext> Campbell M., Ballas D. 2013. " A Spatial Microsimulation Approach to Economic Policy Analysis in Scotland." Regional Science Policy and Practice. 5(3):263–89.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref94" type="bt">8</bibl> <bibtext> Chambers R., Chandra H., Tzavidis N. 2011. " On Bias-robust Mean Squared Error Estimation for Pseudo-linear Small Area Estimators." Survey Methodology. 37:153–70.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref31" type="bt">9</bibl> <bibtext> Chen H., Shen R. 2015. "Variance Estimation for Survey-weighted Data Using Bootstrap Resampling Methods: 2013 Methods-of-payment Survey Questionnaire." Technical Reports 104. Bank of Canada. Retrieved from (https://ideas.repec.org/s/bca/bocatr.html).</bibtext> </blist> <blist> <bibtext> Chin S. F., Harding A. 2006. Regional Dimensions: Creating Synthetic Small-area Microdata and Spatial Microsimulation Models. Online Technical Paper—TP33, NATSEM, University of Canberra.</bibtext> </blist> <blist> <bibtext> Clarke M., Holm E. 1987. " Microsimulation Methods in Spatial Analysis and Planning." Geografiska Annaler, Series B, Human Geography. 69(2):145–64.</bibtext> </blist> <blist> <bibtext> Commonwealth Department of Social Services. 2015. Survey of Disability, Ageing and Carers, 2012: Modelled Estimates for Small Areas, Projected 2015. Prepared by the Regional Statistics National Centre, ABS (Release 1: February 2015). Retrieved August 10, 2018 (https://<ulink href="http://www.health.gov.au/internet/main/publishing.nsf/Content/98DCE47FC10BDD51CA257F15000413F5/$File/SDAC%202012%20Modelled%20Estimates%20for%20Small%20Areas%20projected%202015%5fExplanatory%20Notes%20-%20Release%201.pdf">www.health.gov.au/internet/main/publishing.nsf/Content/98DCE47FC10BDD51CA257F15000413F5/$File/SDAC%202012%20Modelled%20Estimates%20for%20Small%20Areas%20projected%202015%5fExplanatory%20Notes%20-%20Release%201.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Cullinan J., Hynes S., O'Donoghue C. 2006, "The Use of Spatial Microsimulation and Geographic Information Systems (GIS) in Benefit Function Transfer—An Application to Modelling the Demand for Recreational Activities in Ireland." Paper presented at the 8th Nordic Seminar on Microsimulation models, Oslo, Norway, June 7–9.</bibtext> </blist> <blist> <bibtext> D'Arrigo J., Skinner C. 2010. " Linearization Variance Estimation for Generalized Raking Estimators in the Presence of Nonresponse," Survey Methodology. 36(2):181–92.</bibtext> </blist> <blist> <bibtext> Datta G. S., Day B., Basawa I. 1999. " Empirical Best Linear Unbiased and Empirical Bayes Prediction in Multivariate Small Area Estimation." Journal of Statistical Planning and Inference. 75:269–79</bibtext> </blist> <blist> <bibtext> Deville J. C., Särndal C. E. 1992. " Calibration Estimators in Survey Sampling." Journal of the American Statistical Association. 87(418):376–82.</bibtext> </blist> <blist> <bibtext> Dodge Y., Commenges D. eds. 2006. The Oxford Dictionary of Statistical Terms. Oxford, England: Oxford University Press on Demand.</bibtext> </blist> <blist> <bibtext> Edwards K. L., Clarke G. P., Ransley J. K., Cade J. E. 2010. " The Neighbourhood Matters: Studying Exposures Relevant to Childhood Obesity and the Policy Implications in Leeds, UK." Journal of Epidemiology and Community Health. 64(3):194–201.</bibtext> </blist> <blist> <bibtext> Efron B., Tibshirani R. 1993. An Introduction to the Bootstrap. London, England: Chapman and Hall.</bibtext> </blist> <blist> <bibtext> Espuny-Pujol F., Morrissey K., Williamson P. 2018. " A Global Optimization approach to Range-restricted Survey Calibration." Statistics and Computing. 28:427–39.</bibtext> </blist> <blist> <bibtext> Ferrante L., Cameriere R. 2009. " Statistical Methods to Assess the Reliability of Measurements in the Procedures for Forensic Age Estimation." International Journal of Legal Medicine. 123(4):277–83.</bibtext> </blist> <blist> <bibtext> Fuller W.A.2002. " Regression Estimation for Survey Samples." Survey Methodology. 28(1):5–23</bibtext> </blist> <blist> <bibtext> Giusti C., Masserini L., Pratesi M. 2015. " Local Comparisons of Small Area Estimates of Poverty: An Application within the Tuscany Region of Italy." Social Indicators Research. 131(1):235–54.</bibtext> </blist> <blist> <bibtext> Giusti C., Tzavidis N., Pratesi M., Salvati N. 2013. " Resistance to Outliers of m-Quantile and Robust Random Effects Small Area Models." Communications in Statistics: Simulation and Computation. 43:549–68.</bibtext> </blist> <blist> <bibtext> Goedemé T., Van den Bosch K., Salanauskaite L., Verbist G. 2013. " Testing the Statistical Significance of Microsimulation Results: A Plea." International Journal of Microsimulation. 6(3):50–77.</bibtext> </blist> <blist> <bibtext> González-Manteiga W., Lombardía M. J., Molina I., Morales D., Santamaría L. 2008a. " Analytic and Bootstrap Approximations of Prediction Errors under a Multivariate Fay-Herriot Model." Computational Statistics and Data Analysis. 52:5242–52.</bibtext> </blist> <blist> <bibtext> González-Manteiga W., Lombardía M. J., Molina I., Morales D., Santamaría L. 2008b. " Bootstrap Mean Squared Error of a Small-area EBLUP." Journal of Statistical Computation and Simulation. 78:443–62.</bibtext> </blist> <blist> <bibtext> Haagenars A., de Vos K., Zaidi M. A. 1994. Poverty Statistics in the Late 1980s: Research Based on Micro-data. Luxembourg, Europe: Office for Official Publications of the European Communities.</bibtext> </blist> <blist> <bibtext> Horvitz D. G., Thompson D. J. 1952. " A Generalization of Sampling without Replacement from Finite Universe." Journal of the American Statistical Association. 47:663–85.</bibtext> </blist> <blist> <bibtext> Ipsos MORI. 2018. Small Area Estimation of Sport Participation and Activity Using Data from the Active Lives Survey. London, England: Ipsos MORI.</bibtext> </blist> <blist> <bibtext> Johnson A., Chandra F. H., Brown J., Sabu S. P. 2012. " Small Area Estimation for Policy Development: A Case Study of Child Undernutrition in Ghana." Journal of the Indian Society of Agricultural Statistics. 66(1):171–86.</bibtext> </blist> <blist> <bibtext> Klevmarken A. N.2002. " Statistical Inference in Micro-simulation Models: Incorporating External Information." Mathematics and Computers in Simulation. 59(1-3):255–65.</bibtext> </blist> <blist> <bibtext> Koch R.2008. The 80/20 Principle: The Secret of Achieving More with Less. New York: Doubleday.</bibtext> </blist> <blist> <bibtext> Kolenikov S.2014. " Calibrating Survey Data Using Iterative Proportional Fitting (Raking)." The Stata Journal. 14(1):22–59.</bibtext> </blist> <blist> <bibtext> Kott P. S.2009. "Calibration Weighting: Combining Probability Samples and Linear Prediction Models" Pp. 55 in Handbook of Statistics 29B Sample Surveys: Inference and Analysis, edited by Pfeffermann D., Rao C. R. North Holland: Elsevier.</bibtext> </blist> <blist> <bibtext> Lovelace R., Dumont M. 2016. Spatial Microsimulation With R. Boca Raton, FL: Chapman and Hall/CRC.</bibtext> </blist> <blist> <bibtext> Marchetti S., Beręsewicz M., Salvati N., Szymkowiak M., Wawrowski Ł. 2018. " The Use of a Three-level M-quantile Model to Map Poverty at Local Administrative Unit 1 in Poland." Journal of Royal Statistical Society—Series A. 181(4):1077–1104.</bibtext> </blist> <blist> <bibtext> Marshall A. 2010. " Small Area Estimation Using ESDS Government Surveys—An Introductory Guide." Economic and Social Data Service: P. 15.</bibtext> </blist> <blist> <bibtext> Molina I., Nandram B., Rao J. N. K. 2014. " Small Area Estimation of General Parameters with Application to Poverty Indicators: A Hierarchical Bayes Approach." The Annals of Applied Statistics. 8(2):852–85.</bibtext> </blist> <blist> <bibtext> Moretti A., Shlomo N., Sakshaug J. W. 2018. " Parametric Bootstrap Mean Squared Error of a Small Area Multivariate EBLUP." Communications in Statistics—Simulation and Computation. 49(6):1474–86</bibtext> </blist> <blist> <bibtext> Moretti A., Whitworth A. 2019. " Development and Evaluation of an Optimal Composite Estimator in Spatial Microsimulation Small Area Estimation," Geographical Analysis. 52(3):351–70.</bibtext> </blist> <blist> <bibtext> Moretti A., Shlomo N., Sakshaug J. W.2019. Small Area Estimation of Latent Economic Well-being. Sociological Methods & Research.</bibtext> </blist> <blist> <bibtext> Münnich R. 2014. "Small Area Applications: Some Results from a Design-based View." International Small Area Estimation Conference, SAE 2014, Poznan, Poland. Retrieved January 13, 2021 (<ulink href="http://www.sae2014.ue.poznan.pl/presentations/SAE2014%5fRalf%5fMunnich%5fc330a31c0a.pdf">http://www.sae2014.ue.poznan.pl/presentations/SAE2014%5fRalf%5fMunnich%5fc330a31c0a.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Nagle N., Buttenfield B., Leyk S., Spielman S. 2014. " Dasymetric Modeling and Uncertainty." Annals of the Association of American Geographers. 104(1):80–95.</bibtext> </blist> <blist> <bibtext> Office for National Statistics (ONS). 2019. Research Outputs: Small Area Estimation of Fuel Poverty in England, 2013 to 2017. London, England: Office for National Statistics.</bibtext> </blist> <blist> <bibtext> Pratesi M.2016. Analysis of Poverty Data by Small Area Estimation. London, England: Wiley.</bibtext> </blist> <blist> <bibtext> Rahman A., Harding A. 2017. Small Area Estimation and Microsimulation Modeling. Boca Raton, FL: CRC Press Taylor & Francis Group.</bibtext> </blist> <blist> <bibtext> Rahman A., Harding A., Tanton R., Liu S. 2013. " Simulating the Characteristics of Populations at the Small Area Level: New Validation Techniques for a Spatial Microsimulation Model in Australia." Computational Statistics & Data Analysis. 57(1):149–65.</bibtext> </blist> <blist> <bibtext> Rao J. N. K., Molina I. 2015. Small Area Estimation. New York: Wiley.</bibtext> </blist> <blist> <bibtext> Ravulaparthy S., Goulias K. 2011. Forecasting with Dynamic Microsimulation: Design, Implementation, and Demonstration: Final Report on Review, Model Guidelines, and a Pilot Test for a Santa Barbara County Application. Technical Report, May, California: University of California Transportation Center (UCTC).</bibtext> </blist> <blist> <bibtext> Särndal C. E., Swensson B., Wretman J. 1992. Model Assisted Survey Sampling. the Netherlands: Springer.</bibtext> </blist> <blist> <bibtext> Schirripa-Spagnolo F., D'Agostino A., Salvati N. 2018. " Measuring Differences in Economic Standard of Living between Immigrant Communities in Italy." Quality and Quantity. 52:1643–67.</bibtext> </blist> <blist> <bibtext> Simpson L, Tranmer M. 2005. " Combining Sample and Census Data in Small Area Estimates: Iterative Proportional Fitting with Standard Software." The Professional Geographer. 57(2):222–34.</bibtext> </blist> <blist> <bibtext> Statistics Canada. 2009. Statistics Canada Quality Guidelines. Catalogue no. 12-539-X. https://www150.statcan.gc.ca/n1/en/catalogue/12-539-X</bibtext> </blist> <blist> <bibtext> Tanton R., Edwards K. L. 2013. Spatial Microsimulation: A Reference Guide for Users. the Netherlands: Springer.</bibtext> </blist> <blist> <bibtext> Tanton R., McNamara J., Harding A., Morrison T. 2009. "Rich Suburbs, Poor Suburbs? Small Area Poverty Estimates for Australia's Eastern Seaboard in 2006" Pp. 79 in New Frontiers in Microsimulation Modelling, edited by Zaidi A., Harding Williamson P. London, England: Ashgate.</bibtext> </blist> <blist> <bibtext> Tanton R., Williamson P., Harding A. 2014. " Comparing Two Methods of Reweighting a Survey File to Small Area Data." International Journal of Microsimulation. 7(1):76–99.</bibtext> </blist> <blist> <bibtext> Tribby C. P., Zandbergen P. A. 2012. " High-resolution Spatio-temporal Modeling of Public Transit Accessibility." Applied Geography. 34:345–55.</bibtext> </blist> <blist> <bibtext> Whitworth A. (2012). Sustaining evidence-based policing in an era of cuts: Estimating fear of crime at small area level in England. Crime Prevention & Community Safety. 14(1): 48–68.</bibtext> </blist> <blist> <bibtext> Whitworth A. (Ed.). 2013. "Evaluations and Improvements in Small Area Estimation Methodologies."National Centre for Research Methods Methodological Review Paper, Economic and Social Research Council.</bibtext> </blist> <blist> <bibtext> Whitworth A., Carter E. 2015. Understanding Wales at the Small Area Level: Maximising the Performance of Small Area Estimation. Cardiff, Wales: Welsh Government.</bibtext> </blist> <blist> <bibtext> Whitworth A., Carter E., Ballas D., Moon G. 2017. " Estimating Uncertainty in Spatial Microsimulation Approaches to Small Area Estimation: A New Approach to Solving an Old Problem." Computers, Environment and Urban Systems. 63:50–57.</bibtext> </blist> <blist> <bibtext> Williamson P., Birkin M., Rees P. 1998. " The Estimation of Population Microdata by Using Data from Small Area Statistics and Samples of Anonymised Records." Environment and Planning A. 30(5):785–816.</bibtext> </blist> <blist> <bibtext> World Bank. 2018. Small Area Estimation of Poverty Under Structural Change. Washington, DC: World Bank Group: Poverty and Equity Global Practice.</bibtext> </blist> </ref> <ref id="AN0173122298-18"> <title> Footnotes </title> <blist> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibtext> The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the UK Economic and Social Research Council National Centre for Research Methods (grant number ES/N011619/1).</bibtext> </blist> <blist> <bibtext> Angelo Moretti https://orcid.org/0000-0001-6543-9418</bibtext> </blist> <blist> <bibtext> Supplemental material for this article is available online.</bibtext> </blist> </ref> <aug> <p>By Angelo Moretti and Adam Whitworth</p> <p>Reported by Author; Author</p> <p></p> <p>Angelo Moretti is a lecturer (assistant professor) in the Department of Computing and Mathematics, Manchester Metropolitan University, UK. His research interests are in multivariate small area estimation, survey methodology, well-being and poverty measurement, multivariate statistics, and mixed-effects modeling. He has a PhD in social statistics from the University of Manchester, UK, and an MSc in marketing and market research and BSc in economics from the University of Pisa, Italy.</p> <p>Adam Whitworth is a senior lecturer in the Department of Geography, University of Sheffield, UK. He is a spatial social policy scholar with ongoing programs of work focusing on the design and analysis of employment activation policies, local integration, and spatial statistical methodologies. He has a PhD in social policy, MSc in comparative social policy, and BA (Hons) degree in politics, philosophy, and economics all from the University of Oxford, UK.</p> </aug> <nolink nlid="nl1" bibid="bib29" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib49" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib60" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib47" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib38" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib20" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib51" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib31" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib18" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib45" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib46" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib64" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib11" firstref="ref15"></nolink> <nolink nlid="nl14" bibid="bib30" firstref="ref16"></nolink> <nolink nlid="nl15" bibid="bib36" firstref="ref17"></nolink> <nolink nlid="nl16" bibid="bib50" firstref="ref18"></nolink> <nolink nlid="nl17" bibid="bib58" firstref="ref19"></nolink> <nolink nlid="nl18" bibid="bib41" firstref="ref20"></nolink> <nolink nlid="nl19" bibid="bib57" firstref="ref21"></nolink> <nolink nlid="nl20" bibid="bib61" firstref="ref22"></nolink> <nolink nlid="nl21" bibid="bib10" firstref="ref23"></nolink> <nolink nlid="nl22" bibid="bib13" firstref="ref24"></nolink> <nolink nlid="nl23" bibid="bib55" firstref="ref25"></nolink> <nolink nlid="nl24" bibid="bib63" firstref="ref26"></nolink> <nolink nlid="nl25" bibid="bib56" firstref="ref29"></nolink> <nolink nlid="nl26" bibid="bib14" firstref="ref32"></nolink> <nolink nlid="nl27" bibid="bib16" firstref="ref33"></nolink> <nolink nlid="nl28" bibid="bib44" firstref="ref36"></nolink> <nolink nlid="nl29" bibid="bib62" firstref="ref37"></nolink> <nolink nlid="nl30" bibid="bib48" firstref="ref41"></nolink> <nolink nlid="nl31" bibid="bib53" firstref="ref42"></nolink> <nolink nlid="nl32" bibid="bib34" firstref="ref50"></nolink> <nolink nlid="nl33" bibid="bib22" firstref="ref53"></nolink> <nolink nlid="nl34" bibid="bib17" firstref="ref54"></nolink> <nolink nlid="nl35" bibid="bib54" firstref="ref55"></nolink> <nolink nlid="nl36" bibid="bib21" firstref="ref56"></nolink> <nolink nlid="nl37" bibid="bib27" firstref="ref57"></nolink> <nolink nlid="nl38" bibid="bib37" firstref="ref58"></nolink> <nolink nlid="nl39" bibid="bib40" firstref="ref59"></nolink> <nolink nlid="nl40" bibid="bib19" firstref="ref61"></nolink> <nolink nlid="nl41" bibid="bib35" firstref="ref68"></nolink> <nolink nlid="nl42" bibid="bib33" firstref="ref71"></nolink> <nolink nlid="nl43" bibid="bib43" firstref="ref80"></nolink> <nolink nlid="nl44" bibid="bib15" firstref="ref83"></nolink> <nolink nlid="nl45" bibid="bib59" firstref="ref87"></nolink> <nolink nlid="nl46" bibid="bib39" firstref="ref89"></nolink> <nolink nlid="nl47" bibid="bib24" firstref="ref95"></nolink> <nolink nlid="nl48" bibid="bib23" firstref="ref96"></nolink> <nolink nlid="nl49" bibid="bib28" firstref="ref97"></nolink> <nolink nlid="nl50" bibid="bib42" firstref="ref102"></nolink> <nolink nlid="nl51" bibid="bib52" firstref="ref104"></nolink> <nolink nlid="nl52" bibid="bib12" firstref="ref106"></nolink> <nolink nlid="nl53" bibid="bib26" firstref="ref110"></nolink> <nolink nlid="nl54" bibid="bib25" firstref="ref112"></nolink> <nolink nlid="nl55" bibid="bib32" firstref="ref113"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1397535
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Moretti%2C+Angelo%22">Moretti, Angelo</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6543-9418">0000-0001-6543-9418</externalLink>)<br /><searchLink fieldCode="AR" term="%22Whitworth%2C+Adam%22">Whitworth, Adam</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. 2023 52(4):1785-1815.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 31
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2023
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Evaluative
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Simulation%22">Simulation</searchLink><br /><searchLink fieldCode="DE" term="%22Geometric+Concepts%22">Geometric Concepts</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Error+of+Measurement%22">Error of Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Bias%22">Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Income%22">Income</searchLink><br /><searchLink fieldCode="DE" term="%22Municipalities%22">Municipalities</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Italy%22">Italy</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/0049124120986199
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241<br />1552-8294
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Spatial microsimulation encompasses a range of alternative methodological approaches for the small area estimation (SAE) of target population parameters from sample survey data down to target small areas in contexts where such data are desired but not otherwise available. Although widely used, an enduring limitation of spatial microsimulation SAE approaches is their current inability to deliver reliable measures of uncertainty--and hence confidence intervals--around the small area estimates produced. In this article, we overcome this key limitation via the development of a measure of uncertainty that takes into account both variance and bias, that is, the mean squared error. This new approach is evaluated via a simulation study and demonstrated in a practical application using European Union Statistics on Income and Living Conditions data to explore income levels across Italian municipalities. Evaluations show that the approach proposed delivers accurate estimates of uncertainty and is robust to nonnormal distributions. The approach provides a significant development to widely used spatial microsimulation SAE techniques.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2023
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1397535
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1397535
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/0049124120986199
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 31
        StartPage: 1785
    Subjects:
      – SubjectFull: Simulation
        Type: general
      – SubjectFull: Geometric Concepts
        Type: general
      – SubjectFull: Computation
        Type: general
      – SubjectFull: Measurement
        Type: general
      – SubjectFull: Error of Measurement
        Type: general
      – SubjectFull: Bias
        Type: general
      – SubjectFull: Income
        Type: general
      – SubjectFull: Municipalities
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: Italy
        Type: general
    Titles:
      – TitleFull: Estimating the Uncertainty of a Small Area Estimator Based on a Microsimulation Approach
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Moretti, Angelo
      – PersonEntity:
          Name:
            NameFull: Whitworth, Adam
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
            – Type: issn-electronic
              Value: 1552-8294
          Numbering:
            – Type: volume
              Value: 52
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1