Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility

Saved in:
Bibliographic Details
Title: Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility
Language: English
Authors: Yoosoon Chang (ORCID 0000-0003-2919-0349), Steven N. Durlauf (ORCID 0000-0001-8699-7056), Bo Hu (ORCID 0009-0001-0552-0617), Joon Y. Park
Source: Sociological Methods & Research. 2025 54(4):1505-1531.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 27
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Nonparametric Statistics, Social Mobility, Parent Influence, Markov Processes, Race, Educational Attainment, Parent Background, Mothers, Birth, Age, Family Income, Models, Probability
Assessment and Survey Identifiers: Panel Study of Income Dynamics
DOI: 10.1177/00491241251339654
ISSN: 0049-1241
1552-8294
Abstract: This article proposes a fully nonparametric model to investigate the dynamics of intergenerational income mobility for discrete outcomes. In our model, an individual's income class probabilities depend on parental income in a manner that accommodates nonlinearities and interactions among various individual and parental characteristics, including race, education, and parental age at childbearing, and so generalizes Markov chain mobility models. We show how the model may be estimated using kernel techniques from machine learning. Utilizing data from the panel study of income dynamics, we show how race, parental education, and mother's age at birth interact with family income to determine mobility between generations.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1485923
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFVriokXj8aK23ObzmA1bJ9AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDGDyEvzsTt_SOPPU8QIBEICBm-rwd_XH7aWPcBY3wX43a0oMb-afzrhiHj5JH5MUBjMS0K7WE5EvVsdxaPvGfBcVgiNDm0yVT2XPIbp-9QH5kOtktGJKsnaZeNvtz7mkuCHwkvvtxft0aDjtM4sDipps0Xl6mpd3cqFOris7nCZvFAddZ0TNEEQOIEJbT0cMoV12nEiNic0594jF2TBkvTCxpGupNIBjIenxX9AA
Text:
  Availability: 1
  Value: <anid>AN0188424696;som01nov.25;2025Oct06.06:36;v2.2.500</anid> <title id="AN0188424696-1">Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility </title> <p>This article proposes a fully nonparametric model to investigate the dynamics of intergenerational income mobility for discrete outcomes. In our model, an individual's income class probabilities depend on parental income in a manner that accommodates nonlinearities and interactions among various individual and parental characteristics, including race, education, and parental age at childbearing, and so generalizes Markov chain mobility models. We show how the model may be estimated using kernel techniques from machine learning. Utilizing data from the panel study of income dynamics, we show how race, parental education, and mother's age at birth interact with family income to determine mobility between generations.</p> <p>Keywords: intergenerational income mobility; ordered multinomial probability model; nonparametric estimation; heterogeneous treatment effects; reproducing kernel Hilbert space; effects of parental education; C10; C50; D30; I24</p> <hd id="AN0188424696-2">Introduction</hd> <p>This article proposes an approach to studying intergenerational income mobility based on a generalization of Markov chain mobility models. We group households into different income categories and assume that the probabilities of an individual from different family backgrounds being in different income categories are determined by various characteristics of a person's parents. Our generalization allows transition probabilities between parent and child to depend nonparametrically on parental characteristics. This type of dependence is natural from the perspective of theories of intergenerational mobility, whether models in which parental income determines investment, models in which parental income determines the neighborhoods and/or schools of children, models in which discrimination creates persistent Black/White mobility differences, or models in which parental skills as determined by education or experience affect the productivity of education on children's human capital formation; see [<reflink idref="bib17" id="ref1">17</reflink>] for the elaboration of the many paths by which Markov transition probabilities will differ across families. This framework enables us to study the joint distribution of parental-child income pairs and therefore income mobility dynamics at the aggregate level.</p> <p>Previous research has extensively explored heterogeneity in transition processes, with race being a standard dimension of analysis; see [<reflink idref="bib14" id="ref2">14</reflink>] and [<reflink idref="bib19" id="ref3">19</reflink>] for older classic studies and [<reflink idref="bib5" id="ref4">5</reflink>] and [<reflink idref="bib6" id="ref5">6</reflink>] for more recent contributions. Our aim is to provide tools that capture this heterogeneity in richer ways than previous studies. To achieve this, we develop a fully nonparametric ordered multinomial probability model that can accommodate general nonlinear relationships between parental income status and the income class into which children move. In addition, we incorporate factors such as race, parental education, and parental age at childbearing to influence the conditional probability structure linking parental and offspring income statuses without relying on functional form assumptions linking offspring income and parental characteristics. There are other approaches to studying heterogeneous effects. For example, [<reflink idref="bib8" id="ref6">8</reflink>] consider group-specific treatment effects, [<reflink idref="bib1" id="ref7">1</reflink>] introduces quantile treatment effects, and one may estimate the joint distribution of the factor and outcome using copulas as by [<reflink idref="bib10" id="ref8">10</reflink>]. These approaches usually study levels instead of class probabilities, and they assume specific functional forms for the interaction structure.</p> <p>The flexibility of our model presents challenges in terms of estimation. It has long been understood that fully nonparametric estimation of a nonlinear model can be exceedingly difficult, especially when dealing with a large number of factors and/or a large sample size. To overcome these difficulties, we employ kernel methods from the machine learning literature coupled with regularization through principal component analysis. These tools have become increasingly prevalent for addressing high-dimensional problems.[<reflink idref="bib6" id="ref9">6</reflink>] By leveraging these methods, we introduce a new approach to estimate our multinomial model, which is robust in environments with large samples and/or a large number of covariates, thereby circumventing the curse of dimensionality while maintaining computational efficiency.</p> <p>We apply our general multinomial probability model to the panel study of income dynamics (PSID) data to examine how gender, race, parental education, parental age at childbirth, and parental income status interact to influence offspring income status. Our analysis reveals significant racial disparities, with Black individuals more likely to fall into the low-income category and less likely to belong to the middle- and high-income categories, particularly among those raised in middle-income families. We also find that parental college education substantially reduces the likelihood of a child being in the low-income category and increases the chances of belonging to the middle- and high-income categories. This positive effect of parental education is particularly pronounced for individuals with middle-to-high-income parents and those born when their parents are in their late twenties to mid-thirties, maximizing the predictive probability of a child attaining high-income status. Collectively, race, parental education, and parental age at childbirth can influence the probabilities of low-income status for children by about 20 percent, given a certain parental income level. This provides compelling evidence of the ways in which heterogeneity in downward mobility can occur for middle-income families.</p> <p>Relationships between family level characteristics and future incomes have of course been extensively studied. [<reflink idref="bib12" id="ref10">12</reflink>] demonstrate nonlinearities in the intergenerational income transmission process. [<reflink idref="bib21" id="ref11">21</reflink>] and [<reflink idref="bib2" id="ref12">2</reflink>] demonstrate racial differences in intergenerational educational mobility. [<reflink idref="bib20" id="ref13">20</reflink>] and [<reflink idref="bib7" id="ref14">7</reflink>] show how a family structure affects intergenerational income effects. Our overall findings are corroborative of past work. Our contribution is to demonstrate the robustness of previous claims when the analysis allows for a much more general structure than past studies. In particular, we are able to show that particular family level factors retain their predictive power when simultaneously considered with other factors and are able to do this by allowing nonlinear relationships between each factor and offspring income category probabilities. Beyond this, our methods give additional substantive insights. In particular, we show that there are stark differences between the children of non-Black families with college-educated parents, and born when their parents are around 30 versus children who are Black, born to parents at the age of 18, and without college degrees. Thus one set of background conditions makes it particularly difficult for some children to achieve socioeconomic success in income and education while another makes it hard to fail.</p> <p>The article is organized as follows: We begin with a brief motivation for our work in the "Beyond Linearity in Intergenerational Mobility Analysis" section. The "Methodology" section introduces our ordered multinomial model for income class probabilities and discusses its estimation and inference. The "Data" section describes the PSID data we use. The "Empirical Results" section applies our methods to explore the effects of various factors on the relationship between parental and offspring's income status. The "Conclusions" section concludes the article. Technical details and robustness checks are provided in four online Appendices, supplemental materials.</p> <hd id="AN0188424696-3">Beyond Linearity in Intergenerational Mobility Analysis</hd> <p>The workhorse model to study intergenerational income mobility is <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>log</mi><mspace width="0.2em" /><mo stretchy="false">(</mo><msub><mi>y</mi><mi>c</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mi>α</mi><mo>+</mo><mi>β</mi><mi>log</mi><mspace width="0.2em" /><mo stretchy="false">(</mo><msub><mi>y</mi><mi>p</mi></msub><mo stretchy="false">)</mo><mo>+</mo><mi>ε</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>c</mi></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>p</mi></msub></math> </ephtml> are specific measures of the child's income and the parental income, respectively, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ε</mi></math> </ephtml> represents individual-specific heterogeneity. The parameter β is the intergenerational elasticity (IGE) of income and has become the primary measure of the persistence of income across generations.</p> <p>As such, this workhorse model has nothing to say about the heterogeneity in intergenerational persistence either across individuals or across time, since β is a constant. Researchers have, therefore, augmented this model to include additional factors. A frequently used regression model takes the form of <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>log</mi><mspace width="0.2em" /><mo stretchy="false">(</mo><msub><mi>y</mi><mi>c</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mi>α</mi><mo>+</mo><mi>β</mi><mi>log</mi><mspace width="0.2em" /><mo stretchy="false">(</mo><msub><mi>y</mi><mi>p</mi></msub><mo stretchy="false">)</mo><mo>+</mo><msup><mi>γ</mi><mrow><mi mathvariant="normal">′</mi></mrow></msup><mi>s</mi><mo>+</mo><mi>ε</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <emph>s</emph> is a vector of factors beyond parental income that are believed to shape a child's income.</p> <p>Although this augmented model enables one to study the effects of factors beyond parental income on intergenerational mobility, it is still very restrictive since it does not allow interactions between different factors in determining the income level of the child. As such, it preserves an implicit dichotomy between the measure of intergenerational mobility, β, and other mechanisms. Social science theory does not justify this independence. For example, it could be the case that the effect of parental income on children's future income is affected by discrimination or by parental education. This has led to a literature that allows β to differ by categories such as race. More generally, products of variables are often taken to capture the interactions of different determinants of offspring outcomes. However, this is not an entirely satisfactory solution, since it amounts to employing a second-order Taylor series approximation of the true interactions of different variables, and there is no basis, from the vantage point for formal theoretical models of intergenerational mobility for thinking such an approximation will be particularly accurate.</p> <p>Our proposed framework can accommodate rich interactions and nonlinearities. We propose a fully nonparametric model to link these factors to the probabilities of a child belonging to different relative income classes. Unlike the IGE model, which focuses on levels of income, we consider probabilities that link the income classes of parents and children. We choose this outcome variable for several reasons. First, our model permits a natural integration of interactions by making income class probabilities functions of various factors. Second, income categories such as the middle class hold a distinct substantive interest as compared to absolute income levels. Third, many of the publicly available income data contain left and/or right-censored observations and might contain zero/negative income figures. Estimating an IGE with censored data might lead to biased estimates, and taking logs with zero and/or negative values could be problematic, even with some of the usual transformation techniques such as adding one before taking logs ([<reflink idref="bib9" id="ref15">9</reflink>]). Our approach, by analyzing income classes instead of income levels, remains robust in the presence of such data issues.</p> <p>In the next section, we propose a fully nonparametric multinomial model that can be used to study the link between various factors and an individual's probabilities of membership into different income classes.</p> <hd id="AN0188424696-4">Methodology</hd> <p></p> <hd id="AN0188424696-5">Multinomial Model for Income Class Probabilities</hd> <p>The evolution of income distributions over time is evident. A pertinent inquiry arises: What factors propel these changes, and how exactly do they impact income distribution dynamics? To address this, we propose a nonparametric ordered multinomial model for mobility across income classes.</p> <p>To be specific, let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>m</mi></math> </ephtml> denote the <emph>m</emph> income classes. In our study, we shall set <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>m</mi><mo>=</mo><mn>3</mn></math> </ephtml> , and let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mn>3</mn></math> </ephtml> represent the low-, middle-, and high-income classes, respectively. We use subscript <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>n</mi></math> </ephtml> to index individuals in our sample and use <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></math> </ephtml> to denote the probability of belonging to class <emph>j</emph> for the individual with covariates <emph>x</emph>. We shall call these probabilities the income class probabilities hereafter. Evidently, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></munderover><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mn>1</mn></math> </ephtml> for all <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>n</mi></math> </ephtml> , indicating that each individual's probabilities across all income classes sum up to one.</p> <p>Individuals' characteristics <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> are related to their income class probabilities by the functions <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mo>⋅</mo><mo stretchy="false">)</mo></math> </ephtml> in the form of an ordered multinomial model <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mrow><mi mathvariant="double-struck">P</mi></mrow></mrow><mrow><mo>{</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>−</mo><mn>1</mn></mrow></msub><mo><</mo><msub><mi>y</mi><mi>i</mi></msub><mo>*</mo><mo>≤</mo><msub><mi>τ</mi><mi>j</mi></msub><mrow><mo maxsize="1.2em" minsize="1.2em">|</mo></mrow><msub><mi>x</mi><mi>i</mi></msub><mo>=</mo><mi>x</mi><mo>}</mo></mrow><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>m</mi></math> </ephtml> , with the convention <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mn>0</mn></msub><mo>=</mo><mo>−</mo><mi mathvariant="normal">∞</mi></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mi>m</mi></msub><mo>=</mo><mi mathvariant="normal">∞</mi></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>i</mi></msub><mo>*</mo></math> </ephtml> is a latent variable that represents the unobservable permanent income of the individual <emph>i</emph>, which depends on the individual covariates <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>x</mi><mi>i</mi></msub></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>τ</mi><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msub></math> </ephtml> are constant income thresholds that determine the categories of permanent income. For convenience, from now on we shall simply call <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>i</mi></msub><mo>*</mo></math> </ephtml> the permanent income. We set the permanent income <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo>*</mo><mo stretchy="false">)</mo></math> </ephtml> to be determined by the covariates <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> through <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>i</mi></msub><mo>*</mo><mo>=</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>+</mo><msub><mi>u</mi><mi>i</mi></msub><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <emph>g</emph> is a nonparametric function to be estimated, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>u</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> is the random component that represents the heterogeneity in permanent income not captured by the covariates <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> . We shall estimate the distribution of the random component nonparametrically.</p> <p>Our framework offers substantial flexibility and generality. We depart from the linear setting by allowing for a general nonlinear form of <emph>g</emph>, taking values in a rich function space. The function space employed in our analysis allows for a precise approximation of any continuous function over a compact subset of its domain. This departure from linearity is not solely about freedom in functional forms; rather, it empowers us to explore the heterogeneous impacts of factors on income distribution, and thus intergenerational mobility. In addition, it facilitates the exploration of intricate interactions among various factors that influence income distributions and intergenerational mobility, far beyond those allowed in conventional linear discrete choice models. The generalities of our approach will be explained further in the "Heterogeneous Effects of Factors" section. Moreover, unlike parametric models such as logit or probit, which assume a Gaussian or logistic distribution for the random component <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>u</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> , our model does not confine the random component to any specific distribution family. This grants us greater adaptability in mirroring real income distributions, for example, allowing the presence of fat tails in the income distribution.</p> <p>The widely used rank-rank model for intergenerational mobility research can be interpreted as a special case of our framework. Specifically, by dividing income into a sufficiently large number of classes (e.g., 100), each class can approximate a specific income rank with minimal error. By using parental income rank as a covariate instead of raw income, the rank-rank model is essentially embedded within our more general setup. Our model, however, offers more than the rank-rank regression: while the rank-rank approach estimates the conditional mean of a child's income rank given the parent's rank, our model estimates the entire conditional distribution of child outcomes by modeling income class probabilities. From this, the conditional mean can be derived, but richer insights—such as distributional shifts or inequality patterns—can also be explored.[<reflink idref="bib7" id="ref16">7</reflink>]</p> <p>To identify the effects of various factors in our model, we may, for unrestricted <emph>g</emph>, set one of the parameters in <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>τ</mi><mn>1</mn></msub><mo>,</mo><msub><mi>τ</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>τ</mi><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> to a fixed number, say <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mn>1</mn></msub><mo>=</mo><mn>0</mn></math> </ephtml> , or we may restrict the level of <emph>g</emph> and allow all the parameters in <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> to vary freely. In the article, we set <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><munder><mo form="prefix" movablelimits="true">inf</mo><mrow><mi>x</mi><mo>∈</mo><mi>D</mi></mrow></munder><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>=</mo><mn>0</mn><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <emph>D</emph> is the support of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> , which imposes a level restriction on the function <emph>g</emph>, so that all the parameters in <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>τ</mi><mn>1</mn></msub><mo>,</mo><msub><mi>τ</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>τ</mi><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msub><mo stretchy="false">)</mo></math> </ephtml> are identified without any further restriction.</p> <p>If the random component <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>u</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> has an invertible cumulative distribution function (CDF) <emph>F</emph>, it follows that <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mi>j</mi></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>=</mo><msup><mi>F</mi><mrow><mo>−</mo><mn>1</mn></mrow></msup><mrow><mo>(</mo><munderover><mo>∑</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>j</mi></munderover><msub><mi>π</mi><mi>k</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>)</mo></mrow><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>m</mi><mo>−</mo><mn>1</mn></math> </ephtml> . This implies that, for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">)</mo></math> </ephtml> given, parametric models such as logit and probit specifying <emph>F</emph> as the CDF of standard Gaussian distribution and logistic distribution, respectively, impose unintended and uninterpretable restrictions on <emph>g</emph> whenever <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>m</mi><mo>></mo><mn>2</mn></math> </ephtml> . This problem does not arise in our model, where we allow the distribution of the random component to be fully nonparametric.</p> <p>We need to introduce an appropriate identification condition to separately identify the unknown function <emph>g</emph> in the systematic component and the distribution of the random component.[<reflink idref="bib8" id="ref17">8</reflink>] In this article, however, they are not separately identified, since our analysis will be focused only on various probabilities. As a consequence, our model should be viewed as one that studies relative mobility.[<reflink idref="bib9" id="ref18">9</reflink>] We leave for our future work the structural analysis based on the function <emph>g</emph> in the systematic component, which is identified by an appropriate identifying restriction.</p> <hd id="AN0188424696-6">Heterogeneous Effects of Factors</hd> <p>The study of income intergenerational mobility is based on the belief that the income status of parents is linked to the adult income status of their children. One intriguing quantitative question is: if one family has a higher income than another by a certain margin, how does this difference affect the likelihood of their offspring belonging to a specific income class in their adulthood?</p> <p>While all multinomial models can offer insights into this question, their efficacy varies. To articulate this more formally, let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></math> </ephtml> represent the probability of an offspring's income falling within class <emph>j</emph>, where <emph>x</emph> denotes the logged parental income—the only factor considered at present for illustration. The partial effect <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">∂</mi><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>/</mo><mi mathvariant="normal">∂</mi><mi>x</mi></math> </ephtml> serves to answer our question by quantifying the increased likelihood of an offspring being in income class <emph>j</emph> if their parents' income were increased by 1% from level <emph>x</emph>. This partial effect, contingent upon the functional form of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub></math> </ephtml> , is potentially heterogeneous across families with different parental income levels. If we employ a linear probability model as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>=</mo><mi>x</mi><mi>β</mi></math> </ephtml> , the partial effect implied is β, which is identical across all families with different parental income levels. If we employ the ordered probit or logit model with a linear <emph>g</emph> function, the partial effect is given by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><mrow><mi mathvariant="normal">∂</mi><mi>x</mi></mrow></mfrac></mrow><mo>=</mo><mrow><mo maxsize="1.2em" minsize="1.2em">[</mo></mrow><mspace width=".1em" /><mi>f</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>−</mo><mi>x</mi><mi>β</mi><mo stretchy="false">)</mo><mo>−</mo><mi>f</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi></mrow></msub><mo>−</mo><mi>x</mi><mi>β</mi><mo stretchy="false">)</mo><mrow><mo maxsize="1.2em" minsize="1.2em">]</mo></mrow><mi>β</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <emph>f</emph> is the probability density function (PDF) of the standard normal distribution or the logistic distribution. This partial effect, although heterogeneous in <emph>x</emph>, depends heavily on the shape of the PDF <emph>f</emph> under consideration. It could be the case that partial effects as functions of <emph>x</emph> with certain shapes cannot be generated from the probit or logit model. In contrast, our approach gives a partial effect <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><mrow><mi mathvariant="normal">∂</mi><mi>x</mi></mrow></mfrac></mrow><mo>=</mo><mo stretchy="false">[</mo><mi>f</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo stretchy="false">)</mo><mo>−</mo><mi>f</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi></mrow></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo stretchy="false">]</mo><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><mrow><mi mathvariant="normal">∂</mi><mi>x</mi></mrow></mfrac></mrow><mo>.</mo></math> </ephtml></p> <p>Graph</p> <p>By allowing for flexible forms of <emph>f</emph> and <emph>g</emph>, we are able to generate heterogeneous partial effects with no restrictions on their shapes if it is viewed as a function of the given covariates <emph>x</emph>.</p> <p>Usually, the covariates consist of multiple factors, denoted as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>x</mi><mi>i</mi></msub><mo>=</mo><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><msup><mo stretchy="false">)</mo><mrow><mi mathvariant="normal">′</mi></mrow></msup></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mi>i</mi></msub></math> </ephtml> is the factor whose heterogeneous impact is of primary interest, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>w</mi><mi>i</mi></msub></math> </ephtml> consists of all other factors considered under our study. The conditional average partial effect (CAPE) of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mi>i</mi></msub></math> </ephtml> on income class probabilities <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub></math> </ephtml> may be formally defined as follows: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mi mathvariant="normal">CAPE</mi></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mrow><mi mathvariant="double-struck">E</mi></mrow></mrow><mrow><mo>[</mo><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mrow><mi mathvariant="normal">∂</mi><msub><mi>z</mi><mi>i</mi></msub></mrow></mfrac></mrow><mrow /><mo>|</mo><mrow /><msub><mi>z</mi><mi>i</mi></msub><mo>=</mo><mi>z</mi><mo>]</mo></mrow><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>evaluated at a particular point <emph>z</emph>. As we vary the evaluation point <emph>z</emph>, we get the CAPE as a function of <emph>z</emph>. Once we obtain an estimator <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></math> </ephtml> for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></math> </ephtml> , we may estimate the heterogeneous average partial effect by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mrow><mi mathvariant="normal">CAPE</mi></mrow><mo>^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mfrac><mn>1</mn><mi>n</mi></mfrac></mrow><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>ρ</mi><mi>i</mi></msub><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mrow><mi mathvariant="normal">∂</mi><mi>z</mi></mrow></mfrac></mrow><msub><mi>K</mi><mi>h</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>−</mo><msub><mi>z</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>ρ</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> are the survey weights,[<reflink idref="bib10" id="ref19">10</reflink>] and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>K</mi><mi>h</mi></msub><mo stretchy="false">(</mo><mo>⋅</mo><mo stretchy="false">)</mo><mo>=</mo><mo stretchy="false">(</mo><mn>1</mn><mo>/</mo><mi>h</mi><mo stretchy="false">)</mo><mi>K</mi><mo stretchy="false">(</mo><mo>⋅</mo><mo>/</mo><mi>h</mi><mo stretchy="false">)</mo></math> </ephtml> is defined with a kernel function <emph>K</emph> and bandwidth parameter <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>h</mi><mo>></mo><mn>0</mn></math> </ephtml> . The kernel function is introduced here to take the local average of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">∂</mi><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>/</mo><mi mathvariant="normal">∂</mi><mi>z</mi></math> </ephtml> in a neighborhood of any given <emph>z</emph>. The standard normal density function is commonly used for the kernel function in this context.[<reflink idref="bib11" id="ref20">11</reflink>]</p> <p>The same idea can be applied to study the heterogeneous treatment effect of certain treatments. For instance, consider a treatment such as a college degree for the parents. It is expected that there exists a disparity in the probability of belonging to a specific income class between children whose parents have or have not obtained a college degree. Moreover, it is plausible that such discrepancy in probabilities might exhibit variations among children raised in families with diverse parental income levels. Exploring these variations can offer insights into how parental college education and other family background factors can interact with each other to determine mobility.</p> <p>To conduct such an analysis, we first partition our covariates into <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>x</mi><mi>i</mi></msub><mo>=</mo><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo>,</mo><msub><mi>d</mi><mi>i</mi></msub><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> . Here, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>z</mi><mi>i</mi></msub></math> </ephtml> represents the contingency variable under investigation (in our example, logged parental income), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>d</mi><mi>i</mi></msub></math> </ephtml> is the treatment variable (1 if at least one of the parents has a college degree, and 0 otherwise), and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>w</mi><mi>i</mi></msub></math> </ephtml> includes all other factors considered in our study. We then compute the conditional average treatment effect (CATE) in the probability gap given a particular parental income level <emph>z</emph>, formally defined as follows: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mi mathvariant="normal">CATE</mi></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mrow><mi mathvariant="double-struck">E</mi></mrow></mrow><mrow><mo>[</mo><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo>,</mo><mn>1</mn><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>−</mo><msub><mi>π</mi><mi>j</mi></msub><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo>,</mo><mn>0</mn><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mrow><mo maxsize="1.2em" minsize="1.2em">|</mo></mrow><msub><mi>z</mi><mi>i</mi></msub><mo>=</mo><mi>z</mi><mo>]</mo></mrow><mo>.</mo></math> </ephtml></p> <p>Graph</p> <p>With a properly estimated income class probability function <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></math> </ephtml> , we may estimate the heterogeneous average treatment effect by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mover><mrow><mi mathvariant="normal">CATE</mi></mrow><mo>^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mfrac><mn>1</mn><mi>n</mi></mfrac></mrow><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>ρ</mi><mi>i</mi></msub><mo stretchy="false">[</mo><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>,</mo><mn>1</mn><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>−</mo><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>,</mo><mn>0</mn><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo stretchy="false">]</mo><msub><mi>K</mi><mi>h</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>−</mo><msub><mi>z</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where we use a kernel function again for local averaging over <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> around a given <emph>z</emph>.</p> <p>To sum up, our general and flexible framework enables us to design analytical tools to capture complex interactions among the factors without imposing functional forms or predefined interaction terms commonly employed in traditional regression techniques. Moreover, as will be shown later, machine learning techniques and tools empower us to estimate and conduct statistical inference in a framework as general and flexible as ours, even when we face a large number of potential factors and a large sample size.</p> <hd id="AN0188424696-7">Maximum Likelihood Estimation</hd> <p>The nonparametric function <emph>g</emph> in the systematic component, the PDF <emph>f</emph> of the random component, and the threshold values <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>τ</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>τ</mi><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msub><msup><mo stretchy="false">)</mo><mrow><mi mathvariant="normal">′</mi></mrow></msup></math> </ephtml> , can be jointly estimated by maximum likelihood estimation for our nonparametric ordered multinomial model. We define the maximum likelihood estimators <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>g</mi><mo stretchy="false">^</mo></mover></mrow><mo>,</mo><mrow><mover><mi>f</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>τ</mi><mo stretchy="false">^</mo></mover></mrow></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>τ</mi><mo stretchy="false">^</mo></mover></mrow><mo>=</mo><mo stretchy="false">(</mo><msub><mrow><mover><mi>τ</mi><mo stretchy="false">^</mo></mover></mrow><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mrow><mover><mi>τ</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msub><msup><mo stretchy="false">)</mo><mrow><mi mathvariant="normal">′</mi></mrow></msup></math> </ephtml> , for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>g</mi><mo>,</mo><mi>f</mi></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>(</mo><mrow><mover><mi>g</mi><mo stretchy="false">^</mo></mover></mrow><mo>,</mo><mrow><mover><mi>f</mi><mo stretchy="false">^</mo></mover></mrow><mo>,</mo><mo stretchy="false">(</mo><msub><mrow><mover><mi>τ</mi><mo stretchy="false">^</mo></mover></mrow><mi>j</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msubsup><mo>)</mo></mrow><mo>=</mo><msub><mrow><mi mathvariant="normal">arg</mi><mspace width=".1em" /><mi mathvariant="normal">max</mi></mrow><mrow><mstyle scriptlevel="1"><mtable columnspacing="0em" displaystyle="false" rowspacing="0.1em"><mtr><mtd><mi>g</mi><mo>∈</mo><mrow><mi mathvariant="script">G</mi></mrow><mo>,</mo><mi>f</mi><mo>∈</mo><mrow><mi mathvariant="script">F</mi></mrow><mo>,</mo></mtd></mtr><mtr><mtd><mi>τ</mi><mo>∈</mo><msup><mrow><mrow><mi mathvariant="double-struck">R</mi></mrow></mrow><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msup></mtd></mtr></mtable></mstyle></mrow></msub><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>ρ</mi><mi>i</mi></msub><mi>ℓ</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><mi>θ</mi><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> contains the parameters <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mi>g</mi><mo>,</mo><mi>f</mi><mo>,</mo><mi>τ</mi><mo stretchy="false">)</mo></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>ρ</mi><mi>i</mi></msub></math> </ephtml> is the survey weight for the <emph>i</emph>-th observation introduced earlier, the log-likelihood function <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ℓ</mi></math> </ephtml> is given by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ℓ</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><mi>θ</mi><mo stretchy="false">)</mo><mo>=</mo><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></munderover><mn>1</mn><mo fence="false" stretchy="false">{</mo><msub><mi>y</mi><mi>i</mi></msub><mo>=</mo><mi>j</mi><mo fence="false" stretchy="false">}</mo><mrow><mo maxsize="1.2em" minsize="1.2em">[</mo></mrow><mi>F</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mi>j</mi></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo stretchy="false">)</mo><mo>−</mo><mi>F</mi><mo stretchy="false">(</mo><msub><mi>τ</mi><mrow><mspace width=".1em" /><mi>j</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo stretchy="false">)</mo><mrow><mo maxsize="1.2em" minsize="1.2em">]</mo></mrow><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>with the convention <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mn>0</mn></msub><mo>=</mo><mo>−</mo><mi mathvariant="normal">∞</mi></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mi>m</mi></msub><mo>=</mo><mi mathvariant="normal">∞</mi></math> </ephtml> , and <emph>F</emph> is the CDF of the distribution given by <emph>f</emph>. In the main text, we estimate the PDF <emph>f</emph> of the random component fully nonparametrically, as will be explained in detail below.[<reflink idref="bib12" id="ref21">12</reflink>]</p> <p>The function <emph>g</emph> in the systematic component of our model is assumed to belong to the class <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="script">G</mi></mrow></math> </ephtml> of functions that are given by any linear combination of a set of basis functions <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>K</mi><mo stretchy="false">(</mo><mo>⋅</mo><mo>,</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mi>exp</mi><mspace width="0.2em" /><mrow><mo maxsize="1.623em" minsize="1.623em">(</mo></mrow><mo>−</mo><mi>κ</mi><mo fence="false" stretchy="false">‖</mo><mo>⋅</mo><mo>−</mo><msub><mi>x</mi><mi>i</mi></msub><msup><mo fence="false" stretchy="false">‖</mo><mn>2</mn></msup><mrow><mo maxsize="1.623em" minsize="1.623em">)</mo></mrow><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>for <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>n</mi></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>κ</mi><mo>></mo><mn>0</mn></math> </ephtml> is a scale parameter and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo fence="false" stretchy="false">‖</mo><mi>z</mi><msup><mo fence="false" stretchy="false">‖</mo><mn>2</mn></msup><mo>=</mo><msup><mi>z</mi><mrow><mi mathvariant="normal">′</mi></mrow></msup><mi>z</mi></math> </ephtml> denotes the squared norm, in the so-called the <emph>reproducing kernel Hilbert space</emph> defined by <emph>K</emph>. The function <emph>K</emph> we use here to generate a functional basis is referred to as a <emph>kernel function</emph>.[<reflink idref="bib13" id="ref22">13</reflink>] The scale parameter <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>κ</mi></math> </ephtml> in the kernel function <emph>K</emph> is a tuning parameter and has to be set a priori. We use a particular kernel function given in (<reflink idref="bib5" id="ref23">5</reflink>), which is most commonly used and called the <emph>radial kernel</emph>, though other choices are also possible. The class <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="script">G</mi></mrow></math> </ephtml> of functions is known to be large enough to approximate any continuous function <emph>g</emph> arbitrarily well over any compact subset of its domain uniformly. Using a linear combination of the basis functions given by (<reflink idref="bib5" id="ref24">5</reflink>) to estimate the function <emph>g</emph> means that we obtain our estimate for <emph>g</emph> essentially by a linear combination of normal densities centered at <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup></math> </ephtml> with the same variance <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>σ</mi><mn>2</mn></msup><mo>=</mo><mn>1</mn><mo>/</mo><mo stretchy="false">(</mo><mn>2</mn><mi>κ</mi><mo stretchy="false">)</mo></math> </ephtml> .</p> <p>We let <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>g</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>=</mo><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>c</mi><mi>j</mi></msub><mi>K</mi><mo stretchy="false">(</mo><mi>x</mi><mo>,</mo><msub><mi>x</mi><mi>j</mi></msub><mo stretchy="false">)</mo></math> </ephtml></p> <p>Graph</p> <p>with a set of coefficients <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>c</mi><mi>j</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup></math> </ephtml> . It is clear that there exists a set of coefficients <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>c</mi><mi>j</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup></math> </ephtml> such that <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>=</mo><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>c</mi><mi>j</mi></msub><mi>K</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>j</mi></msub><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>for all <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mi>n</mi></math> </ephtml> . Indeed, if we define <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>g</mi><mo stretchy="false">]</mo><mo>=</mo><mo stretchy="false">(</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mn>1</mn></msub><mo stretchy="false">)</mo><mo>,</mo><mo>...</mo><mo>,</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>n</mi></msub><mo stretchy="false">)</mo><msup><mo stretchy="false">)</mo><mrow><mi mathvariant="normal">′</mi></mrow></msup></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo><mo>=</mo><mo stretchy="false">(</mo><mi>K</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>j</mi></msub><mo stretchy="false">)</mo><msubsup><mo stretchy="false">)</mo><mrow><mi>i</mi><mo>,</mo><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></msubsup></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>c</mi><mo>=</mo><mo stretchy="false">(</mo><msub><mi>c</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>c</mi><mi>n</mi></msub><msup><mo stretchy="false">)</mo><mrow><mi mathvariant="normal">′</mi></mrow></msup></math> </ephtml> , then we have <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>g</mi><mo stretchy="false">]</mo><mo>=</mo><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo><mi>c</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>from which we may easily obtain such <emph>c</emph> as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>c</mi><mo>=</mo><mo stretchy="false">[</mo><mi>K</mi><msup><mo stretchy="false">]</mo><mrow><mo>−</mo><mn>1</mn></mrow></msup><mo stretchy="false">[</mo><mi>g</mi><mo stretchy="false">]</mo></math> </ephtml> , since <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> is invertible.</p> <p>However, estimating the function <emph>g</emph> as in (<reflink idref="bib6" id="ref25">6</reflink>) with such <emph>c</emph> yields overfitting, and we need to reduce the dimension of <emph>c</emph> through an appropriate regularization method. Note that <emph>c</emph> includes <emph>n</emph> unknown parameters, that is, as many as the sample size. To avoid the problem, we simply set <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>c</mi><mo>=</mo><msub><mi>V</mi><mo>∙</mo></msub><mi>β</mi></math> </ephtml></p> <p>Graph</p> <p>with <emph>p</emph>-dimensional parameter vector β, where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>V</mi><mo>∙</mo></msub></math> </ephtml> is an <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi><mo>×</mo><mi>p</mi></math> </ephtml> matrix whose columns are leading principal components of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> , which are the <emph>p</emph> eigenvectors of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> corresponding to its <emph>p</emph> largest eigenvalues <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>λ</mi><mi>i</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>p</mi></msubsup></math> </ephtml> . This amounts to approximating (<reflink idref="bib7" id="ref26">7</reflink>) as follows: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>g</mi><mo stretchy="false">]</mo><mo>≈</mo><mo fence="false" stretchy="false">⟮</mo><mo stretchy="false">(</mo><mi>g</mi><mo stretchy="false">)</mo><mo fence="false" stretchy="false">⟯</mo><mo>=</mo><msub><mi>V</mi><mo>∙</mo></msub><msub><mi mathvariant="normal">Λ</mi><mo>∙</mo></msub><mi>β</mi><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">Λ</mi><mo>=</mo><mstyle displaystyle="false" scriptlevel="0"><mtext>diag</mtext></mstyle><mspace width=".1em" /><mo stretchy="false">(</mo><msub><mi>λ</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>λ</mi><mi>p</mi></msub><mo stretchy="false">)</mo></math> </ephtml> .[<reflink idref="bib14" id="ref27">14</reflink>] Our approach here is often used in machine learning. See online Appendix A, supplemental materials, for a more detailed discussion.</p> <p>For the PDF <emph>f</emph> of the random component, we follow [<reflink idref="bib18" id="ref28">18</reflink>] and choose <emph>f</emph> in the class <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="script">F</mi></mrow></math> </ephtml> of density functions given by the following equation: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>f</mi><mo stretchy="false">(</mo><mi>u</mi><mo stretchy="false">)</mo><mo>=</mo><mrow><mfrac><mn>1</mn><mi>w</mi></mfrac></mrow><msup><mrow><mo>(</mo><mn>1</mn><mo>+</mo><munderover><mo>∑</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>q</mi></munderover><msub><mi>α</mi><mi>k</mi></msub><msup><mi>u</mi><mi>k</mi></msup><mo>)</mo></mrow><mn>2</mn></msup><mi>ϕ</mi><mo stretchy="false">(</mo><mi>u</mi><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>α</mi><mi>k</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>q</mi></msubsup></math> </ephtml> are the coefficients of polynomial terms, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ϕ</mi></math> </ephtml> is the standard normal PDF, and <emph>w</emph> is a normalization constant given as a function of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>α</mi><mi>k</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>q</mi></msubsup></math> </ephtml> introduced to make <emph>f</emph> a proper PDF. A wide variety of densities can be approximated arbitrarily well by a function of the form in (<reflink idref="bib9" id="ref29">9</reflink>). The class <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="script">F</mi></mrow></math> </ephtml> of PDFs we consider here is broad and includes, for instance, all Hermite polynomials of finite order. Hermite polynomial approximation of the PDF <emph>f</emph> is particularly suitable in our model, where we let <emph>f</emph> have unbounded support. We impose the mean zero restriction <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mo>∫</mo><mrow><mo>−</mo><mi mathvariant="normal">∞</mi></mrow><mi mathvariant="normal">∞</mi></msubsup><mi>u</mi><mi>f</mi><mo stretchy="false">(</mo><mi>u</mi><mo stretchy="false">)</mo><mi>d</mi><mi>u</mi><mo>=</mo><mn>0</mn><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>for the PDF <emph>f</emph>. Our specification of <emph>f</emph> in (<reflink idref="bib9" id="ref30">9</reflink>) with the zero mean restriction in (<reflink idref="bib10" id="ref31">10</reflink>) defines the likelihood function <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>ℓ</mi></math> </ephtml> in (<reflink idref="bib4" id="ref32">4</reflink>) explicitly as a function of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>α</mi><mi>k</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>k</mi><mo>=</mo><mn>1</mn></mrow><mi>q</mi></msubsup></math> </ephtml> . The interested reader is referred to online Appendix B, supplemental materials, for more details.</p> <p>Consequently, our problem of maximizing likelihood function in (<reflink idref="bib3" id="ref33">3</reflink>) reduces to <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover><mi>θ</mi><mo stretchy="false">^</mo></mover></mrow><mo>=</mo><msub><mrow><mi mathvariant="normal">arg</mi><mspace width=".1em" /><mi mathvariant="normal">max</mi></mrow><mrow><mstyle scriptlevel="1"><mtable columnspacing="0em" displaystyle="false" rowspacing="0.1em"><mtr><mtd><mi>β</mi><mo>∈</mo><msup><mrow><mrow><mi mathvariant="double-struck">R</mi></mrow></mrow><mi>p</mi></msup><mo>,</mo><mi>α</mi><mo>∈</mo><msup><mrow><mrow><mi mathvariant="double-struck">R</mi></mrow></mrow><mi>q</mi></msup><mo>,</mo></mtd></mtr><mtr><mtd><mi>τ</mi><mo>∈</mo><msup><mrow><mrow><mi mathvariant="double-struck">R</mi></mrow></mrow><mrow><mi>m</mi><mo>−</mo><mn>1</mn></mrow></msup></mtd></mtr></mtable></mstyle></mrow></msub><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover><msub><mi>ρ</mi><mi>i</mi></msub><mi>ℓ</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><mi>θ</mi><mo stretchy="false">)</mo><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>θ</mi></math> </ephtml> contains <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>+</mo><mi>q</mi><mo>+</mo><mo stretchy="false">(</mo><mi>m</mi><mo>−</mo><mn>1</mn><mo stretchy="false">)</mo></math> </ephtml> parameters in <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><mi>β</mi><mo>,</mo><mi>α</mi><mo>,</mo><mi>τ</mi><mo stretchy="false">)</mo></math> </ephtml> . This is a completely standard problem. Our approach is thus able to handle without extra difficulty the situation when the dimension of the covariate is large and/or the sample size is large. Regardless of how large the dimension of the covariate <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> is, the dimensionality of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> does not pose any problem to our approach. Note that we only need the covariate <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> in the evaluation of the kernel <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>K</mi><mo stretchy="false">(</mo><mo>⋅</mo><mo>,</mo><mo>⋅</mo><mo stretchy="false">)</mo></math> </ephtml> in our approach, and the value of the kernel function is dependent only on the norm <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo fence="false" stretchy="false">‖</mo><msub><mi>x</mi><mi>i</mi></msub><mo>−</mo><msub><mi>x</mi><mi>j</mi></msub><mo fence="false" stretchy="false">‖</mo></math> </ephtml> of the data pairs <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>x</mi><mi>i</mi></msub><mo>,</mo><msub><mi>x</mi><mi>j</mi></msub><mo stretchy="false">)</mo></math> </ephtml> .</p> <p>To select the tuning parameters including the dimension <emph>p</emph> of the parameter β and the dimension <emph>q</emph> of the parameter <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>α</mi></math> </ephtml> , which are needed to regularize our estimator for <emph>g</emph> and estimate the error density function <emph>f</emph>, respectively, we use the cross-validation, which is a standard method for selecting tuning parameters in nonparametric statistics. In the <emph>i</emph>-th iteration of the cross-validation procedure, we construct a sub-sample by leaving out the <emph>i</emph>-th observation, estimate the model with this sub-sample, make predictions for the <emph>i</emph>-th observation based on the estimated model, obtain the predicted probabilities <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub><msubsup><mo stretchy="false">)</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></msubsup></math> </ephtml> for the <emph>m</emph> classes we consider, and finally calculate the sum of squared errors of the predicted probabilities across <emph>m</emph> classes as follows: <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><munderover><mo>∑</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></munderover><msup><mrow><mo>(</mo><msub><mrow><mover><mi>π</mi><mo stretchy="false">^</mo></mover></mrow><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>−</mo><msub><mi>π</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>)</mo></mrow><mn>2</mn></msup><mo>,</mo></math> </ephtml></p> <p>Graph</p> <p>where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>π</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><msubsup><mo stretchy="false">)</mo><mrow><mspace width=".1em" /><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></msubsup></math> </ephtml> is a degenerate distribution that reflects the true class probabilities of the <emph>i</emph>-th observation. Then we obtain the average of the above squared loss computed for each <emph>i</emph> and select the parameter combination that yields the smallest average loss. We search within the range of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mn>10</mn></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>...</mo><mo>,</mo><mn>5</mn></math> </ephtml> and end up with <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>=</mo><mn>7</mn></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>q</mi><mo>=</mo><mn>2</mn></math> </ephtml> . We set the scale parameter <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>κ</mi><mo>=</mo><mn>1</mn><mo>/</mo><mn>2</mn></math> </ephtml> for the kernel function. This seems to be a reasonable choice, given that we follow the usual practice of standardizing the covariates so that they have mean zero and variance one. Setting <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>κ</mi><mo>=</mo><mn>1</mn><mo>/</mo><mn>2</mn></math> </ephtml> means that we use the standard normal densities as basis functions to estimate <emph>g</emph>. Finally, we use a bootstrap procedure to obtain confidence intervals/bands of our estimates. Details of our bootstrap procedure are presented in online Appendix C, supplemental materials. The asymptotic distribution of our nonparametric estimator is not available.</p> <hd id="AN0188424696-8">Data</hd> <p>Our sample is constructed from the PSID. PSID is a comprehensive longitudinal household survey in the United States, tracking individuals and their descendants over several decades and containing variables on the economic, health, educational, and social behavior of individuals and families. Given the survey's time span and the fact that it tracks families across generations, it is one of the most widely used data sets in the study of intergenerational mobility.</p> <p>The PSID was initiated in 1968 and contains annual data from the years 1968 to 1997. Data is available biannually after 1997. We focus on a sample from the years 1968 to 1997 to avoid any inconsistency due to the change in the survey design. Our sample includes individuals who reached an age between 30 and 35 years old (inclusive) during any of our sample periods (1968–1997). We also track their parents and ultimately we consider child-parent pairs in conducting our analysis. To reflect the fact that we are studying such pairs, we shall refer to the individuals in our study as the child from now on.</p> <p>Due to sample size constraints, we use the logged average household income of the head and spouse within the age range of 30–35 (inclusive) for the child as a measure of the child's overall economic status during adulthood. We use household income instead of personal income due to the economic partnership and risk-sharing function of marriage, by which we think that an individual's economic status is better reflected by the household income instead of his or her personal income. We adopt the Pew Research Center's methodology to categorize children into low-, middle-, and high-income classes ([<reflink idref="bib23" id="ref34">23</reflink>]). Specifically, we calculate the median income of the children and set the threshold for low income at two-thirds of this median income and for high income at twice the median income.[<reflink idref="bib15" id="ref35">15</reflink>] The highest income for a low-income family is $14,147 and the lowest income for a high-income family is $42,442, both in 1977 dollars.</p> <p>The factors we propose that may affect the economic status of an individual are parental income, parental age at childbirth, parental college education, child gender, child race, child college education, and child health condition at birth. We measure parental income by the parents' logged average income during the period when the child is between 15 and 20 years old (inclusive). We choose this time span for two reasons. First, it reflects the parents' economic situation during the child's period of dependency, which significantly influences the child's economic status rather than the parents' economic condition after the child becomes economically independent. Given that many children leave home for college or work after adolescence, we focus on income up until the child reaches the age of 20. Second, due to the sample size considerations, we are not able to use the parents' income throughout the entire childhood of the individual. We therefore strike a balance and consider this specific time span.[<reflink idref="bib16" id="ref36">16</reflink>]</p> <p>We also consider parents' average age when the child was born. Reasons we consider this include parental maturity and location of parents in the life cycle of income and overall family resources.</p> <p>It should be noted that income is top-coded in the PSID. Also, there are instances of zero and negative incomes in the data, which could potentially indicate measurement errors. Our method is robust in handling this censored data issue for the dependent variable as we categorize children into income classes rather than analyzing specific income levels. For parental income used as one of the covariates, we simply set the zero or negative incomes to one in our empirical analysis, as often done in the studies of intergenerational mobility.</p> <p>Our benchmark sample consists of a total of 962 child-parent pairs. Survey weights provided by PSID are also used to adjust for sample selection (oversampling of low-income families) and non-random attrition in the PSID survey. Income is deflated by the Consumer Price Index for All Urban Consumers (CPI-U-RS, 1977 = 100), following the usual practice. Table 1 provides the summary statistics of the variables we use in our study.</p> <p>Table 1. Summary Statistics of Variables.</p> <p>Graph</p> <p> <ephtml> <table><colgroup><col align="left" /><col align="left" /><col align="char" /><col align="char" /><col align="char" /><col align="char" /></colgroup><thead><tr><th align="left">Variable</th><th align="center">Type</th><th align="center">Mean</th><th align="center">Std</th><th align="center">Min</th><th align="center">Max</th></tr></thead><tbody><tr><td>Log child income</td><td>Continuous</td><td>9.833</td><td>0.904</td><td>0.000</td><td>11.724</td></tr><tr><td> Low income class</td><td>Dummy</td><td>0.260</td><td>0.438</td><td>0.000</td><td>1.000</td></tr><tr><td> Middle income class</td><td>Dummy</td><td>0.664</td><td>0.472</td><td>0.000</td><td>1.000</td></tr><tr><td> High income class</td><td>Dummy</td><td>0.076</td><td>0.265</td><td>0.000</td><td>1.000</td></tr><tr><td>Log parental income</td><td>Continuous</td><td>9.813</td><td>1.275</td><td>0.000</td><td>13.016</td></tr><tr><td>Male</td><td>Dummy</td><td>0.508</td><td>0.500</td><td>0.000</td><td>1.000</td></tr><tr><td>Child college degree</td><td>Dummy</td><td>0.393</td><td>0.488</td><td>0.000</td><td>1.000</td></tr><tr><td>Parental college degree</td><td>Dummy</td><td>0.312</td><td>0.463</td><td>0.000</td><td>1.000</td></tr><tr><td>Black</td><td>Dummy</td><td>0.060</td><td>0.237</td><td>0.000</td><td>1.000</td></tr><tr><td>Parental age at birth</td><td>Continuous</td><td>28.030</td><td>5.354</td><td>12.000</td><td>46.000</td></tr><tr><td>Underweight at birth</td><td>Dummy</td><td>0.049</td><td>0.216</td><td>0.000</td><td>1.000</td></tr></tbody></table> </ephtml> </p> <hd id="AN0188424696-9">Empirical Results</hd> <p></p> <hd id="AN0188424696-10">Effects of Single Factors</hd> <p>We first employ our framework to analyze heterogeneous effects on children by considering gender, race, and parental education as separate factors, conditioning on different parental income levels. We do this by estimating a single nonparametric multinomial model and then integrating out each variable except income and the factor under consideration using (<reflink idref="bib2" id="ref37">2</reflink>).</p> <p>Figure 1 plots the probability differentials between male and female children, as functions of parental income. In each figure, the vertical axis measures the probability differential between male and female children, and the horizontal axis measures the parental income. Each parental income/gender pair produces a probability differential for membership, by the child, in each of the income classes. The upper-left panel plots the three probability differentials of the child's membership in the low-, middle-, and high-income classes for comparisons based on their parents' incomes. The next three panels plot the three differentials separately with confidence bands. The light and dark gray areas correspond to the 95% and 90% confidence bands, respectively.</p> <p>Graph: Figure 1. Probability differentials between males and females, as functions of parental income.</p> <p>The plots show that males, compared to females, exhibit slightly lower probabilities of entering the low-income class and slightly higher probabilities of entering the high-income class. However, these differentials are generally not statistically significant at the 0.1 significance level, except for children from the poorest families. For those, we observe that males are less (more) likely than females to be in the low (middle) category when they were born into a very poor family as shown in the upper-right and lower-left panels in Figure 1. The effects are small, but statistically significant. For the probability differentials of being in the high-income category, the gender effect is positive for all parental income levels but the magnitude is bigger for those with richer parents. However, the effect is not statistically significant at the 0.1 level. Overall, at best we find weak evidence that parental income has differential effects on the income class of offspring of different genders.</p> <p>Figure 2 illustrates the probability differentials between Black and non-Black children across different parental income levels. The plots show that Black individuals have a higher likelihood of falling into the low-income class and face a comparative disadvantage in accessing the middle- and high-income classes compared to their non-Black counterparts. All these racial differentials are statistically significant. The summary of the differences is straightforward. First, for all parental income levels, Black children are substantially more likely to reside in the low-income category than comparable non-Blacks and are less likely to reside in either the middle or higher-income categories than non-Blacks. By implication, Black families are less able to lock in middle or high incomes than non-Black families while the low-income category is harder to escape for Blacks. These results on lower rates of upward mobility for Blacks are qualitatively similar to [<reflink idref="bib5" id="ref38">5</reflink>] and the results on relatively higher rates of downward mobility for Blacks are qualitatively similar to [<reflink idref="bib11" id="ref39">11</reflink>]. Our ability to generate similar findings when one allows for distinct heterogeneity variables across families is an important corroboration of the salience of race as a distinct source of disparities in mobility.</p> <p>Graph: Figure 2. Probability differentials between Blacks and non-Blacks, as functions of parental income.</p> <p>Figure 3 illustrates the probability differentials based on the parents' possession of a college degree, relative to those without, as a function of parental income. We note here that our estimates capture the direct effect of parental education, as well as its interaction with other factors—including the child's education—but they do not account for the indirect effect of parental education that operates through the child's educational attainment as a mediator. Our results reveal a significant contrast in the probability of a child ending up in the lower income category when the parents have not attended college versus when they have. This effect is especially large for families whose incomes lie in the middle-to-high range of income support, with college degrees making low income among children 5 percent less likely than otherwise. This result suggests a complementarity between parental income and parental education.</p> <p>Graph: Figure 3. Probability differentials between children with parents as college and non-college graduates, as functions of parental income.</p> <p>Figure 4 illustrates the probability differentials between children whose parents have or have not obtained a college degree when the average age of parents at childbirth is allowed to vary. Complementing Figure 3, Figure 4 reveals that the disparities in income class probabilities due to parental college education are most pronounced when parents give birth during their late twenties to mid-thirties.</p> <p>Graph: Figure 4. Probability differentials between children with parents as college and non-college graduates, as functions of parental age at childbirth.</p> <hd id="AN0188424696-11">Combining Factors</hd> <p>Combining our single factor analyses in the previous section reveals the existence of family background configurations that make it challenging for a child to avoid the low-income category. Our findings from these analyses indicate the presence of a privileged group of children, originating from non-Black families with college-educated parents, and born when their parents are around 30. This group of children contrasts sharply with children who are Black, born to parents at the age of 18, and without college degrees.</p> <p>Figure 5 presents the probability differentials as a function of parental incomes. Particularly striking are the results in the upper-right panel, where the probability of a child being in the low-income category consistently remains 10 percent higher for our disadvantaged category across all family income levels. For middle-to-high-income categories, this disparity exceeds 20 percent. These findings provide further insight into how the probability of downward mobility varies across different demographic groups.</p> <p>Graph: Figure 5. Probability differentials (multiple treatments), as functions of parental income.</p> <p>All of our empirical results are based on the particular sample we use from the PSID. As with all studies relying on survey data, issues related to representativeness and missing data—particularly when sample selection is present—can introduce bias into the estimates. An additional concern is measurement error, especially in our outcome variable. Due to data availability constraints, we use children's income in their early thirties, which may systematically understate long-term earnings potential and thus permanent income.</p> <p>To mitigate the influence of this measurement issue, we model a latent permanent income variable, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>y</mi><mo>*</mo></math> </ephtml> , and focus on income class—a relative measure—rather than income level. In scenarios where underreporting or transitory income shocks affect all individuals similarly, income class and rank may be less sensitive to such errors than absolute income levels.</p> <p>While our modeling choices help to reduce the impact of measurement error, we acknowledge that such error remains and may still bias our estimates. A more comprehensive treatment of sampling and measurement issues would require a detailed investigation of the PSID's sampling structure and possibly additional methodological tools, which is beyond the scope of the current article. We present these limitations to help contextualize our findings and to suggest directions for future research.</p> <hd id="AN0188424696-12">Conclusions</hd> <p>This article proposes a fully nonparametric multinomial outcome model to study intergenerational income mobility. Our approach effectively captures the nonlinear and interactive effects of various factors on personal income status and societal mobility levels. It demonstrates strong computational efficiency and robustness, particularly suitable for analyzing large datasets with high-dimensional covariates. We affirm race, parental education, and parental childbearing age as crucial determinants influencing intergenerational mobility. Each of these factors significantly impacts the predictive power of parental income for the incomes of children. These findings are all consistent with the overall state of the mobility literature. The robustness of these claims to the general nonparametric framework we set up reinforces the vision that they simultaneously matter.</p> <p>Our findings, which highlight the distinct relationships shaped by race, parental education, and parental childbearing age, underscore the importance of systematically investigating <emph>bottlenecks</emph> in intergenerational mobility. By bottlenecks, we refer to a set of family background variables that perpetuate low incomes across generations, where higher incomes alone may not suffice to break such persistence. The differences between children of non-Black families with college-educated parents, and born to parents in their early thirties as opposed to children who are Black, born to parents around 18, without college degrees are stark. These speak to the idea of bottlenecks in the income dynamics where some set of conditions during childhood makes socioeconomic successes in income and education highly unlikely. These phenomena represent a natural stochastic generalization of poverty trap models. Currently, we are actively pursuing further research on this topic.</p> <p>While our analysis is statistical and, therefore, does not directly identify underlying mechanisms, we note that it does speak to theories of intergenerational mobility. First, if one contrasts the classical economic models of mobility due to [<reflink idref="bib3" id="ref40">3</reflink>], [<reflink idref="bib4" id="ref41">4</reflink>]) with the classic sociological perspective associated with the Wisconsin Status Attainment Model, [<reflink idref="bib25" id="ref42">25</reflink>] and [<reflink idref="bib24" id="ref43">24</reflink>], we think the interactions between educational status, ethnicity and age suggest the importance of placing multiple channels at the heart of the sociology approach. This highlights the value of recent advances in formal economic mobility models that include social and psychological factors, see [<reflink idref="bib17" id="ref44">17</reflink>] for elaboration. Second, the stark contrasts we estimate between the effects of parental income on offspring for parents with different ethnicities, educations, and ages are suggestive of mechanisms that apply to groups rather than individuals. By this, theories of categorical inequality ([<reflink idref="bib22" id="ref45">22</reflink>]; [<reflink idref="bib26" id="ref46">26</reflink>]), or theories in economics based on group memberships ([<reflink idref="bib13" id="ref47">13</reflink>]; [<reflink idref="bib15" id="ref48">15</reflink>], [<reflink idref="bib16" id="ref49">16</reflink>]) are suggestive of the types of patterns we find. To be clear, to say more will require even richer models than we consider here. One natural path involves explicit attention to neighborhood effects along the lines pursued by [<reflink idref="bib27" id="ref50">27</reflink>]. Expanding our tools in these directions is underway.</p> <hd id="AN0188424696-13">Supplemental Material</hd> <p>Graph: Supplemental material, sj-pdf-1-smr-10.1177_00491241251339654 for Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility by Yoosoon Chang, Steven N. Durlauf, Bo Hu and Joon Y. Park in Sociological Methods & Research</p> <hd id="AN0188424696-14">Acknowledgements</hd> <p>The authors gratefully acknowledge comments and suggestions from the editors and four anonymous referees. We would like to thank Kristina Butaeva, Ruli Xiao, Guo Yan, Liyuan Yang, and participants in the <emph>Conference on Inequality and Mobility</emph>, co-organized by the Stone Center for Research on Wealth Inequality and Mobility (SCRWIM) at the University of Chicago and the Institute of Economic Research at Seoul National University, as well as the <emph>Conference on Intergenerational Mobility</emph> jointly organized by SCRWIM at University of Chicago and the Stone Center on Inequality Dynamics at University of Michigan, for their comments. This article is part of the research activities at the Centre for Applied Macroeconomics and Commodity Prices (CAMP) at the BI Norwegian Business School.</p> <ref id="AN0188424696-15"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref7" type="bt">1</bibl> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref12" type="bt">2</bibl> <bibtext> The author(s) received no financial support for the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref33" type="bt">3</bibl> <bibtext> Yoosoon Chang https://orcid.org/0000-0003-2919-0349 Steven N. Durlauf https://orcid.org/0000-0001-8699-7056 Bo Hu https://orcid.org/0009-0001-0552-0617</bibtext> </blist> <blist> <bibl id="bib4" idref="ref32" type="bt">4</bibl> <bibtext> The data, replication code for this article, and a MATLAB package for nonparametric estimation of ordered multinomial models are available online at: https://github.com/bohulab/het-mobility.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref4" type="bt">5</bibl> <bibtext> Supplemental material and and Appendix for this article are available online.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref5" type="bt">6</bibl> <bibtext> Methodologically, our article is a multinomial extension of [28], which develops a framework to analyze fully nonparametric large dimensional binary models.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref14" type="bt">7</bibl> <bibtext> We note that in practice, our model may not serve as a simple drop-in replacement for the rank-rank regression. Accurately estimating a large number of conditional probabilities nonparametrically requires substantial amounts of data, whereas estimating a conditional mean is generally less data-intensive.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref6" type="bt">8</bibl> <bibtext> The reader is referred to, for example, [28] for a detailed discussion on the required identification condition for discrete outcome models.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref15" type="bt">9</bibl> <bibtext> We note that the general methodology we use for nonparametric estimation can be applied to study absolute mobility, but under a different model. If a reasonable measure or proxy for the child's permanent income <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>c</mi></msub></math> </ephtml> is available, we may specify a model such as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>c</mi></msub><mo>=</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>p</mi></msub><mo>,</mo><mi>x</mi><mo stretchy="false">)</mo><mo>+</mo><mi>u</mi></math> </ephtml> where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>y</mi><mi>p</mi></msub></math> </ephtml> is the parental permanent income and <emph>x</emph> is a set of covariates, <emph>g</emph> is a function to be estimated parametrically, and <emph>u</emph> is an error term with distribution function <emph>F</emph> to be estimated nonparametrically. The quantity <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML">1<mo>−</mo><mi>F</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>p</mi></msub><mo>−</mo><mi>g</mi><mo stretchy="false">(</mo><msub><mi>y</mi><mi>p</mi></msub><mo>,</mo><mi>x</mi><mo stretchy="false">)</mo><mo stretchy="false">)</mo></math> </ephtml> measures absolute mobility, representing the proportion of children whose income exceeds that of their parents.</bibtext> </blist> <blist> <bibtext> These weights are provided by our data set and used in our empirical study to adjust for sample selection and non-random attrition, as will be explained later in the "Data" section.</bibtext> </blist> <blist> <bibtext> However, the uniform kernel, which is given by <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>K</mi><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo>1<mo fence="false" stretchy="false">{</mo><mo fence="false" stretchy="false">|</mo><mi>z</mi><mo fence="false" stretchy="false">|</mo><mo>≤</mo>1<mo>/</mo>2<mo fence="false" stretchy="false">}</mo></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>K</mi><mi>h</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>=</mo><mo stretchy="false">(</mo>1<mo>/</mo><mi>h</mi><mo stretchy="false">)</mo><mo fence="false" stretchy="false">{</mo><mo fence="false" stretchy="false">|</mo><mi>z</mi><mo fence="false" stretchy="false">|</mo><mo>≤</mo><mi>h</mi><mo>/</mo>2<mo fence="false" stretchy="false">}</mo></math> </ephtml> , makes it more clear what the kernel function does here. If it is used, we take the local average of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">∂</mi><msub><mrow><mi>π</mi><mo stretchy="false">^</mo></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo>,</mo><msub><mi>w</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>/</mo><mi mathvariant="normal">∂</mi><mi>z</mi></math> </ephtml> to estimate <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mrow><mrow><mi mathvariant="normal">CAPE</mi></mrow><mo>^</mo></mrow><mi>j</mi></msub><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo></math> </ephtml> over the values of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>z</mi><mi>i</mi></msub><mo stretchy="false">)</mo></math> </ephtml> such that <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>z</mi><mo>−</mo><mi>h</mi><mo>/</mo>2<mo>≤</mo><msub><mi>z</mi><mi>i</mi></msub><mo>≤</mo><mi>z</mi><mo>+</mo><mi>h</mi><mo>/</mo>2</math> </ephtml> for a small value of <emph>h</emph>.</bibtext> </blist> <blist> <bibtext> For comparison, we also consider the case with <emph>f</emph> being the PDF of the standard normal distribution and the standard logistic distribution. See online Appendix D.1, supplemental materials, for the details. Their results are similar, quantitatively as well as qualitatively, to the benchmark results based on a fully nonparametric approach. Nevertheless, our fully nonparametric approach yields narrower confidence bands in most cases.</bibtext> </blist> <blist> <bibtext> The kernel function is used here to generate a space of functions defined as a reproducing kernel Hilbert space, and it is totally different from the kernel function we introduce earlier for local averaging.</bibtext> </blist> <blist> <bibtext> Since <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> is a symmetric matrix, we may represent it as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo><mo>=</mo><mi>V</mi><mi mathvariant="normal">Λ</mi><mi>V</mi><mrow><mi mathvariant="normal">′</mi></mrow></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">Λ</mi></math> </ephtml> is the diagonal matrix of the eigenvalues <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>λ</mi><mi>i</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>i</mi><mo>=</mo>1</mrow><mi>n</mi></msubsup></math> </ephtml> of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> and <emph>V</emph> is the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi><mo>×</mo><mi>n</mi></math> </ephtml> -orthogonal matrix of the eigenvectors of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">[</mo><mi>K</mi><mo stretchy="false">]</mo></math> </ephtml> associated with the eigenvalues <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mo stretchy="false">(</mo><msub><mi>λ</mi><mi>i</mi></msub><msubsup><mo stretchy="false">)</mo><mrow><mi>i</mi><mo>=</mo>1</mrow><mi>n</mi></msubsup></math> </ephtml> . The matrices <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>V</mi><mo>∙</mo></msub></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi mathvariant="normal">Λ</mi><mo>∙</mo></msub></math> </ephtml> introduced here are <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi><mo>×</mo><mi>p</mi></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi><mo>×</mo><mi>p</mi></math> </ephtml> leading submatrices of <emph>V</emph> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="normal">Λ</mi></math> </ephtml> , respectively.</bibtext> </blist> <blist> <bibtext> In online Appendix D.3, supplemental materials, we also present results using tertiles, quartiles, and quintiles to classify child income groups. Classification based on the quantiles implies constant shares of income groups across generations, whereas the classification used here allows the shares of income groups to vary across generations.</bibtext> </blist> <blist> <bibtext> We also present in online Appendix D.2, supplemental materials, the results using the age 10–15 logged parental income as an alternative measure of parental income. The two measures of parental income generate similar results.</bibtext> </blist> <blist> <bibtext> [4]</bibtext> </blist> </ref> <ref id="AN0188424696-16"> <title> References </title> <blist> <bibtext> Abadie A., Angrist J., Imbens G. 2002. " Instrumental Variables Estimates of the Effect of Subsidized Training on the Quantiles of Trainee Earnings." Econometrica. 70(1): 91–117.</bibtext> </blist> <blist> <bibtext> Bacic R., Zheng A. 2024. " Race and the Income-achievement Gap." Economic Inquiry. 62(1): 5–23.</bibtext> </blist> <blist> <bibtext> Becker G.S., Tomes N. 1979. " An Equilibrium Theory of the Distribution of Income and Intergenerational Mobility." Journal of Political Economy. 87(6): 1153–1189.</bibtext> </blist> <blist> <bibtext> Becker G.S., Tomes N. 1986. " Human Capital and the Rise and Fall of Families." Journal of Labor Economics. 4(3): 1–39.</bibtext> </blist> <blist> <bibtext> Bhattacharya D., Mazumder B. 2011. " A Nonparametric Analysis of Black–White Differences in Intergenerational Income Mobility in the United States." Quantitative Economics. 2: 335–379.</bibtext> </blist> <blist> <bibtext> Bloome D. 2014. " Racial Inequality Trends and the Intergenerational Persistence of Income and Family." American Sociological Review. 79(6): 1196–1225.</bibtext> </blist> <blist> <bibtext> Bloome D. 2017. " Childhood Family Structure and Intergenerational Income Mobility in the United States." Demography. 54(2): 541–569.</bibtext> </blist> <blist> <bibtext> Brand J.E., Xie Y. 2010. " Who Benefits Most From College?: Evidence for Negative Selection in Heterogeneous Economic Returns to Higher Education." American Sociological Review. 75(2): 273–302.</bibtext> </blist> <blist> <bibtext> Chen J., Roth J. 2024. " Logs with Zeros? Some Problems and Solutions." Quarterly Journal of Economics. 139(2): 891–936.</bibtext> </blist> <blist> <bibtext> Chetty R., Grusky D., Hell M., Hendren N., Manduca R., Narang J. 2017. " The Fading American Dream: Trends in Absolute Income Mobility Since 1940." Science. 356: 398–406.</bibtext> </blist> <blist> <bibtext> Chetty R., Hendren N., Jones M., Porter S. 2020. " Race and Economic Opportunity in the United States: An Intergenerational Perspective." Quarterly Journal of Economics. 135(2): 711–783.</bibtext> </blist> <blist> <bibtext> Dahl G., Lochner L. 2012. " The Impact of Family Income on Child Achievement: Evidence From the Earned Income Tax Credit." American Economic Review. 102(5): 1927–1956.</bibtext> </blist> <blist> <bibtext> Darity W. 2022. " Position and Possession: Stratification Economics and Intergroup Inequality." Journal of Economic Literature. 60: 713–736.</bibtext> </blist> <blist> <bibtext> Duncan O.D. 1968. " Patterns of Occupational Mobility Among Negro Men." Demography. 5: 11–22.</bibtext> </blist> <blist> <bibtext> Durlauf S.N. 1999. " The Memberships Theory of Inequality: Ideas and Implications. " In Elites, Minorities, and Economic Growth. edited by Brezis E., Temin P., North Holland.</bibtext> </blist> <blist> <bibtext> Durlauf S.N. 2006. " Groups, Social Influences, and Inequality: A Memberships Theory Perspective on Poverty Traps. " In Poverty Traps. edited by Bowles S., Durlauf S., Hoff K., Princeton University Press.</bibtext> </blist> <blist> <bibtext> Durlauf S.N., Kourtellos A., Tan C.M. 2022. " The Great Gatsby Curve." Annual Review of Economics. 14: 571–605.</bibtext> </blist> <blist> <bibtext> Gallant A.R., Nychka D.W. 1987. " Semi-Nonparametric Maximum Likelihood Estimation." Econometrica. 55(2): 363–390.</bibtext> </blist> <blist> <bibtext> Hout M. 1984. " Occupational Mobility of Black Men: 1962 to 1973." American Sociological Review. 49: 308–322.</bibtext> </blist> <blist> <bibtext> Lopoo L.M., DeLeire T. 2014. " Family Structure and the Economic Wellbeing of Children in Youth and Adulthood." Social Science Research. 43: 30–44.</bibtext> </blist> <blist> <bibtext> Maralani V. 2013. " The Demography of Social Mobility: Black-White Differences in the Process of Educational Reproduction." American Journal of Sociology. 118(6): 1509–1558.</bibtext> </blist> <blist> <bibtext> Massey D. 2007. Categorically Unequal. New York: Russell Sage Foundation Press.</bibtext> </blist> <blist> <bibtext> Pew Research Center. 2020. Most Americans say there is too much economic inequality in the U.S., but fewer than half call it a top priority. Research Report.</bibtext> </blist> <blist> <bibtext> Sewell W., Haller A., Ohlendorf G. 1970. " The Educational and Early Occupational Status Attainment Process: Replication and Revision." American Sociological Review. 35: 1014–1024.</bibtext> </blist> <blist> <bibtext> Sewell W., Haller A., Portes A. 1969. " The Educational and Early Occupational Attainment Process." American Sociological Review. 34: 82–92.</bibtext> </blist> <blist> <bibtext> Tilly C. 1998. Durable Iequality. Berkeley: University of California Press.</bibtext> </blist> <blist> <bibtext> Wodtke G., Harding D., Elwert F. 2016. " Neighborhood Effect Heterogeneity by Family Income and Developmental Period." American Journal of Sociology. 121: 1168–1222.</bibtext> </blist> <blist> <bibtext> Yan G. 2023. Nonparametric Estimation of Large Dimensional Binary Choice Models. Manuscript.</bibtext> </blist> </ref> <aug> <p>By Yoosoon Chang; Steven N. Durlauf; Bo Hu and Joon Y. Park</p> <p>Reported by Author; Author; Author; Author</p> <p></p> <p>Yoosoon Chang is Professor of Economics and Director of the Economic Machine Learning Lab at Indiana University. Her research includes econometrics, machine learning, finance, energy, inequality, and intergenerational mobility.</p> <p>Steven N. Durlauf is Frank P. Hixon Distinguished Service Professor, Harris School of Public Policy and Director of the Stone Center for Research on Wealth Inequality and Mobility, University of Chicago. He has done extensive research on inequality and intergenerational mobility.</p> <p>Bo Hu is Visiting Assistant Professor of Economics at Indiana University. His research includes econometrics, machine learning, functional data analysis, inequality, and intergenerational mobility.</p> <p>Joon Y. Park is Wisnewsky Professor of Human Studies and Professor of Economics at Indiana University. His research includes econometric theory, time series, financial econometrics, machine learning, and functional data analysis.</p> </aug> <nolink nlid="nl1" bibid="bib17" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib14" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib19" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib10" firstref="ref8"></nolink> <nolink nlid="nl5" bibid="bib12" firstref="ref10"></nolink> <nolink nlid="nl6" bibid="bib21" firstref="ref11"></nolink> <nolink nlid="nl7" bibid="bib20" firstref="ref13"></nolink> <nolink nlid="nl8" bibid="bib11" firstref="ref20"></nolink> <nolink nlid="nl9" bibid="bib13" firstref="ref22"></nolink> <nolink nlid="nl10" bibid="bib18" firstref="ref28"></nolink> <nolink nlid="nl11" bibid="bib23" firstref="ref34"></nolink> <nolink nlid="nl12" bibid="bib15" firstref="ref35"></nolink> <nolink nlid="nl13" bibid="bib16" firstref="ref36"></nolink> <nolink nlid="nl14" bibid="bib25" firstref="ref42"></nolink> <nolink nlid="nl15" bibid="bib24" firstref="ref43"></nolink> <nolink nlid="nl16" bibid="bib22" firstref="ref45"></nolink> <nolink nlid="nl17" bibid="bib26" firstref="ref46"></nolink> <nolink nlid="nl18" bibid="bib27" firstref="ref50"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1485923
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yoosoon+Chang%22">Yoosoon Chang</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2919-0349">0000-0003-2919-0349</externalLink>)<br /><searchLink fieldCode="AR" term="%22Steven+N%2E+Durlauf%22">Steven N. Durlauf</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-8699-7056">0000-0001-8699-7056</externalLink>)<br /><searchLink fieldCode="AR" term="%22Bo+Hu%22">Bo Hu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0001-0552-0617">0009-0001-0552-0617</externalLink>)<br /><searchLink fieldCode="AR" term="%22Joon+Y%2E+Park%22">Joon Y. Park</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. 2025 54(4):1505-1531.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 27
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Nonparametric+Statistics%22">Nonparametric Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Social+Mobility%22">Social Mobility</searchLink><br /><searchLink fieldCode="DE" term="%22Parent+Influence%22">Parent Influence</searchLink><br /><searchLink fieldCode="DE" term="%22Markov+Processes%22">Markov Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Race%22">Race</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Attainment%22">Educational Attainment</searchLink><br /><searchLink fieldCode="DE" term="%22Parent+Background%22">Parent Background</searchLink><br /><searchLink fieldCode="DE" term="%22Mothers%22">Mothers</searchLink><br /><searchLink fieldCode="DE" term="%22Birth%22">Birth</searchLink><br /><searchLink fieldCode="DE" term="%22Age%22">Age</searchLink><br /><searchLink fieldCode="DE" term="%22Family+Income%22">Family Income</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Probability%22">Probability</searchLink>
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: <searchLink fieldCode="SU" term="%22Panel+Study+of+Income+Dynamics%22">Panel Study of Income Dynamics</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/00491241251339654
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241<br />1552-8294
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This article proposes a fully nonparametric model to investigate the dynamics of intergenerational income mobility for discrete outcomes. In our model, an individual's income class probabilities depend on parental income in a manner that accommodates nonlinearities and interactions among various individual and parental characteristics, including race, education, and parental age at childbearing, and so generalizes Markov chain mobility models. We show how the model may be estimated using kernel techniques from machine learning. Utilizing data from the panel study of income dynamics, we show how race, parental education, and mother's age at birth interact with family income to determine mobility between generations.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1485923
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1485923
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/00491241251339654
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 27
        StartPage: 1505
    Subjects:
      – SubjectFull: Nonparametric Statistics
        Type: general
      – SubjectFull: Social Mobility
        Type: general
      – SubjectFull: Parent Influence
        Type: general
      – SubjectFull: Markov Processes
        Type: general
      – SubjectFull: Race
        Type: general
      – SubjectFull: Educational Attainment
        Type: general
      – SubjectFull: Parent Background
        Type: general
      – SubjectFull: Mothers
        Type: general
      – SubjectFull: Birth
        Type: general
      – SubjectFull: Age
        Type: general
      – SubjectFull: Family Income
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Probability
        Type: general
      – SubjectFull: Panel Study of Income Dynamics
        Type: general
    Titles:
      – TitleFull: Accounting for Individual-Specific Heterogeneity in Intergenerational Income Mobility
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yoosoon Chang
      – PersonEntity:
          Name:
            NameFull: Steven N. Durlauf
      – PersonEntity:
          Name:
            NameFull: Bo Hu
      – PersonEntity:
          Name:
            NameFull: Joon Y. Park
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
            – Type: issn-electronic
              Value: 1552-8294
          Numbering:
            – Type: volume
              Value: 54
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1