Moving beyond Linear Regression: Implementing and Interpreting Quantile Regression Models with Fixed Effects

Saved in:
Bibliographic Details
Title: Moving beyond Linear Regression: Implementing and Interpreting Quantile Regression Models with Fixed Effects
Language: English
Authors: Fernando Rios-Avila, Michelle Lee Maroto (ORCID 0000-0002-7506-0046)
Source: Sociological Methods & Research. 2024 53(2):639-682.
Availability: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
Peer Reviewed: Y
Page Count: 44
Publication Date: 2024
Document Type: Journal Articles
Reports - Descriptive
Descriptors: Regression (Statistics), Research Methodology, Alternative Assessment, Models, Scores, Mothers, Income, Data Analysis
DOI: 10.1177/00491241211036165
ISSN: 0049-1241
1552-8294
Abstract: Quantile regression (QR) provides an alternative to linear regression (LR) that allows for the estimation of relationships across the distribution of an outcome. However, as highlighted in recent research on the motherhood penalty across the wage distribution, different procedures for conditional and unconditional quantile regression (CQR, UQR) often result in divergent findings that are not always well understood. In light of such discrepancies, this paper reviews how to implement and interpret a range of LR, CQR, and UQR models with fixed effects. It also discusses the use of Quantile Treatment Effect (QTE) models as an alternative to overcome some of the limitations of CQR and UQR models. We then review how to interpret results in the presence of fixed effects based on a replication of Budig and Hodges's work on the motherhood penalty using NLSY79 data.
Abstractor: As Provided
Entry Date: 2024
Accession Number: EJ1422473
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFj73TL_DwU-n4lkH_u-qExAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDHvgL-N8kIPwDa3TwIBEICBm7ZR5GBGhMBlnv2YzehoD7-49yOYgxx9nlwl8Tp4E__VBhfh8Ax42wFJoFuZ-ed_JbmJ4M9UFjMO79utOtWSVf7xXioetRG4vr0N0zW086MYSaTiG0_6zq3TkF0VGsc64D-5qQZRiRGcQaoQZu5t_xDfGVjLBtKJwXpfmiCdurhJyww3LpZNcOJKzAZZqCDD355kI7Hpb1zeRhTP
Text:
  Availability: 1
  Value: <anid>AN0176861662;som01may.24;2024Apr30.03:38;v2.2.500</anid> <title id="AN0176861662-1">Moving Beyond Linear Regression: Implementing and Interpreting Quantile Regression Models With Fixed Effects </title> <p>Quantile regression (QR) provides an alternative to linear regression (LR) that allows for the estimation of relationships across the distribution of an outcome. However, as highlighted in recent research on the motherhood penalty across the wage distribution, different procedures for conditional and unconditional quantile regression (CQR, UQR) often result in divergent findings that are not always well understood. In light of such discrepancies, this paper reviews how to implement and interpret a range of LR, CQR, and UQR models with fixed effects. It also discusses the use of Quantile Treatment Effect (QTE) models as an alternative to overcome some of the limitations of CQR and UQR models. We then review how to interpret results in the presence of fixed effects based on a replication of Budig and Hodges's work on the motherhood penalty using NLSY79 data.</p> <p>Keywords: quantile regression; motherhood penalty; quantitative methods; longitudinal data; fixed effects</p> <hd id="AN0176861662-2">Introduction</hd> <p> <emph>How does motherhood affect women's earnings?</emph> This has been an important question for social scientists, as child-rearing is a large contributing factor for the gender earnings gap. The effects of motherhood on earnings are particularly hard to isolate because of selection biases, confounders, and unobserved factors. Not all women become mothers, and motherhood is associated with a range of other factors, many of which cannot be directly observed. Although the literature has converged around an estimate of a 5–8 percent motherhood penalty for each additional child, recent studies show that this penalty likely varies across a mother's earnings distribution and by education levels (Avellar and Smock 2003; [<reflink idref="bib8" id="ref1">8</reflink>]; [<reflink idref="bib6" id="ref2">6</reflink>], [<reflink idref="bib7" id="ref3">7</reflink>]; [<reflink idref="bib13" id="ref4">13</reflink>]; [<reflink idref="bib20" id="ref5">20</reflink>]; [<reflink idref="bib22" id="ref6">22</reflink>]; [<reflink idref="bib40" id="ref7">40</reflink>]). Discrepancies in findings are partially explained by the use of different methodologies that implicitly answer different sets of questions.</p> <p>In light of these discrepancies, how should a researcher go about answering the question, how does motherhood affect women's earnings? Or, more precisely, how do women's earnings differ when comparing earnings distributions between mothers and nonmothers? Or, how does the female wage distribution change when more women become mothers?</p> <p>Traditionally, linear regression (LR) analysis has been the conventional method for answering questions such as these among social scientists. Under the classical linear model assumptions (CLMA; i.e., linearity, random sampling, no perfect collinearity, zero conditional mean, homoscedasticity, and normality), linear regression provides the best unbiased estimator for the expected change in a dependent variable, <emph>y</emph>, associated with a unit change in the independent variable, <emph>x</emph>, and conditional on all other controls remaining constant ([<reflink idref="bib43" id="ref8">43</reflink>]). This is also known as the marginal or partial effect of <emph>x</emph> on <emph>y</emph> because all other factors that affect earnings (observed and unobserved) are assumed constant. Under the CLMA, the estimated change can even be interpreted as a causal effect <emph>x</emph> has on <emph>y.</emph>[<reflink idref="bib5" id="ref9">5</reflink>] Although the classical assumptions are often hard to meet in practice, even with some deviations from these, namely, under heteroscedasticity and nonnormality of errors, linear regression models provide unbiased estimates of the average effects for the relationship between <emph>x</emph> and <emph>y</emph>.[<reflink idref="bib6" id="ref10">6</reflink>] As a result, linear regression analysis is often the starting point in most quantitative analyses and much of what social scientists "know" has been built on studies focused on the so-called average person.</p> <p>With the ever-growing availability of longitudinal or panel data, the use of linear models that control for one or more high dimensional fixed effects has increased. This is important as it allows researchers to control for otherwise unobserved heterogeneity, making causal interpretations more reasonable.[<reflink idref="bib7" id="ref11">7</reflink>] For the analysis of earnings and motherhood, for example, individual fixed effects control for unobserved time-constant characteristics, including factors like skill or desire to be a parent. Linear regression models allow for the inclusion of fixed effects by explicitly adding dummy variable sets in the model specification or partialling out the fixed effects before implementing the data analysis.</p> <p>Even when accounting for individual-level fixed effects, however, LR does not necessarily provide the best estimator for summarizing the relationship between two variables in settings where assumptions of homoscedasticity and normality do not hold.[<reflink idref="bib8" id="ref12">8</reflink>] For example, studies often involve the analysis of skewed variables with many outliers. This is especially true for earnings, which often includes a small number of very high values. In these studies, the median, which is less influenced by outliers than the mean, potentially offers a better summary of the data. Furthermore, even if LR can still be used to identify average effects, it ignores any heterogeneity of the relationships of interest, leaving out important variation across the distribution of <emph>y,</emph> and masking potential inequalities within the data.</p> <p>When the relationship between <emph>x</emph> and <emph>y</emph> differs at high and low levels of <emph>y</emph>, researchers might be more interested in understanding the heterogeneity of these relationships, in addition to the relationship at the mean and median. Researchers may, for example, be interested in analyzing how motherhood affects high or low earning women who are otherwise similar in observed characteristics; how the distribution of earnings would change if a larger share of the population had children; or how earnings of mothers and women without children at the top or bottom of the distribution compare. This type of analysis can be accomplished using quantile regression (QR) methods that analyze heterogeneous relationships between dependent and independent variables across the conditional or unconditional distribution of the dependent variable ([<reflink idref="bib16" id="ref13">16</reflink>]; [<reflink idref="bib17" id="ref14">17</reflink>]; [<reflink idref="bib23" id="ref15">23</reflink>]).</p> <p>Although quantile regression constitutes a powerful methodological tool that allows researchers to analyze effects beyond the mean and across an entire distribution, there are still misunderstandings regarding what quantile regression models do and how to interpret them. Most notable have been discussions about when to apply conditional quantile regression (CQR) versus unconditional quantile regression (UQR) models and how to interpret the results. This is especially true in settings that require the inclusion of individual fixed effects using longitudinal data.</p> <p>QR models became part of the motherhood penalty debate through an exchange between [<reflink idref="bib6" id="ref16">6</reflink>], [<reflink idref="bib7" id="ref17">7</reflink>]) and [<reflink idref="bib22" id="ref18">22</reflink>]. In addition to estimating the effects of motherhood on women's wages across the wage distribution, these papers had an added challenge of controlling for unobserved characteristics using individual fixed effects in their analyses. [<reflink idref="bib6" id="ref19">6</reflink>] first used CQR to analyze the motherhood penalty across the distribution, adjusting for individual fixed effects, and finding larger penalties for mothers at the lower end of the wage distribution. [<reflink idref="bib22" id="ref20">22</reflink>] responded to this analysis by reestimating the models using UQR, finding the largest penalty for mothers at the middle of the distribution. Finally, responding again, [<reflink idref="bib7" id="ref21">7</reflink>], reestimated their models using UQR and different specifications for fixed effects with results similar to their first set of models. More recently, [<reflink idref="bib13" id="ref22">13</reflink>] incorporated UQR with updated data to examine how motherhood penalties vary across different combinations of skill and wage levels with slightly different findings.</p> <p>These debates extend beyond sociology. For instance, in the developmental literature, [<reflink idref="bib31" id="ref23">31</reflink>] attempted to provide an introduction to CQR, but inadvertently suggested misleading interpretations that were later addressed by [<reflink idref="bib41" id="ref24">41</reflink>]. [<reflink idref="bib41" id="ref25">41</reflink>] distinguished between CQR and UQR and laid additional assumptions required for the type of interpretation usually given to these methodologies. In health economics, [<reflink idref="bib3" id="ref26">3</reflink>] also provided a discussion on the differences between CQR and UQR models, but misinterpreted [<reflink idref="bib17" id="ref27">17</reflink>] UQR methods with the [<reflink idref="bib16" id="ref28">16</reflink>] estimation of quantile treatment effects (QTE). More recently, [<reflink idref="bib5" id="ref29">5</reflink>] have also emphasized the need to address differences in estimating individual-level and population-level effects.</p> <p>In light of these ongoing debates, this paper provides an accessible review to quantile regression models that incorporate fixed effects with an emphasis on application with social science data. Specifically, we discuss the estimation and interpretation of three types of quantile regression models, comparing them to the standard LR model and addressing the application of these models to longitudinal data. As an empirical example for our discussion, we provide a replication of [<reflink idref="bib7" id="ref30">7</reflink>], using the alternative methods.[<reflink idref="bib9" id="ref31">9</reflink>] This allows us to contribute to broader discussions regarding the magnitude of the motherhood penalty on the wage distribution.</p> <p>We begin with a brief description of LR and review its interpretation under the classical assumptions. We then review standard conditional quantile regression (CQR), introduced by [<reflink idref="bib23" id="ref32">23</reflink>], emphasizing the connections to the LR model and the assumptions required for its interpretation. Next, we discuss unconditional quantile regression (UQR), introduced by [<reflink idref="bib17" id="ref33">17</reflink>] as a special case of Recentered Influence Function (RIF) regressions. We then discuss the use of QTE models following the work of [<reflink idref="bib16" id="ref34">16</reflink>] and [<reflink idref="bib15" id="ref35">15</reflink>], and expanding on the use of RIF regressions. QTE can be considered as a compromise between CQR and UQR that focuses on analyzing the distributional impact of a single binary variable on the outcome of interest, holding the distribution of other characteristics constant. Finally, we conclude with a guide of best practices for applying quantile regression models that should be useful to researchers with different levels of statistical expertise.</p> <p>In addition to addressing the interpretation of different models for estimating the motherhood penalty across the earnings distribution, this paper offers several broad contributions to the literature on QR models. First, we provide a simple and straightforward comparison of the most common QR models, emphasizing the differences between CQR and UQR. We describe how each model is capable of answering different sets of questions, depending on the interests of the researcher, and how they relate to each other and standard LR.</p> <p>Second, in discussing the estimation of QTE models, we suggest an alternative method for QR analysis that has received less attention in the literature. The QTE model is essentially a compromise between CQR and UQR models that allows researchers to examine differences across two distributions caused by a single variable of interest, based on the methodology proposed by [<reflink idref="bib16" id="ref36">16</reflink>] and [<reflink idref="bib15" id="ref37">15</reflink>].</p> <p>UQR models are able to identify the effect of marginal changes in the distribution of all controlled characteristics, <emph>x,</emph> on the unconditional distribution of the outcome, <emph>y</emph>.[<reflink idref="bib10" id="ref38">10</reflink>] However, they cannot provide estimates of changes in the distribution of <emph>y</emph> when considering large changes in the distribution of the independent variables, <emph>x</emph> ([<reflink idref="bib38" id="ref39">38</reflink>]), especially when the variable of interest is discrete.[<reflink idref="bib11" id="ref40">11</reflink>] In contrast, QTE models allow researchers to estimate and identify distributional effects considering a large change in the distribution of a single discrete characteristic while controlling for other factors.</p> <p>Third, we focus on how to estimate and interpret UQR and QTE models in the presence of individual-level fixed effects, specifically with longitudinal data, using the newly developed Stata command <emph>rifhdreg</emph> ([<reflink idref="bib36" id="ref41">36</reflink>]), which simplifies all intermediate steps necessary for the estimation of UQR and QTE models.[<reflink idref="bib12" id="ref42">12</reflink>] We also contrast the analysis with CQR in the presence of individual-level fixed effects using the methodology proposed by [<reflink idref="bib10" id="ref43">10</reflink>]. As shown in our empirical example of the motherhood penalty, examining within-person change over time through QR models has been a complicated endeavor with conflicting results. We further note how this debate is tied to issues over incorporating fixed effects into QR models and assessing relationships across a distribution over time.</p> <p>While we often use language that makes causal interpretations of the motherhood, as well as other control variables, across all the models discussed here, it should be emphasized that the identification of causal effects is often difficult. Problems of endogeneity, sample selection, omitted variables, and model specification, among others, are also present when estimating quantile regression models. Many of these problems have yet to be addressed in the literature. Nevertheless, the discussion we provide should help to clarify the kind of effects one is able to estimate using quantile regression methodologies, under ideal scenarios. Furthermore, while emphasizing the interpretation of the effect motherhood on wages, the main variable of interest, we also provide some discussion on other factors for the sake of completeness, but those should be better interpreted as correlational, rather than causal effects.</p> <hd id="AN0176861662-3">Linear Regression (LR)</hd> <p>In order to understand what quantile regression does and the difference between conditional and unconditional partial effects, it is useful to first review what LR does, especially when using longitudinal or panel data. For example, assume that a researcher has access to panel data on earnings information for a fixed number of women (<emph>N</emph>) across time (<emph>T</emph>). The interest lies in analyzing the effect of motherhood, measured by the number of children (<emph>x</emph>) on earnings (<emph>y</emph>), while controlling for other observed characteristics, such as marital status, education, occupation, and work experience (<emph>z</emph>). With panel data, it is possible to differentiate between unobserved individual characteristics that are constant across time ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> ), and an unobserved error ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> ) that is the product of an independent and identically distributed (i.i.d.) error term ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> ) and a strictly positive function, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>σ</mtext><mfenced><mrow><mi>x</mi><mo>,</mo><mi>z</mi></mrow></mfenced></mrow></math> </ephtml> . If <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>σ</mtext><mfenced><mrow><mi>x</mi><mo>,</mo><mi>z</mi></mrow></mfenced></mrow></math> </ephtml> is a constant, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> is homoscedastic, otherwise it varies with respect to <emph>x</emph> or <emph>z</emph>, suggesting a heteroskedastic error.[<reflink idref="bib13" id="ref44">13</reflink>] Finally, assume that the population or "true" model that describes their relationship is given by:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>×</mo><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>∼</mo><mi>i</mi><mi>i</mi><mi>d</mi><mfenced><mrow><mn>0</mn><mo>,</mo><mn>1</mn></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Under the CLMA the coefficients of this model can be estimated via OLS using two strategies ([<reflink idref="bib43" id="ref45">43</reflink>]). The traditional approach is to assume the unobserved effect <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> is a parameter that needs to be estimated for each individual <emph>i</emph>. This is done by including a set of N dummy variables <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mfenced><mrow><msub><mi>D</mi><mi>i</mi></msub></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> [<reflink idref="bib14" id="ref46">14</reflink>] one for each cross-section observation, and finding the set of parameters <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>β</mtext><mo>=</mo><mfenced close="]" open="["><mrow><msub><mi>b</mi><mn>0</mn></msub><mo>,</mo><msub><mi>b</mi><mi>x</mi></msub><mo>,</mo><mo /><msub><mi>b</mi><mi>z</mi></msub></mrow></mfenced></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>Γ</mtext><mo>=</mo><mfenced close="]" open="["><mrow><msub><mi>δ</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>δ</mi><mi>N</mi></msub></mrow></mfenced></mrow></math> </ephtml> that best fit the data by minimizing the sum of squared differences between the observed outcome, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> , and the estimated conditional mean of the model <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mi>E</mi><mo>^</mo></mover><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>,</mo><mi>z</mi><mo>,</mo><mi>D</mi></mrow></mfenced></mrow></math> </ephtml> , highlighted in equation (<reflink idref="bib2" id="ref47">2</reflink>):</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>β</mi><mo>^</mo></mover><mrow><mi>o</mi><mi>l</mi><mi>s</mi></mrow></msub><mo>,</mo><msub><mover accent="true"><mtext>Γ</mtext><mo>^</mo></mover><mrow><mi>o</mi><mi>l</mi><mi>s</mi></mrow></msub><mo>=</mo><munder><mrow><mtext>min</mtext></mrow><mrow><mover accent="true"><mi>β</mi><mo>^</mo></mover><mo>,</mo><mover accent="true"><mtext>Γ</mtext><mo>^</mo></mover></mrow></munder><munderover><mstyle displaystyle="true" mathsize="140%"><mo>∑</mo></mstyle><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>N</mi></munderover><munderover><mstyle displaystyle="true" mathsize="140%"><mo>∑</mo></mstyle><mrow><mi>t</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></munderover><msup><mrow><mfenced><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>b</mi><mo>^</mo></mover><mn>0</mn></msub><mo>−</mo><msub><mover accent="true"><mi>b</mi><mo>^</mo></mover><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>b</mi><mo>^</mo></mover><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>δ</mi><mo>^</mo></mover><mi>i</mi></msub><msub><mi>D</mi><mi>i</mi></msub></mrow></mfenced></mrow><mn>2</mn></msup><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Because the parameters <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> are usually not of interest, a simpler approach is to use a within transformation process (Frisch-Waugh Theorem), partialling out the effect of the fixed effects across all variables before using OLS to estimate <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>β</mtext></math> </ephtml> .[<reflink idref="bib15" id="ref48">15</reflink>] In panel data, this is equivalent to reexpressing the independent variables as deviations from their individual-specific means in equation 3a,[<reflink idref="bib16" id="ref49">16</reflink>] and applying OLS to the model expressed in equation (3b).</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>y</mi><mo>¯</mo></mover><mi>i</mi></msub><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>x</mi><mo>¯</mo></mover><mi>i</mi></msub></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mrow><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>z</mi><mo>¯</mo></mover><mi>i</mi></msub></mrow></mfenced><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>−</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><mfenced><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mi>u</mi><mo>¯</mo></mover><mi>i</mi></msub></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>y</mi><mo>~</mo></mover><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mover accent="true"><mi>x</mi><mo>~</mo></mover><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mover accent="true"><mi>z</mi><mo>~</mo></mover><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mover accent="true"><mi>u</mi><mo>~</mo></mover><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Once the coefficients <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>β</mtext></math> </ephtml> are estimated, LR can be used to answer at least three different questions, resulting in three interpretations when explaining the effect that a change in the independent variable of interest, number of children (<emph>x</emph>), has on the dependent variable, earnings (<emph>y</emph>).</p> <p>First, how much would the earnings for person <emph>i</emph> at time <emph>t</emph> change, if they had an additional child, keeping everything else constant? Under the assumption of homoscedasticity and using the model defined in equation (1a), this effect is equal to <emph>b<subs>x</subs></emph> and constant across every woman:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow><mrow><mtext>′</mtext></mrow></msubsup><mo>−</mo><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>σ</mi><mi>u</mi></msub><mtext /><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>σ</mi><mi>u</mi></msub><mtext /><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Δ</mi><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><msub><mi>b</mi><mi>x</mi></msub><mo>→</mo><mfrac><mrow><mi>Δ</mi><msub><mi>y</mi><mi>i</mi></msub></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub></mrow></mfrac><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>However, if the model is heteroskedastic with respect to number of children (<emph>x</emph>), this effect will not be constant (unobserved heterogeneity) and will depend on an unobserved component <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> .</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow><mtext>′</mtext></msubsup><mo>−</mo><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><mo>,</mo><mo>.</mo></mrow></mfenced><mtext /><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mo>−</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><mo>.</mo></mrow></mfenced><mtext /><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Δ</mi><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><msub><mi>b</mi><mi>x</mi></msub><mo>+</mo><mfenced><mrow><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub><mo>,</mo><mo>.</mo></mrow></mfenced><mo>−</mo><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><mo>.</mo></mrow></mfenced></mrow></mfenced><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>→</mo><mfrac><mrow><mi>Δ</mi><msub><mi>y</mi><mi>i</mi></msub></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi></mrow></msub></mrow></mfrac><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mo>+</mo><mi>Δ</mi><msub><mi>σ</mi><mi>u</mi></msub><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><mo>.</mo></mrow></mfenced><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Because <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> is i.i.d. with an expected value equal to zero, we can abstract from this unknown by averaging the effect of an additional child among women who have the same characteristics (e.g., same number of children, years of education, and marital status). This is the average conditional effect, which helps answer a second question: how much would earnings change on average among, for instance, married women with one child and 12 years of education ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>=</mo><mi>X</mi></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>z</mi><mo>=</mo><mi>Z</mi><mo stretchy="false">)</mo></mrow></math> </ephtml> if they have an additional child?</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mfenced><mrow><mfrac><mrow><mi>Δ</mi><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfrac><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi></mrow></mfenced><mo>=</mo><mfrac><mrow><mi>Δ</mi><mi>E</mi><mfenced><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi></mrow></mfenced></mrow><mrow><mi>Δ</mi><mi>X</mi></mrow></mfrac><mo>=</mo><mi>E</mi><mfenced><mrow><msub><mi>b</mi><mi>x</mi></msub><mo>+</mo><mfrac><mrow><mi>Δ</mi><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfrac><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Finally, it is also possible to make inferences for the population as a whole, by averaging the individual-level effect across all women in the sample.[<reflink idref="bib17" id="ref50">17</reflink>] This average effect would answer a third question: how much would average earnings in the population change if every woman had an additional child?</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mfenced><mrow><mfrac><mrow><mi>Δ</mi><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mn>1</mn><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfrac></mrow></mfenced><mo>=</mo><mi>E</mi><mfenced><mrow><msub><mi>b</mi><mi>x</mi></msub><mo>+</mo><mfrac><mrow><mi>Δ</mi><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced></mrow><mrow><mi>Δ</mi><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfrac><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>In summary, under the CLMA, there is no unobserved heterogeneity when estimating the individual, conditional, and unconditional levels effects. Because we assume a model that is linear in parameters and variables, the effect is constant regardless of individual characteristics.[<reflink idref="bib18" id="ref51">18</reflink>] If the model is heteroskedastic, the effects at the individual level will depend on unobserved factors. In this case, average effects conditional on characteristics and unconditional population average effects can still be used to abstract from the unobserved components and make inferences from the model. The presence of unobserved heterogeneous effects opens the possibility of analyzing the effects of changes in the independent variables across the distribution of the dependent variable, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo /><mi>y</mi></mrow></math> </ephtml> . This type of relationship, however, cannot be estimated using LR models.</p> <hd id="AN0176861662-4">Quantile Regression (QR)</hd> <p>QR models can be used to obtain a richer characterization of the relationships between independent and dependent variables that go beyond the mean. We review three QR methods that can be used for the analysis of effects across the conditional and unconditional distribution of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>y</mi><mo>.</mo></mrow></math> </ephtml> These include conditional quantile regression, unconditional quantile regression, and quantile treatment effect models.</p> <hd id="AN0176861662-5">Conditional Quantile Regression (CQR)</hd> <p>[<reflink idref="bib23" id="ref52">23</reflink>] introduced conditional quantile regression into the econometrics toolbox over 40 years ago as an extension of the least absolute deviation estimator, which focuses on quantiles as a set of statistic that better describes the distribution of the outcome. Whereas LR models aim to explain how the expected outcome of a person changes in relation to a change in their characteristics, CQR tries to explain how the outcome of a person who is ranked above a specified quantile <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo stretchy="false">(</mo><mi>%</mi><mtext>τ</mtext><mo stretchy="false">)</mo></mrow></math> </ephtml> among people with the same characteristics changes in relation to a change in their characteristics. However, this is assuming that their outcome is still above % <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>τ</mtext></math> </ephtml> of the new group of individuals with the same (but new) set of characteristics.</p> <p>Intuitively, this strategy takes advantage of the unobserved heterogeneity described in (4d), quantifying the size of the unobserved effect. In the context of our question, LR models can identify the average change in earnings if women with characteristics <emph>X</emph> and <emph>Z</emph> have an additional child (average effects conditional characteristics). In contrast, CQR can identify heterogeneous effects by quantifying changes in the distribution of earnings, measured in quantile differences, among women with characteristics <emph>X</emph> and <emph>Z</emph> if they were to have an additional child.[<reflink idref="bib19" id="ref53">19</reflink>] In other words, CQR can be used to answer the question—how does an additional child affect the conditional distribution of earnings, <emph>y</emph>, for a woman with observed characteristics <emph>X</emph> and <emph>Z</emph>?</p> <p>To better understand the meaning of coefficients estimated through CQR, it is useful to start with the setup used for LR, explicitly lifting the homoscedasticity assumption, so that <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>δ</mi><mi>i</mi></msub></mrow></mfenced></mrow></math> </ephtml> is a strictly positive function of <emph>x</emph>, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>z</mi><mo>,</mo></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> . We maintain the assumption that all explanatory variables are exogenous. Furthermore, following [<reflink idref="bib28" id="ref54">28</reflink>], we can write <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>δ</mi><mi>i</mi></msub></mrow></mfenced><mo>=</mo><msub><mi>a</mi><mn>0</mn></msub><mo>+</mo><msub><mi>a</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>a</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>a</mi><mi mathsize="small">δ</mi></msub><msub><mi>δ</mi><mi>i</mi></msub></mrow></math> </ephtml> . Finally, we assume the population model that describes their relationship as:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>δ</mi><mi>i</mi></msub></mrow></mfenced><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>In this setup, we can apply a quantile function <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mtext>τ</mtext></msub><mfenced><mo>.</mo></mfenced></mrow></math> </ephtml> , conditioning on <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><mi>X</mi></mrow></math> </ephtml> , <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><mi>Z</mi></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>δ</mi><mi>i</mi></msub><mo>=</mo><mi>δ</mi></mrow></math> </ephtml> , to both sides of equation (<reflink idref="bib7" id="ref55">7</reflink>), to obtain an expression for the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub><mo /></mrow></math> </ephtml> quantile of the conditional distribution of <emph>y</emph>:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>=</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><mi>σ</mi><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>δ</mi><mi>i</mi></msub></mrow></mfenced><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mi>X</mi><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mi>Z</mi><mo>+</mo><mi>δ</mi><mo>+</mo><mfenced><mrow><msub><mi>a</mi><mn>0</mn></msub><mo>+</mo><msub><mi>a</mi><mi>x</mi></msub><mi>X</mi><mo>+</mo><msub><mi>a</mi><mi>z</mi></msub><mi>Z</mi><mo>+</mo><msub><mi>a</mi><mi mathsize="small">δ</mi></msub><mi>δ</mi></mrow></mfenced><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mi>τ</mi></mfenced><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>=</mo><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>a</mi><mn>0</mn></msub><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mi>τ</mi></mfenced></mrow></mfenced><mo>+</mo><mfenced><mrow><msub><mi>b</mi><mi>x</mi></msub><mo>+</mo><msub><mi>a</mi><mi>x</mi></msub><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mi>τ</mi></mfenced></mrow></mfenced><mi>X</mi><mo>+</mo><mfenced><mrow><msub><mi>b</mi><mi>z</mi></msub><mo>+</mo><msub><mi>a</mi><mi>z</mi></msub><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mi>τ</mi></mfenced></mrow></mfenced><mi>Z</mi><mo>+</mo><mfenced><mrow><mn>1</mn><mo>+</mo><msub><mi>a</mi><mi mathsize="small">δ</mi></msub><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mi>τ</mi></mfenced></mrow></mfenced><mi>δ</mi><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>X</mi><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Z</mi><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><mi>δ</mi><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>F</mi><mi>v</mi><mrow><mo>−</mo><mn>1</mn></mrow></msubsup><mfenced><mtext>τ</mtext></mfenced></mrow></math> </ephtml> . represents the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile of the unconditional distribution of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> .</p> <p>Although this setup is somewhat restrictive, it has implications that translate to the more general case of CQR.[<reflink idref="bib20" id="ref56">20</reflink>] First, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mtext>δ</mtext></mrow></mfenced></mrow></math> </ephtml> represents the value that is larger than <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>%</mi><mtext>τ</mtext></mrow></math> </ephtml> of all outcomes among observations with observed characteristics <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>X</mi><mo>,</mo><mo /><mi>Z</mi></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> . Second, the conditional quantile <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mtext>τ</mtext></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mo>.</mo></mrow></mfenced></mrow></math> </ephtml> is modeled as a linear function of <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>X</mi><mo>,</mo><mi>Z</mi><mo /></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> . This assumes that the difference between the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mtext>τ</mtext></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mtext>δ</mtext></mrow></mfenced></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mtext>τ</mtext></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>X</mi><mo>'</mo><mo>,</mo><mi>Z</mi><mo>,</mo><mtext>δ</mtext></mrow></mfenced></mrow></math> </ephtml> , is given by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>b</mi><mi>x</mi></msub><mfenced><mtext>τ</mtext></mfenced><mfenced><mrow><msup><mi>X</mi><mo>′</mo></msup><mo>−</mo><mi>X</mi></mrow></mfenced></mrow></math> </ephtml> . This relationship is constant for the same quantile across individuals with different characteristics but may vary by quantile.</p> <p>Third, the set of coefficients <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><mo /><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><mo /><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><mo /></mrow></math> </ephtml> will vary across quantiles if the model is heteroskedastic with respect to <emph>X</emph>, Z, or <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> . Conversely, if the error is homoscedastic, the coefficients will be constant across quantiles, and will be the same as the average effect from a LR model.[<reflink idref="bib21" id="ref57">21</reflink>]</p> <p>Lastly, CQR cannot be used for individual level interpretations because effects at the individual level depend on an unknown factor, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> , or that person's position in the conditional distribution, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>τ</mtext></math> </ephtml> . Doing so requires the rank invariance assumption (that a person's ranking remains constant even if their characteristics change), and that we observe <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> or <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> .</p> <p>When the model does not include individual fixed effects, for instance when using cross-sectional data, [<reflink idref="bib23" id="ref58">23</reflink>] show that the set of coefficients in this model <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>β</mi><mfenced><mi>τ</mi></mfenced><mo>=</mo><mfenced close="]" open="["><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>t</mi></mfenced><mo>,</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced></mrow></mfenced></mrow></math> </ephtml> ,[<reflink idref="bib22" id="ref59">22</reflink>] corresponding to the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile, can be found by minimizing the weighted sum of absolute deviations:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mrow><mi>β</mi><mtext /></mrow><mo stretchy="true">^</mo></mover><mfenced><mi>τ</mi></mfenced><mo>=</mo><mtext>min</mtext><munder><mstyle displaystyle="true" mathsize="100%"><mo>∑</mo></mstyle><mrow><mi>y</mi><mo>≥</mo><msub><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi></msub><mo stretchy="false">(</mo><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>,</mo><mi>z</mi><mo stretchy="false">)</mo></mrow></munder><mi>τ</mi><mfenced close="|" open="|"><mrow><mi>y</mi><mo>−</mo><msub><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi></msub><mo stretchy="false">(</mo><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>,</mo><mi>z</mi><mo stretchy="false">)</mo></mrow></mfenced><mo>+</mo><munder><mstyle displaystyle="true" mathsize="100%"><mo>∑</mo></mstyle><mrow><mi>y</mi><mo><</mo><msub><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi></msub><mo stretchy="false">(</mo><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>,</mo><mi>z</mi><mo stretchy="false">)</mo></mrow></munder><mfenced><mrow><mn>1</mn><mo>−</mo><mi>τ</mi></mrow></mfenced><mfenced close="|" open="|"><mrow><mi>y</mi><mo>−</mo><msub><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi></msub><mo stretchy="false">(</mo><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>,</mo><mi>z</mi><mo stretchy="false">)</mo></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>In contrast to the LR model, CQR cannot account for individual fixed effects by simply including sets of dummy variables representing each cross-sectional unit. Doing so creates an incidental parameter problem, affecting the consistent estimation of all coefficients in the model.[<reflink idref="bib23" id="ref60">23</reflink>] This problem has been discussed extensively in the literature (see [<reflink idref="bib24" id="ref61">24</reflink>]; [<reflink idref="bib33" id="ref62">33</reflink>]; [<reflink idref="bib10" id="ref63">10</reflink>]; and [<reflink idref="bib28" id="ref64">28</reflink>]), providing various procedures to obtain consistent estimates for other parameters of the CQR, under different assumptions. A simple approach for estimating conditional quantile regressions with fixed effects is discussed by [<reflink idref="bib10" id="ref65">10</reflink>].</p> <p>As previously described, if the error <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> is homoscedastic with respect to an independent variable, the effect of that variable is constant across quantiles and is equal to the LR estimator. Based on this premise, the [<reflink idref="bib10" id="ref66">10</reflink>] estimator assumes that the model error is homoscedastic with respect to the individual effect <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> . If this is the case, a consistent estimator for the parameters <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>β</mi><mfenced><mi>τ</mi></mfenced></mrow></math> </ephtml> for each <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>τ</mi><mrow><mi>t</mi><mi>h</mi></mrow></msub><mo /></mrow></math> </ephtml> conditional quantile can be obtained using a two-step procedure:</p> <p></p> <ulist> <item> Estimate the model <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> using LR methods and obtain a consistent estimate for the individual effect <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> .</item> <p></p> <item> Obtain <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>y</mi><mo>~</mo></mover><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>=</mo><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>−</mo><msub><mover accent="true"><mtext>δ</mtext><mo>^</mo></mover><mi>i</mi></msub></mrow></math> </ephtml> , and use CQR to estimate the model</item> </ulist> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mo stretchy="false">(</mo><mover accent="true"><mi>y</mi><mo>~</mo></mover><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo stretchy="false">)</mo><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>X</mi><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Z</mi><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Although equation (<reflink idref="bib10" id="ref67">10</reflink>) can be directly estimated using the objective function described in equation (<reflink idref="bib9" id="ref68">9</reflink>), standard errors need to be estimated using other approaches such as bootstrap resampling methods. Assuming the true model is given by equation (<reflink idref="bib7" id="ref69">7</reflink>), CQR coefficients can also be estimated using the method of moments proposed by [<reflink idref="bib28" id="ref70">28</reflink>].[<reflink idref="bib24" id="ref71">24</reflink>]</p> <p>The natural interpretation for CQR can be obtained by answering the question: how much would the earnings distribution change, measured by changes in the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub><mo /></mrow></math> </ephtml> quantile, if every woman with observed characteristics <emph>X, Z,</emph> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> had an additional child, holding everything else constant?, or how much would the earnings of a woman, who earns more than <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>τ</mtext><mi>%</mi></mrow></math> </ephtml> of other women with the same characteristics, change if she had an additional child, but remained in the same position among her (new) peers? [<reflink idref="bib25" id="ref72">25</reflink>] This is the conditional quantile effect, as illustrated in equation (<reflink idref="bib11" id="ref73">11</reflink>).</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>+</mo><mi>Δ</mi><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>−</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>t</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mfenced><mrow><mi>X</mi><mo>+</mo><mi>Δ</mi><mi>X</mi></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Z</mi><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><mi>δ</mi><mo>−</mo><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>t</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>X</mi><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Z</mi><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><mi>δ</mi></mrow></mfenced></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Δ</mi><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><msub><mi>δ</mi><mi>i</mi></msub></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Δ</mi><mi>X</mi><mo>→</mo><mfrac><mrow><mi>Δ</mi><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mtext>|</mtext><mi>X</mi><mo>,</mo><mi>Z</mi><mo>,</mo><mi>δ</mi></mrow></mfenced></mrow><mrow><mi>Δ</mi><mi>X</mi></mrow></mfrac><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Three remarks are worth describing here. First, based on equation (<reflink idref="bib11" id="ref74">11</reflink>), we cannot make interpretations at the individual level because we do not know that person's position in the conditional distribution ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>τ</mtext></math> </ephtml> ). The alternative is to compare the same <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile across different conditional distributions, or assuming a person remains in the same ranking, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> after the change in characteristics occurs. Second, the interpretation does not change if <emph>X</emph> is continuous or discrete. Third, because the conditional quantile effects depend on the conditioning characteristics, it is common practice to estimate and interpret these effects for the average person (at mean characteristics), or report the "average" conditional quantile effect across the whole population, as the representative effects.</p> <p>Researchers, however, may be more interested in answering policy-related questions such as: how does motherhood affect the overall unconditional distribution of earnings? Or, how would the distribution of earnings change if every woman had an additional child? Answering this question requires a different type of analysis that incorporates the concept of unconditional quantile regression.</p> <hd id="AN0176861662-6">Unconditional Quantile Regression (UQR) based on the Recentered Influence Function (RIF)</hd> <p>As described in the previous section, the most important feature of CQR is its capacity to identify otherwise heterogeneous effects of changes in independent variables across conditional distributions. Often, however, researchers may be more interested in identifying the effects of changes of independent variables on the overall or unconditional distribution of the outcome. For example, instead of trying to identify how an additional child affects earnings for single mothers with one child, researchers may be interested in analyzing the effect of every woman in the population having an additional child on the unconditional distribution of earnings. In this sense, unconditional distributions are affected by changes in the distribution of other characteristics. There is, however, a close link between analyzing changes on conditional distributions and unconditional distributions.</p> <p>[<reflink idref="bib17" id="ref75">17</reflink>] show that unconditional quantile effects[<reflink idref="bib26" id="ref76">26</reflink>] can be derived as a weighted average of all conditional quantile partial effects. [<reflink idref="bib29" id="ref77">29</reflink>] and [<reflink idref="bib30" id="ref78">30</reflink>] use a similar principle to simulate unconditional distributions of outcomes based on CQR to estimate unconditional quantile effects. In principle, the procedure requires the estimation of a large set of quantile regressions, for example, estimating separate models from the 1st through 99th conditional quantiles, to characterize the whole distribution of the dependent variable. After the models are estimated, simulation methods can be used to identify the effect of a change in the distribution of characteristics, for example, every woman having an additional child, on the distribution of the dependent variable. This process, however, is not always practical.</p> <p>Addressing these impracticalities, [<reflink idref="bib17" id="ref79">17</reflink>] proposed a computationally simpler strategy to identify unconditional quantile effects that does not require reconstructing the entire distribution of the dependent variable. This methodology, unconditional quantile regression, uses the recentered influence function (RIF) to provide a first order approximation of the marginal effect of small location shift changes in the distribution of independent variables on any unconditional quantile. In practice, as described in [<reflink idref="bib36" id="ref80">36</reflink>], this small shift changes should be understood as changes in the average of the independent variables.</p> <p>In contrast with LR and CQR, RIF regressions in general, and UQR in particular, can only be used to draw inferences in terms of unconditional effects.[<reflink idref="bib27" id="ref81">27</reflink>] This implies that UQR can be used to analyze what would happen to the population <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile when there is a small change in the distribution of an independent variable, but not what would happen to a specific person or a specific group of individuals when <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo /></mrow></math> </ephtml> changes. Furthermore, because UQR is based on local approximations, particular care is needed when interpreting effects of discrete and categorical variables.</p> <p>To better understand UQR, it is useful to understand what influence functions (IF) are, how they are constructed, and how they link to the unconditional distribution of the outcome. Assume women's earnings (<emph>y</emph>) across time (<emph>t</emph>), is a random variable with cumulative distribution function <emph>F<subs>y</subs></emph> and a density distribution <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>f</mi><mi>y</mi></msub><mo>=</mo><mi>d</mi><msub><mi>F</mi><mi>y</mi></msub></mrow></math> </ephtml> . Any statistic of interest <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>v</mi><mo /></mrow></math> </ephtml> can be written as a function of the cumulative distribution function (CDF):</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>F</mi><mi>y</mi></mrow></msub><mo>=</mo><mi>v</mi><mfenced><mrow><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></math> </ephtml> </p> <p>Graph</p> <p>Assume there is a second distribution <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>G</mi><mi>y</mi><mn>0</mn></msubsup></mrow></math> </ephtml> that is the result of a new person, with earnings equal to <emph>y</emph><subs>0</subs>, entering the sample, which was originally of size <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="script">N</mi><mo>.</mo></mrow></math> </ephtml> This new distribution could be written as:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>G</mi><mi>y</mi><mn>0</mn></msubsup><mo>=</mo><mfrac><mi mathvariant="script">N</mi><mrow><mi mathvariant="script">N</mi><mo>+</mo><mn>1</mn></mrow></mfrac><msub><mi>F</mi><mi>y</mi></msub><mo>+</mo><mfrac><mn>1</mn><mrow><mi mathvariant="script">N</mi><mo>+</mo><mn>1</mn></mrow></mfrac><mtext>*</mtext><mi>Δ</mi><mfenced><mrow><mi>y</mi><mo>≥</mo><msub><mi>y</mi><mn>0</mn></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Using these two distributions, the change in the distributional statistic caused by the added person is simply the difference between <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><mi>F</mi><mi>y</mi></mrow></msub></mrow></math> </ephtml> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>v</mi><mrow><msubsup><mi>G</mi><mi>y</mi><mn>0</mn></msubsup></mrow></msub></mrow></math> </ephtml> . If we rescale this difference by the relative change in the population size, we obtain the Influence Function of that observation. Mathematically, if we define <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>ϵ</mi><mo>=</mo><msup><mrow><mfenced><mrow><mi mathvariant="script">N</mi><mo>+</mo><mn>1</mn></mrow></mfenced></mrow><mrow><mo>−</mo><mn>1</mn></mrow></msup></mrow></math> </ephtml> , the influence function (IF) of a person with earnings <emph>y<subs>i</subs></emph> on the statistic <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>v</mi><mo /></mrow></math> </ephtml> is defined as <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mo>:</mo></math> </ephtml></p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>=</mo><munder><mrow><mtext>lim</mtext></mrow><mrow><mi>ϵ</mi><mo>↓</mo><mn>0</mn></mrow></munder><mfrac><mrow><mi>v</mi><mfenced><mrow><mfenced><mrow><mn>1</mn><mo>−</mo><mi>ϵ</mi></mrow></mfenced><msub><mi>F</mi><mi>y</mi></msub><mo>+</mo><mi>ϵ</mi><mtext>*</mtext><mi>Δ</mi><mfenced><mrow><mi>y</mi><mo>≥</mo><msub><mi>y</mi><mi>i</mi></msub></mrow></mfenced></mrow></mfenced><mo>−</mo><mi>v</mi><mfenced><mrow><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow><mi>ϵ</mi></mfrac><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>This expression is also known as a directional or Gateaux derivative, and represents a first order (linear) approximation of the rate of change, or influence, an observation with earnings <emph>y<subs>i</subs></emph> has on the distributional statistic <emph>v</emph>. To complement the idea of IF, [<reflink idref="bib17" id="ref82">17</reflink>] introduced the concept of the Recentered Influence Function, which is defined as:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>=</mo><mtext /><mi>v</mi><mfenced><mrow><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>+</mo><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>In this case, the RIF is better understood as the linear approximation of the contribution of a single observation on the construction of the distributional statistic, <emph>v</emph>. The RIF has two important properties that have been discussed by [<reflink idref="bib39" id="ref83">39</reflink>], [<reflink idref="bib12" id="ref84">12</reflink>], and [<reflink idref="bib17" id="ref85">17</reflink>]:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mfenced><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></mfenced><mo>=</mo><mi>v</mi><mfenced><mrow><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mi>v</mi><mfenced><mrow><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></mfenced><mo>=</mo><mfrac><mn>1</mn><mi>N</mi></mfrac><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></mfenced><mo>=</mo><mfrac><mn>1</mn><mi>N</mi></mfrac><mi>V</mi><mi>a</mi><mi>r</mi><mfenced><mrow><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>The first property (16a) implies that the unconditional expectation of the RIF function is equal to the distributional statistic of interest. This property is the basis for interpreting RIF regression in the framework of unconditional effects. The second states that influence functions can be used to obtain the variance of distributional statistics <emph>v</emph> ([<reflink idref="bib12" id="ref86">12</reflink>]).</p> <p>Although RIF functions can be used to analyze a large set of distributional statistics,[<reflink idref="bib28" id="ref87">28</reflink>][<reflink idref="bib17" id="ref88">17</reflink>] concentrate on the analysis of unconditional quantiles, for which the RIF is defined as follows:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mo>.</mo></mfenced><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>=</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced><mo>+</mo><mfrac><mrow><mi>τ</mi><mo>−</mo><mi>Δ</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>≤</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></mfenced></mrow><mrow><msub><mi>f</mi><mi>y</mi></msub><mfenced><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></mfenced></mrow></mfrac><mo>,</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></math> </ephtml> is the unconditional quantile, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> is the rank or percentile of interest, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Δ</mi><mfenced><mo>.</mo></mfenced></mrow></math> </ephtml> is an indicator function for whether observation <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>y</mi><mi>i</mi></msub><mo /></mrow></math> </ephtml> is below <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></math> </ephtml> , and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>f</mi><mi>y</mi></msub><mfenced><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></mfenced></mrow></math> </ephtml> is the density function of distribution of <emph>y</emph> evaluated at the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow></math> </ephtml> .</p> <p>Once RIF's have been obtained, UQR can be estimated through standard LR (RIF-OLS) using the corresponding RIF as the dependent variable, instead of <emph>y</emph>.[<reflink idref="bib29" id="ref89">29</reflink>] For instance, returning to the setup of panel data, where earnings (<emph>y</emph>) is a function of number of children (<emph>x</emph>), other explanatory variables (<emph>Z</emph>), individual fixed effect ( <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> ), we would estimate the following model using standard OLS:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mo>.</mo></mfenced><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>e</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Even though quantile functions are inherently nonlinear, one of the main advantages of estimating UQR using OLS is that they can be easily adapted to include fixed effects, as discussed in [<reflink idref="bib22" id="ref90">22</reflink>] and [<reflink idref="bib4" id="ref91">4</reflink>]. In other words, it is possible to use the within transformation procedure described in section 2, to control for the individual effects <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> , without suffering from the incidental parameter problem. Similar to CQR, we can obtain different sets of parameters, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>β</mi><mfenced><mi>τ</mi></mfenced><mo>=</mo><mfenced close="]" open="["><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><mo>,</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced></mrow></mfenced></mrow></math> </ephtml> , for any <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>τ</mtext></math> </ephtml> between 0 and 1. Similar to the LR model, UQR requires the assumption that the unobserved component <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> is distributed independently from <emph>x</emph>, <emph>z</emph> and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>δ</mtext></math> </ephtml> , so that its influence on the conditional expectation of RIF can be average out.</p> <p>Three important aspects differentiate UQR from standard LR. First, equation (<reflink idref="bib18" id="ref92">18</reflink>) models how changes in the number of children relate to changes in the influence function person <emph>i</emph> has on the distributional statistic. The average of those changes can be interpreted as an effect on the unconditional quantile. Second, UQR has no interpretation in terms of individual or conditional effects. This happens because the RIF is constructed using the overall unconditional distribution, and thus can only be used to analyze influences on unconditional distribution statistics.[<reflink idref="bib30" id="ref93">30</reflink>] When using panel data, the unconditional distribution corresponds to the outcome distribution observed across all observations (individuals across time).</p> <p>Third, because this strategy aims to analyze unconditional effects, all interpretations have to be made based on changes in the distribution of independent variables. If the model specification is set as in equation (<reflink idref="bib18" id="ref94">18</reflink>), without interactions or higher order polynomials, we assume that the only distributional changes of <emph>x</emph> that affect the unconditional quantile are changes in their means.[<reflink idref="bib31" id="ref95">31</reflink>]</p> <p>The natural interpretation of UQR can be obtained by answering the question: how much would the observed distribution of women's earnings (across individuals and time) change, measured by the change in the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile, if all women had, on average, an additional child, holding everything else constant? This question can be answered by estimating the difference between the observed <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> earnings quantile among all women, and the predicted quantile assuming every woman has an additional child:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi><mtext>′</mtext></msubsup><mfenced><mi>y</mi></mfenced><mo>=</mo><mi>E</mi><mfenced><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mtext>′</mtext><mfenced><mrow><msub><mi>y</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>,</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mo>.</mo></mfenced><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></mfenced><mo>=</mo><mi>E</mi><mo>(</mo><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><mi>Δ</mi><mi>x</mi><mtext /></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>e</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow><mo>)</mo><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi><mo>'</mo></msubsup><mfenced><mi>y</mi></mfenced><mo>=</mo><mi>E</mi><mfenced><mrow><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mfenced><mrow><msub><mi>x</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>+</mo><msub><mi>b</mi><mi>z</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>z</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub><mo>+</mo><msub><mi>b</mi><mi mathsize="small">δ</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>δ</mi><mi>i</mi></msub><mo>+</mo><msub><mi>e</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></mfenced><mo>+</mo><mi>E</mi><mfenced><mrow><msub><mi>b</mi><mi>x</mi></msub><mi>Δ</mi><mi>x</mi></mrow></mfenced><mo>=</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Δ</mi><mover accent="true"><mi>x</mi><mo>¯</mo></mover><mo /><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mover accent="true"><mi>Q</mi><mo>^</mo></mover><mi>τ</mi><mo>'</mo></msubsup><mfenced><mi>y</mi></mfenced><mo>−</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>Δ</mi><mover accent="true"><mi>x</mi><mo>¯</mo></mover><mo>→</mo><mfrac><mrow><mi>Δ</mi><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mi>y</mi></mfenced></mrow><mrow><mi>Δ</mi><mover accent="true"><mi>x</mi><mo>¯</mo></mover></mrow></mfrac><mo>=</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>This partial effect is what [<reflink idref="bib17" id="ref96">17</reflink>] describe as the unconditional quantile partial effect, which is defined as the change in the unconditional quantile caused by a change in the distribution of <emph>x</emph>, approximated by a change in its mean value <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Δ</mi><mover accent="true"><mi>x</mi><mo>¯</mo></mover></mrow></math> </ephtml> .</p> <p>When a variable is continuous with a wide range of values, using the thought experiment of "an additional unit change" is appropriate because such a change can be considered a "small" location shift effect, which UQR can approximate well. For example, an additional year of education is a small location shift compared to the average number of years of education in the population. However, an additional child likely results in a larger location shift, because a typical family has less than two children, and this may result in an incorrect interpretation of UQR estimates. In cases like this, instead of referring to one additional child per woman, one could refer to an increase in the fertility rate that results, for example, in a 0.2 increase in the average number of children.</p> <p>When the variable of interest is binary, for instance, being a mother (<reflink idref="bib1" id="ref97">1</reflink>) or not (0), the thought experiment described above is not adequate. On the one hand, only a fraction of the sample could change from "nonmothers" to "mothers." On the other hand, comparing scenarios where every woman was a mother versus not would imply a large change in the distribution of motherhood, which UQR does not approximate well. Often, researchers incorrectly interpret coefficients of binary variables as if they were treatment effects, but they are not. A better interpretation is to treat the location shift as a change in the "incidence rate," for example referring to a <emph>p</emph> percentage point increase in the share of mothers in the sample.</p> <p>If a researcher is interested in analyzing distributional treatment effects, namely comparing the outcome distributions of two groups of individuals with the same distribution of characteristics, but who belong to different groups (e.g., treated and untreated group, mothers versus women without children), the most appropriate approach is to use a methodology known as quantile treatment effects.</p> <hd id="AN0176861662-7">Quantile Treatment Effects (QTE) via RIF</hd> <p>UQR models provide linear approximations of changes in how unconditional quantiles of the dependent variable change when there is a small change in the distribution of independent characteristics. Such approximations, however, may not be appropriate for analyzing large changes in the distribution of characteristics. This may happen when the characteristic of interest has a limited range of values (e.g., number of children), or when it is a binary variable (e.g., mothers and women without children).</p> <p>When the variable of interest is binary, [<reflink idref="bib16" id="ref98">16</reflink>] and [<reflink idref="bib15" id="ref99">15</reflink>] propose estimators that identify what they refer to as quantile or distributional treatment effects under the assumption of exogeneity.[<reflink idref="bib32" id="ref100">32</reflink>] These estimators use an inverse probability weights (IPW) to control for differences in the distribution of characteristics across two groups. Once such differences are controlled for, treatment effects of the variable interest are estimated by calculating differences in statistics across groups. The literature, however, is not clear regarding how to control for individual fixed effects when using this strategy.</p> <p>[<reflink idref="bib19" id="ref101">19</reflink>] present the command <emph>ivqte</emph> to implement the [<reflink idref="bib16" id="ref102">16</reflink>] estimator. This command uses a semiparametric approach to estimate the IPW combined with a standard CQR procedure to identify quantile treatment effects. In contrast, the estimation procedure we propose relies on using RIF functions in combination with a probit or logit model for the estimation of the IPW weights, in the spirit of [<reflink idref="bib15" id="ref103">15</reflink>] and [<reflink idref="bib18" id="ref104">18</reflink>]. Using RIF function enables us to estimate a more flexible model that controls for differences in distribution of characteristics using the IPW, but also controls for them directly, by including them in the model specification, as it is done with UQR models.</p> <p>Consider a situation where motherhood is measured as a binary variable <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo stretchy="false">(</mo><mi>x</mi></mrow></math> </ephtml> ) and denotes a potential treatment that any woman may receive. Assume that all women face two potential outcomes (earnings), <emph>y</emph><subs>0</subs> or <emph>y</emph><subs>1</subs>, depending on their motherhood status, and that we only observe <emph>y</emph>:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>y</mi><mo>=</mo><msub><mi>y</mi><mn>1</mn></msub><mi>x</mi><mo>+</mo><msub><mi>y</mi><mn>0</mn></msub><mfenced><mrow><mn>1</mn><mo>−</mo><mi>x</mi></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>If we were able to observe both potential outcomes, quantile treatment effects could be estimated by calculating the differences in the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantiles of both distributions.</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Δ</mi><mi>τ</mi></msub><mo>=</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><msub><mi>y</mi><mn>1</mn></msub></mrow></mfenced><mo>−</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><msub><mi>y</mi><mn>0</mn></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>However, this is an unfeasible estimator. If the treatment <emph>motherhood</emph> were to be assigned at random, independent of all observed (<emph>z</emph>) and unobserved characteristics, quantile treatment effects could be estimated directly by calculating the difference in the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile between the observed earnings distribution of mothers and nonmothers:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Δ</mi><mi>τ</mi></msub><mo>=</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>1</mn></mrow></mfenced><mo>−</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>This is possible because, under the random assignment assumption, the observed distributions of earnings among the treated (or untreated) group <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>F</mi><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>=</mo><mn>1</mn></mrow></msub><mfenced><mrow><msub><mi>F</mi><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>=</mo><mn>0</mn></mrow></msub></mrow></mfenced></mrow></math> </ephtml> is the same as the distribution of the potential outcomes <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>F</mi><mrow><mi>y</mi><mn>1</mn></mrow></msub><mfenced><mrow><msub><mi>F</mi><mrow><mi>y</mi><mn>0</mn></mrow></msub></mrow></mfenced></mrow></math> </ephtml> . In this framework, the treatment effect can be estimated using the CQR, and the observed outcome distributions, by estimating equation (22c):</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi></mrow></mfenced><mo>=</mo><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>1</mn></mrow></mfenced><mtext>*</mtext><msub><mi>x</mi><mi>i</mi></msub><mo>+</mo><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mtext>*</mtext><mfenced><mrow><mn>1</mn><mo>−</mo><msub><mi>x</mi><mi>i</mi></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi></mrow></mfenced><mo>=</mo><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mo>+</mo><mfenced><mrow><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>1</mn></mrow></mfenced><mo>−</mo><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>0</mn></mrow></mfenced></mrow></mfenced><msub><mi>x</mi><mi>i</mi></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi></mrow></mfenced><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><mi>x_i</mi><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>An alternative strategy to estimate the treatment effect is applying an extended use of RIF regressions. Specifically, rather than estimating equation (22c) via CQR, we can estimate this equation using RIF functions to approximate the conditional quantile of <emph>y</emph>.[<reflink idref="bib33" id="ref105">33</reflink>] Starting with equation (22b), we reorder the terms on the right side, and obtain:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>1</mn></mrow></mfenced><msub><mi>x</mi><mi>i</mi></msub><mo>+</mo><mi>Q</mi><mfenced><mrow><mi>y</mi><mtext>|</mtext><mi>x</mi><mo>=</mo><mn>0</mn></mrow></mfenced><mfenced><mrow><mn>1</mn><mo>−</mo><msub><mi>x</mi><mi>i</mi></msub></mrow></mfenced><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>Next, we obtain the RIF for equation (<reflink idref="bib23" id="ref106">23</reflink>), and use it as the dependent variable, in a regression similar to (22c) but that can be estimated using OLS:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mo>.</mo></mfenced><mo>,</mo><msub><mi>F</mi><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>=</mo><mn>1</mn></mrow></msub></mrow></mfenced><msub><mi>x</mi><mi>i</mi></msub><mo>+</mo><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><msub><mi>Q</mi><mi mathsize="small">τ</mi></msub><mfenced><mo>.</mo></mfenced><mo>,</mo><msub><mi>F</mi><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>x</mi><mo>=</mo><mn>0</mn></mrow></msub></mrow></mfenced><mfenced><mrow><mn>1</mn><mo>−</mo><msub><mi>x</mi><mi>i</mi></msub></mrow></mfenced><mtext /><mo>=</mo><msub><mi>b</mi><mn>0</mn></msub><mfenced><mi>τ</mi></mfenced><mo>+</mo><msub><mi>b</mi><mi>x</mi></msub><mfenced><mi>τ</mi></mfenced><msub><mi>x</mi><mi>i</mi></msub><mo>+</mo><msub><mi>e</mi><mi>i</mi></msub><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>According to [<reflink idref="bib16" id="ref107">16</reflink>] and [<reflink idref="bib15" id="ref108">15</reflink>], if the treatment assignment is not random, even if it depends only on observed characteristics <emph>z</emph> (conditionally exogenous),[<reflink idref="bib34" id="ref109">34</reflink>] neither of the strategies above will correctly identify the treatment effects. This happens because the differences in quantiles captured by <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>b</mi><mi>x</mi></msub><mfenced><mtext>τ</mtext></mfenced></mrow></math> </ephtml> may also account for differences in the distribution of characteristics among the treated and untreated group, and the distribution of the potential outcome will be different from the distribution of the observed outcome among the treated, that is, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>F</mi><mrow><mi>y</mi><mn>1</mn></mrow></msub><mo>≠</mo><msub><mi>F</mi><mrow><mi>y</mi><mo stretchy="false">|</mo><mi>X</mi><mo>=</mo><mn>1</mn></mrow></msub></mrow></math> </ephtml> .</p> <p>Under the additional assumption that there are individuals with characteristics <emph>z</emph> among both the treated and untreated group (common support), [<reflink idref="bib16" id="ref110">16</reflink>] and [<reflink idref="bib15" id="ref111">15</reflink>] suggest that distributional treatment effects can be identified using a two-stage procedure. In the first stage a set of weights are constructed to equalize the observed distribution of characteristics <emph>Z</emph> among the treated group and untreated groups. To do so, we first estimate a propensity score <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mi>p</mi><mo>^</mo></mover><mfenced><mi>z</mi></mfenced><mo /></mrow></math> </ephtml> that reflects the probability of an observation belonging to the treated group, conditional on <emph>z</emph>. Then, we construct the weights, also known as inverse probability weights (IPW), as follows:</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>ω</mi><mn>0</mn></msub><mo>=</mo><mfrac><mrow><mn>1</mn><mo>−</mo><mi>x</mi></mrow><mrow><mn>1</mn><mo>−</mo><mover accent="true"><mi>p</mi><mo>^</mo></mover><mfenced><mi>z</mi></mfenced></mrow></mfrac><mtext> and </mtext><msub><mi>ω</mi><mn>1</mn></msub><mo>=</mo><mfrac><mi>x</mi><mrow><mover accent="true"><mi>p</mi><mo>^</mo></mover><mfenced><mi>z</mi></mfenced></mrow></mfrac><mo>.</mo></mrow></math> </ephtml> </p> <p>Graph</p> <p>In the second stage, the constructed weights are used to estimate the quantile average treatment effects using equation (22c) with a weighted quantile regression ([<reflink idref="bib16" id="ref112">16</reflink>]), or they can be used to construct the appropriate reweighted RIF functions and estimate equation (<reflink idref="bib24" id="ref113">24</reflink>) using weighted RIF regressions via OLS ([<reflink idref="bib18" id="ref114">18</reflink>]).[<reflink idref="bib35" id="ref115">35</reflink>]</p> <p>In principle, IPW are used to <emph>reshape</emph> the observed distribution of characteristics <emph>z</emph>, so that the treated (or untreated) groups resemble the whole population. In doing so, the reweighted distribution of the observed outcome among the treated (or untreated) will be similar to the distribution of the potential outcome under treatment for the full sample. This estimation of the potential outcome distributions can then be used to identify the quantile treatment effects.[<reflink idref="bib36" id="ref116">36</reflink>] Despite the intuitive appeal of using these methods, IPW methods can be very sensitive when the propensity score is close to 0 or 1 ([<reflink idref="bib26" id="ref117">26</reflink>]). In such cases, the recommendation is to trim the weights, or propensity scores, in order to reduce the sensitivity of the estimates.</p> <p>Unlike the CQR approach, the RIF regression approach (eq. 24) allows researchers to directly control for differences in the distribution of characteristics by including control variables in the model specification. This approach may be equivalent to the IPW regression adjustment estimator of the inequality treatment effects ([<reflink idref="bib42" id="ref118">42</reflink>]). In addition, it also allows researchers to obtain estimates for the covariates <emph>z</emph> that can be interpreted in the same way as variables in a UQR.</p> <p>In contrast to other QR model estimation strategies, the literature has not yet reached a consensus regarding the estimation of fixed effects in the framework of QTE with panel data.[<reflink idref="bib37" id="ref119">37</reflink>] However, we suggest that it is feasible to use fixed effects as controls directly in the model specification making use of the properties of RIF-OLS regressions. Nevertheless, we concur that the estimation of QTE remains an open question when using panel data. Furthermore, when using panel data, like UQR, the QTE estimators compare distributions for the treated and untreated group pooling the information across years and individuals.</p> <p>The interpretation of QTE estimations depends on whether the goal is to estimate average treatment effects or treatment effects on the treated (or untreated). For our example, if the treatment variable is motherhood, the interpretation can be obtained by answering the question: how much would the distribution of childless women's earnings change, measured by the change in the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile, if they become mothers, holding everything else constant? (treatment effect on the untreated), or how much would the distribution of mothers' earnings change because they became mothers (treatment effect on treated)? Finally, how large is the difference in wage distributions when comparing mothers to women without children, everything else assumed constant (average treatment effect)?</p> <p>For other variables, the precise interpretation differs from UQR. Because the RIF is constructed for a distributional statistic conditional on <emph>x</emph>, <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>β</mtext><mi>z</mi></msub></mrow></math> </ephtml> no longer captures the unconditional effect of <emph>z</emph> on the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>τ</mtext><mrow><mi>t</mi><mi>h</mi></mrow></msub></mrow></math> </ephtml> quantile of earnings among all women. Instead, it captures the weighted average of the unconditional effects on earnings quantiles across the conditioning groups, in this case mothers and nonmothers. In practice, however, both UQR and QTE via RIF provide similar coefficient estimates for nonconditioning variables.</p> <p>The proposed QTE estimator can be implemented with the Stata community-contributed command -<emph>rifhdreg</emph>- utilizing the option -<emph>over</emph>-, to define the treatment variable, and the options -<emph>rwprobit</emph>- or -<emph>rwlogit</emph>- for the estimation of propensity scores. It also provides the option -<emph>trim-</emph> to exclude observations with estimated propensity scores that are close to 0 or 1. This command also allows for the estimation of different treatment effects using the options <emph>att</emph>, <emph>ate,</emph> or <emph>atu</emph>, as described in [<reflink idref="bib36" id="ref120">36</reflink>].</p> <hd id="AN0176861662-8">Application: Reexamining the Motherhood Penalty</hd> <p>Isolating the effects of motherhood on earnings has proven quite complicated for social scientists because it requires accounting for many confounding factors, some of which, such as work productivity and personal preferences, are hard to observe in survey data. Selection bias is also an issue, as only certain women may choose motherhood. As a result, researchers often rely on longitudinal data, most notably, the 1979 cohort of the National Longitudinal Survey of Youth (NLSY79), and fixed effects models to account for unobserved time-invariant individual-level factors in models that examine earnings over time. Results from these studies show that mothers experience an average wage penalty of 5–8 percent per child in the United States (Avellar and Smock 2003; [<reflink idref="bib8" id="ref121">8</reflink>]; [<reflink idref="bib20" id="ref122">20</reflink>]; [<reflink idref="bib40" id="ref123">40</reflink>]). More recent work ([<reflink idref="bib6" id="ref124">6</reflink>], [<reflink idref="bib7" id="ref125">7</reflink>]; [<reflink idref="bib13" id="ref126">13</reflink>]; [<reflink idref="bib22" id="ref127">22</reflink>]) confirmed these findings, expanding the analysis beyond the mean, and indicating that the motherhood penalty varies across the wage distribution. However, due to the use of different QR models, these recent studies have led to conflicting findings regarding the motherhood penalty.</p> <p>As illustrated by [<reflink idref="bib7" id="ref128">7</reflink>], [<reflink idref="bib22" id="ref129">22</reflink>], and [<reflink idref="bib13" id="ref130">13</reflink>], several issues arise when QR is used to analyze longitudinal data. First, the development of estimators for CQR with fixed effects is relatively new. None of the proposed methodologies has gained widespread usage because of the computational complexity of the method and the restrictive assumptions required for estimation. In contrast, because UQR can be estimated using OLS, this methodology can more easily account for fixed effects, which has gained attention in the applied literature.</p> <p>Second, previous attempts to account for individual fixed effects were either computationally difficult to implement because they required a large number of parameters (dummy inclusion approach), or because they were not appropriate outside LR analysis (incidental parameter problem).</p> <p>Third, the specific interpretation of the estimated effects depends on whether CQR or UQR is used, as these models capture conceptually different aspects of the relationships between dependent and independent variables. For instance, partial effects obtained using CQR are only valid when conditioning on all observed characteristics and cannot be easily generalized as effects on the unconditional distribution. On the other hand, partial effects obtained using UQR focus on explaining potential effects that may affect the distribution as a whole (the unconditional distribution), but they cannot be interpreted as effects for specific individuals. Additionally, because UQR relies on local approximations, the analysis of events like motherhood, may be biased if considered as a change that affects the motherhood status of all women in the sample.</p> <p>In this section, we revisit [<reflink idref="bib6" id="ref131">6</reflink>] to show how best to interpret findings related to the motherhood effect on earnings. To do so, we contrast insights from the LR model and the three types of quantile regression models discussed in Section 3, CQR, UQR and QTE, while accounting for individual fixed effects. We use the same model specifications and sample selection used in [<reflink idref="bib6" id="ref132">6</reflink>]; [<reflink idref="bib7" id="ref133">7</reflink>]). The sample contains 36,361 observations on 3,293 non-Hispanic white women, covering the years 1979–2004, extracted from the NLSY79. Following [<reflink idref="bib6" id="ref134">6</reflink>], we examine the effect of number of children, as the proxy for motherhood, on logged hourly wages for women with wages between $1 and $200 per hour. The model specifications include sets of family structure, work effort, human capital, job characteristics, and demographic variables.[<reflink idref="bib38" id="ref135">38</reflink>]</p> <p>We assume that there are no issues involving endogeneity, self-selection, omitted variables, or incorrect model function forms, once individual fixed effects are accounted for.[<reflink idref="bib39" id="ref136">39</reflink>] We do this because we aim to replicate [<reflink idref="bib7" id="ref137">7</reflink>], contrasting the findings across different methodologies. Furthermore, the specification is common across studies on the motherhood penalty. Nevertheless, it is important to consider that controlling for individual fixed effects only controls for time-fixed unobservable characteristics. If preferences for children change across time, using individual fixed effects may not be enough to control for those unobserved effects, which might generate inconsistent results. Because these assumptions are strong, causal interpretation is unrealistic in all cases. Nonetheless, they can be used as examples for the appropriate interpretation across different models.</p> <p>Given the differences between CQR and UQR with QTE models, we provide two sets of results. Table 1 provides estimates for the effect of the number of children on logged hourly wages across three model types with fixed effects—linear regression (Model 1, LR), conditional quantile regression (Model 2, CQR) and unconditional quantile regression (Model 3, UQR), while controlling for individual fixed effects. Table 2 uses the same models (Models 1–3) to show the effect of motherhood as a binary variable that takes the value of one if a woman has any children and incorporates quantile treatment effects (Model 4, QTE). Tables provide results from models that include all covariates and fixed effects. Additional model results are available in the Online Appendix, which can be found at <ulink href="http://smr.sagepub.com/supplemental/">http://smr.sagepub.com/supplemental/</ulink>. Expanding on Tables 1 and 2, which provide estimates at the 10th, 25th, 50th, 75th, and 90th percentiles, Figures 1 and 2 present results from each model across the earnings distribution.</p> <p>Graph</p> <p>Table 1. Results From Linear Regression (LR), Conditional Quantile Regression (UQR), Unconditional Quantile Regression (UQR) Models Predicting Logged Hourly Wages With Fixed Effects (FE) for Selected Quantiles Associated With Number of Children.</p> <p> <ephtml> <table><thead><tr><th rowspan="2">Quantile</th><th colspan="2">Model 1: LR With FE</th><th colspan="2">Model 2: CQR With FE</th><th colspan="2">Model 3: UQR With FE</th></tr><tr><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th></tr></thead><tbody><tr><td>q10</td><td>−.028*** (.006)</td><td>−.027</td><td>−.026** (.008)</td><td>−.025</td><td>−.023*** (.007)</td><td>−.023</td></tr><tr><td>q25</td><td>−.028*** (.006)</td><td>−.027</td><td>−.022*** (.006)</td><td>−.022</td><td>−.041*** (.008)</td><td>−.040</td></tr><tr><td>q50</td><td>−.028*** (.006)</td><td>−.027</td><td>−.020** (.006)</td><td>−.020</td><td>−.033*** (.009)</td><td>−.032</td></tr><tr><td>q75</td><td>−.028*** (.006)</td><td>−.027</td><td>−.019** (.006)</td><td>−.018</td><td>−.023* (.010)</td><td>−.023</td></tr><tr><td>q90</td><td>−.028*** (.006)</td><td>−.027</td><td>−.018* (.007)</td><td>−.018</td><td>.015 (.018)</td><td>.015</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Source</emph>: NLSY79, 1979–2004, non-Hispanic white women with earnings, <emph>N</emph> = 36,361 observations and 3,293 individuals.</p> <ulist> <item>2 <emph>Note</emph>: Relationship between number of children and logged hourly wages. Clustered standard errors at the individual level in parentheses. Models include all control variables from BH (2010). CQR model fixed effects obtained following [<reflink idref="bib10" id="ref138">10</reflink>], with standard errors based on 100 bootstrapped samples, clustered at the individual level. UQR with fixed effects estimated using the within transformation estimator and implemented using the command -rifhdreg- with standard errors based on 100 bootstrapped samples, clustered at the individual level. UQR estimates differ from BH (2014) because -rifhdreg- uses a different bandwidth compared to the command -rifreg-. Because some of these coefficients exceed 0.1, we use the following formula to determine the percent change in wages for a one-unit change in each predictor variable: %Δ(y) = 100*(e<sups>b</sups> − 1).</item> <item>3 *<emph>p</emph> <.05. **<emph>p</emph> <.01. ***<emph>p</emph> <.001.</item> </ulist> <p>Graph</p> <p>Table 2. Results From Linear Regression (LR), Conditional Quantile Regression (UQR), Unconditional Quantile Regression (UQR), and Quantile Treatment Effects (QTE) Models Predicting Logged Hourly Wages With Fixed Effects (FE) for Selected Quantiles Associated With Motherhood.</p> <p> <ephtml> <table><thead><tr><th rowspan="2">Quantile</th><th colspan="2">Model 1: LR With FE</th><th colspan="2">Model 2: CQR With FE</th><th colspan="2">Model 3: UQR With FE</th><th colspan="2">Model 4: QTE With FE</th></tr><tr><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th><th><italic>b (SE)</italic></th><th>e<sup>b</sup> −1</th><th><italic>b (SE)</italic></th><th>e<sup>b</sup> − 1</th></tr></thead><tbody><tr><td>q10</td><td>−.044*** (.011)</td><td>−.043</td><td>−.043*** (.013)</td><td>−.042</td><td>−.045*** (.012)</td><td>−.044</td><td>−.050** (.024)</td><td>−.049</td></tr><tr><td>q25</td><td>−.044*** (.011)</td><td>−.043</td><td>−.033*** (.010)</td><td>−.032</td><td>−.085*** (.017)</td><td>−.081</td><td>−.071*** (.025)</td><td>−.069</td></tr><tr><td>q50</td><td>−.044*** (.011)</td><td>−.043</td><td>−.030*** (.009)</td><td>−.030</td><td>−.054*** (.015)</td><td>−.053</td><td>−.039** (.019)</td><td>−.038</td></tr><tr><td>q75</td><td>−.044*** (.011)</td><td>−.043</td><td>−.030*** (.009)</td><td>−.030</td><td>−.034* (.018)</td><td>−.033</td><td>.011 (.026)</td><td>.011</td></tr><tr><td>q90</td><td>−.044*** (.011)</td><td>−.043</td><td>−.033*** (.011)</td><td>−.032</td><td>.063* (.033)</td><td>.065</td><td>.096** (.038)</td><td>.101</td></tr></tbody></table> </ephtml> </p> <ulist> <item>4 <emph>Source</emph>: NLSY79, 1979–2004, non-Hispanic white women with earnings, N = 36,361 observations and 3,293 individuals.</item> <item>5 <emph>Note</emph>: Relationship between motherhood and logged hourly wages. Motherhood takes the value of 1 if a woman has any children. Clustered standard errors at the individual level in parentheses. Models include all control variables from BH (2010). CQR model fixed effects obtained following [<reflink idref="bib10" id="ref139">10</reflink>], with standard errors based on 100 bootstrapped samples, clustered at the individual level. UQR with fixed effects estimated using the within transformation estimator and implemented using the command -rifhdreg- with standard errors based on 100 bootstrapped samples, clustered at the individual level. QTE with fixed effects estimated using the within transformation, and an inverse probability weighting (IPW) approach, and reports average treatment effects. The IPW is constructed using propensity scores derived from a logit model that includes the full set of explanatory variables, except for the individual fixed effects. The propensity score is trimmed to be between 0.025 and 0.975. QTE estimations are implemented using the command -rifhdreg- with standard errors based on 100 bootstrapped samples, clustered at the individual level. Because some of these coefficients exceed 0.1, we use the following formula to determine the percent change in wages for a one-unit change in each predictor variable: %Δ(y) = 100*(eb − 1).</item> <item>6 *<emph>p</emph> <.05. **<emph>p</emph> <.01 ***<emph>p</emph> <.001.</item> </ulist> <p>Graph: Figure 1. Partial effects of number of children on log wages: Comparison across LR, CQR and UQR. Source : NLSY79, 1979–2004, non-Hispanic white women with earnings, N = 36,361 observations and 3,293 individuals. Note: Results from linear regression (LR), conditional quantile regression (UQR), unconditional quantile regression (UQR) models include individual and year fixed effects (FE), and the full set of explanatory variables.</p> <p>Graph: Figure 2. Partial effects of motherhood on log wages: Comparison across LR, CQR, UQR, and QTE. Source : NLSY79, 1979–2004, non-Hispanic white women with earnings, N = 36,361 observations and 3,293 individuals. Note : Results from linear regression (LR), conditional quantile regression (UQR), unconditional quantile regression (UQR), and quantile treatment effects (QTE) models include individual and year fixed effects (FE), and the full set of explanatory variables. QTE estimates correspond to the average treatment effects.</p> <hd id="AN0176861662-9">LR Estimates</hd> <p>According to the results in Model 1 in Table 1, there is a significant and negative wage penalty associated with motherhood, when measured as the discrete variable, number of children. Based on the LR model, mother's wages are expected to decrease by 2.8 percent per additional child on average, a finding in line with Avellar and Smock (2003) and [<reflink idref="bib13" id="ref140">13</reflink>] but lower than that of [<reflink idref="bib8" id="ref141">8</reflink>].[<reflink idref="bib40" id="ref142">40</reflink>] When measured as a binary variable (Table 2 Model 1), the effect of motherhood on wages is larger; it is associated with a wage reduction of about 4.3 percent. Although this is the most common way in which researchers interpret the estimated effects from LR models, we can also provide other interpretations to the partial effects dependent on certain model assumptions.</p> <p>First, under the assumption of homoscedasticity, everyone in the population will experience the same effect. Consider, for example, three random women, one with no children, one with one child, and one with two children, the first model predicts each of these women will experience a decline in their wages of 2.8 percent if they had an additional child.[<reflink idref="bib41" id="ref143">41</reflink>] Alternatively, in the second model, upon becoming a mother, a woman would experience a wage decline of 4.3 percent. This interpretation can be considered the individual level partial effect.</p> <p>Second, if we relax the assumption of homoscedasticity, we can no longer assume that every woman will experience the same effects associated with the birth of additional child. Consider all women who are identical in terms of their characteristics. For example, they are all 30 years old, married, and have one child, 12 years of education, 10 years of work experience, and so on. If they all have an additional child, each woman is not going to experience an identical wage penalty. Some will experience a larger decline in wages while others may experience no change at all. On average, however, having an additional child would reduce their wages by 2.8 percent. Alternatively, becoming a mother would reduce wages on average by 4.3 percent. This interpretation can be considered the conditional partial effect.</p> <p>Finally, because the LR model specification assumes a linear relationship between the dependent and independent variables, we can also make inferences regarding the unconditional changes in wages across the population. For example, if every woman in the population has an additional child, assuming other characteristics remain constant, we would expect average wages among all women to decline by 2.8 percent. In this situation, it may be more accurate to discuss this effect in relation to a smaller change in the average number of children. For instance, if the average number of children increased by 0.5, we would expect average wages to decrease by 1.4 percent (2.8 percent × 0.5). This gives us the unconditional partial effect. In the case of motherhood, the interpretation refers to how much lower women's wages would be if all women were mothers, compared to if they had no children. A more reasonable interpretation would indicate that if the share of mothers in the sample increased by 10 percent, average wages would decline in 0.43 percent.</p> <p>The differences between interpreting LR results as individual, conditional, or unconditional partial effects are very fine distinctions that few studies make. This is understandable because such distinctions are not always central to most LR analyses. In fact, under the homoscedasticity assumption, no distinction is needed, because everyone in the population is affected in the same way. As this paper shows, however, understanding the differences between individual, conditional, and unconditional partial effects is crucial for interpreting results from conditional and unconditional quantile regression models.</p> <hd id="AN0176861662-10">CQR Estimates</hd> <p>Depending on the research question of interest, LR estimates can be used to draw inferences about how changes in the independent variables translate into changes in the dependent variable at the individual level (individual effect), on average for individuals with the same characteristics (conditional effects), or on the whole population average (unconditional effects). In the case of conditional quantile regression only the conditional effect interpretation is feasible for all independent variables, unless stronger assumptions regarding ranking are applied.</p> <p>Model 2 in Tables 1 and 2 provides estimates of the motherhood penalty measured as number of children (Table 1) and as any children (Table 2) applying CQR with fixed effects using the [<reflink idref="bib10" id="ref144">10</reflink>] estimator.[<reflink idref="bib42" id="ref145">42</reflink>]Figures 1 and 2 also provide an illustration of the effects across the different quantiles. According to the CQR model results, motherhood has varying effects on earnings across the wage distribution. As shown with the LR results in Model 1, a typical woman experiences an average wage penalty of 2.8 percent per additional child and a 4.3 percent average wage penalty for being a mother and this is true at every quantile. Based on the CQR estimates in Model 2 (Table 1), however, a typical woman would experience a wage penalty that ranges between 2.6 percent if she is ranked at the 10th percentile to 1.8 percent if she is ranked at the 90th percentile, assuming her rank remain constant. The trend of the impact of motherhood (Table 2 Model 2), provides qualitatively similar results, suggesting that a typical woman may experience a larger wage penalty at the 10th percentile (4.2 percent), but a smaller penalty near the 90th percentile of the distribution (3.2 percent). Similar to [<reflink idref="bib6" id="ref146">6</reflink>] CQR models, Figures 1 and 2 suggest that the motherhood penalty is larger for women at the bottom of the conditional distribution, although differences across the distribution are small.</p> <p>We use the term "typical" woman to refer to a nonspecific woman who has average or median characteristics, thus stating the conditionality of the effects. This woman would be unmarried with no children and 12 years of education (approximately a high school diploma). We are also interpreting the coefficients under two strict assumptions. We assume that we know a specific woman's position among other women with identical observed characteristics, and we assume that her position does not change after she has a new child (rank invariance assumption). Although both assumptions are consistent with the model as described in equation (8d), population rankings are never observed in empirical settings and the assumption that a person's conditional ranking does not change is strong. This implies that we cannot make inferences for a specific individual because we do not know their ranking in the conditional distribution or if that ranking will remain constant after a change in the number of children.</p> <p>A second interpretation, which is more fitting in the framework of CQR, refers to the conditional effect, implicitly acknowledging that we cannot be certain about individuals' rankings. This can be accomplished in two ways. The first avoids references to specific individuals, and refers instead to groups of women who have the same characteristics. For example, we could say, if women with average characteristics had an additional child, their wage distribution would decline between 2.5 percent at the 10th percentile that declines to about 1.8 percent at the 90th percentile (Table 1 Model 2). Or, if the focus in on motherhood, the wage penalty after transitioning to motherhood would range from 4.2 percent at the 10th percentile to 3.2 percent at the 90th percentile (Table 2 Model 2).</p> <p>An alternative way to understand the results is to consider the distribution of wages across two groups of women who are identical in every way except that one group comprises mothers with one child and the other comprises women with no children. The difference in average wages would be 2.8 percent for women with one additional child, according to the LR model. However, based on the CQR model, we should expect wages among women with one child at the 10th percentile to be 2.5 percent lower than wages for women without children, but only 1.8 percent lower at the 90th percentile (Table 1 Model 2). We would also expect wages among all mothers at the 10th percentile to be 4.2 percent lower than wages for women without children, but only 3.2 percent lower at the 90th percentile (Table 2 Model 2). This interpretation emphasizes how the two distributions compare to each other.</p> <hd id="AN0176861662-11">UQR Estimates</hd> <p>Model 3 in Tables 1 and 2 provides UQR results for selected quantiles for the effect of number of children and motherhood. In contrast to CQR, UQR models analyze how changes in the distribution of characteristics affect the unconditional distribution of the outcome, but these models do not provide information about how those changes influence individual experiences. This means that coefficients from UQR models must be interpreted in relation to how changes in the average number of children per woman or in the share of mothers affect the overall (across time) earnings distribution.</p> <p>According to the LR results in Table 1 Model 1, if the average number of children per woman were to increase by one child between 1979 and 2004, average wages among all women would have been 2.8 percent lower. However, looking at the effects throughout the distribution, the UQR results in Model 3 suggest that a change in the distribution of the number of children will have a heterogeneous impact across the distribution of wages. If the average number of children per woman was to increase by one, we would expect wages at the bottom of the distribution to decrease by 2.3 percent, but we would also observe an increase in wages at the top of the distribution (90th percentile) by approximately 1.5 percent, although this last coefficient is not statistically significant. Figure 1 further illustrates this relationship. Although an increase in the average number of children per woman is negatively associated with wages for most of the distribution, the negative effect shrinks above the 60th percentile, turning positive and increasing around the 90th percentile.</p> <p>When measured as a binary variable, the results in Table 2 Model 3 suggest a similar trend in terms of the wage penalty of motherhood, with an estimated impact that is positive and statistically significant at the top of the wage distribution. However, the most appropriate way to interpret the coefficients associated with motherhood relies on describing the effect associated with a marginal increase in the share of mothers in the sample. Based on the UQR results, if the share of mothers in the sample were to increase by 10 percentage points, wages at the bottom of the distribution would decrease faster than wages at the top. For instance, we would expect the 10th quantile of wages to decrease by 0.44 percent, the 50th quantile to decrease 0.53 percent. The results in Figure 2 suggests that above the 60th percentile, the negative effect of motherhood shrinks, with a positive impact above the 80th percentile. At the 90th percentile a 10 percentage point increase in the share of mothers may increase wages by 0.65%.</p> <p>As indicated by our discussion, the interpretation of UQR results should be kept in terms of unconditional statistics. It is also important to consider that, in contrast with conditional effects, unconditional quantile effects are not isolated. For example, if each woman were to have an additional child, this would cause a change in earnings distribution that would be observed across all quantiles simultaneously. Because of this, plotting coefficients may be useful to appreciate the full extent of the distributional changes. Furthermore, while we use the language of "an additional child" effect, this may not be adequate because less than 10% of women in the sample have more than 3 children. This kind of interpretation is more appropriate when considering the motherhood as a binary variable. As the proportion of mothers in the sample increases, the whole distribution shifts, affecting all quantiles simultaneously.</p> <hd id="AN0176861662-12">QTE Estimates</hd> <p>Expanding on these more common models, Model 4 in Table 2 provides the QTE results for selected quantiles and at the mean. This model controls for differences in characteristics using IPW and includes these controls and individual fixed effects in the model specification. The IPW are constructed by trimming the propensity score to be between 0.025 and 0.975.</p> <p>We again measure the effect of motherhood using a binary variable that takes the value of one if a woman has any children and zero otherwise. For the estimation of the QTE, a logit model is used for the estimation of the propensity score using the same variables used in [<reflink idref="bib6" id="ref147">6</reflink>], except for the individual fixed effects. To reduce the sensitivity of the estimations to extreme propensity scores, the QTE are estimated excluding observations with a predicted propensity score below 0.025 and above 0.975. For comparison, we report average treatment effects (ATE), with ATE on the mean being reported in the Online Appendix, which can be found at <ulink href="http://smr.sagepub.com/supplemental/">http://smr.sagepub.com/supplemental/</ulink>.</p> <p>The estimated coefficients should be interpreted as the impact motherhood would have on the distribution of wages of all women, assuming that the distribution of other characteristics remain constant. In other words, the estimated effects measure how different the distribution of wages would be if we were to compare the wage distribution assuming all women are mothers, against all women not being mothers.</p> <p>When controlling for observed covariates as part of the IPW and directly in the model specification and unobserved time-invariant characteristics using individual fixed effects in QTE models (Table 2 Model 4), we observe a trend in the average QTE similar to the one observed for the UQR (Table 2 Model 3). Based on these estimates, motherhood would have an heterogenous effect on the wage distribution, potentially reducing wages at the bottom of the distribution in as much as 5 percent at the 10th percentile and reducing wages by 3.8 percent in the middle of the distribution, but increasing wages at the 90th percentile by 10.1 percent. This average effect is larger than the effect estimated using UQR.</p> <p>Importantly, QTE models should be interpreted by comparing the potential wages if all women were not mothers to a scenario where all women are mothers, and assuming the distribution of other characteristics remaining constant. While not discussed in this example, a researcher may also choose to interpret the average treatment effects on the treated (or untreated). This could facilitate the interpretation by analyzing potential effects on the treated or untreated group, rather than on the population as a whole.</p> <hd id="AN0176861662-13">Discussion</hd> <p>What is the relationship between <emph>x</emph> and <emph>y</emph>? Because social scientists typically use LR to answer this question, most research presents an answer and interpretation related to the mean that people have come to expect. This is a good example of how the methods we use to answer a research question shape the answers we find. Due to the ubiquity of linear regression models, much of the accumulated quantitative social science knowledge is based on the mean, which is not necessarily problematic. Often, the mean presents a good summary of the outcome, and, hence, a good description of the relationship between two variables, <emph>x</emph> and <emph>y</emph>. But, in many situations, when the relationship between <emph>x</emph> and <emph>y</emph> varies across the distribution of <emph>y,</emph> the mean might not offer the best summary available. Focusing on only the mean can obscure results, especially when researchers are interested in issues like gender inequality ([<reflink idref="bib2" id="ref148">2</reflink>]). When this happens, researchers must expand their toolkits to test new methods for studying this relationship.</p> <p>Quantile regression provides a framework for analyzing heterogeneous effects beyond what LR can provide. However, in applying quantile regression, researchers must take care in defining a research question and interpreting the results. This comes down to determining what it means to ask and answer the questions—How does a change in <emph>x</emph> affect the outcome <emph>y</emph> for any individual in the data? How does a change in <emph>x</emph> affect the conditional distribution of <emph>y</emph>? Or, how does a change in <emph>x</emph> affect the unconditional distribution of <emph>y</emph>? Although QR models do not provide an answer to the first question like LR models do, CQR can be used to answer the second question, and UQR can be used to answer the third. In this paper, we also discuss a fourth strategy, QTE, as a middle ground between what CQR and UQR can estimate. This strategy answers the second question, with respect to a single variable of interest, by identifying a distributional treatment effect. Using RIF regressions to identify QTE, also allows a researcher to answer the third question for other variables in the model. Each method, therefore, provides a slightly different interpretation for the relationship between <emph>x</emph> and <emph>y</emph> across the distribution of <emph>y.</emph></p> <p>CQR can be used to study the relationship between variables across the conditional distribution of the outcome variable. However, the interpretation should be framed as effects experienced by groups that are defined by a set of characteristics (conditional effects). CQR models are easy to interpret in the case of a single variable, but interpretation problems typically arise once additional covariates are added. These controls essentially change an observation's place in the distribution, implicitly creating subgroups defined by covariates. Because of this, the interpretation of the estimations should be considered as local effects, given a set of individuals with specific characteristics (e.g., women with the same years of education, hours of work, and marital status). These results cannot be generalized as effects that would affect the unconditional statistic of interest.</p> <p>Overcoming this limitation, UQR provides an additional framework for analyzing heterogeneous effects across a distribution where the definitions of quantiles are not affected by individual values of model covariates, as they describe a characteristic of the distribution of <emph>y</emph> as a whole. However, because of this, UQR can only be used to identify effects on the unconditional distribution of the outcome. Consequently, coefficients can only be interpreted as in relation to how changes in the distribution of independent characteristics <emph>x</emph>, usually approximated by changes in the unconditional mean, affect the unconditional quantile of the outcome. Importantly, inferences are only valid when analyzing small changes in the distribution. Because of this, special care is needed when interpreting the effects of categorical and discrete data, such as motherhood and the number of children. UQR may not provide consistent estimates when the changes in the distribution of characteristics are large. Furthermore, it remains to the researcher to evaluate if the unconditional distribution, in a panel data framework, is appropriate for their research question.</p> <p>QTE presents an important compromise between CQR and UQR. By constructing the RIF's across groups defined by a single variable of interest, distributional treatment effects, measured via changes in quantiles, can be estimated and discussed. This allows researchers to examine how the distribution of the dependent variable, <emph>y</emph>, changes as the main conditioning variable changes, after controlling for differences in the distribution of other characteristics. If the model is estimated using RIF regressions, the interpretation of all other variables in the model is similar to the one for UQR. The interpretation of QTE will also depend on whether the researcher is interested in analyzing average treatment effects that apply to the population as a whole or treatment effects on the treated or untreated groups. Similar to UQR, it needs to be evaluated if unconditional distributions pooling data across time are of interest in a given research question.</p> <p>What does this mean for the debate over using QR models to research the motherhood penalty? In this case, the use of different methods has led to varying results for the effects of children on women's wages and a large debate over the "true" penalty of motherhood. Our replication of [<reflink idref="bib6" id="ref149">6</reflink>] shows that neither type of QR model specification is essentially "right" or "wrong." Each, however, offers a different interpretation and understanding of the motherhood penalty, as indicated in Tables 1 and 2, and Figures 1 and 2.</p> <p>According to the LR model, controlling for time-varying covariates and unobserved individual-level time-invariant factors, there is a wage penalty associated with motherhood. Specifically, this model indicates that on average, having an additional child will reduce women's wages by 2.8 percent, and if the average number of children per woman in the population increases by one, holding all other characteristics constant, women's wages will decline by 2.8 percent. Using a different definition of motherhood, a binary variable, also suggests a larger effect of a 4.3 percent average penalty associated with motherhood. Certain QR models, however, indicate that this relationship varies across the conditional and unconditional wage distribution.</p> <p>Estimates from the CQR model suggest that an additional child, or motherhood more generally, has a mostly homogenous effect for women across the conditional distribution. The CQR estimates also present a somewhat smaller wage penalty, predicting that among women with the same characteristics an additional child would lead to a wage decline of about 1.8–2.6 percent with larger penalties below the 50th percentile, and a wage decline of 3.2–4.2 percent when using the binary definition of motherhood.</p> <p>Although having an additional child has a relatively stable negative effect on wages across the conditional wage distribution, the UQR results suggest a generalized increase in the number of children per woman in the population will have a more heterogenous impact on the unconditional distribution of wages observed across all years. If the average number of children per woman were to increase by one, the UQR estimates suggest that up to the 60th quantile, unconditional quantiles of wages would decline between 2.3 percent to 4.0 percent. For the upper section of the distribution, the decline shrinks to zero, and even small increases above the 85th quantile are estimated. Although the point estimates when using a binary definition of motherhood are larger when using UQR, the interpretation of the coefficients need to be rescaled to consider a marginal increase in the share of mothers in the sample.</p> <p>Overall, the UQR results suggest that an exogenous increase in the number of children, or an increase in the share of mothers in the sample, will increase inequality, widening the wage gap between the top and the bottom of the distribution. However, this result should be considered cautiously for two reasons. First, because we are using panel data, UQR is measuring the impact of an additional child on the overall distribution of wages observed across all the survey waves. It is up to the researcher to decide the merits of analyzing changes on the overall wage distribution across multiple years. Second, half of the person-years in the sample have no children, and only 6 percent of those who are mothers have more than three children, thus the thought experiment of one additional child per women may not be appropriate in the context of UQR. Similar concerns are raised if we try to analyze the binary variable of motherhood as a treatment effect in the context of UQR.</p> <p>In addition to replicating findings regarding a key question in the sociological literature, this paper incorporates an additional model based on quantile treatment effects that provides researchers with an alternative to CQR and UQR models. Based on our estimates, QTE, which aims to compare the distribution of women's earnings based on their motherhood status, presents a varying relationship across the distribution, from a negative effect at the bottom, to a large and positive effect at the top of the distribution (up to 10 percent effect). These are similar to the point estimations obtained using the UQR model. Similar to the UQR analysis, however, it is up to the researcher to decide if analyzing a treatment effect on the overall distribution of wages across years is appropriate for their research question.</p> <p>Results further support [<reflink idref="bib6" id="ref150">6</reflink>], [<reflink idref="bib7" id="ref151">7</reflink>]) who found that mothers at the top of the wage distribution do not experience a motherhood penalty after controlling for key job characteristics, especially a change in work hours. In terms of mechanisms behind this relationship, it is important to note that women with wages this high have access to support in the form of nannies, chefs, and cleaning services that other women do not. The large wage premiums we find at the top of the wage distribution also suggest that there may be other factors, like individual time varying variables, we are not controlling for, that may be driving this effect. For instance, it is possible that women's preferences for having children may change over time.</p> <hd id="AN0176861662-14">Conclusion</hd> <p>With several potential models, what's a researcher who wants to study the relationship between <emph>x</emph> and <emph>y</emph> across the distribution of <emph>y</emph> to do? Thankfully, there are several options for approaching this problem through quantile regression. Below, we list a set of best practices for applying quantile regression models. These best practices are not exhaustive, but they do provide a roadmap toward approaching a study that requires QR analysis. Most suggestions also relate to steps that come before data analysis—an extremely important part of creating a research project that can often go overlooked.</p> <hd id="AN0176861662-15">Best Practices for QR</hd> <p></p> <ulist> <item> 1. <emph>Clearly identify the research question and variables of interest.</emph></item> </ulist> <p>As in any research article, the choice of the econometric approach will depend on the research question and relationship of interest. It is useful to remember that QR allows identifying heterogeneous effects with respect to the conditional (CQR) and unconditional distributions (UQR) of the dependent variable. QTE can be used if the interest lies on analyzing distributional treatment effects. If the interest lies in analyzing heterogeneity with respect to an independent variable, other approaches not discussed here may be required.</p> <p>"How does motherhood affect women's earnings?" is a broad question that can be answered in many ways. Perhaps the simplest answer comes from linear regression, which shows that, on average, each additional child is associated with a 2.8 percent decrease in women's earnings and motherhood more generally is associated with a 4.3 percent decrease. However, that still leaves open questions of heterogeneity. Is the motherhood penalty larger for low-wage or high-wage women? Asking this question creates a need to examine the relationship between motherhood and earnings across the wage distribution, but it can be answered in multiple ways that examine either the conditional or unconditional wage distributions. Here, it is important to identify if the interest falls in analyzing the effects that any particular woman will experience or if the interest falls in analyzing how changes in the overall number of children in the population or motherhood status. In the case of the former, LR and CQR models are more appropriate, while UQR and QTE present better approaches for the latter.</p> <p></p> <ulist> <item> 2. <emph>Determine the most useful or adequate interpretation for answering the research question, especially in terms of interpreting partial effects.</emph></item> </ulist> <p>As discussed from the beginning of this paper, LR provides a good approximation of the average effects between dependent and independent variables that is applicable for most research questions. If the interest lies in analyzing heterogeneous effects, CQR can be used to analyze local effects of how changes in an independent variable affect the conditional distribution of the dependent variable, which complement the average effects identified using LR. If the interest lies in analyzing global distribution effects, UQR can be used to identify how small changes in the distribution of independent variables affect the distribution of the dependent variable, measured by changes in the unconditional quantiles. Finally, if the interest lies in analyzing distributional effects of policies or large changes in the independent variables, thus analyzing how two or more distributions compare across quantiles, after controlling for the distribution of other factors, QTE may be the most appropriate approach.</p> <p>In the case of the motherhood penalty, CQR model results indicate that wages among mothers at the 10th percentile would be 4.2 percent lower than wages for women without children, but only 3.2 percent lower at the 90th percentile (Table 2 Model 2). UQR results indicate that if the share of mothers in the sample were to increase by 10 percentage points, the 10th quantile of wages would decrease by 0.44 percent and the 50th quantile would decrease by 0.53 percent, but a 10 percentage point increase in the share of mothers at the 90th percentile may increase wages by 0.65 percent (Table 2 Model 3). Finally, QTE models indicate that if all women became mothers, this would decrease wages at the bottom of the distribution by 4.9 percent at the 10th percentile and 3.8 percent in the middle of the distribution, but increase wages at the 90th percentile by 10.1 percent (Table 2 Model 4).</p> <p></p> <ulist> <item> 3. <emph>Note the assumptions for the interpretations of different QR model types.</emph></item> </ulist> <p>As described earlier, each QR model has its own advantages and limitations. While CQR can be used to estimate local effects in the distribution caused by changes in a single the independent variables, the effects cannot be described as an individual effect, unless rank invariance is assumed, nor extrapolated as an effect on the overall/unconditional distribution. UQR, on the other hand, can be used to draw global distributional effects caused by small changes in the distribution of characteristics, but cannot be used to draw inferences of individual or local effects. QTE, when estimated via RIF regressions, is a compromise between CQR and UQR. It allows us to estimate unconditional like effects for all characteristics except for a single conditioning variable for which a distributional treatment effects are obtained.</p> <p></p> <ulist> <item> 4. <emph>Compare results across multiple model types.</emph></item> </ulist> <p>Although it is important to choose a method that best fits your research question, it is equally important to understand that all these methodologies complement each other, as they identify different aspects regarding the relationships between dependent and independent variables. Observing how results vary with a different model and interpretation may reveal patterns that are otherwise overlooked and hidden. The addition of QTE models to the motherhood penalty debate provides further support for studies that have relied on UQR models, as these models present similar results with more flexible specifications.</p> <p>These best practices do not just apply to analyses using QR. They include questions that all researchers should consider before embarking on a new project. However, following such guidelines becomes even more important in situations where researchers have many choices for their models that imply different interpretations, such as with quantile regression.</p> <hd id="AN0176861662-16">Supplemental Material</hd> <p>Graph: Supplemental Material, sj-do-1-smr-10.1177_00491241211036165 for Moving Beyond Linear Regression: Implementing and Interpreting Quantile Regression Models With Fixed Effects by Fernando Rios-Avila and Michelle Lee Maroto in Sociological Methods & Research</p> <ref id="AN0176861662-17"> <title> Notes </title> <blist> <bibl id="bib1" idref="ref97" type="bt">1</bibl> <bibtext> The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref47" type="bt">2</bibl> <bibtext> The author(s) received no financial support for the research, authorship, and/or publication of this article.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref26" type="bt">3</bibl> <bibtext> Michelle Lee Maroto https://orcid.org/0000-0002-7506-0046</bibtext> </blist> <blist> <bibl id="bib4" idref="ref91" type="bt">4</bibl> <bibtext> Supplemental material for this article is available online</bibtext> </blist> <blist> <bibl id="bib5" idref="ref9" type="bt">5</bibl> <bibtext> It should be noted that when the variable of interest <emph>x</emph> is binary (treatment status), under the stated assumptions, the causal effect can also be interpreted as a treatment effect or policy effect of how the distribution of <emph>y</emph> changes when comparing the treated and untreated group.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref2" type="bt">6</bibl> <bibtext> In practice, due to problems of omitted variables, endogenous selection, or model misspecification, assuming that all control variables are exogenous is often unrealistic. Thus, researchers must be cautious in regards to providing causal interpretation to all explanatory variables in LR or QR models alike.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref3" type="bt">7</bibl> <bibtext> There is recent discussion in the literature in regards to the limitation of fixed effects models and the estimation of causal effects, which may apply not only to LR models, but also QR models. See [21] for the explanation of the arguments regarding the identification of treatment effects in a difference in difference setting.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref1" type="bt">8</bibl> <bibtext> While LR models still provide unbiased and consistent estimates, they are no longer efficient and may hide some heterogeneity.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref31" type="bt">9</bibl> <bibtext> We want to thank Michelle J. Budig and Melissa J. Hodges for kindly sharing the datasets and programs to replicate their 2014 paper.</bibtext> </blist> <blist> <bibtext> This identification is based on a local linear approximation via recentered influence functions (RIF).</bibtext> </blist> <blist> <bibtext> Treatment effects are an example of large changes in the distribution of <emph>x</emph> when one aims to analyze distribution changes between a treated (<emph>x=1</emph>) and an untreated (<emph>x=0</emph>) population.</bibtext> </blist> <blist> <bibtext> While the discussion focuses on individual fixed effect, a single high dimensional fixed effect, many of the methodologies discussed can also be adapted for the inclusion of multiple high dimensional fixed effects. Nevertheless, as recently discussed in [25] for LR models, using multiple sets of fixed effects may provide estimations that are difficult to interpret.</bibtext> </blist> <blist> <bibtext> For simplicity, we assume the function <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>σ</mtext><mfenced><mo>.</mo></mfenced></mrow></math> </ephtml> does not depend on the individual effect <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> , but that assumption can be lifted as explored in [28] for the estimation of conditional quantile regressions via method of moments (see next section). While the estimation of LR models does not require any assumption regarding the idiosyncratic error <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>u</mi><mrow><mi>i</mi><mi>t</mi></mrow></msub></mrow></math> </ephtml> , using a specification that explicitly states heteroskedasticity allows to transition from LR models to conditional quantile models.</bibtext> </blist> <blist> <bibtext> Or N-1 if an intercept is in the model is considered.</bibtext> </blist> <blist> <bibtext> When including 2 or more fixed effects, iterative within transformations, as discussed in [11] and [35], can be applied. In Stata, OLS with fixed effects can be estimated using the using the official commands <emph>areg, absorb()</emph> or <emph>xtreg, fe,</emph> or the community-contributed command <emph>reghdfe, absorb()</emph> ([11]), which allows for the inclusion of multiple fixed effects. Nevertheless, one should be aware of the limitations of identifying treatment effects in this setting as discussed in [21]</bibtext> </blist> <blist> <bibtext> Alternative methods include the estimation of random effects to panel data, which requires stronger assumption to obtain consistent estimates for <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mtext>β</mtext></math> </ephtml> , and estimation correlated random effect models ([44]), which provide estimates equivalent to fixed effect estimator for LR models.</bibtext> </blist> <blist> <bibtext> In Stata, the estimation of the conditional effect can be done with the postestimation command <emph>margins</emph> using the <emph>at()</emph> or <emph>at means</emph> option. By default, however, the margins command estimates the equivalent to the average marginal effect.</bibtext> </blist> <blist> <bibtext> This conclusion does not change if variables are continuous or categorical.</bibtext> </blist> <blist> <bibtext> For a more technical discussion of CQR see [23] and [32].</bibtext> </blist> <blist> <bibtext> A more general form of expressing CQR is assuming the <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>y</mi><mo>=</mo><mi>x</mi><mi>β</mi><mfenced><mi>τ</mi></mfenced><mo>+</mo><mi>F</mi><mrow><mo>−</mo>1</mrow><mfenced><mi>τ</mi></mfenced></mrow></math> </ephtml> , where <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>F</mi><mrow><mo>−</mo>1</mrow><mfenced><mi>τ</mi></mfenced></mrow></math> </ephtml> is the inverse CDF for an iid distribution, and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mtext>β</mtext><mfenced><mtext>τ</mtext></mfenced></mrow></math> </ephtml> are functions of <emph>τ∼</emph>uniform <emph>(0,1).</emph></bibtext> </blist> <blist> <bibtext> It is useful to notice that in [28], <emph>b<subs>k</subs></emph> represents what is known as the location shift effect, whereas <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>a</mi><mi>k</mi></msub><mi>F</mi><mrow><mo>−</mo>1</mrow><mfenced><mi>τ</mi></mfenced></mrow></math> </ephtml> represents the scale shift effect.</bibtext> </blist> <blist> <bibtext> In this case <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>b</mi>0</msub><mfenced><mtext>τ</mtext></mfenced><mo /></mrow></math> </ephtml> corresponds to the CQR constant.</bibtext> </blist> <blist> <bibtext> The incidental parameter problem arises because the number of parameters <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mtext>δ</mtext><mi>i</mi></msub></mrow></math> </ephtml> that require estimation generally increases at the same speed as the sample size. In panel data this implies the number of effective observations used for its identification of each individual fixed effect is small (T), and the random nature effect of the individual effect remains. When using OLS, individual fixed effects can be "averaged out," but that is not possible with nonlinear estimation methods, like the one used for conditional quantile regressions.</bibtext> </blist> <blist> <bibtext> This estimator can be implemented in Stata using the community-contributed command <emph>xtqreg</emph> ([27]) or <emph>mmqreg</emph> ([37]).</bibtext> </blist> <blist> <bibtext> The assumption also implies that the internal rankings (unobserved factors) remain constant ([32]).</bibtext> </blist> <blist> <bibtext> More specifically, [17] call this an unconditional quantile partial effect.</bibtext> </blist> <blist> <bibtext> As described in [18], and discussed in [36], RIF regressions are meant to provide inferences in regards to unconditional effects on the distribution, with two exceptions. First, the standard LR model can also be considered as a special case of RIF regression, but one can still draw inferences and the individual and conditional levels. Second, RIF regressions involving FGT poverty measures can also be used to obtain both individual and conditional level effects. On the other hand, similar analysis can done pooling RIFs from different conditional distributions, as [18] do for their proposed decomposition analysis.</bibtext> </blist> <blist> <bibtext> RIFs have been derived for a large set of distributional statistics.[14] and more recently [36] provide a comprehensive list of RIF's for a large set distributional statistics developed in the literature.</bibtext> </blist> <blist> <bibtext> [17] also discuss the estimation of UQR using logit/probit models and other nonlinear/nonparametric models. The results, however, were comparable to the ones based on OLS. Nevertheless, using OLS for the estimation of UQR has drawbacks similar to linear probability models (LPM) when compared to other binomial models like probit or logit. The community-contributed command <emph>uqreg</emph>, companion to <emph>rifhdreg</emph>, can be used to estimate UQR using both logit and probit models.</bibtext> </blist> <blist> <bibtext> Interpretations for individual and conditional effects can be obtained if <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>R</mi><mi>I</mi><mi>F</mi><mfenced><mrow><msub><mi>y</mi><mi>i</mi></msub><mo>,</mo><mi>v</mi><mo>,</mo><msub><mi>F</mi><mi>y</mi></msub></mrow></mfenced></mrow></math> </ephtml> is a function of <emph>y<subs>i</subs></emph> only. That is the case for the mean, Watts poverty index, and Foster-Greer-Thorbecke poverty indices. Additionally, it is possible to make conditional interpretations for key variables, by constructing RIF's based on conditional, rather than unconditional, distributions. This is discussed in the next section.</bibtext> </blist> <blist> <bibtext> [36] suggests that centered quadratic terms could be used to control for changes in the variance of the independent variable.</bibtext> </blist> <blist> <bibtext> [38] also proposes a strategy for a more general case that includes large changes in the distribution of continues variables, or variables with limited range, based on a reweighting strategy.</bibtext> </blist> <blist> <bibtext> The use of RIF regressions for the estimation of statistics based on conditional distributions, rather than unconditional distributions, is not new.[18] use the same strategy for their application of oaxaca-type decomposition. This is also described in [36] for the estimation of distributional treatment effects, following [15]. Empirically, this is implemented using the option <emph>-over(tvar)-</emph> in the <emph>-rifhdreg-</emph> command, where <emph>-tvar-</emph> is a single variable that conditions the distribution of the outcome.</bibtext> </blist> <blist> <bibtext> Formally, this condition is also known as the conditional independence assumption of unconfoundedness. This assumption is similar to the one described for UQR via RIF's where the errors are assumed to be independently distributed from the independent variables.</bibtext> </blist> <blist> <bibtext> It is also possible to use weights defined as <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>ω</mi>0<mo>'</mo></msubsup><mo>=</mo><msub><mi>ω</mi>0</msub><mi>p</mi><mo>^</mo><mfenced><mi>z</mi></mfenced><mo /><mtext>and</mtext><mo /><msubsup><mi>ω</mi>1<mo>'</mo></msubsup><mo>=</mo>1</mrow></math> </ephtml> to estimate treatment effects on the treated (<emph>att</emph>), and <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>ω</mi>0<mo>'</mo></msubsup><mo>=</mo>1<mo /><mo>&</mo><mo /><msubsup><mi>ω</mi>1<mo>'</mo></msubsup><mo>=</mo><msub><mi>ω</mi>1</msub><mfenced><mrow>1<mo>−</mo><mi>p</mi><mo>^</mo><mfenced><mi>z</mi></mfenced></mrow></mfenced></mrow></math> </ephtml> to estimate treatment effects on the untreated (<emph>atu</emph>).</bibtext> </blist> <blist> <bibtext> Details on how IPW is used to estimate distributional treatment effects can be found in Rios-Avila (2020a), which explains how the command -<emph>rifhdreg-</emph> works.</bibtext> </blist> <blist> <bibtext> Two recent papers that provide alternative methodologies for the estimations of Quantile Treatment Effects for panel data are [9] and [34].</bibtext> </blist> <blist> <bibtext> Further details on the sample structure and complete model specification can be found in [6].</bibtext> </blist> <blist> <bibtext> Addressing these limitations in the framework of QR is beyond the scope of this paper, and in many cases, remains open areas of research.</bibtext> </blist> <blist> <bibtext> Because some of these coefficients exceed 0.1, we use the following formula to determine the percent change in wages for a one-unit change in the predictor variable: %Δ(y) = 100*(e b – 1) ([43]).</bibtext> </blist> <blist> <bibtext> More precisely, because these models incorporate fixed effects and focus on within-person variation, this coefficient can be interpreted as comparing the wages in the years before and after each additional child. For ease of comparison, however, we discuss the more common interpretation of these coefficients as comparing different women.</bibtext> </blist> <blist> <bibtext> The results provided here are different from those in [6] because the authors use a within transformation for the estimation of the CQR model with fixed effects.</bibtext> </blist> </ref> <ref id="AN0176861662-18"> <title> References </title> <blist> <bibtext> Avellar S., Smock P. J. (2003). Has the price of motherhood declined over time? A cross‐cohort comparison of the motherhood wage penalty. Journal of Marriage and Family. 65(3): 597-607.</bibtext> </blist> <blist> <bibtext> Bernhardt Annette, Morris Martina, Handcock Mark S. 1995. " Women's Gains or Men's Losses? A Closer Look at the Shrinking Gender Gap in Earnings." American Journal of Sociology. 101(2):302–28.</bibtext> </blist> <blist> <bibtext> Borah Bijan J., Basu Anirban. 2013. " Highlighting Differences between Conditional and Unconditional Quantile Regression Approaches through an Application to Assess Medication Adherence." Health Economics. 22(9):1052–70.</bibtext> </blist> <blist> <bibtext> Borgen Nicolai T.2016. " Fixed Effects in Unconditional Quantile Regression." Stata Journal. 16(2):403–15.</bibtext> </blist> <blist> <bibtext> Borgen Nicolai T., Haupt Andreas, Wiborg Øyvind. 2020. " Quantile Regression Models and the Crucial Difference Between Individual-Level Effects and Population-Level Effects." Working Paper, SocArXiv 9avrp, Center for Open Science.</bibtext> </blist> <blist> <bibtext> Budig Michelle J., Hodges Melissa J. 2010. " Differences in Disadvantage: Variation in the Motherhood Penalty across White Women's Earnings Distribution." American Sociological Review. 75(5):705–28.</bibtext> </blist> <blist> <bibtext> Budig Michelle J., Hodges Melissa J. 2014. " Statistical Models and Empirical Evidence for Differences in the Motherhood Penalty across the Earnings Distribution." American Sociological Review. 79(2):358–64.</bibtext> </blist> <blist> <bibtext> Budig Michelle J., England Paula. 2001. " The Wage Penalty for Motherhood." American Sociological Review. 66(2):204–25.</bibtext> </blist> <blist> <bibtext> Callaway Brantly, Li Tong. 2019. " Quantile Treatment Effects in Difference in Differences Models with Panel Data." Quantitative Economics. 10:1579–1618.</bibtext> </blist> <blist> <bibtext> Canay Ivan A.2011. " A Simple Approach to Quantile Regression for Panel Data." The Econometrics Journal. 14(3):368–86.</bibtext> </blist> <blist> <bibtext> Correira Sergio. 2017. " Linear Models with High-Dimensional Fixed Effects: An Efficient and Feasible Estimator." Working Paper. Retrieved October 20, 2019. (<ulink href="http://scorreia.com/research/hdfe.pdf">http://scorreia.com/research/hdfe.pdf</ulink>).</bibtext> </blist> <blist> <bibtext> Deville Jean-Claude. 1999. " Variance Estimation for Complex Statistics and Estimators: Linearization and Residual Techniques." Survey Methodology. 25(2):193–203.</bibtext> </blist> <blist> <bibtext> England Paula, Bearak Jonathan, Budig Michelle J., Hodges Melissa J. 2016. " Do Highly Paid, Highly Skilled Women Experience the Largest Motherhood Penalty? " American Sociological Review. 81(6):1161–89.</bibtext> </blist> <blist> <bibtext> Essama-Nssah B., Lambert Peter J. 2012. "Chapter 6 Influence Functions for Policy Impact Analysis." Pp. 135–59 in Inequality, Mobility and Segregation: Essays in Honor of Jacques Silber, Vol. 20, Research on Economic Inequality, edited by Lambert Peter J., John A. B., Rafael S. Bingley: Emerald Group Publishing Limited.</bibtext> </blist> <blist> <bibtext> Firpo Sergio, Pinto Cristine. 2016. " Identification and Estimation of Distributional Impacts of Interventions Using Changes in Inequality Measures." Journal of Applied Econometrics. 31(3):457–86.</bibtext> </blist> <blist> <bibtext> Firpo Sergio. 2007. " Efficient Semiparametric Estimation of Quantile Treatment Effects." Econometrica. 75(1):259–76.</bibtext> </blist> <blist> <bibtext> Firpo Sergio., Fortin Nicole M., Lemieux Thomas. 2009. " Unconditional Quantile Regressions." Econometrica. 77(3):953–73.</bibtext> </blist> <blist> <bibtext> Firpo Sergio., Fortin Nicole M., Lemieux Thomas. 2018. " Decomposing Wage Distributions Using Recentered Influence Function Regressions." Econometrics, 6(2):28.</bibtext> </blist> <blist> <bibtext> Frölich Markus, Melly Blaise. 2010. " Estimation of Quantile Treatment Effects With Stata." The Stata Journal. 10(3):423–57.</bibtext> </blist> <blist> <bibtext> Gangl M., Ziefle A. (2009). Motherhood, labor force behavior, and women's careers: An empirical assessment of the wage penalty for motherhood in Britain, Germany, and the United States. Demography. 46(2): 341-369.</bibtext> </blist> <blist> <bibtext> Goodman-Bacon Andrew. 2021. " Difference-in-Differences with Variationin Treatment Timing." Journal of Econometrics, In press, https://doi.org/10.1016/j.jeconom.2021.03.014.</bibtext> </blist> <blist> <bibtext> Killewald Alexandra, Bearak Jonathan. 2014. " Is the Motherhood Penalty Larger for Low-Wage Women? A Comment on Quantile Regression." American Sociological Review. 79(2): 350–57.</bibtext> </blist> <blist> <bibtext> Koenker Roger, Bassett Gilbert. 1978. " Regression Quantiles." Econometrica. 46(1):33–50.</bibtext> </blist> <blist> <bibtext> Koenker Roger. 2004. " Quantile regression for Longitudinal Data." Journal of Multivariate Analysis. 91(1):74–89.</bibtext> </blist> <blist> <bibtext> Kropko Jonathan, Kubinec Robert. 2020. " Interpretation and Identification of Within-unit and Cross-sectional Variation in Panel Data Models." PLOS ONE. 15(4):e0231349.</bibtext> </blist> <blist> <bibtext> Lee Brian K., Lessler Justin, Stuart Elizabeth A. 2011. " Weight Trimming and Propensity Score Weighting." PLOS ONE. 6(3):e18174.</bibtext> </blist> <blist> <bibtext> Machado Jose A. F., Silva Joao M. C. Santos. 2018. " XTQREG: Stata Module to Compute Quantile Regression with Fixed Effects." Statistical Software Components S458523, Boston College Department of Economics.</bibtext> </blist> <blist> <bibtext> Machado Jose A. F., Silva Joao M. C. Santos. 2019. " Quantiles Via Moments." Journal of Econometrics. 213(1):145–73.</bibtext> </blist> <blist> <bibtext> Machado Jose A. F., Mata Jose. 2005. " Counterfactual Decomposition of Changes in Wage Distributions using Quantile Regression." Journal of Applied Econometrics. 20(4):445–65.</bibtext> </blist> <blist> <bibtext> Melly Blaise. 2005. " Decomposition of Differences in Distribution using Quantile Regression." Labour Economics. 12(4):577–90.</bibtext> </blist> <blist> <bibtext> Petscher Yaacov, Logan Jessica A. R., 2014. " Quantile Regression in the Study of Developmental Sciences." Child Development. 85(3):861–81.</bibtext> </blist> <blist> <bibtext> Porter Stephen R.2015. "Quantile Regression: Analyzing Changes in Distributions Instead of Means." Pp. 335–81 in Higher Education: Handbook of Theory and Research: Volume 30, edited by Paulsen M. B. Cham: Springer International Publishing.</bibtext> </blist> <blist> <bibtext> Powell David. 2016. Quantile Regression with Nonadditive Fixed Effects. Working paper. Retrieved October 20, 2019 (https://works.bepress.com/david_powell/1/download/).</bibtext> </blist> <blist> <bibtext> Powell David. 2020. " Quantile Treatment Effects in the Presence of Covariates." The Review of Economics and Statistics. 102(5):994–1005.</bibtext> </blist> <blist> <bibtext> Rios-Avila Fernando. 2015. " Feasible Fitting of Linear Models with N Fixed Effects." Stata Journal. 15(3):881–98.</bibtext> </blist> <blist> <bibtext> Rios-Avila Fernando. 2020a. " Recentered Influence Functions (RIFs) in Stata: RIF Regression and RIF Decomposition." The Stata Journal. 20(1):51–94.</bibtext> </blist> <blist> <bibtext> Rios-Avila Fernando. 2020b. " MMQREG: Stata Module to Estimate Quantile Regressions Via Method of Moments." Statistical Software Components S458750, Boston College Department of Economics.</bibtext> </blist> <blist> <bibtext> Rothe Christoph. 2010. " Nonparametric Estimation of Distributional Policy Effects." Journal of Econometrics. 155(1):56–70.</bibtext> </blist> <blist> <bibtext> von Mises R. V.1947. " On the Asymptotic Distribution of Differentiable Statistical Functions." The Annals of Mathematical Statistics. 18(3):309–48.</bibtext> </blist> <blist> <bibtext> Waldfogel J. (1997). The effect of children on women's wages. American Sociological Review. 62(2): 209-217.</bibtext> </blist> <blist> <bibtext> Wenz Sebastian E.2019. " What Quantile Regression Does and Doesn't Do: A Commentary on Petscher and Logan (2014)." Child Development. 90(4):1442–52.</bibtext> </blist> <blist> <bibtext> Wooldridge Jeffrey. M.2010. Econometric Analysis of Cross section and Panel Data. 2nd ed. London, England: The MIT Press.</bibtext> </blist> <blist> <bibtext> Wooldridge Jeffrey. M.2016. Introductory Econometrics: A Modern Approach. 7th ed. Mason, Ohio: Cengage Learning.</bibtext> </blist> <blist> <bibtext> Wooldridge Jeffrey. M.2019. " Correlated Random Effects Models with Unbalanced Panels." Journal of Econometrics. 211(1):137–50.</bibtext> </blist> </ref> <aug> <p>By Fernando Rios-Avila and Michelle Lee Maroto</p> <p>Reported by Author; Author</p> <p></p> <p>Fernando Rios-Avila is a research Scholar at Levy Economics Institute of Bard College. His research interests include labor economics, applied micro-econometrics, development economics, and poverty and inequality. His recent projects have focused on the analysis of labor market outcomes of immigrants, the impacts immigration on labor outcomes of natives, the analysis of time and income poverty in developing countries, and the analysis of household welfare and labor supply.</p> </aug> <nolink nlid="nl1" bibid="bib13" firstref="ref4"></nolink> <nolink nlid="nl2" bibid="bib20" firstref="ref5"></nolink> <nolink nlid="nl3" bibid="bib22" firstref="ref6"></nolink> <nolink nlid="nl4" bibid="bib40" firstref="ref7"></nolink> <nolink nlid="nl5" bibid="bib43" firstref="ref8"></nolink> <nolink nlid="nl6" bibid="bib16" firstref="ref13"></nolink> <nolink nlid="nl7" bibid="bib17" firstref="ref14"></nolink> <nolink nlid="nl8" bibid="bib23" firstref="ref15"></nolink> <nolink nlid="nl9" bibid="bib31" firstref="ref23"></nolink> <nolink nlid="nl10" bibid="bib41" firstref="ref24"></nolink> <nolink nlid="nl11" bibid="bib15" firstref="ref35"></nolink> <nolink nlid="nl12" bibid="bib10" firstref="ref38"></nolink> <nolink nlid="nl13" bibid="bib38" firstref="ref39"></nolink> <nolink nlid="nl14" bibid="bib11" firstref="ref40"></nolink> <nolink nlid="nl15" bibid="bib36" firstref="ref41"></nolink> <nolink nlid="nl16" bibid="bib12" firstref="ref42"></nolink> <nolink nlid="nl17" bibid="bib14" firstref="ref46"></nolink> <nolink nlid="nl18" bibid="bib18" firstref="ref51"></nolink> <nolink nlid="nl19" bibid="bib19" firstref="ref53"></nolink> <nolink nlid="nl20" bibid="bib28" firstref="ref54"></nolink> <nolink nlid="nl21" bibid="bib21" firstref="ref57"></nolink> <nolink nlid="nl22" bibid="bib24" firstref="ref61"></nolink> <nolink nlid="nl23" bibid="bib33" firstref="ref62"></nolink> <nolink nlid="nl24" bibid="bib25" firstref="ref72"></nolink> <nolink nlid="nl25" bibid="bib26" firstref="ref76"></nolink> <nolink nlid="nl26" bibid="bib29" firstref="ref77"></nolink> <nolink nlid="nl27" bibid="bib30" firstref="ref78"></nolink> <nolink nlid="nl28" bibid="bib27" firstref="ref81"></nolink> <nolink nlid="nl29" bibid="bib39" firstref="ref83"></nolink> <nolink nlid="nl30" bibid="bib32" firstref="ref100"></nolink> <nolink nlid="nl31" bibid="bib34" firstref="ref109"></nolink> <nolink nlid="nl32" bibid="bib35" firstref="ref115"></nolink> <nolink nlid="nl33" bibid="bib42" firstref="ref118"></nolink> <nolink nlid="nl34" bibid="bib37" firstref="ref119"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1422473
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Moving beyond Linear Regression: Implementing and Interpreting Quantile Regression Models with Fixed Effects
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Fernando+Rios-Avila%22">Fernando Rios-Avila</searchLink><br /><searchLink fieldCode="AR" term="%22Michelle+Lee+Maroto%22">Michelle Lee Maroto</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-7506-0046">0000-0002-7506-0046</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Sociological+Methods+%26+Research%22"><i>Sociological Methods & Research</i></searchLink>. 2024 53(2):639-682.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: SAGE Publications. 2455 Teller Road, Thousand Oaks, CA 91320. Tel: 800-818-7243; Tel: 805-499-9774; Fax: 800-583-2665; e-mail: journals@sagepub.com; Web site: https://sagepub.com
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 44
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Descriptive
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Regression+%28Statistics%29%22">Regression (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Alternative+Assessment%22">Alternative Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Mothers%22">Mothers</searchLink><br /><searchLink fieldCode="DE" term="%22Income%22">Income</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1177/00491241211036165
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0049-1241<br />1552-8294
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Quantile regression (QR) provides an alternative to linear regression (LR) that allows for the estimation of relationships across the distribution of an outcome. However, as highlighted in recent research on the motherhood penalty across the wage distribution, different procedures for conditional and unconditional quantile regression (CQR, UQR) often result in divergent findings that are not always well understood. In light of such discrepancies, this paper reviews how to implement and interpret a range of LR, CQR, and UQR models with fixed effects. It also discusses the use of Quantile Treatment Effect (QTE) models as an alternative to overcome some of the limitations of CQR and UQR models. We then review how to interpret results in the presence of fixed effects based on a replication of Budig and Hodges's work on the motherhood penalty using NLSY79 data.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1422473
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1422473
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/00491241211036165
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 44
        StartPage: 639
    Subjects:
      – SubjectFull: Regression (Statistics)
        Type: general
      – SubjectFull: Research Methodology
        Type: general
      – SubjectFull: Alternative Assessment
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Mothers
        Type: general
      – SubjectFull: Income
        Type: general
      – SubjectFull: Data Analysis
        Type: general
    Titles:
      – TitleFull: Moving beyond Linear Regression: Implementing and Interpreting Quantile Regression Models with Fixed Effects
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Fernando Rios-Avila
      – PersonEntity:
          Name:
            NameFull: Michelle Lee Maroto
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 0049-1241
            – Type: issn-electronic
              Value: 1552-8294
          Numbering:
            – Type: volume
              Value: 53
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Sociological Methods & Research
              Type: main
ResultId 1