5 Decomposition of Generative AI Exposure: Methods
We next decompose changes in aggregate generative AI exposure to identify the margins through which labor demand changes over time. Our central distinction is between two forms of adjustment. Firms may change where they hire by reallocating postings across different types of jobs. They may also change what jobs contain by revising the task content of comparable jobs. The first margin is hiring reallocation; the second is job redesign.
Our main analysis uses a three-fold extension of the Kitagawa decomposition (Kitagawa, 1955). This decomposition expresses changes in aggregate exposure as the sum of three components: a composition effect, a within-cell exposure effect, and an interaction effect. The composition effect captures changes in the mix of posted jobs. The within-cell exposure effect captures changes in exposure within comparable jobs. The interaction effect captures the additional change that arises when these two adjustments occur simultaneously.
We complement the Kitagawa analysis with a regression-based Oaxaca–Blinder decomposition (Oaxaca, 1973; Blinder, 1973; Oaxaca and Sierminska, 2025). The two decompositions are related but serve different purposes. The Kitagawa decomposition provides a cell-based accounting of aggregate exposure changes. The Oaxaca–Blinder decomposition instead asks which observable job characteristics, such as occupation, industry, seniority, location, remote-work status, internship status, and employment type, are most associated with the exposure gap between the pre- and post-GPT periods. Thus, Kitagawa identifies the adjustment margins, while Oaxaca–Blinder helps characterize the observable dimensions along which compositional change occurs.
5.1 Cell-Level Representation of Aggregate Exposure
Our exposure measure is defined at the posting level. To study aggregate labor-demand adjustment, we aggregate posting-level exposure to job cells. A cell is defined by occupation, seniority, and industry. This cell structure matches the main dimensions of our sampling strategy and captures three sources of heterogeneity central to our research design: the type of work being performed, the sector in which it is performed, and the career stage of the job.
Let $\beta_{p}$ denote the posting-level exposure measure for posting $p$, and let $t$ index time periods. Aggregate exposure in period $t$ is the average exposure across postings in that period:
$$ \bar{E}_{t}=\mathbb{E}[\beta_{p}\mid t]. $$
Because every posting belongs to one mutually exclusive cell $c$, aggregate exposure can be written as a weighted average of cell-level exposure:
$$ \bar{E}_{t}=\sum_{c}\mathbb{E}[\beta_{p}\mid c,t]\Pr(c\mid t). $$
Define $E_{c,t}\equiv\mathbb{E}[\beta_{p}\mid c,t]$ as mean exposure within cell $c$ in period $t$, and $w_{c,t}\equiv\Pr(c\mid t)$ as the share of postings in period $t$ that belong to cell $c$. Then aggregate exposure can be written as
$$ \bar{E}_{t}=\sum_{c}w_{c,t}E_{c,t}. $$
Equation (6) shows that aggregate exposure can change through two basic margins. First, firms may change the mix of jobs they post, which changes the weights $w_{c,t}$. Second, the task content of comparable jobs may change over time, which changes $E_{c,t}$. In our interpretation, the first margin captures hiring reallocation across job cells, while the second captures job redesign within cells.
Let period $0$ denote the baseline period. In our main implementation, period $0$ is the full year 2021. We use year 2021 as the baseline because it precedes the broad public diffusion of generative AI tools and provides a stable pre-GPT benchmark. Our goal is to decompose the change in aggregate exposure between period $0$ and period $t$:
$$ \Delta\bar{E}_{t}=\bar{E}_{t}-\bar{E}_{0}. $$
5.2 Three-Fold Kitagawa Decomposition
We decompose the change in aggregate exposure into three counterfactual components:
$$ \Delta\bar{E}_{t}=\underbrace{\sum_{c}(w_{c,t}-w_{c,0})E_{c,0}}_{\text{Composition effect}}+\underbrace{\sum_{c}w_{c,0}(E_{c,t}-E_{c,0})}_{\text{Within-cell exposure effect}}+\underbrace{\sum_{c}(w_{c,t}-w_{c,0})(E_{c,t}-E_{c,0})}_{\text{Interaction effect}}.\vskip 9.0pt $$
The first term is the composition effect. It measures how aggregate exposure would change if posting shares shifted from their 2021 distribution to their period-$t$ distribution, while exposure within each cell remained fixed at its 2021 level. Substantively, this term captures hiring reallocation across job cells. A positive value means that postings are shifting toward cells that were more exposed to generative AI in the baseline period. A negative value means that postings are shifting away from those more exposed cells.
The second term is the within-cell exposure effect. It measures how aggregate exposure would change if cell-level exposure shifted from its 2021 level to its period-$t$ level, while posting shares remained fixed at their 2021 distribution. Substantively, this term captures job redesign within comparable jobs. A negative value indicates that the tasks described within a cell have become less exposed to generative AI over time, holding the baseline distribution of postings fixed. A positive value indicates movement toward more exposed task content within cells.
The third term is the interaction effect. It captures the additional change that arises because posting shares and within-cell exposure move at the same time. This term indicates whether hiring reallocation and job redesign reinforce or offset each other. For example, the interaction term is negative when postings shift toward cells whose exposure is falling, or away from cells whose exposure is rising. It is positive when the two movements push aggregate exposure in the same direction.
This three-fold formulation is useful in our setting because it separately identifies changes in where firms hire, changes in what comparable jobs contain, and the joint movement of the two. The classic two-fold Kitagawa decomposition allocates the interaction term into the composition and within-cell components. We report the corresponding two-fold decomposition in Appendix G.
5.3 Common Support and Robustness of the Decomposition
The set of observed cells can vary over time. Some occupation-by-seniority-by-industry cells may appear in the baseline period but not in a later period, or vice versa. To ensure that the decomposition compares comparable cells, we define the period-specific common support between the baseline year and period $t$ as
$$ S_{t}=\{c:w_{c,0}>0\text{ and }w_{c,t}>0\}. $$
Cells in $S_{t}$ are observed in both 2021 and period $t$. We re-normalize posting shares within this common-support sample before applying the decomposition. This approach focuses the analysis on changes among persistent job cells, rather than mechanically attributing changes to cell entry or exit. Because the re-normalized common-support aggregate is not identical to the raw aggregate, Appendix F reports overlap diagnostics, reconstruction checks, and the gap between the raw and re-normalized aggregate series. These diagnostics show that the common-support series closely tracks the raw aggregate and that the reconstruction gap is negligible.
Our main text reports the three-fold decomposition in Equation (8). Appendix G reports two additional exercises. First, we implement the symmetric two-fold Kitagawa decomposition, which separates aggregate change into composition and within-cell components by allocating the interaction term evenly across the two. Second, we repeat the analysis using a balanced-cell sample restricted to cells observed throughout the full sample period. These analyses confirm that the main patterns are not driven by the treatment of the interaction term or by changes in cell support over time.
5.4 Oaxaca–Blinder Decomposition
The Kitagawa decomposition separates aggregate exposure changes into reallocation, redesign, and interaction components. It does not, however, identify which observable job characteristics are most associated with the exposure gap. For example, a large composition effect may reflect shifts across occupations, industries, seniority groups, locations, remote-work arrangements, or employment types. To examine these observable dimensions, we complement the Kitagawa analysis with a regression-based Oaxaca–Blinder decomposition (Oaxaca, 1973; Blinder, 1973; Oaxaca and Sierminska, 2025).
The Oaxaca–Blinder decomposition compares two aggregate periods: the pre-GPT period and the post-GPT period. It decomposes the change in mean exposure into an explained component and an unexplained component. The explained component captures the portion associated with changes in observable job characteristics. The unexplained component captures the remaining portion associated with changes in the relationship between those characteristics and exposure, as well as unobserved factors.
We implement the decomposition at the cell level for computational tractability. Each cell is defined by a unique combination of occupation, seniority, and industry. The regression is weighted by the number of postings in each cell-period so that the estimates reproduce posting-level moments. The estimating equation is
$$ Y_{c,g}=\alpha_{g}+X_{c,g}^{\prime}\beta_{g}+\varepsilon_{c,g}, $$
where $g\in\{A,B\}$ indexes the pre- and post-GPT periods, $Y_{c,g}$ is mean exposure in cell $c$ and period group $g$, and $X_{c,g}$ is a vector of observable job characteristics. The covariates are organized into blocks: occupation, industry, seniority, state, remote-work arrangement, internship status, and employment type. The occupation block includes 923 O*NET–SOC categories observed in our sample; the industry block includes 20 two-digit NAICS sectors; the seniority block includes junior, intermediate, and senior postings; the state block includes 51 state categories; the remote-work block includes remote, hybrid, and not-remote postings; the internship block distinguishes intern from non-intern postings; and the employment-type block includes full-time, part-time, and part-time/full-time postings. After omitting one reference category per block, the specification includes approximately 1,000 regressors.
Using the pre-GPT period as the reference structure, the two-fold decomposition is
$$ \bar{Y}_{B}-\bar{Y}_{A}=\underbrace{(\bar{X}_{B}-\bar{X}_{A})^{\prime}\hat{\beta}_{A}}_{\text{Explained}}+\underbrace{(\hat{\alpha}_{B}-\hat{\alpha}_{A})+\bar{X}_{B}^{\prime}(\hat{\beta}_{B}-\hat{\beta}_{A})}_{\text{Unexplained}}.\vskip 9.0pt $$
The explained component answers a counterfactual question: how much would average exposure change if the distribution of observable job characteristics shifted from the pre-GPT period to the post-GPT period, while the association between those characteristics and exposure remained fixed at its pre-GPT level? It therefore captures the part of the exposure gap attributable to shifts in observed job characteristics.
This component is conceptually related to the Kitagawa composition effect because both involve changes in composition. However, the two objects are not identical in our implementation. The Kitagawa composition effect is a cell-based accounting term computed over the full set of occupation-by-industry-by-seniority cells. The Oaxaca–Blinder explained component is a regression-based decomposition that enters observable characteristics as additive covariate blocks. If the Oaxaca–Blinder specification were fully saturated with indicators for the exact same cells used in the Kitagawa decomposition, the explained component would closely correspond to the Kitagawa composition effect (Oaxaca and Sierminska, 2025). Our implementation instead uses the Oaxaca–Blinder decomposition as a complementary diagnostic tool: it identifies which observable characteristics account for the exposure change associated with compositional differences.
The unexplained component captures the remaining exposure gap after accounting for changes in observed job characteristics. Formally, it reflects changes in the intercept and in the estimated coefficients between the pre- and post-GPT periods. We do not interpret this component as a direct analog of the Kitagawa within-cell effect. Its magnitude depends on the choice of reference group, the included covariates, and the functional form of the regression (Fortin et al., 2011). We therefore interpret it as the portion of the exposure gap not accounted for by shifts in observable job characteristics, rather than as a separate measure of job redesign.
We report detailed contributions only for the explained component and only at the level of covariate blocks. This choice follows standard practice in decomposition analysis. For categorical variables, the total contribution of a block, such as occupation or industry, is invariant to the choice of omitted reference category. In contrast, contributions of individual categories within a block depend on which category is omitted; changing the reference category can redistribute contributions across categories without changing the block total (Fortin et al., 2011). We therefore report block-level explained contributions, such as the total contribution of occupation, industry, or seniority. For the unexplained component, category-level contributions are even more sensitive to the choice of reference group because they measure coefficient changes relative to the omitted category. We therefore interpret the unexplained component only in the aggregate.