Transformers, the neural-network architecture behind much of modern AI, combine flexible nonlinear approximation with an attention mechanism that learns how information is combined across inputs. Can neural attention reveal which parts of an agent distribution matter for aggregate dynamics? I represent economic states as tokens—such as wealth bins, capital, debt, and productivity—and embed a transformer in a recursive Krusell–Smith-style solution method. The transformer forecasts the continuation tokenized distribution entering individual Bellman equations. Its neural-network component approximates nonlinear aggregate policy functions, while attention learns how distributional information is aggregated and yields interpretable weights. In Krusell–Smith benchmarks, the transformer rediscovers approximate aggregation, assigning nearly uniform attention across the distribution when aggregate capital is nearly sufficient. In a heterogeneous-firm economy with endogenous default, attention instead concentrates near the default margin and rises during crises for financially constrained firms when their default risk becomes more consequential for aggregate dynamics. Attention therefore serves simultaneously as a solution device and a diagnostic of economic aggregation.
Neural Attention in Heterogeneous-Agent Economies