PolicyEngine

Implied top rates

One fixed optimal-tax calibration applied to every model's elicited taxable-income elasticity. The exercise maps elicited parameters through a disclosed formula; it computes no optimum of its own and takes no position on what the top rate should be.

Each model's pooled ETI median enters the Saez top-rate formula τ* = (1 − ḡ) / (1 − ḡ + a·e) with a threshold-normalized log-utility welfare weight ḡ = 0.618 and a Pareto tail a = 1.621 calibrated from PolicyEngine US microdata. The bar spans the rate implied by each model's own 90 percent ETI band — the models' stated uncertainty, propagated. Among frontier models the medians run from 29.8% (Qwen 3.7 Max) to 38.8% (Kimi K3).

29 of 29 models
Qwen 3.7 Max
GPT-5.4 nano
Claude Haiku 4.5
Claude Sonnet 4.6
Grok 4.20
Kimi K2.6
GLM-5.2
GPT-5.6 Terra
DeepSeek V4 Pro
Claude Sonnet 5
Grok 4.3
MiniMax M3
GPT-5.4 mini
Claude Fable 5
GPT-5.6 Sol
Gemini 3.6 Flash
GPT-5.4
Grok 4.5
Claude Opus 4.7
Gemini 3 Flash
Grok 4.1 Fast
Gemini 3.1 Flash-Lite
Claude Opus 4.8
GPT-5.6 Luna
Kimi K3
GPT-5.5
Gemini 3.5 Flash
Gemini 3.1 Pro
Claude Opus 5

Dot: implied top rate at the pooled ETI median. Bar: rates implied by the model's own 90 percent ETI interval. Sorted lowest to highest; color = provider family. Filters change which models render; the axis stays fixed to the full panel.

Full mapping

The revenue-max column reports the ḡ → 0 Diamond–Saez benchmark 1 / (1 + a·e) at the same ETI median. Paper Table 4 documents the calibration; Appendix Table A13 re-runs the mapping under alternative Pareto tails and CRRA curvature — the cross-model ordering survives every variant.

ModelETI median [90%]Implied top rate [90%]Revenue-max90% width (pp)
Qwen 3.7 Max· frontier0.555 [0.16, 1.42]29.8% [14.3%, 59.4%]52.6%45.1
GPT-5.4 nano0.546 [0.15, 1.46]30.1% [13.9%, 60.9%]53.1%47.0
Claude Haiku 4.50.502 [0.14, 1.30]31.9% [15.3%, 61.9%]55.1%46.6
Claude Sonnet 4.60.500 [0.14, 1.28]32.0% [15.5%, 63.1%]55.2%47.6
Grok 4.200.500 [0.17, 1.24]32.0% [15.9%, 57.5%]55.2%41.6
Kimi K2.60.499 [0.15, 1.24]32.1% [15.9%, 61.7%]55.3%45.8
GLM-5.2· frontier0.495 [0.16, 1.23]32.2% [16.1%, 60.1%]55.5%44.0
GPT-5.6 Terra0.492 [0.17, 1.18]32.4% [16.6%, 58.7%]55.6%42.1
DeepSeek V4 Pro· frontier0.479 [0.07, 1.81]32.9% [11.5%, 76.3%]56.3%64.7
Claude Sonnet 50.471 [0.14, 1.15]33.3% [17.0%, 61.9%]56.7%44.9
Grok 4.30.439 [0.13, 1.09]34.9% [17.8%, 64.0%]58.4%46.2
MiniMax M3· frontier0.438 [0.12, 1.33]34.9% [15.0%, 65.9%]58.5%50.9
GPT-5.4 mini0.437 [0.12, 1.38]35.0% [14.6%, 67.1%]58.6%52.5
Claude Fable 5· frontier0.437 [0.15, 1.08]35.0% [17.9%, 60.9%]58.6%42.9
GPT-5.6 Sol· frontier0.431 [0.15, 1.17]35.3% [16.8%, 61.1%]58.9%44.3
Gemini 3.6 Flash· frontier0.423 [0.11, 0.90]35.8% [20.8%, 67.7%]59.3%46.9
GPT-5.40.420 [0.15, 1.00]35.9% [19.1%, 61.1%]59.5%42.0
Grok 4.5· frontier0.410 [0.10, 1.13]36.5% [17.3%, 69.3%]60.1%52.0
Claude Opus 4.70.400 [0.12, 0.99]37.0% [19.2%, 65.7%]60.7%46.5
Gemini 3 Flash0.400 [0.13, 1.04]37.0% [18.5%, 64.3%]60.7%45.8
Grok 4.1 Fast0.400 [0.11, 1.19]37.0% [16.5%, 68.4%]60.7%51.9
Gemini 3.1 Flash-Lite0.389 [0.10, 0.95]37.7% [19.8%, 69.9%]61.3%50.1
Claude Opus 4.80.383 [0.10, 0.99]38.0% [19.2%, 69.1%]61.7%49.9
GPT-5.6 Luna0.377 [0.07, 1.20]38.4% [16.4%, 77.5%]62.0%61.1
Kimi K3· frontier0.371 [0.08, 0.99]38.8% [19.2%, 74.4%]62.4%55.2
GPT-5.50.369 [0.12, 0.99]39.0% [19.1%, 65.9%]62.6%46.8
Gemini 3.5 Flash0.357 [0.10, 0.83]39.8% [22.0%, 69.8%]63.4%47.7
Gemini 3.1 Pro0.351 [0.10, 0.87]40.2% [21.3%, 69.5%]63.8%48.2
Claude Opus 50.337 [0.09, 0.96]41.2% [19.7%, 73.2%]64.7%53.5

The welfare weight is a normalization choice, not an implication of the elicited data, and the within-model bands run far wider than the cross-model spread in medians: the models report far more parameter uncertainty than their disagreement. Elasticity source: elasticity of taxable income.

11,310 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Code and dataElicited April and July 2026 · 29 models · v4 prompts