Implied top rates
One fixed optimal-tax calibration applied to every model's elicited taxable-income elasticity. The exercise maps elicited parameters through a disclosed formula; it computes no optimum of its own and takes no position on what the top rate should be.
Each model's pooled ETI median enters the Saez top-rate formula τ* = (1 − ḡ) / (1 − ḡ + a·e) with a threshold-normalized log-utility welfare weight ḡ = 0.618 and a Pareto tail a = 1.621 calibrated from PolicyEngine US microdata. The bar spans the rate implied by each model's own 90 percent ETI band — the models' stated uncertainty, propagated. Among frontier models the medians run from 29.8% (Qwen 3.7 Max) to 38.8% (Kimi K3).
Dot: implied top rate at the pooled ETI median. Bar: rates implied by the model's own 90 percent ETI interval. Sorted lowest to highest; color = provider family. Filters change which models render; the axis stays fixed to the full panel.
Full mapping
The revenue-max column reports the ḡ → 0 Diamond–Saez benchmark 1 / (1 + a·e) at the same ETI median. Paper Table 4 documents the calibration; Appendix Table A13 re-runs the mapping under alternative Pareto tails and CRRA curvature — the cross-model ordering survives every variant.
| Model | ETI median [90%] | Implied top rate [90%] | Revenue-max | 90% width (pp) |
|---|---|---|---|---|
| Qwen 3.7 Max· frontier | 0.555 [0.16, 1.42] | 29.8% [14.3%, 59.4%] | 52.6% | 45.1 |
| GPT-5.4 nano | 0.546 [0.15, 1.46] | 30.1% [13.9%, 60.9%] | 53.1% | 47.0 |
| Claude Haiku 4.5 | 0.502 [0.14, 1.30] | 31.9% [15.3%, 61.9%] | 55.1% | 46.6 |
| Claude Sonnet 4.6 | 0.500 [0.14, 1.28] | 32.0% [15.5%, 63.1%] | 55.2% | 47.6 |
| Grok 4.20 | 0.500 [0.17, 1.24] | 32.0% [15.9%, 57.5%] | 55.2% | 41.6 |
| Kimi K2.6 | 0.499 [0.15, 1.24] | 32.1% [15.9%, 61.7%] | 55.3% | 45.8 |
| GLM-5.2· frontier | 0.495 [0.16, 1.23] | 32.2% [16.1%, 60.1%] | 55.5% | 44.0 |
| GPT-5.6 Terra | 0.492 [0.17, 1.18] | 32.4% [16.6%, 58.7%] | 55.6% | 42.1 |
| DeepSeek V4 Pro· frontier | 0.479 [0.07, 1.81] | 32.9% [11.5%, 76.3%] | 56.3% | 64.7 |
| Claude Sonnet 5 | 0.471 [0.14, 1.15] | 33.3% [17.0%, 61.9%] | 56.7% | 44.9 |
| Grok 4.3 | 0.439 [0.13, 1.09] | 34.9% [17.8%, 64.0%] | 58.4% | 46.2 |
| MiniMax M3· frontier | 0.438 [0.12, 1.33] | 34.9% [15.0%, 65.9%] | 58.5% | 50.9 |
| GPT-5.4 mini | 0.437 [0.12, 1.38] | 35.0% [14.6%, 67.1%] | 58.6% | 52.5 |
| Claude Fable 5· frontier | 0.437 [0.15, 1.08] | 35.0% [17.9%, 60.9%] | 58.6% | 42.9 |
| GPT-5.6 Sol· frontier | 0.431 [0.15, 1.17] | 35.3% [16.8%, 61.1%] | 58.9% | 44.3 |
| Gemini 3.6 Flash· frontier | 0.423 [0.11, 0.90] | 35.8% [20.8%, 67.7%] | 59.3% | 46.9 |
| GPT-5.4 | 0.420 [0.15, 1.00] | 35.9% [19.1%, 61.1%] | 59.5% | 42.0 |
| Grok 4.5· frontier | 0.410 [0.10, 1.13] | 36.5% [17.3%, 69.3%] | 60.1% | 52.0 |
| Claude Opus 4.7 | 0.400 [0.12, 0.99] | 37.0% [19.2%, 65.7%] | 60.7% | 46.5 |
| Gemini 3 Flash | 0.400 [0.13, 1.04] | 37.0% [18.5%, 64.3%] | 60.7% | 45.8 |
| Grok 4.1 Fast | 0.400 [0.11, 1.19] | 37.0% [16.5%, 68.4%] | 60.7% | 51.9 |
| Gemini 3.1 Flash-Lite | 0.389 [0.10, 0.95] | 37.7% [19.8%, 69.9%] | 61.3% | 50.1 |
| Claude Opus 4.8 | 0.383 [0.10, 0.99] | 38.0% [19.2%, 69.1%] | 61.7% | 49.9 |
| GPT-5.6 Luna | 0.377 [0.07, 1.20] | 38.4% [16.4%, 77.5%] | 62.0% | 61.1 |
| Kimi K3· frontier | 0.371 [0.08, 0.99] | 38.8% [19.2%, 74.4%] | 62.4% | 55.2 |
| GPT-5.5 | 0.369 [0.12, 0.99] | 39.0% [19.1%, 65.9%] | 62.6% | 46.8 |
| Gemini 3.5 Flash | 0.357 [0.10, 0.83] | 39.8% [22.0%, 69.8%] | 63.4% | 47.7 |
| Gemini 3.1 Pro | 0.351 [0.10, 0.87] | 40.2% [21.3%, 69.5%] | 63.8% | 48.2 |
| Claude Opus 5 | 0.337 [0.09, 0.96] | 41.2% [19.7%, 73.2%] | 64.7% | 53.5 |
The welfare weight is a normalization choice, not an implication of the elicited data, and the within-model bands run far wider than the cross-model spread in medians: the models report far more parameter uncertainty than their disagreement. Elasticity source: elasticity of taxable income.
11,310 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)