PolicyEngine

Generations

How elicited elasticities move as labs ship newer models: same-product-line successor pairs, then every multi-model lab's trajectory. Descriptive only — elicitation wave and harness configuration are confounded with model generation (see Methods), so read shifts as observations, not effects.

Same-product-line successors

Each pair compares a model with its direct successor in the same product line. Ranks run 1 (most elastic, or tightest) to 31; Δ is newer minus older, so positive Δ = the newer model ranks less elastic (or wider).

PairLabor-tax rank (old → new)ΔMacro-trade rank (old → new)ΔWidth rank (old → new)Δ
Claude Opus 4.7Claude Opus 4.815.815.8+0.126.319.8-6.59.810.5+0.7
Claude Opus 4.8Claude Opus 515.816.8+1.019.88.2-11.710.514.5+4.0
Claude Sonnet 4.6Claude Sonnet 58.318.1+9.814.716.8+2.215.88.7-7.2
GPT-5.4GPT-5.515.123.2+8.19.715.8+6.216.39.5-6.8
GPT-5.5GPT-5.6 Sol23.217.8-5.415.814.3-1.59.512.8+3.3
Gemini 3 FlashGemini 3.5 Flash14.826.5+11.721.810.5-11.310.011.7+1.7
Gemini 3.5 FlashGemini 3.6 Flash26.521.2-5.310.517.3+6.811.76.5-5.2
Grok 4.3Grok 4.512.810.0-2.85.811.3+5.517.717.2-0.5
Kimi K2.6Kimi K316.914.6-2.312.017.8+5.820.318.5-1.8
Qwen 3.7 MaxQwen 3.8 Max10.816.6+5.724.025.0+1.023.819.5-4.3

Counted across the 10 pairs: the newer model ranks less elastic on labor-and-tax in 6, more elastic in 4; on macro-and-trade, less elastic in 6, more elastic in 4; and wider in 4, tighter in 6. No adjustment for multiple comparisons — these are counts, not tests.

Lab trajectories

Every lab with more than one model in the panel, ordered by elicitation wave. Single-model labs (DeepSeek, MiniMax, Thinking Machines, Zhipu AI) have no within-lab comparison yet.

OpenAI

ModelWaveLabor-taxMacroWidth
GPT-5.4April 202615.19.716.3
GPT-5.4 miniApril 202615.68.324.3
GPT-5.4 nanoApril 202616.27.719.5
GPT-5.5July 2026 frontier23.215.89.5
GPT-5.6 LunaJuly 2026 GPT-5.619.97.526.8
GPT-5.6 SolJuly 2026 GPT-5.617.814.312.8
GPT-5.6 TerraJuly 2026 GPT-5.612.319.317.3

Anthropic

ModelWaveLabor-taxMacroWidth
Claude Haiku 4.5April 202613.324.511.7
Claude Opus 4.7April 202615.826.39.8
Claude Sonnet 4.6April 20268.314.715.8
Claude Fable 5July 2026 frontier15.521.76.2
Claude Opus 4.8July 2026 frontier15.819.810.5
Claude Sonnet 5July 2026 frontier18.116.88.7
Claude Opus 5July 2026 late16.88.214.5

Google

ModelWaveLabor-taxMacroWidth
Gemini 3 FlashApril 202614.821.810.0
Gemini 3.1 Flash-LiteApril 202618.419.211.3
Gemini 3.1 ProApril 202617.514.07.3
Gemini 3.5 FlashJuly 2026 frontier26.510.511.7
Gemini 3.6 FlashJuly 2026 late21.217.36.5

xAI

ModelWaveLabor-taxMacroWidth
Grok 4.1 FastApril 202618.020.026.3
Grok 4.20April 20269.86.823.8
Grok 4.3July 2026 frontier12.85.817.7
Grok 4.5July 2026 late10.011.317.2

Alibaba

ModelWaveLabor-taxMacroWidth
Qwen 3.7 MaxJuly 2026 independent labs10.824.023.8
Qwen 3.8 MaxAugust 202616.625.019.5

Moonshot AI

ModelWaveLabor-taxMacroWidth
Kimi K2.6July 2026 independent labs16.912.020.3
Kimi K3July 2026 late14.617.818.5

Ranks are averages of within-quantity ranks over each subpanel (labor-and-tax and macro-and-trade), 1 = most elastic; width rank 1 = tightest pooled 90 percent interval. Within-lab comparisons share the lab's serving path in most cases, but elicitation wave, output mechanism, and completion budget still differ across generations — the paper's harness-disclosure and cross-mechanism ablation appendices bound those effects.

12,090 successful runs · elicited April through August 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Code and dataElicited April through August 2026 · 31 models · v4 prompts