PolicyEngine

Generations

How elicited elasticities move as labs ship newer models: same-product-line successor pairs, then every multi-model lab's trajectory. Descriptive only — elicitation wave and harness configuration are confounded with model generation (see Methods), so read shifts as observations, not effects.

Same-product-line successors

Each pair compares a model with its direct successor in the same product line. Ranks run 1 (most elastic, or tightest) to 29; Δ is newer minus older, so positive Δ = the newer model ranks less elastic (or wider).

PairLabor-tax rank (old → new)ΔMacro-trade rank (old → new)ΔWidth rank (old → new)Δ
Claude Opus 4.7Claude Opus 4.814.614.7+0.125.018.8-6.29.29.8+0.7
Claude Opus 4.8Claude Opus 514.715.8+1.218.87.8-11.09.813.8+4.0
Claude Sonnet 4.6Claude Sonnet 57.817.1+9.314.315.8+1.514.78.2-6.5
GPT-5.4GPT-5.513.921.7+7.89.315.5+6.215.38.8-6.5
GPT-5.5GPT-5.6 Sol21.716.6-5.115.513.7-1.88.811.8+3.0
Gemini 3 FlashGemini 3.5 Flash13.824.7+10.920.810.5-10.39.711.2+1.5
Gemini 3.5 FlashGemini 3.6 Flash24.719.8-4.810.516.3+5.811.26.0-5.2
Grok 4.3Grok 4.512.19.5-2.65.810.7+4.816.716.0-0.7
Kimi K2.6Kimi K315.813.4-2.311.716.8+5.219.317.5-1.8

Counted across the 9 pairs: the newer model ranks less elastic on labor-and-tax in 5, more elastic in 4; on macro-and-trade, less elastic in 5, more elastic in 4; and wider in 4, tighter in 5. No adjustment for multiple comparisons — these are counts, not tests.

Lab trajectories

Every lab with more than one model in the panel, ordered by elicitation wave. Single-model labs (Alibaba, DeepSeek, MiniMax, Zhipu AI) have no within-lab comparison yet.

OpenAI

ModelWaveLabor-taxMacroWidth
GPT-5.4April 202613.99.315.3
GPT-5.4 miniApril 202614.48.023.2
GPT-5.4 nanoApril 202615.37.318.3
GPT-5.5July 2026 frontier21.715.58.8
GPT-5.6 LunaJuly 2026 GPT-5.618.87.225.2
GPT-5.6 SolJuly 2026 GPT-5.616.613.711.8
GPT-5.6 TerraJuly 2026 GPT-5.611.518.316.5

Anthropic

ModelWaveLabor-taxMacroWidth
Claude Haiku 4.5April 202612.523.211.0
Claude Opus 4.7April 202614.625.09.2
Claude Sonnet 4.6April 20267.814.314.7
Claude Fable 5July 2026 frontier14.420.75.8
Claude Opus 4.8July 2026 frontier14.718.89.8
Claude Sonnet 5July 2026 frontier17.115.88.2
Claude Opus 5July 2026 late15.87.813.8

Google

ModelWaveLabor-taxMacroWidth
Gemini 3 FlashApril 202613.820.89.7
Gemini 3.1 Flash-LiteApril 202617.118.210.7
Gemini 3.1 ProApril 202616.313.36.8
Gemini 3.5 FlashJuly 2026 frontier24.710.511.2
Gemini 3.6 FlashJuly 2026 late19.816.36.0

xAI

ModelWaveLabor-taxMacroWidth
Grok 4.1 FastApril 202616.719.324.5
Grok 4.20April 20269.56.822.3
Grok 4.3July 2026 frontier12.15.816.7
Grok 4.5July 2026 late9.510.716.0

Moonshot AI

ModelWaveLabor-taxMacroWidth
Kimi K2.6July 2026 independent labs15.811.719.3
Kimi K3July 2026 late13.416.817.5

Ranks are averages of within-quantity ranks over each subpanel (labor-and-tax and macro-and-trade), 1 = most elastic; width rank 1 = tightest pooled 90 percent interval. Within-lab comparisons share the lab's serving path in most cases, but elicitation wave, output mechanism, and completion budget still differ across generations — the paper's harness-disclosure and cross-mechanism ablation appendices bound those effects.

11,310 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Code and dataElicited April and July 2026 · 29 models · v4 prompts