Generations
How elicited elasticities move as labs ship newer models: same-product-line successor pairs, then every multi-model lab's trajectory. Descriptive only — elicitation wave and harness configuration are confounded with model generation (see Methods), so read shifts as observations, not effects.
Same-product-line successors
Each pair compares a model with its direct successor in the same product line. Ranks run 1 (most elastic, or tightest) to 29; Δ is newer minus older, so positive Δ = the newer model ranks less elastic (or wider).
| Pair | Labor-tax rank (old → new) | Δ | Macro-trade rank (old → new) | Δ | Width rank (old → new) | Δ |
|---|---|---|---|---|---|---|
| Claude Opus 4.7 → Claude Opus 4.8 | 14.6 → 14.7 | +0.1 | 25.0 → 18.8 | -6.2 | 9.2 → 9.8 | +0.7 |
| Claude Opus 4.8 → Claude Opus 5 | 14.7 → 15.8 | +1.2 | 18.8 → 7.8 | -11.0 | 9.8 → 13.8 | +4.0 |
| Claude Sonnet 4.6 → Claude Sonnet 5 | 7.8 → 17.1 | +9.3 | 14.3 → 15.8 | +1.5 | 14.7 → 8.2 | -6.5 |
| GPT-5.4 → GPT-5.5 | 13.9 → 21.7 | +7.8 | 9.3 → 15.5 | +6.2 | 15.3 → 8.8 | -6.5 |
| GPT-5.5 → GPT-5.6 Sol | 21.7 → 16.6 | -5.1 | 15.5 → 13.7 | -1.8 | 8.8 → 11.8 | +3.0 |
| Gemini 3 Flash → Gemini 3.5 Flash | 13.8 → 24.7 | +10.9 | 20.8 → 10.5 | -10.3 | 9.7 → 11.2 | +1.5 |
| Gemini 3.5 Flash → Gemini 3.6 Flash | 24.7 → 19.8 | -4.8 | 10.5 → 16.3 | +5.8 | 11.2 → 6.0 | -5.2 |
| Grok 4.3 → Grok 4.5 | 12.1 → 9.5 | -2.6 | 5.8 → 10.7 | +4.8 | 16.7 → 16.0 | -0.7 |
| Kimi K2.6 → Kimi K3 | 15.8 → 13.4 | -2.3 | 11.7 → 16.8 | +5.2 | 19.3 → 17.5 | -1.8 |
Counted across the 9 pairs: the newer model ranks less elastic on labor-and-tax in 5, more elastic in 4; on macro-and-trade, less elastic in 5, more elastic in 4; and wider in 4, tighter in 5. No adjustment for multiple comparisons — these are counts, not tests.
Lab trajectories
Every lab with more than one model in the panel, ordered by elicitation wave. Single-model labs (Alibaba, DeepSeek, MiniMax, Zhipu AI) have no within-lab comparison yet.
OpenAI
| Model | Wave | Labor-tax | Macro | Width |
|---|---|---|---|---|
| GPT-5.4 | April 2026 | 13.9 | 9.3 | 15.3 |
| GPT-5.4 mini | April 2026 | 14.4 | 8.0 | 23.2 |
| GPT-5.4 nano | April 2026 | 15.3 | 7.3 | 18.3 |
| GPT-5.5 | July 2026 frontier | 21.7 | 15.5 | 8.8 |
| GPT-5.6 Luna | July 2026 GPT-5.6 | 18.8 | 7.2 | 25.2 |
| GPT-5.6 Sol | July 2026 GPT-5.6 | 16.6 | 13.7 | 11.8 |
| GPT-5.6 Terra | July 2026 GPT-5.6 | 11.5 | 18.3 | 16.5 |
Anthropic
| Model | Wave | Labor-tax | Macro | Width |
|---|---|---|---|---|
| Claude Haiku 4.5 | April 2026 | 12.5 | 23.2 | 11.0 |
| Claude Opus 4.7 | April 2026 | 14.6 | 25.0 | 9.2 |
| Claude Sonnet 4.6 | April 2026 | 7.8 | 14.3 | 14.7 |
| Claude Fable 5 | July 2026 frontier | 14.4 | 20.7 | 5.8 |
| Claude Opus 4.8 | July 2026 frontier | 14.7 | 18.8 | 9.8 |
| Claude Sonnet 5 | July 2026 frontier | 17.1 | 15.8 | 8.2 |
| Claude Opus 5 | July 2026 late | 15.8 | 7.8 | 13.8 |
| Model | Wave | Labor-tax | Macro | Width |
|---|---|---|---|---|
| Gemini 3 Flash | April 2026 | 13.8 | 20.8 | 9.7 |
| Gemini 3.1 Flash-Lite | April 2026 | 17.1 | 18.2 | 10.7 |
| Gemini 3.1 Pro | April 2026 | 16.3 | 13.3 | 6.8 |
| Gemini 3.5 Flash | July 2026 frontier | 24.7 | 10.5 | 11.2 |
| Gemini 3.6 Flash | July 2026 late | 19.8 | 16.3 | 6.0 |
xAI
| Model | Wave | Labor-tax | Macro | Width |
|---|---|---|---|---|
| Grok 4.1 Fast | April 2026 | 16.7 | 19.3 | 24.5 |
| Grok 4.20 | April 2026 | 9.5 | 6.8 | 22.3 |
| Grok 4.3 | July 2026 frontier | 12.1 | 5.8 | 16.7 |
| Grok 4.5 | July 2026 late | 9.5 | 10.7 | 16.0 |
Ranks are averages of within-quantity ranks over each subpanel (labor-and-tax and macro-and-trade), 1 = most elastic; width rank 1 = tightest pooled 90 percent interval. Within-lab comparisons share the lab's serving path in most cases, but elicitation wave, output mechanism, and completion budget still differ across generations — the paper's harness-disclosure and cross-mechanism ablation appendices bound those effects.
11,310 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)