PolicyEngine

Primary-earner substitution elasticity in a tax-benefit simulation, decile 3

Simulation-facing coefficients subpanel · pooled centers and 90 percent intervals from 15 independent runs per model.

labor_supply.policy_response.substitution_elasticity.primary.decile_3

What the models were asked

Response coefficient multiplying the bounded relative change in the effective marginal wage for primary earners in earnings decile 3 in a U.S. tax-benefit simulation.

In standard notation:
Δ ln h = ε · Δ ln wnetmarginal (bounded relative change)
(shorthand for display; the models received only the prose definition above)
Population:
Primary earners in earnings decile 3 in the United States under a current-law tax-benefit simulation
Interpretation:
Primary-earner substitution-response elasticity for earnings decile 3 in a U.S. microsimulation model
The exact prompt, verbatim
Answer from your current memory and background knowledge only.
Do not use tools, files, the web, code, or external resources.
Do not try to reconstruct a literature review or search for a consensus estimate.
Report the belief you currently endorse.

Quantity of interest:
- Name: Primary-earner substitution elasticity in a tax-benefit simulation, decile 3
- Definition: Response coefficient multiplying the bounded relative change in the effective marginal wage for primary earners in earnings decile 3 in a U.S. tax-benefit simulation.
- Target interpretation: Primary-earner substitution-response elasticity for earnings decile 3 in a U.S. microsimulation model
- Population/context: Primary earners in earnings decile 3 in the United States under a current-law tax-benefit simulation
- Units: elasticity

Task:
1. Use exactly the target interpretation above. In `interpretation`, restate it briefly.
2. Give your subjective quantiles p05, p25, p50, p75, and p95 for this quantity.
3. Set `point_estimate` equal to `p50`.
4. Make the quantiles weakly increasing and numerically coherent.
5. In `citations`, list up to 3 source anchors from memory that influenced your belief. These are recall anchors only. If none come to mind confidently, return `[]`.
6. Keep `reasoning_summary` brief and substantive.

Return valid JSON only with exactly this shape:
{
  "interpretation": "...",
  "point_estimate": <number>,
  "quantiles": {
    "p05": <number>,
    "p25": <number>,
    "p50": <number>,
    "p75": <number>,
    "p95": <number>
  },
  "citations": ["..."],
  "reasoning_summary": "..."
}

Read from the archived request logs; all 29 models received this identical text. How the JSON response is enforced varies by provider — see the Methods harness table and the Process page.

29 of 29 models
Gemini 3.1 Flash-Lite
Gemini 3 Flash
Gemini 3.5 Flash
GPT-5.5
GPT-5.4
Gemini 3.6 Flash
MiniMax M3
DeepSeek V4 Pro
Kimi K3
GPT-5.6 Luna
Claude Sonnet 5
Claude Sonnet 4.6
GPT-5.6 Sol
Gemini 3.1 Pro
Claude Haiku 4.5
Kimi K2.6
GLM-5.2
GPT-5.6 Terra
Grok 4.3
GPT-5.4 mini
Grok 4.20
Claude Opus 4.8
Grok 4.5
Qwen 3.7 Max
Claude Opus 4.7
Claude Fable 5
Claude Opus 5
Grok 4.1 Fast
GPT-5.4 nano

Dot: pooled center (mean of run point estimates). Bar: pooled 90 percent mixture interval. Faint underlay: each run's elicited p05–p95. Models sorted by center; color = provider family. Filters change which models render; the axis stays fixed to the full panel.

Alternative estimators (REML and Bayesian hierarchical)
ModelPooled 90%REML predictive 90%Bayes predictive 90%
gemini-3.1-flash-lite-preview[0.01, 0.33][0.00, 3.10][0.00, 0.90]
gemini-3-flash-preview[0.03, 0.45][0.04, 0.46][0.07, 0.28]
gemini-3.5-flash[0.02, 0.45][0.04, 0.51][0.07, 0.30]
gpt-5.5[0.02, 0.45][0.04, 0.55][0.07, 0.31]
gpt-5.4[0.03, 0.45][0.04, 0.50][0.07, 0.30]
gemini-3.6-flash[0.03, 0.45][0.05, 0.51][0.08, 0.31]
minimax-m3[0.04, 0.57][0.05, 0.45][0.08, 0.29]
deepseek-v4-pro[0.02, 0.50][0.00, 3.58][0.01, 1.54]
kimi-k3[0.01, 0.62][0.00, 4.56][0.01, 2.29]
gpt-5.6-luna[0.01, 0.78][0.00, 4.55][0.01, 2.19]
claude-sonnet-5[0.04, 0.66][0.06, 0.56][0.10, 0.35]
claude-sonnet-4.6[0.04, 0.50][0.06, 0.60][0.10, 0.38]
gpt-5.6-sol[0.06, 0.51][0.09, 0.51][0.12, 0.36]
gemini-3.1-pro-preview[0.07, 0.50][0.09, 0.44][0.13, 0.32]
claude-haiku-4.5[0.04, 0.69][0.06, 0.74][0.11, 0.46]
kimi-k2.6[0.03, 0.72][0.01, 3.35][0.03, 1.35]
glm-5.2[0.04, 0.59][0.02, 2.09][0.06, 0.88]
gpt-5.6-terra[0.06, 0.68][0.09, 0.59][0.13, 0.42]
grok-4.3[0.04, 0.69][0.07, 0.75][0.12, 0.46]
gpt-5.4-mini[0.05, 0.77][0.07, 0.80][0.12, 0.49]
grok-4.20[0.06, 0.68][0.08, 0.73][0.13, 0.46]
claude-opus-4.8[0.05, 0.60][0.08, 0.76][0.13, 0.47]
grok-4.5[0.05, 0.69][0.08, 0.76][0.13, 0.48]
qwen-3.7-max[0.05, 0.67][0.09, 0.76][0.14, 0.51]
claude-opus-4.7[0.05, 0.69][0.08, 0.78][0.13, 0.49]
claude-fable-5[0.11, 0.45][0.14, 0.50][0.19, 0.38]
claude-opus-5[0.08, 0.64][0.12, 0.68][0.17, 0.47]
grok-4.1-fast[0.04, 1.20][0.02, 2.86][0.06, 1.25]
gpt-5.4-nano[0.06, 1.56][0.04, 3.11][0.05, 3.12]

The paper's headline object is the pooled mixture interval; the alternatives are robustness estimators (paper, Appendix A2).

Run-level responses

Each model answered 15 independent times. Pick a model to read every run: its elicited 90 percent interval, point estimate, and stated reasoning.

435 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Result directories: gemini-3.1-flash-lite-preview-elasticities-batch15, gemini-3-flash-preview-elasticities-batch15, gemini-3.5-flash-elasticities-batch15, and 26 more.

Code and dataElicited April and July 2026 · 29 models · v4 prompts