PolicyEngine

Capital gains realizations elasticity

Labor and tax subpanel · pooled centers and 90 percent intervals from 15 independent runs per model · shaded region marks the review range [-1, -0.2].

tax.capital_gains_realizations.elasticity

What the models were asked

Elasticity of long-term capital gains realizations with respect to the capital gains marginal tax rate in a U.S. tax model. A 1 percent increase in the capital-gains marginal tax rate changes realizations by this elasticity.

In standard notation:
ετ = ∂ ln R / ∂ ln τ
(shorthand for display; the models received only the prose definition above)
Population:
Individuals with long-term capital gains in the United States
Interpretation:
Medium-run elasticity of long-term capital gains realizations with respect to the capital-gains marginal tax rate
The exact prompt, verbatim
Answer from your current memory and background knowledge only.
Do not use tools, files, the web, code, or external resources.
Do not try to reconstruct a literature review or search for a consensus estimate.
Report the belief you currently endorse.

Quantity of interest:
- Name: Capital gains realizations elasticity
- Definition: Elasticity of long-term capital gains realizations with respect to the capital gains marginal tax rate in a U.S. tax model. A 1 percent increase in the capital-gains marginal tax rate changes realizations by this elasticity.
- Target interpretation: Medium-run elasticity of long-term capital gains realizations with respect to the capital-gains marginal tax rate
- Population/context: Individuals with long-term capital gains in the United States
- Units: elasticity

Sign convention for this quantity:
- An elasticity of ε means that a 1 percent increase in the marginal capital-gains tax rate changes long-term realizations by ε percent; for example, if ε = 0.5, a 1 percent increase in the marginal capital-gains tax rate changes long-term realizations by 0.5 percent (not 50 percent).
- ε > 0 if and only if a higher capital-gains tax rate raises realizations.
- ε < 0 if and only if a higher capital-gains tax rate reduces realizations.
- This is the elasticity with respect to the tax rate itself, not with respect to the net-of-tax rate (1 - τ).

Task:
1. Use exactly the target interpretation above. In `interpretation`, restate it briefly.
2. Give your subjective quantiles p05, p25, p50, p75, and p95 for this quantity.
3. Set `point_estimate` equal to `p50`.
4. Make the quantiles weakly increasing and numerically coherent.
5. In `citations`, list up to 3 source anchors from memory that influenced your belief. These are recall anchors only. If none come to mind confidently, return `[]`.
6. Keep `reasoning_summary` brief and substantive.

Return valid JSON only with exactly this shape:
{
  "interpretation": "...",
  "point_estimate": <number>,
  "quantiles": {
    "p05": <number>,
    "p25": <number>,
    "p50": <number>,
    "p75": <number>,
    "p95": <number>
  },
  "citations": ["..."],
  "reasoning_summary": "..."
}

Read from the archived request logs; 22 of 29 models received exactly this text, and the other 7 an earlier v4 wording — every model's prompt is archived verbatim, and the two-wording comparison below shows the four models elicited under both. How the JSON response is enforced varies by provider — see the Methods harness table and the Process page.

29 of 29 models
GPT-5.4 mini
Gemini 3 Flash
GPT-5.4
Claude Haiku 4.5
Grok 4.1 Fast
Gemini 3.1 Flash-Lite
Claude Opus 5
Grok 4.3
Qwen 3.7 Max
Gemini 3.1 Pro
Claude Fable 5
Claude Opus 4.7
Claude Sonnet 4.6
Claude Opus 4.8
GPT-5.5
Gemini 3.6 Flash
GLM-5.2
Grok 4.5
Claude Sonnet 5
GPT-5.6 Sol
GPT-5.6 Luna
Grok 4.20
Kimi K3
Kimi K2.6
GPT-5.6 Terra
DeepSeek V4 Pro
GPT-5.4 nano
MiniMax M3
Gemini 3.5 Flash

Dot: pooled center (mean of run point estimates). Bar: pooled 90 percent mixture interval. Faint underlay: each run's elicited p05–p95. Models sorted by center; color = provider family. Filters change which models render; the axis stays fixed to the full panel.

Review-range sources: Dowd, McClelland, and Muthitacharoen 2015; Burman and Randolph 1994; CBO/JCT medium-run convention. These are hand-coded literature anchors, not benchmark truths.

Same model, two clarifier wordings

The sign clarifier for this quantity was revised two days into the April 2026 wave: plain conditionals with the conventional direction first became symmetric if-and-only-if clauses. Seven April models keep the original wording (the split disclosed above), while the four April premium models were re-elicited in full under the revision — so those four answered this quantity under both wordings. Their superseded April 19 runs remain in git history and pool to:

ModelApril 19 center (original wording)April 21 center (revised wording)Change
Claude Opus 4.7-0.70-0.700.000
Claude Sonnet 4.6-0.70-0.700.000
Gemini 3.1 Pro-0.67-0.71-0.039
Grok 4.20-0.55-0.47+0.080

Pooled centers under the paper's piecewise-uniform construction, 15 runs per cell on both sides. The comparison is not a pure wording experiment — the April 21 re-elicitation also moved to the per-quantity harness that added request logging, and two days elapsed — so wording is confounded with harness path and time (paper, Appendix Tables A18–A19). Rerunning the paper's implied-tax-rate convention audit per wording moves Claude Sonnet 4.6 from 0.123 (plausible sign, outside bands) to 0.167 (LTCG-rate consistent) and Gemini 3.1 Pro from 0.545 (ordinary-income-rate consistent) to 0.259 (LTCG-rate consistent), while Claude Opus 4.7 and Grok 4.20 stay in their bands.

Alternative estimators (REML and Bayesian hierarchical)
ModelPooled 90%REML predictive 90%Bayes predictive 90%
gpt-5.4-mini[-3.43, 0.14][-2.02, 0.07][-1.53, -0.26]
gemini-3-flash-preview[-1.49, -0.20][-1.38, -0.27][-1.11, -0.49]
gpt-5.4[-2.67, 0.16][-1.91, 0.12][-1.38, -0.23]
claude-haiku-4.5[-1.50, -0.13][-1.47, -0.19][-1.15, -0.44]
grok-4.1-fast[-2.77, 0.09][-1.74, 0.08][-1.27, -0.24]
gemini-3.1-flash-lite-preview[-2.18, -0.04][-1.57, -0.05][-1.19, -0.33]
claude-opus-5[-1.74, -0.21][-1.38, -0.18][-1.08, -0.41]
grok-4.3[-2.32, -0.10][-1.67, 0.05][-1.23, -0.26]
qwen-3.7-max[-2.15, -0.10][-1.47, -0.07][-1.12, -0.33]
gemini-3.1-pro-preview[-1.20, -0.20][-1.21, -0.27][-0.98, -0.45]
claude-fable-5[-1.57, -0.25][-1.26, -0.22][-1.00, -0.42]
claude-opus-4.7[-1.73, -0.20][-1.38, -0.13][-1.07, -0.37]
claude-sonnet-4.6[-1.50, -0.10][-1.45, -0.07][-1.10, -0.33]
claude-opus-4.8[-1.57, -0.20][-1.34, -0.15][-1.05, -0.38]
gpt-5.5[-1.58, -0.02][-1.33, -0.05][-1.02, -0.28]
gemini-3.6-flash[-1.46, -0.11][-1.21, -0.14][-0.95, -0.35]
glm-5.2[-3.64, 0.15][-1.15, 0.37][-1.04, 0.29]
grok-4.5[-1.76, -0.05][-1.39, 0.00][-1.04, -0.26]
claude-sonnet-5[-1.40, -0.08][-1.25, -0.07][-0.96, -0.30]
gpt-5.6-sol[-1.45, -0.08][-1.11, -0.08][-0.89, -0.26]
gpt-5.6-luna[-3.50, 0.33][-1.72, 0.49][-1.16, 0.16]
grok-4.20[-1.65, 0.16][-1.19, 0.13][-0.86, -0.12]
kimi-k3[-1.27, -0.01][-0.88, 0.08][-0.78, -0.02]
kimi-k2.6[-1.50, 0.70][-1.29, 0.43][-1.26, 0.38]
gpt-5.6-terra[-1.36, 0.07][-0.86, 0.10][-0.66, -0.05]
deepseek-v4-pro[-1.45, -0.01][-0.69, 0.17][-0.69, 0.16]
gpt-5.4-nano[-0.98, 0.01][-0.67, -0.07][-0.55, -0.17]
minimax-m3[-1.56, 0.11][-0.59, 0.23][-0.60, 0.20]
gemini-3.5-flash[-1.14, 1.30][-1.97, 1.29][-2.16, 1.34]

The paper's headline object is the pooled mixture interval; the alternatives are robustness estimators (paper, Appendix A2).

Run-level responses

Each model answered 15 independent times. Pick a model to read every run: its elicited 90 percent interval, point estimate, and stated reasoning.

435 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Result directories: gpt-5.4-mini-elasticities-batch15, gemini-3-flash-preview-elasticities-batch15, gpt-5.4-elasticities-batch15, and 26 more.

Code and dataElicited April and July 2026 · 29 models · v4 prompts