Capital gains realizations elasticity (net-of-tax-rate convention)
Simulation-facing coefficients subpanel · pooled centers and 90 percent intervals from 15 independent runs per model.
What the models were asked
Elasticity of long-term capital gains realizations with respect to the net-of-tax rate (1 - tau) on capital gains. A 1 percent increase in the net-of-tax rate changes realizations by this elasticity. Sibling of tax.capital_gains_realizations.elasticity, which uses the opposite w.r.t.-tax-rate convention; the two are related by epsilon_taxrate = -(tau / (1 - tau)) * epsilon_netoftax.
- In standard notation:
- ε1−τ = ∂ ln R / ∂ ln(1−τ) (shorthand for display; the models received only the prose definition above)
- Population:
- Individuals with long-term capital gains in the United States
- Interpretation:
- Medium-run elasticity of long-term capital gains realizations with respect to the net-of-tax rate (1 - tau)
The exact prompt, verbatim
Answer from your current memory and background knowledge only.
Do not use tools, files, the web, code, or external resources.
Do not try to reconstruct a literature review or search for a consensus estimate.
Report the belief you currently endorse.
Quantity of interest:
- Name: Capital gains realizations elasticity (net-of-tax-rate convention)
- Definition: Elasticity of long-term capital gains realizations with respect to the net-of-tax rate (1 - tau) on capital gains. A 1 percent increase in the net-of-tax rate changes realizations by this elasticity. Sibling of tax.capital_gains_realizations.elasticity, which uses the opposite w.r.t.-tax-rate convention; the two are related by epsilon_taxrate = -(tau / (1 - tau)) * epsilon_netoftax.
- Target interpretation: Medium-run elasticity of long-term capital gains realizations with respect to the net-of-tax rate (1 - tau)
- Population/context: Individuals with long-term capital gains in the United States
- Units: elasticity
Sign convention for this quantity:
- An elasticity of ε means that a 1 percent increase in the net-of-tax rate (1 - τ) changes long-term realizations by ε percent; for example, if ε = 0.5, a 1 percent increase in the net-of-tax rate changes long-term realizations by 0.5 percent (not 50 percent).
- ε > 0 if and only if a higher net-of-tax rate raises realizations.
- ε < 0 if and only if a higher net-of-tax rate reduces realizations.
- This is the elasticity with respect to the net-of-tax rate, not with respect to the tax rate τ itself.
Task:
1. Use exactly the target interpretation above. In `interpretation`, restate it briefly.
2. Give your subjective quantiles p05, p25, p50, p75, and p95 for this quantity.
3. Set `point_estimate` equal to `p50`.
4. Make the quantiles weakly increasing and numerically coherent.
5. In `citations`, list up to 3 source anchors from memory that influenced your belief. These are recall anchors only. If none come to mind confidently, return `[]`.
6. Keep `reasoning_summary` brief and substantive.
Return valid JSON only with exactly this shape:
{
"interpretation": "...",
"point_estimate": <number>,
"quantiles": {
"p05": <number>,
"p25": <number>,
"p50": <number>,
"p75": <number>,
"p95": <number>
},
"citations": ["..."],
"reasoning_summary": "..."
}Read from the archived request logs; 22 of 29 models received exactly this text, and the other 7 an earlier v4 wording — every model's prompt is archived verbatim, and the two-wording comparison below shows the four models elicited under both. How the JSON response is enforced varies by provider — see the Methods harness table and the Process page.
Dot: pooled center (mean of run point estimates). Bar: pooled 90 percent mixture interval. Faint underlay: each run's elicited p05–p95. Models sorted by center; color = provider family. Filters change which models render; the axis stays fixed to the full panel.
Same model, two clarifier wordings
The sign clarifier for this quantity was revised two days into the April 2026 wave: plain conditionals with the conventional direction first became symmetric if-and-only-if clauses, and the definition line's conversion identity — stated backwards in the original wording — was corrected. Seven April models keep the original wording (the split disclosed above), while the four April premium models were re-elicited in full under the revision — so those four answered this quantity under both wordings. Their superseded April 19 runs remain in git history and pool to:
| Model | April 19 center (original wording) | April 21 center (revised wording) | Change |
|---|---|---|---|
| Claude Opus 4.7 | 0.70 | 0.70 | 0.000 |
| Claude Sonnet 4.6 | 4.60 | 3.37 | -1.233 |
| Gemini 3.1 Pro | 0.37 | 2.08 | +1.703 |
| Grok 4.20 | 0.67 | 0.65 | -0.020 |
Pooled centers under the paper's piecewise-uniform construction, 15 runs per cell on both sides. The comparison is not a pure wording experiment — the April 21 re-elicitation also moved to the per-quantity harness that added request logging, and two days elapsed — so wording is confounded with harness path and time (paper, Appendix Tables A18–A19). Rerunning the paper's implied-tax-rate convention audit per wording moves Claude Sonnet 4.6 from 0.123 (plausible sign, outside bands) to 0.167 (LTCG-rate consistent) and Gemini 3.1 Pro from 0.545 (ordinary-income-rate consistent) to 0.259 (LTCG-rate consistent), while Claude Opus 4.7 and Grok 4.20 stay in their bands.
Alternative estimators (REML and Bayesian hierarchical)
| Model | Pooled 90% | REML predictive 90% | Bayes predictive 90% |
|---|---|---|---|
| gpt-5.4-nano | [-1.01, 0.92] | [-0.49, 0.39] | [-0.53, 0.45] |
| claude-sonnet-5 | [0.16, 1.38] | [0.11, 1.16] | [0.32, 0.91] |
| gpt-5.6-terra | [0.07, 1.93] | [-0.03, 1.33] | [0.21, 1.00] |
| grok-4.20 | [0.05, 1.95] | [0.01, 1.39] | [0.27, 1.05] |
| claude-opus-4.8 | [0.20, 1.57] | [0.14, 1.23] | [0.35, 0.97] |
| gpt-5.6-luna | [-0.19, 2.76] | [-0.34, 1.99] | [0.05, 1.36] |
| claude-opus-4.7 | [0.13, 1.78] | [0.10, 1.42] | [0.35, 1.09] |
| qwen-3.7-max | [0.11, 2.23] | [0.06, 1.47] | [0.32, 1.12] |
| grok-4.5 | [0.15, 2.79] | [-0.01, 1.54] | [0.28, 1.15] |
| gemini-3.6-flash | [0.11, 1.78] | [0.11, 1.48] | [0.37, 1.14] |
| kimi-k2.6 | [0.07, 3.29] | [-0.15, 1.65] | [0.14, 1.23] |
| grok-4.3 | [0.14, 2.49] | [0.02, 1.71] | [0.32, 1.28] |
| gemini-3.5-flash | [0.11, 1.94] | [0.12, 1.60] | [0.39, 1.23] |
| gemini-3-flash-preview | [0.30, 1.91] | [0.22, 1.50] | [0.46, 1.18] |
| deepseek-v4-pro | [0.11, 3.27] | [-0.03, 1.69] | [0.26, 1.28] |
| kimi-k3 | [0.03, 3.23] | [-0.11, 1.55] | [-0.10, 1.53] |
| gemini-3.1-flash-lite-preview | [0.13, 3.55] | [-0.05, 1.92] | [0.29, 1.43] |
| claude-haiku-4.5 | [0.22, 2.87] | [0.15, 1.65] | [0.41, 1.29] |
| gpt-5.4 | [0.20, 3.19] | [-0.13, 2.27] | [0.28, 1.64] |
| gpt-5.6-sol | [0.15, 3.57] | [-0.15, 1.94] | [0.18, 1.44] |
| gpt-5.4-mini | [0.13, 3.89] | [-0.07, 2.24] | [0.30, 1.69] |
| minimax-m3 | [0.06, 6.14] | [-0.52, 2.08] | [-0.37, 1.81] |
| gpt-5.5 | [0.15, 5.18] | [-0.32, 2.54] | [-0.01, 2.02] |
| glm-5.2 | [0.24, 5.82] | [-0.25, 2.67] | [0.08, 2.12] |
| claude-opus-5 | [0.31, 4.37] | [0.12, 3.23] | [0.64, 2.44] |
| grok-4.1-fast | [0.29, 5.83] | [-1.77, 8.59] | [-1.17, 5.94] |
| claude-fable-5 | [0.42, 4.29] | [0.32, 3.52] | [0.87, 2.72] |
| gemini-3.1-pro-preview | [0.37, 4.97] | [0.16, 3.86] | [0.41, 3.53] |
| claude-sonnet-4.6 | [1.01, 8.65] | [-0.20, 7.34] | [1.05, 5.74] |
The paper's headline object is the pooled mixture interval; the alternatives are robustness estimators (paper, Appendix A2).
Run-level responses
Each model answered 15 independent times. Pick a model to read every run: its elicited 90 percent interval, point estimate, and stated reasoning.
435 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)
Result directories: gpt-5.4-nano-elasticities-batch15, claude-sonnet-5-elasticities-batch15, gpt-5.6-terra-elasticities-batch15, and 26 more.