PolicyEngine

Working paper

How large language models answer questions about economic elasticities

A repeated-elicitation study of prompt-conditioned response distributions: 31 models from ten organizations, 15 runs per model-quantity cell over 26 US-scoped quantities, pooled into predictive distributions and mapped through a fixed optimal-tax calibration. This page embeds the manuscript snapshot that matches the live site — the 31-model panel with capability correlates pinned to PolicyBench release dashboard-data-20260805 — and every number in it is verified against the committed tables by a 131-check prose gate.

31-model panel · 2026-08-08 · Max Ghenis, PolicyEngine

12,090 successful runs · elicited April through August 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Code and dataElicited April through August 2026 · 31 models · v4 prompts