PolicyEngine

GPT-5.5

OpenAI · elicited July 2026 frontier · average interval-width rank #14.2 of 29 across the canonical panel · $5.42 total elicitation cost.

gpt-5.5

Canonical panel profile

One row per canonical quantity, each on its own scale: this model's pooled center and 90 percent interval, with the 29-model panel median as a gray tick.

Generation harness

From the paper's per-model disclosure (Appendix A16). The prompt and repeated-run design are identical across models; the output mechanism follows each provider's API surface.

Provider path
OpenAI Chat Completions
Output mechanism
strict JSON schema
Completion budget
1200 (8000 for the 40 re-elicited runs)
Sampling
temperature 1.0, batched n <= 8
Reasoning
provider default effort
API identifier
gpt-5.5 (alias)

195 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)

Code and dataElicited April and July 2026 · 29 models · v4 prompts