Kimi K2.6
Moonshot AI · elicited July 2026 independent labs · average interval-width rank #21.3 of 29 across the canonical panel · — (cost untracked).
Canonical panel profile
One row per canonical quantity, each on its own scale: this model's pooled center and 90 percent interval, with the 29-model panel median as a gray tick.
Generation harness
From the paper's per-model disclosure (Appendix A16). The prompt and repeated-run design are identical across models; the output mechanism follows each provider's API surface.
- Provider path
- LiteLLM via OpenRouter
- Output mechanism
- forced JSON object (schema validated locally)
- Completion budget
- 8000
- Sampling
- temperature 1.0
- Reasoning
- provider default
- API identifier
- openrouter/moonshotai/kimi-k2.6 (alias)
195 successful runs · elicited April and July 2026 · v4 prompts · 15 runs per model-quantity cell. Code · Raw responses · Paper (PDF)
Code and dataElicited April and July 2026 · 29 models · v4 prompts