Xyntherium
Model Scorecards
A permanent record of how each model performs on the panel — how often it answers, how fast, how closely it tracks the verified consensus, and where it’s strongest. Every number is earned on live verification traffic.
Last updated: Sep 12, 2026
ClaudeActive
- Responded
- 100%
- p95
- 19.2s
- Agreement
- 74%
GPTActive
- Responded
- 100%
- p95
- 15.0s
- Agreement
- 73%
GrokActive
- Responded
- 100%
- p95
- 9.0s
- Agreement
- 71%
PerplexityActive
- Responded
- 100%
- p95
- 17.2s
- Agreement
- 75%
GeminiActive
- Responded
- 100%
- p95
- 10.4s
- Agreement
- 96%
DeepSeekActive
- Responded
- 100%
- p95
- 27.0s
- Agreement
- 71%
KimiBenched
Gathering data — the first numbers appear after the next daily refresh.
View scorecard →Scorecards are rebuilt daily from live verification traffic. The same per-task scores shown here are what the router weighs when it has more capable models than a panel needs — so the strongest model for a task is the one most likely to answer it.
