All scorecards
Xyntherium
Grok
ActiveLast updated: Jul 29, 2026
Routed for
✓ General reasoning✓ Documents & images✓ Live URL fetch✓ Deep research
Grokis routed to questions that play to these strengths. Where a task needs a capability it doesn’t have, the question goes to the models that do — and Grok sits that one out.
Across all tasks
| Metric | 7d | 30d | All-time |
|---|---|---|---|
| Response rate | 100% | 100% | 98% |
| p50 latency | 13.8s | 13.2s | 9.3s |
| p95 latency | 13.8s | 20.5s | 20.7s |
| Avg cost / query | 0.1598¢ | 0.1819¢ | 0.1755¢ |
| Agreement w/ verdict | 100% | 82% | 76% |
| Consensus flip rate | 100% | 67% | 76% |
Routed on 9of the last 30 days’ queries it was eligible for, answering 9.
By task type · 30-day
Score = router weight| Task | Responded | p95 | Agreement | Flip | Score |
|---|---|---|---|---|---|
| General | 100% | 20.5s | 82% | 67% | 0.71 |
The score blends agreement with the verified verdict, response rate, and speed over the last 30 days. When a task has more capable models than a panel needs, the router prefers the higher scores — a soft preference, never a hard exclusion.
Rebuilt daily from live verification traffic. Capability flags are set by hand; the performance numbers are earned.
