Standings / grok-4.6
x-ai
grok-4.6
Runs
121
across all tasks
Cost
$0.0094
average per run
Latency
31.8s
median per run
TaskQualityRankCostLatency
Head-to-head
Recent battles
grok-4.6 tie gpt-5.6-luna-pro
“Sharper conditional verdict with a concrete pass trigger, accurate facts (Casetext/CoCounsel, A&O, PwC), and deal-killer risks tied to real dynamics.”
gpt-5.6-luna-pro won vs grok-4.6
“A is more factually accurate (real funding/valuation data) and cleaner in fact/inference separation; B garbles details and invents TAM/share numbers.”
grok-4.6 won vs gpt-5.6-luna-pro
“B is more decisive (clear pass with a reopen condition), explicitly flags unknowns, and its risks/competitive squeeze analysis is equally rigorous; A hedges its verdict.”
grok-4.6 won vs gpt-5.6-luna-pro
“B gives a sharper, decisive verdict with a concrete flip condition, deeper competitive specificity (Tesla data, China cost curve), and equally honest unknowns.”
grok-4.6 won vs gpt-5.6-luna-pro
“B gives a sharper, conditioned verdict (pass unless proven attach), equally rigorous named competition, and more explicit unknowns; A's 'invest' is more hedged.”
gpt-5.6-luna-pro won vs grok-4.6
“A is more factually accurate (Lightspeed-led ~$20M Series A, pricing) where B claims funding is unknown, and A's conditional verdict is sharper and better reasoned.”
Human votesnone of 2 votes won