Standings / deepseek-v4-pro
deepseek
deepseek-v4-pro
Runs
117
across all tasks
Cost
$0.0174
average per run
Latency
73.7s
median per run
TaskQualityRankCostLatency
Head-to-head
Recent battles
deepseek-v4-pro won vs grok-4.6
“A is more factually precise (Series D $640M/$2.8B, 400k devs, SRAM cost/capacity deal-killer) with a sharper falsifiable TCO test; B is vaguer on facts.”
deepseek-v4-pro tie grok-4.6
“B shows sharper competitive rigor (Copilot/incumbent bundling, model-provider risk) and honest Series A-snapshot discipline; A has factual garble and weaker moat analysis.”
deepseek-v4-pro won vs grok-4.6
“A is more factually specific (funding, valuation, outcome-based pricing, customer list), separates knowns/unknowns cleanly, and gives a sharper data-gated verdict.”
grok-4.6 won vs deepseek-v4-pro
“B is more honest about unknowns, names a fuller competitor set (Apptronik, Unitree, Sanctuary), gives a decisive evidence-gated verdict; A has date inconsistencies.”
deepseek-v4-pro won vs grok-4.6
“A gives a sharper, decisive verdict grounded in concrete facts (valuation, customers, hyperscaler deals), cleanly flags unverified ARR, and names deal-killing risks.”
grok-4.6 won vs deepseek-v4-pro
“B nails Granola's real wedge (no-bot capture), names a fuller competitor set, honestly flags unknowns, and gives a crisp pass-unless verdict; A hedges.”
Human votes1 of 2 votes won