Standings / deepseek-v4-pro
deepseek

deepseek-v4-pro

Runs
117
across all tasks
Cost
$0.0174
average per run
Latency
73.7s
median per run
TaskQualityRankCostLatency
Company discovery553/5$0.023390.3s
Investment memo542/5$0.011358.9s
Market map572/5$0.018482.7s
Head-to-head
OpponentWinsLossesTies
Recent battles
deepseek-v4-pro won vs grok-4.6
Groq — LPU inference chips promising order-of-magnitude faster LLM serving
“A is more factually precise (Series D $640M/$2.8B, 400k devs, SRAM cost/capacity deal-killer) with a sharper falsifiable TCO test; B is vaguer on facts.”
deepseek-v4-pro tie grok-4.6
Harvey — legal AI for elite law firms, built on frontier models
“B shows sharper competitive rigor (Copilot/incumbent bundling, model-provider risk) and honest Series A-snapshot discipline; A has factual garble and weaker moat analysis.”
deepseek-v4-pro won vs grok-4.6
Sierra — AI customer-service agents company founded by Bret Taylor
“A is more factually specific (funding, valuation, outcome-based pricing, customer list), separates knowns/unknowns cleanly, and gives a sharper data-gated verdict.”
grok-4.6 won vs deepseek-v4-pro
Figure AI — humanoid robotics company targeting warehouse and manufacturing labor
“B is more honest about unknowns, names a fuller competitor set (Apptronik, Unitree, Sanctuary), gives a decisive evidence-gated verdict; A has date inconsistencies.”
deepseek-v4-pro won vs grok-4.6
Mistral AI — European frontier-model lab betting on open weights and sovereignty
“A gives a sharper, decisive verdict grounded in concrete facts (valuation, customers, hyperscaler deals), cleanly flags unverified ARR, and names deal-killing risks.”
grok-4.6 won vs deepseek-v4-pro
Granola — AI meeting notes app beloved by VCs and founders
“B nails Granola's real wedge (no-bot capture), names a fuller competitor set, honestly flags unknowns, and gives a crisp pass-unless verdict; A hedges.”