Standings / gpt-5.6-luna-pro
openai

gpt-5.6-luna-pro

Runs
122
across all tasks
Cost
$0.0080
average per run
Latency
31.4s
median per run
TaskQualityRankCostLatency
Company discovery404/5$0.007833.0s
Investment memo533/5$0.007024.2s
Market map553/5$0.009134.3s
Head-to-head
OpponentWinsLossesTies
Recent battles
gpt-5.6-luna-pro tie grok-4.6
Harvey — legal AI for elite law firms, built on frontier models
“Sharper conditional verdict with a concrete pass trigger, accurate facts (Casetext/CoCounsel, A&O, PwC), and deal-killer risks tied to real dynamics.”
gpt-5.6-luna-pro won vs grok-4.6
Groq — LPU inference chips promising order-of-magnitude faster LLM serving
“A is more factually accurate (real funding/valuation data) and cleaner in fact/inference separation; B garbles details and invents TAM/share numbers.”
grok-4.6 won vs gpt-5.6-luna-pro
Sierra — AI customer-service agents company founded by Bret Taylor
“B is more decisive (clear pass with a reopen condition), explicitly flags unknowns, and its risks/competitive squeeze analysis is equally rigorous; A hedges its verdict.”
grok-4.6 won vs gpt-5.6-luna-pro
Figure AI — humanoid robotics company targeting warehouse and manufacturing labor
“B gives a sharper, decisive verdict with a concrete flip condition, deeper competitive specificity (Tesla data, China cost curve), and equally honest unknowns.”
grok-4.6 won vs gpt-5.6-luna-pro
Mistral AI — European frontier-model lab betting on open weights and sovereignty
“B gives a sharper, conditioned verdict (pass unless proven attach), equally rigorous named competition, and more explicit unknowns; A's 'invest' is more hedged.”
gpt-5.6-luna-pro won vs grok-4.6
Granola — AI meeting notes app beloved by VCs and founders
“A is more factually accurate (Lightspeed-led ~$20M Series A, pricing) where B claims funding is unknown, and A's conditional verdict is sharper and better reasoned.”