Marc Lou says the Opus 5.5 gap over OpenAI keeps widening
One founder's side by side on a non coding task, and the complaint is hedging rather than accuracy.
Marc Lou says the more he uses Opus 5.5, the further behind OpenAI's models look, and not just on coding.
OpenAI models right now are like that overthinking guy who can't make a decision.
His example is not a programming task at all. He says he sent GPT-6 Astra his VO2 max lab test results and got a headache back, roughly 200 sentences of your report shows this, I think that, therefore it's inconclusive. He says Opus 5.5 gave him the gist and actionable steps instead. His framing is that OpenAI models right now behave like someone who cannot make a decision, while Opus 5.5 is the high agency friend who tells you what you need to hear in one sentence. He ends by saying he does not know how the gap widened so much in the last few months.
What this is and is not
This is one founder running one prompt on each model and reporting how it felt. It is not a benchmark, there is no scoring, and nothing here says the OpenAI answer was wrong. Hedging and being incorrect are different failures, and a cautious read of medical style lab data is a defensible thing for a model to do. Lou is judging usefulness, not accuracy.
That said, usefulness is what solo operators pay for. When you are the only person in the company, a model that hands you five steps is worth more than a model that hands you a balanced summary you then have to decide on yourself. The cost of a confident wrong answer is real, but so is the cost of reading 200 sentences that end in inconclusive.
Worth noting that Lou is not the only person moving weight onto Opus 5.5 this week. Theo said the model now takes half of all T3 Code prompts, which is a usage number rather than a preference, and it is about coding rather than lab results. Two data points pointing the same way is still two data points.
The business read is that model choice is increasingly a taste call about tone and decisiveness, not a spec sheet comparison, and founders are making it one prompt at a time.
