AI Founder Weekly

Grok 4.7 hits 94% on Next.js evals at 2x to 7x cheaper

Three models tie at the top, and the interesting number is the one underneath them.

Guillermo Rauch posted fresh Next.js evals with a four-model tally. Claude Opus 5.5, GPT-6 Sol and Claude Fable 5.1 all landed at 97%. Grok 4.7 came in at 94%, and Rauch flagged the part that actually decides things: Grok is 2x to 7x cheaper.

Guillermo Rauch
@rauchg
X
Notably, Grok is 2x-7x cheaper
Sep 22, 2026 · View on X

Three points of separation across four frontier models is not much of a spread. If you are picking a model to write your app code, the quality question has quietly stopped being the interesting one.

Cost is now the tiebreaker

The Next.js account added the detail that makes the leaderboard readable. GPT-6 Sol debuted at 97%, tying the highest success rate alongside Opus 5.5 and Fable 5.1. Among those three, Opus ranks first on average cost.

So at the top of the board you have a three-way tie on accuracy and a cost ranking sitting underneath it. Drop three points to Grok 4.7 and the price falls by somewhere between half and seven eighths, depending on which comparison you are running. Nobody in the sources says what the right trade is for a given workload, and that is genuinely a per-founder call.

OpenAI cut prices the same day

OpenAI shipped Sol and Luna as faster, more affordable siblings to GPT-6 Astra, saying both build on the advances behind Astra. The company says it made caching and inference more efficient and is passing the savings on, with 50% lower API prices for Sol and Luna compared with GPT-5.6 promotional pricing.

Simon Willison, who has been building on the cheap tier for a while, called out Luna specifically. He says Luna is half the price of 5.6 Luna, which he already considered astonishingly cheap given its capability.

Simon Willison
@simonw
X
Luna is my favorite model for building product features thanks to its cost (and speed)
Sep 22, 2026 · View on X

The read for a small team

For a solo founder or a two-person shop, the practical upshot is that you are no longer choosing a model on whether it can do the job. All four of these can, at least on Next.js work, at least on this eval. You are choosing on what the bill looks like at the end of the month and how fast the thing comes back.

That also means model loyalty is getting expensive. If the top of the board is a three-way tie and a discount challenger sits three points back, the sensible setup is one you can swap. Rauch's numbers are an eval on a specific framework, not a general verdict, and one leaderboard is not a benchmark suite. But it is the framework a lot of readers here actually ship in, which makes it the number worth watching.

Get the next one by email

Founders building with AI, every day. Real revenue, real pricing, real launches and what they would redo. No hype sludge. Every number links out.