AI Founder Weekly

Lemkin says use the cheap model additively, never as a filter

The cheap model sweeps work nobody was running, and the expensive model still makes the call.

Jason Lemkin has settled on where a cheap fast model fits in his stack, and it is not in place of anything he already runs. In a reply to Tommy Long, Lemkin says the answer is "not as a replacement, and not as a pre-filter either", and that the only use he trusts is additive work. Sweep the volume you currently skip, hand the best candidates to the expensive model.

Jason ✨👾SaaStr.Ai✨ Lemkin
@jasonlk
X
not as a pre-filter either, since a cheap model rejecting silently removes good options before anything better sees them
Sep 20, 2026 · View on X
Jason ✨👾SaaStr.Ai✨ Lemkin
@jasonlk
X
Only additive. Sweep the volume you currently skip, hand the best to the expensive model.
Sep 20, 2026 · View on X

The number that moved him is latency, not price. Lemkin says a cold search at SaaStr.Ai evaluates thousands of candidates and takes minutes, and that 219ms versus 4,188ms per call changes what is possible there. That is a roughly nineteen times gap on a job where the run time is the constraint, which is why he frames the test as whether speed is worth anything to you rather than whether the cheap model is good enough.

Why he rules out pre-filtering

The pre-filter case is the one that looks obviously smart and is not. Lemkin's objection is that a cheap model rejecting silently removes good options before anything better sees them. You never see what you lost, there is no error message, and the expensive model downstream looks like it is performing fine because it only ever sees what survived. On a search that evaluates thousands of candidates, a quiet false negative rate is close to invisible.

Tommy Long, who prompted the exchange, reports the mirror image result. He tried swapping in Jev for Gemini 3.8 Flash to annotate spend categories on transactions and found Jev was not far behind while being much cheaper and faster. He stayed on Gemini anyway, because the job is not time-sensitive and Gemini is already cheap enough. Same conclusion about judgment quality, opposite decision, because the workload has no clock on it.

The business read

These are two founders' own reports from their own workloads, not benchmarks, and the split between them is the useful part. A cheap model earns its place when latency is the thing stopping you from running a job at all, and it earns nothing when your current model is already cheap and the job can take its time. Lemkin's version of the win is new work he was not running before, not the same work at a lower bill, which is a different line entirely. If the pitch you are hearing is cost savings on jobs you already run well, the honest answer in both of these cases was to leave it alone.

Get the next one by email

Founders building with AI, every day. Real revenue, real pricing, real launches and what they would redo. No hype sludge. Every number links out.