AI Founder Weekly

Lemkin says Jev judged 600 candidates for $0.064 and still failed

Roughly 90x cheaper on identical inputs, and still not usable for SaaStr Connect.

Jason Lemkin spent much of a day running Jev against the candidate judging for SaaStr Connect and says it came in at $0.064 for 600 judgments, against roughly $5 on Sonnet. He puts input pricing at $0.042 per million tokens with output free, which works out to roughly 90x cheaper on identical inputs. Then he looked at what the cheap judgments actually said.

On agreement with Sonnet, Jev hit 59.5% on 200 borderline-weighted pairs and 67% on 200 random pairs. Against a blind third-model referee on the random set, Lemkin reports Jev at 70.5% accurate and Sonnet at 77.5%. A seven point gap for a ninety-fold price cut is the kind of trade a lot of founders would take without reading further.

The errors were not symmetric

The split is where it fell apart. Lemkin says Jev produced 47 false admissions out of 148 negatives, a 32% rate, and only 2 false rejections out of 52 positives. Sonnet produced 20 false admissions, 13.5%, and 15 false rejections. Same rough accuracy band, opposite personality. Jev waves weak candidates through. Sonnet is stricter and turns away people it should have let in.

That is the part the headline accuracy number hides. Two models can land a few points apart on a scoreboard and be completely different products depending on which mistake costs you. For an event picking who gets in, letting a third of the weak answers through is not a rounding error, it is the whole job. For a task where a false rejection is the expensive one, the same model might look fine.

Lemkin's read on his own test was blunt.

This lands next to the earlier reports of Jev doing bulk classification work for cents. Those numbers were about throughput and cost. Lemkin's is about the other half, whether the output survives contact with a decision that matters. His conclusion is that it did not, at least for now.

The business read is that price per thousand judgments is the easy metric and the least useful one on its own. Before swapping a cheap model into anything that gates people or spend, run both against a referee, then split the errors into the two directions and ask which one you can afford. Lemkin says he did that in a day. It cost him six cents and saved him a bad launch.

Jason ✨👾SaaStr.Ai✨ Lemkin
@jasonlk
X
Jev unfortunately lets too many weak answers (candidates) through. A 32% false-admission rate is too high.
Sep 19, 2026 · View on X
Jason ✨👾SaaStr.Ai✨ Lemkin
@jasonlk
X
So on the surface, it looked great. At first.
Sep 19, 2026 · View on X

Get the next one by email

Founders building with AI, every day. Real revenue, real pricing, real launches and what they would redo. No hype sludge. Every number links out.