If he doesn't help, tell me and I'll refund you.

I can't stand behind that if I don't know Seth works.

So I test him.

26 prompts. 6 criteria. 30 points.
Written before the results existed.
Fixed since.

The 2026 result both numbers out of 30.

We ran the old version of Seth on the new model to check. 23.1 became 23.2. The model didn’t make Seth better. Our upgrades did.

The average is the least interesting number here.

ChatGPT never once reached the top band.

Ten times it fell below the line where we'd say it isn't working.

Seth never did.

One of the 26 prompts needs a document the reader supplies, so the bare run covers 25 every table prints its own base, and bars are proportions of that base.

ChatGPT is good.
On three prompts it did something Seth didn't.

That isn't the argument. We ran one email through it five times: 13, 17, 21, 24, 24. Three of those are good reviews. One quietly weakened a commitment the sender had made.

All five sounded equally sure of themselves.

You can't hire the good day.

The same email, through ChatGPT, five times.

Three of those are good reviews. one quietly weakened a commitment the sender had made. All five sounded equally sure of themselves.

And the thing it's worst at is the thing you're paying for.

1.92 is the lowest number anywhere in the dataset.

One of the 26 prompts needs a document the reader supplies, so the bare run covers 25 every table prints its own base, and bars are proportions of that base.

One prompt.
Both answers.

What still fails.

C11. Asked whether a message should go yet, Seth reviews the writing and doesn't answer the timing question. Failed twice.

C3. Seven identical failures. We think the prompt is broken, not Seth. We haven't changed it, moving the test to fit the result is how you get a meaningless number.

A3, A4. One sentence Seth is meant to return exactly. Twice it came back with a comma instead of a dash.

The models change underneath Seth.

OpenAI ships changes without telling anyone. Seth runs inside ChatGPT, so those changes reach you before they reach me.

Checked monthly. Published once a year. Changed only when a check finds a reason.

Scored against the published rubric, with AI assistance. We sell the thing being tested. Factor that in.