playbooks

Jev is saving us 97%, and that's not all.

One check, put to two judges A line from a video plan, the hook, and one question about it: how strong is the opening, on five levels from weak to compelling. Our old judge, a language model, answered with a JSON block that we pulled out of the reply and parsed. Jev answers with a level, 4 of 5, a probability and a confidence, and lands within a tenth when asked again. Three results underneath: 97 per cent cheaper than a language model, half a second per answer, and a reported confidence that is low where it cannot see. The answers shown are illustrative. ONE CHECK, PUT TO TWO JUDGES A LINE FROM A VIDEO PLAN Hook: a thumb presses the cap and sunscreen loops over the tube. THE QUESTION How strong is the opening? Five levels, from weak to compelling. OUR OLD JUDGE ANSWERED JEV ANSWERS { "hook_strength": 4, "brand_named": false, "reason": "The opening is concrete and specific..." } JSON WE PULLED OUT AND PARSED Level 4 of 5 Probability 0.62, confidence 0.78 Within a tenth when asked again. A NUMBER, NOTHING TO PARSE 97% cheaper than a language model Half a second per answer, 24 questions Reports confidence low where it cannot see

TypeSafe announced Jev on 15 September. We got our hands on it on 28 September and had it in production the same day, as the judge of the plans InstantClips drafts for sellers' product videos. Jev costs 97 per cent less than the large language model (LLM) judge it replaced and answers almost 20 times faster.

That is not all. Jev is a different kind of model. It answers questions with numbers and writes nothing. Asked twice, it gives nearly the same answer. Its confidence drops where it lacks the evidence. Together they let us run consistent brand checks and automatic quality assessment on every asset a workflow produces. The rest of this post is the evidence.

Brand checks and quality assessment on every asset

A row of identical amber glass bottles on a white conveyor, one lit by a narrow band of warm orange inspection light
The same check on every bottle.

The reason this matters to a brand is what a workflow can now check, and who checks what. Literal rules go to code: is the approved product name present, is a banned word absent. Anything visual goes to a vision model first, which describes the frame. Everything between those is Jev's: is this caption in the brand's voice, does this plan lean on stock adjectives, does this script promise more than the brief, is this opening the kind the brand approved. Those are judgement calls, and on a text plan they now cost a fiftieth of a cent and half a second each. An image or video adds the cost of describing it first.

We are putting that division into the Creative Workflows we run for clients. A check that comes back with low confidence goes to a person. The rest are recorded against the asset, and a client's brand rules, approved once (Approve the workflow once), are enforced on everything the workflow produces.

Where the checks run Briefs, plans, scripts and captions go to Jev directly. Images and video are first described by a vision model, then go to Jev. Jev returns brand checks and quality scores as numbers on every asset at $0.00019 and 0.53 seconds each. A confident answer is recorded against the asset; a low-confidence answer goes to a person. WHERE THE CHECKS RUN Brief, plan, script, caption Image, video described first A vision model describes the frame Jev brand checks and quality scores on every asset $0.00019 · 0.53 S Confident recorded against the asset Low confidence a person decides

What Jev is

Three kinds of question, with one we ask of every plan A Choice, for example which of two plans is better, returns the odds of each option and a confidence: here A 0.81, B 0.19, confidence 0.74. A Score, for example how readable the plan is for a seller, returns a probability over five worded levels from confusing to clear and reads as a weighted level, here 3.5. A Noul, TypeSafe's yes or no, for example does the plan name a language, returns one probability, here 0.12, which counts as yes at 0.5 or more. We ask 24 such questions of every plan. The answers shown are examples. THREE KINDS OF QUESTION, WITH ONE WE ASK OF EVERY PLAN Choice one option from a list Which of two plans is better? A 0.81 B 0.19 CONFIDENCE 0.74 Score a level on a worded scale How readable is the plan for a seller? 1 2 3 4 5 confusing clear WEIGHTED LEVEL: 3.5 Noul TypeSafe's yes or no Does the plan name a language? 0 0.5 1 0.12: no YES AT 0.5 OR MORE We ask 24 of these about every plan: 13 Scores, 11 yes-or-no questions, and one Choice between the new plan and the current one.

Every question takes one of the three forms above, and Jev returns numbers and nothing else: no explanation, no prose. It reads text only, and TypeSafe says images are not supported yet.

How we use it

Twelve vertical product video ads from the InstantClips example library: jewellery, soap, a facial brush, coffee, a music box, a cap, a sweatshirt, a golf iron, touch-up paint, an Italian dress, matcha and a pet fountain
Twelve ads from the InstantClips example library.

InstantClips drafts a plan for each product video, and the seller approves it before anything renders. Jev judges every draft of that plan on 24 questions, from how strong the opening is to whether the seller's brand is named, and picks the better of two versions when we change the prompt that writes the plan. So far that is 1,800 drafts and 900 pairs, with the LLM judge it replaced run alongside for comparison. Day one was possible because on our platform the judge is one step in a workflow, and a step can point at a new model with nothing else changing.

What it costs, and what that changes

Cost and time per judged draft Two bar pairs drawn to scale. Cost: the LLM judge $0.006 to $0.012 a draft, Jev $0.00019, 32 to 63 times less. Time per answer: the LLM judge 5 to 10 seconds, Jev 0.53 seconds at the median, 10 to 19 times faster. COST AND TIME TO JUDGE ONE DRAFT, 24 QUESTIONS Jev LLM judge (Claude Sonnet 5); the lighter part is the measured range Cost LLM judge $0.006 to $0.012 Jev $0.00019, which is 32 to 63 times less Time per answer LLM judge 5 to 10 s Jev 0.53 s at the median, 10 to 19 times faster

Jev costs 97 to 98 per cent less than the LLM judge and answers in half a second against five to ten seconds. A million judgements cost under two hundred dollars, against six to twelve thousand.

At that price a judgement is effectively free, and that changes what gets measured. We judge every draft of every prompt build, ask every pair of plans both ways, and judge drafts twice to check that the answers hold. None of it shows up on a bill. When the judge cost thirty to sixty times more, measuring was something we scheduled. Now it runs on everything, every time.

DeepSeek is nearly free too, but it still writes text.

Three judges on the same job Per million judgements: Jev under $200, measured; DeepSeek V4.1 Flash $1,000 to $4,000 at list price; Claude Sonnet 5 $6,000 to $12,000, measured. Jev returns numbers with probabilities; the other two return text or JSON. Jev landed within a tenth on a repeat in our tests; we did not measure the other two. Jev reports a confidence; the other two self-report. Jev never thinks before answering; DeepSeek's thinking mode is on by default and can be turned off; Sonnet's is optional. THREE JUDGES ON THE SAME JOB Jev DeepSeek V4.1 Flash Claude Sonnet 5 MEASURED LIST PRICE MEASURED A million judgements under $200 $1,000 to $4,000 $6,000 to $12,000 What comes back numbers text or JSON text or JSON Repeat within a tenth yes, measured not measured not measured Reports confidence yes self-reported self-reported Thinks first never on by default optional

DeepSeek V4.1 Flash works out at roughly $0.001 to $0.004 a judgement at list price, nearly free as well, and it can return JSON. What it does not return is a probability over every level and a confidence. Its thinking mode is also on by default, and until it is turned off it spends tokens before the answer starts, in amounts it does not control. InstantClips ran on DeepSeek V4 Flash until July, and for an identical prompt it spent anywhere from zero to 260 tokens thinking before a 30-token answer. We turned thinking off, then switched models.

Nothing to parse

What comes back, and what you do with it An LLM judge returns a verdict written as text, which must be parsed and retried when it fails, before the numbers are stored. Jev returns numbers with probabilities that are stored directly: 0 of 1,800 answers unusable. WHAT COMES BACK, AND WHAT YOU DO WITH IT LLM JUDGE JEV A verdict written as JSON ↓ Parse it, retry when it fails an occasional reply will not parse ↓ Store the numbers Numbers, with probabilities ↓ Store the numbers nothing to parse, nothing to retry 0 of 1,800 answers unusable

Our old judge wrote its scores as a JSON block, which we pulled out of the reply with a pattern match and parsed, and now and then that failed. Schema-constrained output would remove most of that at the same price and speed. What it would not give us is what Jev returns natively: a probability for every level and a confidence, in half a second, for a fiftieth of a cent. Across 1,800 drafts not one Jev answer came back unusable.

Nearly the same answer twice

The same draft, judged twice A grid of 330 flags, 30 drafts by 11 yes-or-no questions, judged twice. Seven cells are marked as changed between the two runs; on each of three repeats 5 to 8 flags changed, all borderline. Scores moved 0.02 to 0.09 on average on a scale of 1 to 5. THE SAME DRAFT, JUDGED TWICE 330 flags: 30 drafts, 11 yes-or-no questions each 7 of 330 flags changed 5 to 8 per repeat, all close calls Scores moved 0.02 to 0.09 on average, on a scale of 1 to 5.

Asked the same 24 questions about the same draft twice, Jev moved each score by 0.02 to 0.09 on a scale of 1 to 5 and changed 2 per cent of the yes-or-no answers, all of them close calls. That is not identical, and it is close enough to use a score as a measurement: change the prompt, judge again, and read any difference bigger than a tenth as the prompt's doing.

Its confidence tracks the evidence

Jev's reported confidence, question by question On the four questions the text alone can answer, whether there is a creative idea, whether the hook fits it, whether you can follow the story and whether a seller can read it, Jev's mean confidence was between 0.7 and 0.8. On the three that need the photo or a cross-check, whether every claim is supported, whether the seller's notes were kept and whether camera choices were left out, its mean confidence was under 0.4, with 86, 92 and 97 per cent of answers under 0.5. JEV'S REPORTED CONFIDENCE, QUESTION BY QUESTION THE TEXT ALONE CAN ANSWER: 0.7 TO 0.8 Is there a creative idea? Does the hook fit the idea? Can you follow the story? Can a seller read it? 0.7 to 0.8 NEEDS THE PHOTO OR A CROSS-CHECK: UNDER 0.4 Every claim supported? 86% of answers under 0.5 Seller's notes kept? 92% of answers under 0.5 Camera choices left out? 97% of answers under 0.5 0 0.5 1

Every Score and Choice answer comes with a confidence, TypeSafe's measure of how concentrated the answer is on one level, from 0 to 1. In our data it tracks the evidence available to Jev. On the four questions the text alone can answer, its reported confidence was 0.7 to 0.8. On the three that need the photo it cannot see, or a check across two parts of the text, it was under 0.4, and 86, 92 and 97 per cent of answers came in under 0.5, the level at which TypeSafe's own guidance says to route to a person. Its misses fall in the same place, so we answer those questions another way, with a text search and with a judge that can see the photo.

What to design around

Three things to design around Order: a single A-or-B choice depends on which comes first, and swapped the same winner held only 63 to 77 per cent of the time; fix, ask both ways. Literal reading: it reads each question field by field, so a missing brand was flagged on 7 to 9 per cent of drafts against 36 to 57 by a string check; fix, string checks. Text only: it cannot see the product photo, so the unsupported-fact flag fired on 70 to 85 per cent of drafts; fix, a model that sees. THREE THINGS TO DESIGN AROUND Order A single A-or-B choice depends on which comes first. Swapped, the same winner only 63 to 77% of the time. Ask both ways Literal reading It reads each question literally, field by field. Missing brand: Jev says 7 to 9%, a text search says 36 to 57%. Text search Text only It cannot see the product photo the drafter saw. It called 70 to 85% of drafts unsupported. A model that sees

Three things, each with a mechanical fix, and each one showed up in Jev's own confidence first. Order: ask every pair both ways and average the two. Absence: a text search answers whether a word is present or missing, and Jev gets the questions that need judgement. Sight: whatever needs the image goes to a model that can see, or to a person.

How we measured

  • Native Jev, jev-1.13.0, over TypeSafe's API, text only.
  • Four test runs on 60 real seller inputs, three drafts each per prompt version: 1,800 judged drafts and 900 pairs of plans.
  • Repeat probes: 30 drafts judged three times over, and 30 pairs asked in both orders three times over.
  • Against the previous judge, Claude Sonnet 5 writing JSON: the same verdict on 62 per cent of 180 pairs, and 92 per cent where Jev's margin was 0.9 or wider. No human labels, so agreement is not accuracy.
  • Latency: median 0.53 s, slowest 5 per cent 1.7 to 7 s by hour. Failures: 0 to 8 network errors per 360 calls, retried. No unusable answers.
  • Production here means the judge that gates every change to the prompt that writes InstantClips' plans, run on real seller inputs. It does not yet score live seller drafts as they are made.

Since July we have said that no single model wins every job (One brief, five image models). Jev is the clearest case yet: a model that cannot write a sentence, doing a job the writing models did slower, dearer and with less consistency. If your brand rules are written down, they can be asked as questions. Talk to us about running them on your own work.

Let's get started

Put a brand check on every asset

We design and run managed Creative Workflows for brands and agencies, with brand checks and quality assessment built into the line. Talk to us about running them on your work.