Kaboom.
That was roughly the sound of Jev arriving in my corner of the internet. I got on the waitlist. I got access. And I immediately started thinking about how I would change my application to use it.
TypeSafe's Jev gives you an appealing programming interface: supply some context, ask questions, and get probabilities back. A Boolean probability. A distribution over choices. A score against a rubric.
I build AuthorCraft, which analyzes fiction manuscripts. I can think of plenty of uses for that. Is this a chapter boundary? Are these two names the same character? Is this passage mostly action, description, or introspection?
Then: wait, wait, wait.
First: can I depend on one provider? I don't know what this company's future looks like. I don't know whether it will have the capacity I need when I need it. Even if everything goes wonderfully, I want to be able to use something better when it comes along.
Second: is it any good at my questions? Fast sounds great. Cheap sounds great. I'm a machine learning guy. Obviously we're going to have to run an evaluation.
A radical idea: an abstraction layer
Here's a blast from the past. Software has this thing called an abstraction layer: it separates how an application talks about the work from how the computer executes it. Groundbreaking stuff, I know.
I've been building one for LLMs, called creAItive LLM. It has three jobs:
- Separate what the application asks for from how it gets executed.
- Handle operations like failover and cost tracking.
- Package the machinery of machine learning quality evaluation so you can compare models on your own workload.
Jev was a perfect test case for all three.
Could I give my application a probability-judgment interface, then execute it on Jev or on an ordinary language model?
Turns out the answer is yes. I created a simple judgments API that lets the application ask for probabilities, and wrote two implementations: one calls Jev; the other translates the questions into instructions and a response format for general-purpose models, including OpenAI, Claude, and my local DeepSeek. Both return the same kind of answer to the application.
One interface. More than one way to execute it.
The application wants a judgment: how likely is this, or which option fits? It shouldn’t have to express that as “generate some text, please follow this schema, and then I’ll parse it.” The API captures the application’s intent. Whether a model answers natively or needs instructions to produce that structure is the execution layer’s job.
I can configure the execution policy to fail over if Jev is down, or switch providers entirely if a better alternative comes along. Come on, people. Put away your punch cards. We live in the future.
A little creAItive LLM code
Here's the application request:
from creAItive import llm
request = llm.Request(
input=passage,
output=llm.OutputSpec.judgments({
"chapter_boundary": llm.BooleanQuestion(
"Does this excerpt contain a chapter heading marking the "
"beginning of a new chapter? Ordinary paragraph breaks do not count."
),
}),
)
Choosing the models is simple:
jev = llm.Policy(
name="judgment.jev",
version=1,
model_attempt_order=[llm.models.TYPESAFE_JEV_1_13_0],
request_requirements=llm.RequestRequirements(
input_parts={llm.InputKind.TEXT},
output_kinds={llm.OutputKind.JUDGMENTS},
),
)
luna = llm.Policy(
name="judgment.luna",
version=1,
model_attempt_order=[llm.models.OPENAI_GPT_6_LUNA],
request_requirements=llm.RequestRequirements(
input_parts={llm.InputKind.TEXT},
output_kinds={llm.OutputKind.JUDGMENTS},
),
)
With those policies installed in a runtime manifest and the corresponding credentials configured, the call site is:
jev_result = runtime.invoke(request, llm_policy="judgment.jev")
jev_probability = jev_result.output.judgments["chapter_boundary"].probability
luna_result = runtime.invoke(request, llm_policy="judgment.luna")
luna_probability = luna_result.output.judgments["chapter_boundary"].probability
How did it do?
I built a small benchmark from two of my own manuscripts, Iceborne and Auramancer. The tasks are things AuthorCraft might reasonably need: spotting chapter boundaries and spoken dialogue, resolving character references, and distinguishing description, action, and introspection. I also made a few controlled sentence edits to test passive voice and clear grammatical errors.
These are useful little decisions. They were not designed to challenge the limits of frontier models.
I asked each model the same 43 questions and compared its answers, cost, and speed.1
| Execution policy | Correct / 43 | API cost for 43 questions (¢) | Median call latency |
|---|---|---|---|
| TypeSafe Jev 1.13.0 | 43 | 0.096¢ | 0.221 s |
| GPT-6 Luna | 41 | 0.293¢ | 1.214 s |
| GPT-6 Luna Flex | 40 | 0.146¢ | 1.162 s |
| GPT-6 Sol | 43 | 5.857¢ | 1.482 s |
| Claude Haiku 4.5 | 42 | 6.595¢ | 0.784 s |
| DeepSeek local | 43 | No API fee* | 2.574 s |
DeepSeek local: v4 Flash Vision Exp on two Sparks. Hardware and electricity costs are not included.
The runs and Jev–Luna comparison come from creAItive LLM’s standard offline evaluation workflow: compare execution policies on your own workload, then inspect the answers, costs, and timings side by side. I added the summary charts for this post.
My takeaway
Jev is a really good choice for this workload. It got every expected decision, with the lowest hosted API cost and the shortest median latency. Luna was close on accuracy, and its absolute cost was tiny too. Our local DeepSeek model also got all 43 right, at a slower 2.57-second median. If you aren't especially latency-sensitive, you have useful alternatives.
And come on, we're not cavemen. Put an interface between your application and the models. You can take advantage of a new capability without committing every call site to the provider that got you excited about it.
These layers aren't rocket science. They're a bit of a pain in the butt to vibe-code and get right.2 Ping me if you want to take a look at creAItive LLM.
Small application benchmark using my own manuscripts, with fixed answers and some shared passages: 28 Boolean questions and 15 choices. One run per case; this measures answer agreement, not probability calibration. Hosted costs use recorded usage and configured prices. OpenAI reasoning was disabled; calls ran with three concurrent jobs. Local API fees exclude hardware and electricity. Full inputs and results accompany the comparison.↩
A friend of mine, a Distinguished Engineer and software architect, disagreed: “I’d say it more strongly: they can’t be vibe-coded and have anything worth having. Abstraction layers are all about the design of the API, and models remain bad at it.”↩