Building a Harness with Jev
All About Jev
Jev is a model released by TypeSafe AI. The company reports up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks, trained using Reinforcement Learning for Calibrated Decisions (RLCD).
Even with tool calling and structured outputs, the agent loop remains slow and costly as every decision requires another model call. Jev addresses this by providing fast, structured decisions without full chat LLM calls.
Jev is not a traditional LLM and does not generate text. It belongs to a class of AI models called System One models:
> System One models are built to make fast, structured decisions that software can use directly. A System One model evaluates a state and returns typed answers and probabilities.
It uses results to guide what an agent does next without a full chat LLM call for each decision.
To invoke a Jev model, send it a state (the context) and questions about that state:
{
"model": "jev-latest",
"state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}Response:
{
"is_urgent": {
"type": "noul",
"noul": 0.999
}
}This represents a 99.9% probability that the message is urgent, which application logic can use to prioritize the ticket.
Supported question types:
- Choice: Pick from a set of options. Returns a probability for each option and an overall confidence score.
- Score: Rate an input against ordered levels (e.g., low, medium, high). Returns a continuous score, the underlying distribution, and a confidence value.
- Noul: Answer a yes-or-no question. Returns the probability that a statement is true.
Multiple questions about the same state can be evaluated in parallel within a single request, keeping response times fast and token costs minimal. Unlike traditional LLMs, Jev is neither constrained by text generation nor sequential decision-making.
How to Use Jev with LangChain
The LangChain integration exposes Jev through TypeSafeClassifier. Pass state and questions to .invoke() to receive classification results directly:
from langchain_typesafe import Noul, TypeSafeClassifier
classifier = TypeSafeClassifier()
response = classifier.invoke(
state=(
"The deploy failed twice and customers are seeing 500s. "
"Can someone look now?"
),
questions={
"urgent": Noul(
instructions="Does this need attention right now?"
),
},
)
urgency = response.nouls["urgent"].noulThe state can consist of text, structured data, or LangChain messages, allowing Jev to be called from a node or middleware hook using existing agent context.
Use Cases
Jev serves as a complement to the primary model driving an agent: use an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way.
Model Routing
Model-routing middleware lets Jev assess requests and select models based on defined criteria—routing straightforward tasks to fast, inexpensive models while reserving complex tasks for more capable models:
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
ModelChoice,
ModelRouterMiddleware,
)
router = ModelRouterMiddleware(
choices={
"fast": ModelChoice(
model="openai:luna",
criteria="Direct lookups, extraction, and localized changes.",
),
"powerful": ModelChoice(
model="openai:sol",
criteria="Architecture and high-stakes decisions.",
),
},
instructions="Choose the least costly model that can complete the task.",
)
agent = create_agent("openai:gpt-5.6-luna", middleware=[router])Auto Mode
Coding harnesses use classifiers to detect dangerous actions before they are executed. Implementing AutoModeMiddleware brings this security pattern to general agents:
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
AutoModeMiddleware,
)
guardrail = AutoModeMiddleware(tools=["bash"])
agent = create_agent("openai:gpt-5.6-luna", middleware=[guardrail])AutoModeMiddleware uses Jev to check tool calls for risky decisions and block them before execution.
Applications
Practical implementations of Jev include powering browser-use agents at Browserbase, executing automated live trading strategies, and processing high-volume email triage.
Thoughts
The agent loop is expensive. Every classification that would otherwise require a full LLM call is now a cheap, parallel decision. That's the real win here.
I'd test the routing logic first. Make sure the fast model doesn't get routed to tasks it can't handle. Keep the dangerous action detection tight.
This is the kind of infra that actually scales agents beyond demos.