Decisions instead of text (Jev)
A normal model writes prose, and then you have to take it apart: look for "yes" in the answer, parse JSON, hope the shape holds. Jev writes no text at all. You hand it a state and a list of typed questions; it returns probabilities, scores and choices.
It is fast (a fraction of a second), very cheap, and it does not drift: the answer always has the shape you asked for, because you defined the shape.
Good for: ticket classification, routing, rubric checks, moderation, filtering — anything where you need a decision, not a paragraph.
Many questions, one request
The model answers every question in a single call. Ten questions in one request cost about the same as one — that is where the savings are.
Quick start
curl https://nordrouter.com/v1/evaluate \
-H "Authorization: Bearer YOUR-KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe-ai/jev",
"state": "Order #1841 never arrived, third message from the customer, harsh tone",
"questions": {
"urgent": {
"type": "boolean",
"question": "Is this urgent?",
"criteria": { "true": "repeat contact or harsh tone", "false": "first calm message" }
},
"team": {
"type": "choice",
"question": "Who should take it",
"criteria": { "delivery": "orders and timing", "billing": "money", "support": "everything else" }
}
}
}'Response:
{
"model": "typesafe-ai/jev",
"answers": {
"urgent": { "type": "boolean", "probability": 0.85 },
"team": {
"type": "choice",
"choice": "delivery",
"probabilities": { "delivery": 1, "billing": 0, "support": 0 },
"confidence": 0.99
}
},
"usage": { "input_tokens": 434, "output_tokens": 82 }
}The answer keys are your own question names. Nothing to parse.
Three kinds of question
Each kind takes a different criteria shape
This is the one real trap in the format. Getting it wrong is easy, and the error comes back immediately and clearly.
boolean — yes or no
criteria is an object with two keys:
"angry": {
"type": "boolean",
"question": "Is the customer upset?",
"criteria": { "true": "harsh tone, demands", "false": "calm tone" }
}Returns probability between 0 and 1. The threshold is yours: 0.5 for ordinary calls, 0.8 when a mistake is expensive.
score — a rating on a scale
criteria is an array, lowest to highest:
"severity": {
"type": "score",
"question": "How serious is this",
"criteria": ["minor", "normal", "critical"]
}Returns score, the probabilities across the steps, and confidence.
choice — pick one option
criteria is an object: option and what it means.
"route": {
"type": "choice",
"question": "Where should this go",
"criteria": { "support": "general questions", "billing": "money", "engineer": "something broke" }
}Returns choice, the spread across all options, and confidence.
Instructions instead of criteria
Any kind accepts instructions as a plain string:
"ok": { "type": "boolean", "question": "Is everything fine?", "instructions": "true when there are no complaints" }One of the two — criteria or instructions — is required. Without either the model refuses.
Python
import httpx
r = httpx.post(
"https://nordrouter.com/v1/evaluate",
headers={"Authorization": "Bearer YOUR-KEY"},
json={
"model": "typesafe-ai/jev",
"state": ticket_text,
"questions": {
"spam": {"type": "boolean", "question": "Is this spam?",
"criteria": {"true": "advertising or gibberish", "false": "a real message"}},
"topic": {"type": "choice", "question": "What is it about",
"criteria": {"payment": "money", "access": "cannot sign in",
"bug": "something broke"}},
},
},
timeout=40,
)
a = r.json()["answers"]
if a["spam"]["probability"] > 0.8:
drop(ticket)
else:
route_to(a["topic"]["choice"])Price
| per 1M tokens | |
|---|---|
| input | $0.0504 |
| output | free |
Only what you send is billed. The answer costs nothing.
The ticket above is 434 tokens, that is $0.0000219 — roughly 46,000 calls per dollar.
The exact charge for a call comes back in the X-Charged-USD header.
Rate limit
The channel sustains 15 requests per second. Anything beyond that gets a 429 immediately — we do not park your request in a queue and decide on your behalf how long it should hang there. The refusal carries a Retry-After header in seconds.
Retries are yours to make, deliberately: you know better than we do whether a given call should wait, be deferred, or be dropped.
import time, httpx
def evaluate(payload, tries=3):
for _ in range(tries):
r = httpx.post(URL, headers=HDRS, json=payload, timeout=40)
if r.status_code != 429:
return r
time.sleep(int(r.headers.get("Retry-After", 1)))
return rPace yourself under the ceiling
Refusals are cheap but useless. Fifteen requests per second at a steady pace gets more work done than a hundred at once, eighty of which come back 429.
Limits and quirks
- Context is 32,000 tokens for
stateand the questions together. - No streaming: the answer arrives whole, usually in 0.3–0.6 seconds.
- This model only. Any other name returns
400. - Plain
/v1/chat/completionsdoes not work with it — this is an evaluation model with its own endpoint,/v1/evaluate.
Key and balance
Your normal sk-nr-… key and your normal balance. Nothing to set up: if you already use NordRouter, the evaluation model is available right away.
Charges appear in your dashboard alongside everything else — there is no separate ledger to reconcile.
Errors
| code | meaning | what to do |
|---|---|---|
401 | key rejected | use your normal sk-nr-…; check it is enabled |
402 | balance empty | top up |
400 | wrong model or question shape | check the criteria shapes above |
429 | rate limit exceeded | wait Retry-After seconds, then retry |
502 | model temporarily unavailable | retry in a few seconds |