Skip to content

Decisions instead of text (Jev) ​

A normal model writes prose, and then you have to take it apart: look for "yes" in the answer, parse JSON, hope the shape holds. Jev writes no text at all. You hand it a state and a list of typed questions; it returns probabilities, scores and choices.

It is fast (a fraction of a second), very cheap, and it does not drift: the answer always has the shape you asked for, because you defined the shape.

Good for: ticket classification, routing, rubric checks, moderation, filtering — anything where you need a decision, not a paragraph.

Many questions, one request

The model answers every question in a single call. Ten questions in one request cost about the same as one — that is where the savings are.

Quick start ​

bash
curl https://nordrouter.com/v1/evaluate \
  -H "Authorization: Bearer YOUR-KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "typesafe-ai/jev",
    "state": "Order #1841 never arrived, third message from the customer, harsh tone",
    "questions": {
      "urgent": {
        "type": "boolean",
        "question": "Is this urgent?",
        "criteria": { "true": "repeat contact or harsh tone", "false": "first calm message" }
      },
      "team": {
        "type": "choice",
        "question": "Who should take it",
        "criteria": { "delivery": "orders and timing", "billing": "money", "support": "everything else" }
      }
    }
  }'

Response:

json
{
  "model": "typesafe-ai/jev",
  "answers": {
    "urgent": { "type": "boolean", "probability": 0.85 },
    "team": {
      "type": "choice",
      "choice": "delivery",
      "probabilities": { "delivery": 1, "billing": 0, "support": 0 },
      "confidence": 0.99
    }
  },
  "usage": { "input_tokens": 434, "output_tokens": 82 }
}

The answer keys are your own question names. Nothing to parse.

Three kinds of question ​

Each kind takes a different criteria shape

This is the one real trap in the format. Getting it wrong is easy, and the error comes back immediately and clearly.

boolean — yes or no ​

criteria is an object with two keys:

json
"angry": {
  "type": "boolean",
  "question": "Is the customer upset?",
  "criteria": { "true": "harsh tone, demands", "false": "calm tone" }
}

Returns probability between 0 and 1. The threshold is yours: 0.5 for ordinary calls, 0.8 when a mistake is expensive.

score — a rating on a scale ​

criteria is an array, lowest to highest:

json
"severity": {
  "type": "score",
  "question": "How serious is this",
  "criteria": ["minor", "normal", "critical"]
}

Returns score, the probabilities across the steps, and confidence.

choice — pick one option ​

criteria is an object: option and what it means.

json
"route": {
  "type": "choice",
  "question": "Where should this go",
  "criteria": { "support": "general questions", "billing": "money", "engineer": "something broke" }
}

Returns choice, the spread across all options, and confidence.

Instructions instead of criteria ​

Any kind accepts instructions as a plain string:

json
"ok": { "type": "boolean", "question": "Is everything fine?", "instructions": "true when there are no complaints" }

One of the two — criteria or instructions — is required. Without either the model refuses.

Python ​

python
import httpx

r = httpx.post(
    "https://nordrouter.com/v1/evaluate",
    headers={"Authorization": "Bearer YOUR-KEY"},
    json={
        "model": "typesafe-ai/jev",
        "state": ticket_text,
        "questions": {
            "spam":  {"type": "boolean", "question": "Is this spam?",
                      "criteria": {"true": "advertising or gibberish", "false": "a real message"}},
            "topic": {"type": "choice", "question": "What is it about",
                      "criteria": {"payment": "money", "access": "cannot sign in",
                                   "bug": "something broke"}},
        },
    },
    timeout=40,
)
a = r.json()["answers"]

if a["spam"]["probability"] > 0.8:
    drop(ticket)
else:
    route_to(a["topic"]["choice"])

Price ​

per 1M tokens
input$0.0504
outputfree

Only what you send is billed. The answer costs nothing.

The ticket above is 434 tokens, that is $0.0000219 — roughly 46,000 calls per dollar.

The exact charge for a call comes back in the X-Charged-USD header.

Rate limit ​

The channel sustains 15 requests per second. Anything beyond that gets a 429 immediately — we do not park your request in a queue and decide on your behalf how long it should hang there. The refusal carries a Retry-After header in seconds.

Retries are yours to make, deliberately: you know better than we do whether a given call should wait, be deferred, or be dropped.

python
import time, httpx

def evaluate(payload, tries=3):
    for _ in range(tries):
        r = httpx.post(URL, headers=HDRS, json=payload, timeout=40)
        if r.status_code != 429:
            return r
        time.sleep(int(r.headers.get("Retry-After", 1)))
    return r

Pace yourself under the ceiling

Refusals are cheap but useless. Fifteen requests per second at a steady pace gets more work done than a hundred at once, eighty of which come back 429.

Limits and quirks ​

  • Context is 32,000 tokens for state and the questions together.
  • No streaming: the answer arrives whole, usually in 0.3–0.6 seconds.
  • This model only. Any other name returns 400.
  • Plain /v1/chat/completions does not work with it — this is an evaluation model with its own endpoint, /v1/evaluate.

Key and balance ​

Your normal sk-nr-… key and your normal balance. Nothing to set up: if you already use NordRouter, the evaluation model is available right away.

Charges appear in your dashboard alongside everything else — there is no separate ledger to reconcile.

Errors ​

codemeaningwhat to do
401key rejecteduse your normal sk-nr-…; check it is enabled
402balance emptytop up
400wrong model or question shapecheck the criteria shapes above
429rate limit exceededwait Retry-After seconds, then retry
502model temporarily unavailableretry in a few seconds