Launch offerup to +40% credits on every packends inClaim →
jagent.
Get a free key

Tested 2026-10-09

Jev rate limits and API errors, tested

Jev rate limits and API errors as the API actually returns them: the request limits, the size limits on a single call, and ten errors across six status codes that we produced by breaking requests on purpose, each with its body and its fix. Three differ from TypeSafe's own error table.

Quick answer · verified 2026-10-09

What are the Jev rate limits?

100,000 input tokens a second and 80 requests a second, and a request over either gets a 429. TypeSafe says the Jev rate limits are adjusting as it adds capacity and can change without notice; higher limits come with custom and enterprise plans.

Tokens100K / s
Requests80 / s
Per request64k tokens — 32k for state + longest question
Over a limit429 — retry with backoff

Which Jev rate limit will you hit first?

Usually the request limit. At 80 requests a second, the token limit only binds once the average request passes 1,250 tokens (100,000 ÷ 80). A support ticket with three questions read 411 tokens in our test, so 80 of them is about 33,000 tokens a second, a third of the token allowance. Long documents flip that: a 26,000-token state uses a quarter of a second's token budget in one call.

TypeSafe does not say whether the Jev rate limits count per key or per account, and responses carry no remaining-quota header, so meter your own requests per second if you run near the line.

What does each Jev API error mean?

Every error we produced, with the response body as it came back (shortened where marked) and what fixes it:

StatusWhat we sentBodyFix
400model: "jev-1.13"Unknown model: jev-1.13Use jev-1.13.0 or jev-latest; typesafe/jev-1.13 is OpenRouter's id.
400A question type of "multi", or a temperature fieldInvalid request.Send only model, state and questions, typed choice, score or noul. No field is named.
400256 options in a ChoiceToo many choices. Must have at most 255 choices.Split into a coarse Choice, then a fine one inside the winner.
40011 levels in a ScoreToo many score levels. Must have at most 10 levels.Merge levels; ten is the ceiling.
400A state over the token budget{"error_type":"max_tokens_exceeded"}Chunk the state; the error gives no count.
401A wrong key, or an OpenRouter keyCannot authenticate with the server.Use a TypeSafe key; the two routes do not share keys.
403No Authorization headerMust supply an API key!Send Bearer <key>. TypeSafe's table lists this case under 401.
422No state, no model, or an empty questions objectField required, with loc ["body","state"]Read loc: it is the path to the missing field.
429Documented: over either rate limitNot reproducedBack off and retry; see below.
529Documented: TypeSafe overloadedNot reproducedBack off and retry, or fail over.

Three rows differ from TypeSafe's table, which lists 401, 422, 429 and 529. A missing key is a 403, not a 401. A malformed question, such as an unknown type, is a 400 with no detail, where the table says a 422 that names the field. And the option, level and token limits come back as 400s.

Parse detail three ways

The detail field changes shape with the error: an object with error_type and message for auth and usage errors, a plain string for the option and level limits, and a list of validation entries for a 422. A handler that reads detail.message crashes on two of the three:

python
def jev_error(body: dict) -> str:
    d = body.get("detail")
    if isinstance(d, list):   # 422: [{"loc": [...], "msg": ...}]
        return "; ".join(f"{'.'.join(map(str, e['loc']))}: {e['msg']}" for e in d)
    if isinstance(d, dict):   # 400/401/403: {"error_type", "message"?}
        return d.get("message") or d.get("error_type", "unknown")
    return str(d)             # 400 limits: a plain string

Which mistakes does the Jev API accept without an error?

Three requests we expected to fail came back 200, with an answer worth nothing:

Check these before sending, because the API will not:

python
assert payload["state"] and str(payload["state"]).strip(), "empty state"
for qid, q in payload["questions"].items():
    if q["type"] == "score":
        assert 2 <= len(q["criteria"]) <= 10, f"{qid}: 2 to 10 levels"
    if q["type"] == "choice":
        assert 2 <= len(q["criteria"]) <= 255, f"{qid}: 2 to 255 options"

How should you retry Jev 429 and 529 errors?

With backoff, and only those. TypeSafe's Python SDK already does it: its default RetryPolicy makes two retries, waits 0.5 seconds and doubles to a 5-second cap with up to 25% jitter, honours Retry-After, and retries 408, 429 and every 5xx, which covers 529. Each attempt times out at 10 seconds and the whole call at 30.

Thirty seconds is a long time on a hot path. If a user is waiting, cap it: the SDK's RetryPolicy(max_retries=1, timeout=3.0), or the same rules over plain HTTP:

python
import random, time, requests

RETRY = {408, 429} | set(range(500, 600))  # what TypeSafe's SDK retries by default

def decide(payload, key, attempts=3, budget=8.0):
    start = time.monotonic()
    for i in range(attempts):
        r = requests.post("https://api.typesafe.ai/v1/systemone", json=payload,
                          headers={"Authorization": f"Bearer {key}"}, timeout=5)
        if r.status_code not in RETRY or i == attempts - 1:
            r.raise_for_status()  # 400, 401, 403 and 422 fail the same way twice
            return r.json()
        wait = min(0.5 * 2**i, 5.0) * (1 - 0.25 * random.random())
        try:
            wait = float(r.headers.get("retry-after", wait))
        except ValueError:
            pass
        if time.monotonic() - start + wait > budget:
            r.raise_for_status()
        time.sleep(wait)

Tested live and on stubbed responses: a 429 then a 200 returns the answer, three 529s raise after about 1.3 seconds, a 422 raises at once. The other option is a second route: this site moves to OpenRouter only on a 429 or a 5xx, as described on Jev on OpenRouter.

How big can one Jev request be?

As documented on tokens, and with no limit we could find on questions. In our calls a request reading 26,290 tokens, nearly all of it state, went through; a state of about 36,000 words was refused with max_tokens_exceeded, as was a request of about 66,000 words spread over sixty long questions. That matches the budgets on the Jev context window: 64k for everything, 32k for the state plus the longest question.

The count of questions has no documented cap, and we found none: 500 yes-or-no questions in one request came back with 500 answers in 555 ms, reading 17,666 tokens. A Choice took 255 options and refused 256, and a Score took 10 levels and refused 11.

Is the state billed once when you ask several questions?

Yes, and there is a floor per request besides. Asked one at a time, our three ticket questions read 330, 307 and 326 tokens, 963 in all; asked together they read 411, so the single request cost 43% as much. Even an empty state with one two-option question read 313 tokens, so every extra request pays roughly 300 tokens before your text counts.

Split a request only when a question depends on another's answer, such as asking about refund policy only if the ticket is billing. Often it is cheaper to ask both anyway and ignore the one that does not apply. What a call costs at each volume is on Jev pricing.

What should you log from each Jev call?

The model field and the x-typesafe-request-id header. The first names the build that answered: jev-latest moves when a release ships, so pin jev-1.13.0 if you tuned thresholds against it. The second is the only per-call id, because the first-party body has none. The SDK exposes it on errors as request_id, and its exception classes follow the codes above; TypeSafeRateLimitError also carries retry_after_ms.

Jev rate limits and API errors: common questions

Can I get higher Jev rate limits?

Yes, on TypeSafe's custom and enterprise plans, through sales@typesafe.ai. The published limits are also moving on their own: TypeSafe says they are adjusting while it adds capacity and can change without notice, so read them from the models page rather than hard-coding them.

Does the Jev API return rate-limit headers?

Not on the successful calls we recorded on 2026-10-09: the only custom header was x-typesafe-request-id. There is no remaining-quota header to read, so count your own requests per second if you run near 80.