Measured 2026-10-09
Jev vs Laya: the hosted model and the open one, on the same items
Jev vs Laya, run on the same 633 labelled items with the same questions: which is more accurate on which task, whose confidence you can act on, how fast Laya is on an ordinary laptop, and the volume at which running it yourself costs less than Jev's API.
Quick answer · verified 2026-10-09
Should you use Jev or Laya?
Jev, unless the data cannot leave your machines. Out of the box Laya matched or beat Jev only on few-option topic and toxicity questions, fell to about half Jev's accuracy with 59 to 77 options, and was more than twice as often wrong when it was sure. If you must self-host, Laya is a real option for short option lists, and its makers say fine-tuning is where its accuracy comes from.
| 4-option news | 91% / 88% — Laya / Jev |
|---|---|
| 59–77 options | 44.5% / 79% — Laya / Jev |
| Sure but wrong | 26.8% / 11.5% — at 0.9 or more |
| Price | Hardware / $0.042 — per 1M input tokens |
What is Laya?
An open-weight decision engine from Convai Innovations, first published on PyPI and GitHub on 2026-09-18 under Apache 2.0. Like Jev it answers typed questions, choice, score and yes-or-no, with probabilities, in one forward pass. Unlike Jev it is small enough to run anywhere: the English checkpoint is a 421M-parameter ModernBERT-large reading 512 tokens, beside a 322M multilingual one and a version fine-tuned on its own decision set.
Every comparison we found repeats Laya's own benchmark, whose Jev column is labelled as published figures, never measured by its authors, who had no TypeSafe access. We have both, so we ran Laya on the exact items Jev had answered for our thirteen-model comparison.
Which is more accurate, Jev or Laya?
Jev on most tasks; Laya on the simplest one. One question per request, no examples, the same wording for both:
| Task | Options | Laya | Jev |
|---|---|---|---|
| News topics (AG News) | 4 | 91% | 88% |
| Tweet emotion | 6 | 60% | 63% |
| Support tickets | 5 | 74.1% | 100% |
| Voice-assistant intents (MASSIVE) | 59 | 45% | 76% |
| Bank questions (Banking77) | 77 | 44% | 82% |
| Toxic comment, yes or no | 2 | 91% | 92% |
| Positive review, yes or no | 2 | 67% | 98% |
Our Laya numbers agree with its own: it reports 95% on AG News and 42.5% on Banking77, against our 91% and 44% on a smaller sample. The many-option collapse is architectural, and Laya's README says so: all of a question's options share one budget of 192 tokens, so 77 labels get a few tokens each, and it advises keeping choices under about 20. The support tickets surprised us more. Laya sent three account requests to technical, two sales questions to spam, and a $14.5 million advance-fee letter to billing.
We tried to give Laya a fairer yes-or-no question
Laya called 33 of the 44 positive product reviews not positive, so we reran the two yes-or-no sets three ways: our wording, the same statement without the true and false labels, and a question in Laya's own style. Reviews scored 67%, 71% and 56%; toxicity 91%, 88% and 88%. The wording was not the problem, and the table keeps the original.
Whose confidence can you act on?
Jev's, by a wide margin. Of the choice answers each gave at 0.9 confidence or more, 26.8% of Laya's were wrong against 11.5% of Jev's. On the support tickets Jev's confidence fell from 0.98 on clear tickets to 0.82 on deliberately ambiguous ones; Laya's sat at 0.53 and 0.45, low on both, so it does not separate easy from hard.
Two cautions if you move a threshold across. Laya's README says its checkpoints ship over-confident and recommends refitting a temperature on your own data. And its confidence field is defined differently from Jev's, one minus the normalised entropy, so a Jev threshold does not carry over; we recomputed Laya's on Jev's definition for the figures above.
How fast is Laya on your own hardware?
On an Apple M2 laptop with no other setup, one question took 163 ms at the median and 573 ms at the 90th percentile, and batching raised it to 17.9 decisions a second on news articles. The first load downloaded 843 MB of weights and took about 70 seconds. Laya reports 33 to 39 ms per question on an NVIDIA T4. For scale, Jev through OpenRouter answered in 184 ms at the median from a server in Washington, D.C., network included.
When is self-hosting Laya cheaper than Jev's API?
Above roughly 20 million decisions a month, on hardware cost alone. Jev bills $0.042 per million input tokens, and a typical request reads about 400 tokens, so a million decisions cost about $16.80. A GPU left running all month at an assumed $0.50 an hour costs $365, the price of about 21.7 million such Jev calls, or eight a second around the clock. Laya reports 103 to 332 questions a second batched on a T4, so one GPU covers that load.
Below that volume Jev is cheaper before anyone's time is counted, and on our items the cheaper model is only good enough for short option lists. Put your own GPU price and token count into the same sum; the per-request token floor matters for small inputs.
Can you switch between Jev and Laya without changing code?
Mostly. Laya's server speaks the same request shape on the same path:
pip install "laya[serve]" laya-serve # 0.0.0.0:8000, POST /v1/systemone, same body as Jev
Point a Jev client's base URL at it and requests go through, with three differences its README lists: the option budget instead of Jev's 255-option cap, a description required for every Score level, and the confidence definition above. Run both on a few hundred of your own items before switching; the table above is why.
When should you choose Laya over Jev?
- The data cannot leave your infrastructure. There is no hosted Laya API; it runs where you install it. Jev's retention terms are on Jev data retention.
- Your questions have a handful of options and you can label data. Its makers report 0.766 after fine-tuning against 0.362 zero-shot on their decision benchmark.
- Volume is in the tens of millions a month and latency matters more than the last points of accuracy.
Otherwise Jev: dozens of options, no training data, a confidence you can route on, and 64k tokens of context against 512 on Laya's English checkpoint and 8,192 on its multilingual one. Other open models, some closer to Jev on accuracy, are on open-source Jev alternatives.
Jev vs Laya: common questions
Is Laya free?
The code and weights are Apache 2.0, so there is no licence fee and no per-call price. You pay for the machine it runs on and for the work of running it; there is no hosted Laya API from its maker.
Does Laya need a GPU?
No. We ran all 633 items on an Apple M2 laptop at 163 ms per question at the median. A GPU is what gets it to the 33 to 39 ms Laya reports on a T4.
Is Laya made by TypeSafe?
No. Laya is from Convai Innovations and was first published on 2026-09-18. It speaks the same request shape as Jev, which is why the two get compared, but the models, training data and companies are separate.