Tested · October 2026
The best LLM for text classification in 2026, tested
We gave 17 models the same 627 labelled items, from news topics to bank intents with 77 classes, and counted what each got right. The best LLM for text classification depends on the task, so here is the winner for each kind, with what it costs per 1,000 items.
Quick answer · verified 2026-10-11
What is the best LLM for text classification?
For text classification, decision models (built to answer a typed question with one option) beat general chat LLMs where it matters most: with dozens of classes the best one scored 92.0% against 80.5% for the best chat model. With a few classes most models land between 90% and 97%, so pick on price and on whether you need a confidence for each answer.
| A few classes | d1, 96.5% — news topics and support tickets |
|---|---|
| Dozens of classes | Clef Flash, 92.0% — 59 and 77 intents |
| Yes or no | Clef, 97.0% — toxicity and review sentiment |
| Best chat LLM | Claude Haiku 5.5, 95.5% — a few classes |
How we tested
The items are a fixed-seed sample of 600 texts from BTZSC, a public zero-shot classification benchmark (news topics, tweet emotions, 59 voice-assistant intents, 77 bank intents, toxic comments, review sentiment), plus 27 support tickets we wrote with one decisive signal each. Every model saw the same instruction and the same options for each item. Chat LLMs ran at temperature 0 with reasoning off and were told to reply with one option; a reply that named no option counted as wrong. Decision models answered through the Decisions API. Scripts and raw answers are in jev-measured.
Every model on the same items
Sorted by accuracy on tasks with a few classes. Cost is what the gateway billed, per 1,000 items; only decision models and Laya return a probability with each answer.
| Model | Type | Few classes | 59–77 classes | Yes / no | Per 1,000 | Confidence |
|---|---|---|---|---|---|---|
| d1 | Decision model | 96.5% | 86.0% | 96.5% | $0.0078 | yes |
| Decider V1.1 27B | Decision model | 96.0% | 85.0% | 95.5% | $0.0072 | yes |
| Mercury Decide (free) | Decision model | 95.6% | 85.0% | 94.0% | free | yes |
| Claude Haiku 5.5 | Chat LLM | 95.5% | 78.5% | 96.0% | $0.034 | no |
| Clef | Decision model | 95.0% | 89.0% | 97.0% | $0.154 | yes |
| Kev 4B | Decision model | 95.0% | 78.5% | 91.0% | $0.011 | yes |
| Qwen3.8 Flash | Chat LLM | 95.0% | 77.0% | 95.5% | $0.023 | no |
| Solar Decide | Decision model | 94.6% | refused | 93.0% | $0.022 | yes |
| GPT-6 Luna Decisions | Decision model | 94.1% | 79.8% | 93.5% | $0.040 | yes |
| Solar Decide Flash | Decision model | 94.1% | refused | 91.5% | $0.022 | yes |
| Jev 1.13 | Decision model | 94.0% | 79.0% | 95.0% | $0.030 | yes |
| DeepSeek V4.1 Flash | Chat LLM | 93.6% | 80.5% | 95.5% | $0.022 | no |
| Clef Flash | Decision model | 92.6% | 92.0% | 97.0% | $0.058 | yes |
| GPT-6 Luna (chat) | Chat LLM | 92.5% | 78.5% | 95.0% | $0.024 | no |
| Tev1 4B Experimental | Decision model | 90.4% | refused | 93.5% | $0.0072 | yes |
| Nemotron 3 Ultra (free) | Chat LLM | 89.4% | 78.0% | 92.5% | free | no |
| Laya (ModernBERT) | Local, open weights | 82.5% | 44.5% | 79.0% | your hardware | yes |
With 100 items a set, models a point or two apart are not reliably different; read the table in tiers. d1 leads the few-class column, and chat LLMs (shaded) sit between 89.4% and 95.5%, inside the same tier as most decision models.
Best LLM for intent classification with many classes
Intent classification is where the field spreads out. With 59 and 77 intents, Clef Flash scored 92.0%; the chat LLMs managed 77.0% to 80.5%. Most chat misses were neighbouring labels: ordering a physical card read as getting one, a calendar entry read as an alarm, a missing transfer read as a slow one. A sentence of description per option fixes most of that. Three decision models refuse more than 20 or 26 options outright, and the small local model falls to 44.5%: long label lists are the hard case for everything small.
Best for yes or no: sentiment and toxicity
Binary questions are the easy end. Every hosted model averaged between 91% and 97% on toxic comments and review sentiment, Clef highest at 97.0%. Here the choice is cost and a usable probability: a gate that blocks a comment wants to know how sure the model is, which a chat reply does not say.
Chat LLM or decision model?
The cleanest comparison is one model family through both doors. GPT-6 Luna through the chat endpoint scored 92.5% on few classes and 78.5% on many, for $0.024 per 1,000 items; GPT-6 Luna Decisions scored 94.1% and 79.8% for $0.040. Accuracy is close and the chat route was cheaper here. What the decision endpoint adds is a probability for every option, so you can act on sure answers and send the rest to a person.
- Pick a chat LLM when the label set is small, you already call that provider, and a wrong label costs little.
- Pick a decision model for many classes, for thresholds that trigger actions, or for millions of items: 13 decision models compared, with calibration and speed.
- Fine-tune a classifier when one task has thousands of stable labels; see Jev vs BERT.
Best free and open-weight options
Two free options held up. Mercury Decide, a decision model with a free tier, scored 95.6% on few classes (treat its emotion score with care: it looks like it saw that dataset in training). Nemotron 3 Ultra, free on OpenRouter, scored 89.4% and 78.0%, though 19 of its replies named no option at all. Clef, Clef Flash and Kev 4B publish their weights. Run locally, Laya matched the hosted models on news topics (91.0%) but not on intents, details in Jev vs Laya.
Best LLM for text classification: questions
Is GPT or Claude better at text classification?
Among the chat models we ran, Claude Haiku 5.5 led on few-option tasks (95.5%) and tied GPT-6 Luna on intents with many classes. Their flagships, GPT-6.1 Sol and Claude Sonnet 5.5, only run with reasoning on; we tried 21 items each and did not measure the full set, so this page makes no claim about them.
Should I fine-tune a model instead of using an LLM?
If you have thousands of labels for one task that will not change, a fine-tuned classifier such as BERT is usually more accurate and far cheaper to run. LLMs and decision models win before you have labels, when classes change, and on the long tail.
How should I test models on my own data?
Label 100 to 200 real items per task, send every model the same instructions and options, and count a reply that names no option as wrong. Read the misses: on our intent sets most errors were near-synonym labels, which better option descriptions fix.