Launch offerup to +40% credits on every packends inClaim →

Tested · October 2026

The best LLM for text classification in 2026, tested

We gave 17 models the same 627 labelled items, from news topics to bank intents with 77 classes, and counted what each got right. The best LLM for text classification depends on the task, so here is the winner for each kind, with what it costs per 1,000 items.

Quick answer · verified 2026-10-11

What is the best LLM for text classification?

For text classification, decision models (built to answer a typed question with one option) beat general chat LLMs where it matters most: with dozens of classes the best one scored 92.0% against 80.5% for the best chat model. With a few classes most models land between 90% and 97%, so pick on price and on whether you need a confidence for each answer.

A few classesd1, 96.5% — news topics and support tickets
Dozens of classesClef Flash, 92.0% — 59 and 77 intents
Yes or noClef, 97.0% — toxicity and review sentiment
Best chat LLMClaude Haiku 5.5, 95.5% — a few classes

How we tested

The items are a fixed-seed sample of 600 texts from BTZSC, a public zero-shot classification benchmark (news topics, tweet emotions, 59 voice-assistant intents, 77 bank intents, toxic comments, review sentiment), plus 27 support tickets we wrote with one decisive signal each. Every model saw the same instruction and the same options for each item. Chat LLMs ran at temperature 0 with reasoning off and were told to reply with one option; a reply that named no option counted as wrong. Decision models answered through the Decisions API. Scripts and raw answers are in jev-measured.

Every model on the same items

Sorted by accuracy on tasks with a few classes. Cost is what the gateway billed, per 1,000 items; only decision models and Laya return a probability with each answer.

ModelTypeFew classes59–77 classesYes / noPer 1,000Confidence
d1Decision model96.5%86.0%96.5%$0.0078yes
Decider V1.1 27BDecision model96.0%85.0%95.5%$0.0072yes
Mercury Decide (free)Decision model95.6%85.0%94.0%freeyes
Claude Haiku 5.5Chat LLM95.5%78.5%96.0%$0.034no
ClefDecision model95.0%89.0%97.0%$0.154yes
Kev 4BDecision model95.0%78.5%91.0%$0.011yes
Qwen3.8 FlashChat LLM95.0%77.0%95.5%$0.023no
Solar DecideDecision model94.6%refused93.0%$0.022yes
GPT-6 Luna DecisionsDecision model94.1%79.8%93.5%$0.040yes
Solar Decide FlashDecision model94.1%refused91.5%$0.022yes
Jev 1.13Decision model94.0%79.0%95.0%$0.030yes
DeepSeek V4.1 FlashChat LLM93.6%80.5%95.5%$0.022no
Clef FlashDecision model92.6%92.0%97.0%$0.058yes
GPT-6 Luna (chat)Chat LLM92.5%78.5%95.0%$0.024no
Tev1 4B ExperimentalDecision model90.4%refused93.5%$0.0072yes
Nemotron 3 Ultra (free)Chat LLM89.4%78.0%92.5%freeno
Laya (ModernBERT)Local, open weights82.5%44.5%79.0%your hardwareyes

With 100 items a set, models a point or two apart are not reliably different; read the table in tiers. d1 leads the few-class column, and chat LLMs (shaded) sit between 89.4% and 95.5%, inside the same tier as most decision models.

Best LLM for intent classification with many classes

Intent classification is where the field spreads out. With 59 and 77 intents, Clef Flash scored 92.0%; the chat LLMs managed 77.0% to 80.5%. Most chat misses were neighbouring labels: ordering a physical card read as getting one, a calendar entry read as an alarm, a missing transfer read as a slow one. A sentence of description per option fixes most of that. Three decision models refuse more than 20 or 26 options outright, and the small local model falls to 44.5%: long label lists are the hard case for everything small.

Best for yes or no: sentiment and toxicity

Binary questions are the easy end. Every hosted model averaged between 91% and 97% on toxic comments and review sentiment, Clef highest at 97.0%. Here the choice is cost and a usable probability: a gate that blocks a comment wants to know how sure the model is, which a chat reply does not say.

Chat LLM or decision model?

The cleanest comparison is one model family through both doors. GPT-6 Luna through the chat endpoint scored 92.5% on few classes and 78.5% on many, for $0.024 per 1,000 items; GPT-6 Luna Decisions scored 94.1% and 79.8% for $0.040. Accuracy is close and the chat route was cheaper here. What the decision endpoint adds is a probability for every option, so you can act on sure answers and send the rest to a person.

Best free and open-weight options

Two free options held up. Mercury Decide, a decision model with a free tier, scored 95.6% on few classes (treat its emotion score with care: it looks like it saw that dataset in training). Nemotron 3 Ultra, free on OpenRouter, scored 89.4% and 78.0%, though 19 of its replies named no option at all. Clef, Clef Flash and Kev 4B publish their weights. Run locally, Laya matched the hosted models on news topics (91.0%) but not on intents, details in Jev vs Laya.

Best LLM for text classification: questions

Is GPT or Claude better at text classification?

Among the chat models we ran, Claude Haiku 5.5 led on few-option tasks (95.5%) and tied GPT-6 Luna on intents with many classes. Their flagships, GPT-6.1 Sol and Claude Sonnet 5.5, only run with reasoning on; we tried 21 items each and did not measure the full set, so this page makes no claim about them.

Should I fine-tune a model instead of using an LLM?

If you have thousands of labels for one task that will not change, a fine-tuned classifier such as BERT is usually more accurate and far cheaper to run. LLMs and decision models win before you have labels, when classes change, and on the long tail.

How should I test models on my own data?

Label 100 to 200 real items per task, send every model the same instructions and options, and count a reply that names no option as wrong. Read the misses: on our intent sets most errors were near-synonym labels, which better option descriptions fix.