Launch offerup to +40% credits on every packends inClaim →

Use case · Mixed

Invoice classification with Jev: expense type, review flags and urgency

Read the invoice text once and get three typed answers: the expense category, whether a person should check it before payment, and how soon it is due.

Run this example

3 questions · no key needed

The problem

Accounts payable teams code every invoice to an expense category by hand or with supplier rules nobody maintains, and a repeated or padded line slips through because nobody re-reads a familiar supplier's bill. An LLM can read the invoice, but its answer is a sentence to parse, and AI invoice classification without a number that says how sure it is cannot tell you which invoices to book unattended.

How Jev handles it

For AI invoice classification with Jev, send the invoice text after OCR as the state with three questions: a Choice over your expense categories, a Noul for "needs a human check", and a Score for payment urgency. Dates and totals are computed in code and stated as results, so Jev judges the text and code does the arithmetic.

python
from datetime import date

days_left = (invoice.due_date - date.today()).days
state = f"{ocr_text}\nReceived: {date.today():%-d %B %Y} (due in {days_left} days)."

response = client.system_one(
    model="jev-latest",
    state=state,
    questions={
        "category": Choice(
            instructions="Which expense category does this invoice belong to?",
            criteria={
                "cloud_hosting": "Servers, cloud platforms, storage and hosting",
                "software":      "Software licences and SaaS subscriptions",
                "contractors":   "Freelancers and professional services",
                "travel":        "Flights, hotels and transport",
                "office":        "Office supplies, equipment and rent",
                "marketing":     "Advertising, events and promotion",
                "other":         "Anything else",
            },
        ),
        "needs_review": Noul(
            instructions="Should a person check this invoice before it is paid, because a line repeats, "
                         "the totals do not add up, or a charge looks unusual?",
            criteria=NoulCriteria(true="Needs a human check before payment", false="Looks routine"),
        ),
        "urgency": Score(
            instructions="How soon must this invoice be paid?",
            criteria=["No date pressure", "Due this month", "Due within a week", "Due now or overdue"],
        ),
    },
)

a = response.answers
if a["category"].confidence < 0.6 or a["needs_review"].noul > 0.5:
    send_to_ap_queue(invoice, a)
else:
    book(invoice, category=a["category"].choice)
Measured outputtypesafe/jev-1.13 · 2026-10-11

State sent the example loaded in Run this example above, as it first appears.

Answers returned

categorycloud_hosting
confidence100%
cloud_hosting100%
software0%
contractors0%
travel0%
office0%
marketing0%
other0%
needs_review0.87
urgencyDue within a week (2.09)
confidence87%
Due within a week87%
Due now or overdue11%
Due this month2%
No date pressure0%
Latency (median of 3)
632.3ms
Minus network floor (259.9ms)
≈372ms
Input tokens
724
Cost
$0.00003041

A cloud invoice with one line charged twice. The due date is stated as days left, computed in code.

Designing the questions for invoice classification

Use your own expense categories as the Choice options, with descriptions that say what lands in each: "Servers, cloud platforms, storage and hosting", not a ledger code. Seven were enough for the measured cloud invoice, which went to cloud_hosting at 1.00 with every other option at zero. Map the answer to account codes in your own code afterwards.

Make the review flag one Noul that names its reasons, not a Choice of anomaly types. The measured invoice charges the same priority support add-on twice, and the flag came back at 0.87. The Noul reads repeated lines and odd charges; whether the totals add up is arithmetic, so check that in code before the call and state the result.

Urgency is a Score because "due in 3 days" sits between levels: the measured answer was 2.09, with 0.87 of the probability on "Due within a week". Round it to place the invoice in a queue, or read the probabilities when two invoices compete for the same payment run.

What to put in the state

Send the OCR text as it is, line items included: a repeated line can only be seen if both lines are in the state. Add a short header your code computes, the date received and the days until due. Jev reads "due in 3 days"; it should not be asked to subtract dates.

Leave out bank details and anything the three questions do not read. The supplier, the lines and the totals are what the decision needs; account numbers only add exposure.

What one decision costs

Invoice volumes are modest, so the bill rarely decides this. First line: the measured invoice, three questions. Second: 2,000 invoices a month at that size. Third: re-classifying a 50,000-invoice archive once after the categories change.

arithmetic
# the measured invoice, three questions
724 tokens × $0.042 / 1M = $0.0000304

# 2,000 invoices a month at that size
2,000 × 724 = 1,448,000 tokens × $0.042 / 1M = $0.0608

# re-classifying a 50,000-invoice archive once
50,000 × 724 = 36,200,000 tokens × $0.042 / 1M = $1.52

The archive run is the one that matters: when finance splits or renames a category, re-running invoice classification over every past invoice costs about a dollar and a half.

Jev's answers on the measured call — category: cloud_hosting; needs review: 0.87; urgency: 2.09 of 3, $30.41 per million calls

When not to use Jev for this

Where it fits in your stack

Invoice classification sits after OCR and before the accounts payable queue: extraction turns the PDF into text, code computes dates and totals, Jev classifies and flags, and the result posts to the ledger or to a review queue.

When a reviewer overrides a category or clears a flag, keep the pair. A few hundred of them show which category descriptions need rewording and whether the 0.5 review threshold is too strict.

Notes from the field

Common questions

Can AI classify invoices into expense categories?

Yes, from the invoice text. On the measured cloud invoice Jev chose cloud hosting at 1.00 out of seven categories. How well it does on yours depends on how clearly the categories are described, so compare a few hundred answers with what your accountants chose before booking automatically.

Does it catch duplicate invoice lines?

It caught the planted one: the same support add-on charged twice raised the review flag to 0.87. It reads the text, so it notices repeated or odd lines; an invoice you already paid last month is a lookup against your ledger, which code does better.

Is this an invoice processing tool?

It covers the AI invoice classification and review step of one. Extracting fields from a PDF, matching purchase orders and paying are separate steps; Jev answers the judgment calls between them, with a probability on each.