Use case · Mixed
CSV data validation with Jev: the checks code cannot write
Let code check formats and types, and ask Jev what code cannot answer, row by row: is this value plausible, does this email belong to this person, what is wrong here.
Run this example
3 questions · no key neededThe problem
Schema validators check types, ranges and patterns, and stop there. "ana.souza@exmaple.com" passes every email pattern, "Brasil" is a valid string, and a mouse priced like a laptop is a valid number. Rows like these reach the warehouse and break a report weeks later.
How Jev handles it
Run the schema checks in code first. Then send each remaining row, with its column names and the table's purpose, as the state, and ask one Noul per semantic check plus, if you want one label for the row, a Choice for its most serious problem.
header = "order_id, customer_name, email, country, product, quantity, unit_price_usd, order_date"
state = (
"TABLE: customer_orders.csv, one row per order.\n"
f"COLUMNS: {header}\n"
f"ROW: {', '.join(row)}"
)
response = client.system_one(
model="jev-latest",
state=state,
questions={
"valid": Noul(
instructions="Does every field in this row look plausible for its column?",
criteria=NoulCriteria(true="Every field looks plausible", false="At least one field looks wrong"),
),
"issue": Choice(
instructions="Which problem, if any, is the most serious in this row?",
criteria={
"none": "No problem",
"email": "The email address looks mistyped or fake",
"country": "The country is misspelled or not in English",
"date": "The date is impossible or malformed",
"product_mismatch": "The product does not fit its price or quantity",
"missing": "A required value is empty",
},
),
"name_matches_email": Noul(
instructions="Does the email address plausibly belong to the named customer?",
criteria=NoulCriteria(true="Plausibly the same person", false="Looks like someone else"),
),
},
)
if response.answers["valid"].noul < 0.5:
quarantine(row, reason=response.answers["issue"].choice)State sent the example loaded in Run this example above, as it first appears.
Answers returned
- Latency (median of 3)
- 633.6ms
- Minus network floor (259.9ms)
- ≈374ms
- Input tokens
- 576
- Cost
- $0.00002419
Three planted problems: a misspelled email domain, a non-English country name, and 30 February.
Which checks belong to CSV data validation with Jev
The measured row hides three problems: a misspelled email domain (exmaple.com), a country written in Portuguese (Brasil) and 30 February. Jev returned 0.03 on "every field looks plausible", so the row is quarantined, and named the date as the most serious problem with all of the probability, 1.00, leaving 0 on the email.
That is the design lesson: a Choice reports one problem, so a row with three gets one label. When every problem has to be caught, ask one Noul per check, "the email domain looks mistyped", "the country is not written in English", and keep the Choice as a summary label for the quarantine queue.
Leave formats to code even though Jev caught the date here. A date parser rejects 30 February every time, exactly and for free, and dates and numbers are among Jev's documented weak spots; a validator should not be probabilistic where it does not need to be.
What to put in the state
Send the column names with every row, and one line on what the table is. "One row per order" tells Jev that quantity is a count of items and unit_price_usd is a price, which is what plausibility depends on.
One row per call keeps answers attached to rows. For rows that must agree with each other, such as an order and its refund, put both in one state and ask about the pair.
What one decision costs
Validation runs per row, so cost scales with the file. First line: the measured row, three questions. Second: a 10,000-row import. Third: 1,000,000 rows a month.
# the measured row, three questions 576 tokens × $0.042 / 1M = $0.0000242 # a 10,000-row import 10,000 × 576 = 5,760,000 tokens × $0.042 / 1M = $0.242 # 1,000,000 rows a month 1,000,000 × 576 = 576,000,000 tokens × $0.042 / 1M = $24.19
Send only the rows that pass the schema checks and trip a cheap rule, such as an unusual domain or a price outside its usual band; most rows in a healthy file never need a model.
When not to use Jev for this
- The rule is a format, a type or a range. A schema validator is exact, free and instant.
- The check spans the file, such as duplicate IDs or totals across rows. That is a query, not a per-row judgment.
- Every row must pass a check of record. Jev returns probabilities; use it to surface suspicious rows for a person.
- The data is sensitive. Send only the columns a check needs and leave out identifiers that do not change the answer.
Where it fits in your stack
CSV data validation with Jev sits after the schema validator and before the load: parse, run type and format checks in code, send the survivors to Jev, quarantine rows under the threshold with their issue label, and load the rest.
Quarantined rows that a person fixes are labels. They show which Noul checks fire usefully and which need rewording.
Notes from the field
- Code first, Jev second: formats are free to check exactly.
- One Noul per check when every problem matters; a Choice names only one.
- Column names and the table's purpose belong in the state.
Common questions
Can AI validate CSV data?
It can judge what a schema cannot: whether an email looks mistyped, a value is plausible for its column, or a name matches an email. On the measured row Jev flagged the row and named the impossible date. Formats and types are still exact and free in code.
Did it find all three planted problems?
It flagged the row and named the date, the one problem a Choice can report. Asked whether the email belongs to the customer it said yes at 0.82, which is right about the name and says nothing about the misspelled domain; that needs a question of its own.
Is this a CSV validation tool?
It is the semantic step of one. Run the file through a schema validator, then send the rows that pass through these questions; the example at the top of this page runs one row live.