We compared four language models on one task: reading purchase orders and extracting the order data. We wanted to choose the model our demo uses by default. The question was: which is the cheapest model that never accepts a problem order on its own? A problem order is, for example, one without a buyer, with a delivery date before the order date, or with a negative price. Such an order must go to a person or be rejected.
#TL;DR
- We built a demo,
order-intake-agent, to find the cheapest language model that reads purchase orders well enough and never accepts a problem order on its own. - The demo takes a purchase order as plain text. A model extracts the order data, and code checks it against a schema and business rules. Then a person reviews it in a queue.
- We compared four models:
claude-haiku-4-5through Anthropic's API, andqwen/qwen3.5-9b,qwen/qwen3.8-27bandgoogle/gemma-4-26b-a4b-qaton a local OpenAI-compatible server. - Every model read the same 29 synthetic purchase orders, with the same prompt, schema and tools.
- No model accepted an order that should have gone to a person or been rejected. The models differed in how often they accepted an order with a wrong value inside: from 0 for
qwen/qwen3.8-27bto 0.1154 for gemma. - We picked
qwen/qwen3.5-9bas the default model.claude-haiku-4-5read best and was fastest, but it costs 0.0051 USD per order, and the local models have no price per token. Among the local models,qwen/qwen3.5-9banswers in 3.47 seconds at the median, against 12.15 forqwen/qwen3.8-27b, and it accepted 2 of 26 orders with a wrong value. - During the comparison, we made the test set harder, added a metric for accepted orders with wrong values and fixed how failed answers were counted.
#How the demo works
The demo is called order-intake-agent. The model is only the first step. It reads a purchase order as plain text and proposes the order data. It can also call two read-only lookup tools, one for customers and one for products. Then the answer goes through three more steps:
- a schema check, which rejects any answer that does not have the shape of an order
- business rules, which check whether the order can go ahead, for example whether the dates make sense, and give it a status:
accepted,needs_revieworrejected - a review queue, where a person approves or corrects the order
The model proposes and the code decides. The model cannot write anything, and the demo has no ERP connection. The data is synthetic and the repository is private, so we cannot link to the numbers, but the method below is complete.
A model can misread a field, and the rules can still catch the order. That is why we compare what comes out of the whole pipeline, not only the model's raw answer.
The mistake we care about most is an order that should go to a person or be rejected, but gets accepted instead. We track it as wrong_accept_rate, and its limit is zero. Every limit in the repository comes with a written reason. This one says: "There is no acceptable rate above zero, so this one has no headroom."
#How we compared the models
Each test case is a JSON file with the order text, a fixed date for "today", the expected order data and the expected status. A runner sends every case to a model, runs the answer through the same checks as the demo and writes a report. The report has six metrics, and each one has a limit. If the default model crosses a limit, the build fails. Four metrics appear in this article:
field_accuracy: the share of all fields that the model read correctlydocument_exact_match: the share of orders with every field correctstatus_agreement: the share of orders that got the expected statusaccepted_with_wrong_field_rate: the share of orders that were accepted with at least one wrong field
The demo supports Anthropic's API and any OpenAI-compatible server behind one interface. Every model gets the same orders, prompt, schema and tools.
Each report carries a fingerprint. It is a hash of the model name, the prompt, the schema, the tool descriptions and the sampling parameters, such as temperature. A CI test recomputes it from the code. If someone changes the prompt and does not run the evaluation again, the build fails. For the same reason, the default model is set in code and not in an environment variable.
The client sends every sampling parameter explicitly, so a local server cannot fall back to its own hidden settings. We set repeat_penalty to 1.0, because the common default of 1.1 punishes repetition, and JSON is repetition. On the default model, repeated runs gave the same result for every order.
#What we changed during the comparison
The first test set was too easy
The first set had 21 orders with variants of a simple table: another date format, a decimal comma, units written as words, OCR noise. claude-haiku-4-5 (hosted) and qwen/qwen3.8-27b (local) both scored 1.0 on every metric that has a limit. google/gemma-4-26b-a4b-qat passed every limit with one order to spare. The set could no longer tell us which model to pick.
We added eight harder orders. They check whether the model sees what in an order is not order data: a totals row, a cancelled line, an old order quoted under a reply, four reference numbers next to the real one, an ambiguous date format, a description wrapped over two rows, a discount written as a negative price and a quantity written with a space.
On the default model, qwen/qwen3.5-9b, document_exact_match fell from 1.0 to 0.8846 and status_agreement from 1.0 to 0.9655. The same three orders failed in every run.
We added a metric for accepted orders with wrong values
An order can come out accepted with a wrong value inside. The rules check whether the order can go ahead. They cannot check whether the model copied the quantity correctly. The status looks fine, so the order goes to a person as a proposal to approve.
Our first five metrics did not show this case. wrong_accept_rate compares statuses, and the status was right. field_accuracy counts one wrong field out of 378. document_exact_match counts the order as wrong, but does not say that it was accepted.
The sixth metric, accepted_with_wrong_field_rate, counts these orders. It is calculated over the same orders as document_exact_match. Together, the two metrics split the orders into three groups: read correctly, read wrongly but not accepted, and read wrongly and accepted.
We fixed how failed answers were counted
When an answer failed the schema check, the runner recorded that the order had no expected fields. So the order disappeared from the calculation of document_exact_match and accepted_with_wrong_field_rate. A model that gave no usable answer scored better than a model that read the order almost right. The number of orders in the calculation also differed between models.
Now a failed answer counts as an order with every field wrong. Only gemma was affected, because it was the only model that ever failed the schema check, once. After the fix and a new run, its field_accuracy went from 0.9694 to 0.9259 and its accepted_with_wrong_field_rate from 0.16 to 0.1154. Every model is now scored on the same 26 orders.
#The results
All four models read the same 29 orders. 26 of them have expected order data and count for the field metrics. The other three have no valid order data on purpose and should be rejected.
| Model | field_accuracy | document_exact_match | accepted_with_wrong_field_rate | status_agreement | Median latency | Tool rounds per order |
|---|---|---|---|---|---|---|
claude-haiku-4-5 | 0.9974 | 0.9615 | 0.0385 | 1.0 | 3.124 s | 2 |
qwen/qwen3.5-9b (default) | 0.9894 | 0.8846 | 0.0769 | 0.9655 | 3.4715 s | 1 |
qwen/qwen3.8-27b | 1.0 | 1.0 | 0 (inferred) | 1.0 | 12.1546 s | 1 |
google/gemma-4-26b-a4b-qat | 0.9259 | 0.8077 | 0.1154 | 0.9655 | 3.5988 s | 1 |
wrong_accept_rate is 0.0 for all four models. Twelve of the 29 orders should not be accepted, and no model accepted any of them. So wrong_accept_rate cannot tell these models apart. accepted_with_wrong_field_rate can.
The qwen/qwen3.8-27b report was made before the sixth metric existed, so its value is inferred, not read. Its document_exact_match is 1.0, so no field was wrong and no order was accepted with a wrong value.
Only haiku has a price per token: 0.0051 USD per order at list price. The other three run on our own hardware. We count their running cost as zero, because the tested models are light enough for that hardware.
#Three orders from the test set
Case 26-ambiguous-date-format contains these two lines:
Order date: 03/09/2026
Requested delivery: 15/09/2026There is no fifteenth month, so the format is day first and the order date is 3 September. claude-haiku-4-5 read the order date as 9 March and the delivery date as 15 September. A March date breaks no rule, so the order came out accepted with the wrong date. It is the only order haiku got wrong.
The default model got the same order wrong in a different way. It returned 0309-03-20 and 1509-09-15. The rules flagged an order date that old, so the order went to needs_review. Both models were wrong, but here the pipeline sent the error to a person as a question.
Case 29-two-currencies-and-spaced-quantity has a note, "our catalogue lists prices in USD. This order is placed in EUR.", and one item line:
1 CHEM-777 Thinner 1 200 LTR 12.40Every model got the currency right. The default model read the quantity as 1, and the order came out accepted. A person would see a proposal for one litre instead of 1 200.
The third order the default model got wrong was 27-wrapped-description. It copied a description that runs over two rows together with the line break and the padding spaces. The status was accepted again.
#Which model we picked
Every model passed the test from the title: none of them accepted a problem order. So that test alone could not pick a model. The choice rested on three other numbers: accepted orders with wrong values, speed and cost.
We made the choice in two steps. First, hosted or local. claude-haiku-4-5 read best, with 1 of 26 orders accepted with a wrong value, and it was the fastest, at 3.124 seconds. But it costs 0.0051 USD per order, and the local models have no price per token. So we chose a local model.
Second, which local model. qwen/qwen3.8-27b accepted no orders with a wrong value, but its median answer takes 12.15 seconds, against 3.47 for qwen/qwen3.5-9b. gemma was about as fast as qwen/qwen3.5-9b, but it accepted 3 orders with a wrong value instead of 2. We chose qwen/qwen3.5-9b and set the limit for accepted_with_wrong_field_rate at 0.08. That allows the two orders we know about. If a third one appears, the build fails.
The harder test set also changed the ranking. On 21 orders, the 27B local model and the hosted model looked the same. On 29, they did not. The results come from 29 synthetic purchase orders in one prompt format, so they tell us which model to use for this task in this demo, not which model is better in general.
AKENA is an engineering studio working across AI, blockchain, and software infrastructure. We take projects from first principles to production and operate what we ship. If you are choosing a model for a production task and want the choice to rest on a measurement, we should talk.

