# 99% Accurate Can Still Mean Half Your Documents Are Wrong

Source: https://www.digiparser.com/blog/document-ai-accuracy-trap

[See all posts](/blog)

Last updated on September 23, 2026

# 99% Accurate Can Still Mean Half Your Documents Are Wrong

Research

Document AI

Data quality

[![Pankaj Patidar](https://avatars.githubusercontent.com/u/17493609?v=4)

Pankaj Patidar

@thepantales



](https://x.com/thepantales)

![99% Accurate Can Still Mean Half Your Documents Are Wrong](https://www.digiparser.com/research/document-accuracy-trap/accuracy-gap.png)

**Two document-processing systems can both report 99% field accuracy while one produces errors in 1% of documents and the other in 50%.**

That is a 50-fold difference in affected documents, hidden behind exactly the same headline score. No questionable survey is needed to demonstrate it. Below are two fully reproducible, constructed datasets with identical document counts, field counts and error counts.

Our position is straightforward: **a document AI accuracy claim without its unit of measurement is an incomplete buying argument.** Finance teams pay to process invoices. Order desks need usable sales orders. Neither buys a basket of mostly correct fields.

This analysis shows the arithmetic, provides downloadable data and sets out the measurements buyers should demand. These are mathematical examples, not observed error rates for DigiParser or any other vendor.

## The same 99%. Fifty times as many affected documents.

Take 1,000 documents with exactly 50 required fields each. There are 50,000 fields to score. Both systems get 49,500 fields right and 500 wrong: **99% field accuracy**.

Now change where the 500 errors land.

Measurement

Errors concentrated

Errors spread out

Documents processed

1,000

1,000

Required fields per document

50

50

Incorrect fields

500

500

Field accuracy

99%

99%

Error placement

All 50 fields wrong in 10 documents

One field wrong in 500 documents

Documents with at least one error

10

500

Share of documents with errors

1%

50%

Completely correct documents

990

500

Both datasets are mathematically valid. The concentrated case represents an extreme pattern of whole-document failure; the spread case distributes errors across many more records. We use the extremes to expose what an average cannot tell you, not to suggest either is typical.

> With 50 required fields per document, the same 99% field accuracy can conceal a 50-fold difference in the number of documents containing errors.

If every affected document needs someone to open it, check it and save a correction, the second system creates far more document-level interruptions. That does not make its total correction cost automatically 50 times higher: repairing 50 fields can take longer than repairing one. **Count affected documents and correction effort separately.**

[Download the two constructed datasets as CSV](/research/document-accuracy-trap/constructed-document-errors.csv).

## More line items mean more opportunities to fail

The extremes above require no independence assumption. For a separate planning scenario, suppose every required field has the same probability of being correct and field errors occur independently.

Then:

**Probability of a completely correct document = field accuracy raised to the number of required fields.**

At 99% field accuracy, a 70-field document has only a **49.5%** chance of being completely correct in this model. A 100-field document has a **36.6%** chance.

Required fields per document

Perfect documents at 99% field accuracy

At 99.5%

At 99.9%

10

90.4%

95.1%

99.0%

25

77.8%

88.2%

97.5%

50

60.5%

77.8%

95.1%

70

49.5%

70.4%

93.2%

100

36.6%

60.6%

90.5%

200

13.4%

36.7%

81.9%

These are calculated probabilities, not measured industry performance. The model follows the independent-trial framework described in the [NIST binomial distribution reference](https://itl.nist.gov/div898/handbook/eda/section3/eda366i.htm).

A constructed purchase order with 10 header fields and 20 line items containing three fields each already has 70 fields. A compact-looking document can therefore have a large error surface. For a manufacturer, those fields might include customer part number, quantity and unit of measure. For an AP team, the corresponding invoice fields could include description, quantity and line total. The workflows differ, but both need more than a correctly read grand total.

[Download all probability scenarios as CSV](/research/document-accuracy-trap/independent-field-model.csv).

## A 99.9% score can still miss a 95% perfect-document target

Suppose a buyer wants 95% of 100-field documents to be completely correct. Under the same independent-field model, the required field accuracy is:

**0.95 raised to the power of 1/100 = 99.9487% field accuracy.**

That is higher than 99.9%. In the model, 99.9% produces only 90.5% perfect 100-field documents.

> The accuracy target belongs to the document and the workflow. The field score is an input to that target, not a substitute for it.

Real errors often cluster. A poor scan can damage several fields at once; one shifted table column can corrupt many rows. Some fields are also harder than others. That is why the model is useful for exposing the gap, while a representative document-level test is needed to measure your actual result.

## The review bill that the accuracy slide leaves out

Use the 70-field, 99% scenario with 10,000 documents per month. It predicts approximately **5,052 documents with at least one incorrect field**.

If every one of those documents is detected and takes two minutes to review and correct, the work totals **168 hours per month**. The volume, handling time and perfect detection are explicit scenario assumptions. This is not a measured labour benchmark.

A review system that catches fewer errors may report a smaller queue while allowing more mistakes through. A system that flags correct documents may create a larger queue without improving the extraction itself.

**A small review queue is not proof of accurate automation.** It can mean strong extraction, weak error detection, or a mix of both.

Google's [Document AI evaluation documentation](https://docs.cloud.google.com/document-ai/docs/evaluate) explains why evaluation must include precision, recall and the confidence threshold: predictions below the threshold are excluded by its evaluation logic even when they are correct. A threshold changes what gets accepted. It does not make uncertainty disappear.

## Five numbers to demand in your next document AI pilot

Ask every provider, including DigiParser, to report these on the same representative test set:

1.  **Required-field accuracy.** Correct required values divided by all required values. Missing values count as failures. Report extra rows and invented fields separately.
2.  **Perfect-document rate.** Documents with every required value correct and no unwanted records or rows, divided by all submitted documents. Failed processing stays in the denominator.
3.  **Verified straight-through rate.** Documents completed without human intervention that also pass an independent correctness check and the destination system's business rules. Audit a sample of accepted documents, not only the review queue.
4.  **Review effort.** Documents reviewed, average handling time and total reviewer minutes. Separate review from active correction when possible.
5.  **Critical errors that escaped review.** Wrong quantities, amounts, identifiers or other high-impact values accepted downstream. State both the count and the denominator.

Do not blend precision, recall, F1, character accuracy and required-field accuracy into a single percentage. They answer different questions. Do not compare a provider tested on five invoice headers with one tested on every row of a 200-field order.

For [customer purchase order processing](/solutions/purchase-order-parser), success means the right customer, products, quantities and delivery details reach the sales-order workflow. For [supplier invoice processing](/solutions/invoice-parser), success means the correct bill data reaches AP with the required validation. Neither outcome can be established by extraction accuracy alone.

## The standard we should hold the industry to

A useful document AI demo should end with a scorecard, not an accuracy slogan.

Show the messy documents. Count the missing fields. Keep failed runs in the results. Measure the time spent fixing what the model got wrong. Check a sample of records that the system accepted automatically.

**Stop buying the highest accuracy percentage. Buy the lowest verified cost of getting a correct document into your business system.**

That is the number that connects an AI demonstration to an operating result.

## How we calculated these results

We built two datasets containing 1,000 documents and 50,000 fields each. Both contain exactly 500 field errors. By changing how those errors are distributed, we demonstrate that the same 99% field accuracy can leave either 1% or 50% of documents with errors.

These are the exact endpoints for this example. At least 10 documents must contain errors because each document has only 50 fields. At most 500 documents can contain errors because each affected document needs at least one incorrect field.

The probability table shows what happens when every field has the same accuracy and errors occur independently. At 99% accuracy per field, the probability that all 70 fields are correct is 0.99⁷⁰, or 49.5%. This means multiplying 0.99 by itself 70 times. We round table percentages to one decimal place; the downloadable data keeps ten decimal places.

For the 100-field example, we calculate the field accuracy needed to make 95% of complete documents error-free: approximately 99.9487%. For the review-hours example, we multiply the expected number of documents with errors by two minutes per review, then divide by 60 to convert minutes to hours. We round only the final result.

The probability model follows [NIST's binomial distribution reference](https://itl.nist.gov/div898/handbook/eda/section3/eda366i.htm). For further reading on evaluation metrics and confidence thresholds, see [Google Document AI's evaluation documentation](https://docs.cloud.google.com/document-ai/docs/evaluate).

Analysis by DigiParser. Published 23 September 2026.

## Download and reuse

Download [the document-level dataset](/research/document-accuracy-trap/constructed-document-errors.csv), [the probability table data](/research/document-accuracy-trap/independent-field-model.csv) and [the calculation script](/research/document-accuracy-trap/reproduce.py) to reproduce the results.

You can reuse [the chart](/research/document-accuracy-trap/accuracy-gap.png) and DigiParser's example datasets with attribution and a link to this article. Keep the chart's labels and assumptions so readers can understand the figures.

Suggested citation: DigiParser, "99% Accurate Can Still Mean Half Your Documents Are Wrong," 23 September 2026.

* * *

[See all posts](/blog)

Automate recurring documents next: [supplier invoice parser](/solutions/invoice-parser), [customer purchase order parser](/solutions/purchase-order-parser), and [extract data from PDF](/solutions/extract-data-from-pdf) hub.

## Transform Your Document Processing

Start automating your document workflows with DigiParser's AI-powered solution.

[Start Free Trial](https://app.digiparser.com/auth/join)[Schedule Demo](/contact)