Skip to main content
Document Accessibility Institute
Menu

How we measure

Everything below is what our software actually does, written so that a records clerk, an IT director or a council member can check it. Where we estimate, we say so and we show the sample.

1. What we look at

Only the public website. We never log in, never request access, never install anything. We fetch the same pages a resident would open, follow the links on them, and note every document they point to. We identify ourselves in every request, so a system administrator can see us in the logs and block us. We keep to about two requests per second and we respect the site's robots file.

2. What we count

A document is a PDF file linked from the site. We stop crawling at a fixed number of pages, so on very large sites the count is a floor, not a ceiling, and we say so on the city page. Word, Excel and PowerPoint files are on the roadmap and are not counted in the pilot edition.

3. What we test

We download a random sample of the documents found, usually two hundred, and check each one for:

  • A tag tree. The internal structure that tells a screen reader what is a heading, a paragraph, a list, a table. Without it the reader gets a flat stream of words, or nothing.
  • Real text. A scanned page is a picture. A screen reader cannot read a picture. We check whether the first pages contain text at all.
  • A declared language and a title. Small things that decide whether a reader pronounces the document correctly and announces it by name.
  • Whether it is still online. Dead links are removed from the estimate.

These are the checks that can be run by software with confidence. They are necessary conditions for WCAG 2.1 Level AA and PDF/UA (ISO 14289-1), not the whole of either standard. A document that passes all of them can still fail a human review. A document that fails the first one cannot pass.

3a. The second opinion

Every sampled document is also run through veraPDF, the open-source reference validator maintained by the PDF Association, against the PDF/UA-1 profile. We publish its failure rate next to ours. In the pilot the two agreed on every document we called untagged; veraPDF additionally fails most of the documents that do have a tag tree, for reasons a tag tree alone cannot show. That gap is real, and it is why our repair work ends with a person and a screen reader, not with a checker.

4. How we grade

The grade uses one number: the share of live documents in the sample that have no tag tree. The scale is fixed and applies to every city the same way.

AUnder 10% of documents lack screen-reader structure
B10% to under 30%
C30% to under 55%
D55% to under 80%
F80% or more

5. What we do not claim

  • We do not claim that a document we repaired is “compliant” in a legal sense. We state which checks it passes, and a person reviews every file we return.
  • We do not claim that our count is your legal obligation. The Title II rule exempts documents that were already published and are not used to obtain a service. Only a review of your own site can settle which documents that covers.
  • We are not a government body and we have no authority over anyone. Nobody is required to buy anything from us.

6. Corrections

If a figure on a city page is wrong, write to reports@documentaccessibilityinstitute.org. We re-run the measurement, publish the new result with the date, and keep the old one visible underneath it.

See the cities