AI alt text accuracy: what our test of 124 alt texts found

AI can judge alt text well, but not flawlessly, and its setup matters as much as the model. On 124 labeled alt texts, the vision check we run in production left 97.8% of correct alt text alone and flagged 98.4% of planted errors. Most of its misses were look-alike patterns and ambiguous details.

Published . Last reviewed .

This is original research from GotAlt. We built a test set of product photos with alt text whose right answer was known in advance, ran it through the same code that powers our alt text checker, the deep audit and the paid weekly audit, and recorded every verdict. This page reports what was measured, the exact numbers, the mistakes the model made, and the limits of a test this size. Every figure below comes from our own evaluation records of October 10 and 11, 2026.

Key findings

  • Correct alt text was mostly left alone. Under the production setting, 182 of 186 verdicts on true alt text (97.8%) were accurate: 62 alt texts, three runs each.
  • Planted errors were mostly caught. 183 of 186 verdicts on planted problems (98.4%) flagged them, including all 23 wrong-colorway alt texts in every run (69 of 69 verdicts).
  • Settings moved the results more than the choice of model. The same model at a lower reasoning setting wrongly accused 8.6% of correct alt text. Showing it the alt text before the image cut planted errors caught to 90.3%.
  • Most misses were the hard cases a person would also pause on: a leopard print versus a tortoiseshell, a cartoon headband that reads as cat ears or a crown, which part of a case counts as "black". Two were plain slips: an empty alt on a product photo and a description on a decorative bar, each let through once in three runs.
  • This test measured judging alt text, not writing it. We did not measure how accurate AI-generated alt text is, including from our own alt text generator.

Can AI write good alt text?

Two different jobs hide in that question. Writing alt text means looking at an image with no answer in hand and producing a description. Checking alt text means comparing a description someone already wrote against the image and deciding whether it is true. Our test measured the second job, because that is what our checker does, and because it can be graded precisely: each alt text was either true of its photo or deliberately wrong.

The results say a current vision model, set up carefully, is a strong judge of whether a description matches a picture. They also show the kinds of things it gets wrong, and those are the same visual confusions a model can make when it writes a description from scratch. So the practical answer is: AI can draft good alt text, and AI can catch a lot of bad alt text, but neither is a reason to stop looking. Alt text exists to meet WCAG 1.1.1 Non-text Content (see our plain-English explainer), and that criterion asks for a text alternative that serves the same purpose as the image. Only someone who knows why the image is on the page can confirm that.

Results at a glance

Production setting: Anthropic's Claude Haiku 5.5 vision model at medium reasoning effort, with each image introduced by its number and its alt text given after the image. That is the configuration our alt text checker, deep audit and paid weekly audit use today. Every page of up to six images ran three times.

Summary: GotAlt alt text verdict test, production setting, October 2026 (counts are verdicts pooled over three runs unless labeled otherwise)
Measure Result Count 95% interval
Correct alt text left alone97.8%182 of 18695 to 99%
Planted problems flagged98.4%183 of 18695 to 99%
Exact verdict (right kind of problem)98.1%365 of 37296 to 99%
Same verdict in all three runs96.0%119 of 124 alt texts91 to 98%
Pages with no answer, a refusal or an error00 of 63 page runsNot applicable

"Correct alt text left alone" is the measure we care about most. A check that accuses true alt text wastes the site owner's time and teaches them to ignore it, so our prompt tells the model to mark an image accurate when it is genuinely unsure. "Planted problems flagged" counts any flag on a bad alt text; "exact verdict" also requires the right label, for example calling a filename inaccurate rather than vague.

Results by kind of alt text

The 124 alt texts fell into nine kinds. For true alt text the table shows the share left alone; for planted problems, the share flagged. Small groups have wide intervals: two decorative images, three runs each, give only six chances.

Production setting results by kind of alt text (three runs each)
Kind of alt text Should be Result Count (verdicts over three runs)
True, full description (30)Left alone95.6%86 of 90
True, short description (30)Left alone100%90 of 90
Empty alt on a decorative image (2)Left alone100%6 of 6
Wrong color, copied from another colorway (23)Flagged100%69 of 69
Wrong pattern, copied from another design (7)Flagged95.2%20 of 21
A filename as alt text (10)Flagged100%30 of 30
Vague label, such as "Phone case" or "Phone accessory" (10)Flagged100%30 of 30
Empty alt on a real product photo (10)Flagged96.7%29 of 30
A description on a purely decorative image (2)Flagged83.3%5 of 6

The pattern is worth noticing. Under the production setting, short true alt text ("Matte pale pink phone case") was never accused. Every miss on true alt text came from the long, detailed descriptions, which make more claims and give the model more to disagree with. That is an argument for the advice in our alt text guide: say what matters, not everything you can see.

Settings changed the results more than the model did

We ran the same 124 alt texts through seven configurations. Differences of one or two verdicts out of 186 are within noise; the intervals overlap. Larger gaps are real.

How configuration changed accuracy on the same 124 alt texts (three runs each)
Configuration Correct alt text left alone Planted problems flagged Used in production?
Medium effort, number before each image, alt text after it97.8% (182 of 186)98.4% (183 of 186)Yes, since October 11, 2026
Medium effort, number and alt text both after each image (previous layout)98.9% (184 of 186)97.8% (182 of 186)Replaced (see below)
Medium effort, alt text given before the image94.6% (176 of 186)90.3% (168 of 186)No
Low effort, previous layout91.4% (170 of 186)94.6% (176 of 186)No
Medium effort, previous layout, extra rule to flag alt text that omits information98.4% (183 of 186)95.7% (178 of 186)No
A second vendor's vision model, medium effort, previous layout99.5% (185 of 186)95.7% (178 of 186)No
A second vendor's vision model, low effort, previous layout98.9% (184 of 186)93.5% (174 of 186)No

Three lessons came out of this table.

  • Lower effort settings accuse more. At low effort the model rejected "Black crocodile leather phone case" because "the image cannot confirm the material is genuine leather", and rejected "Yellow phone case with a pink flamingo and palm tree print" for naming only two of the print's motifs. Both alt texts were fine.
  • Telling the model the answer first makes it agree. When the alt text came before the image, the model tended to see what it had just been told. Planted errors caught fell from 97.8% (previous layout) to 90.3%.
  • Models differ in what they miss, not only how often. The second vendor's model at medium effort was as reluctant to accuse true alt text, but it accepted "Brown leopard print phone case" on a tortoiseshell case in 3 of 3 runs, and caught 17 of 21 wrong-pattern verdicts (7 alt texts, three runs) against 21 of 21 for our production model on the previous layout.

For anyone buying or building an AI accessibility tool, the takeaway is that "uses AI" says very little. The same model, on the same images, ranged from 91.4% to 98.9% on correct alt text depending on two settings most users never see.

Where the AI went wrong

Under the production setting the check made seven wrong verdicts out of 372. Here are all of them, with the model's own description of the image where it explains the miss.

Every wrong verdict under the production setting
Alt text on the page Expected What happened Runs
"Pink phone case printed with a cartoon girl with a blonde bob and cat-ear headband, wearing a white blouse and black skirt, waving"AccurateFlagged: the model read the headband as "a black tiara-like headband with three pointed spikes, not cat ears". The drawing can be read either way.2 of 3
"Black square phone case with an iridescent abalone shell back in gold, pink and violet tones, shown beside the front of the phone"AccurateFlagged because the frame, not the back, is black. In one run the model's own explanation ended "this is accurate" while its verdict said inaccurate.2 of 3
"Brown leopard print phone case" (the case is tortoiseshell)FlaggedAccepted. Its description even said "tortoiseshell-style pattern" and still passed the leopard alt text.1 of 3
Empty alt on a photo of an ivory tortoiseshell caseFlaggedAccepted as decorative, though the model described the case correctly.1 of 3
"Decorative gradient divider graphic" on a plain gradient barFlaggedAccepted. A description on a decorative image is noise for screen reader users; the model let it through.1 of 3

Four failure modes stand out, and they apply to AI-written alt text as much as to AI-checked alt text:

  • Look-alike patterns and near colors. Across all seven configurations we ran, "Brown leopard print phone case" on a tortoiseshell case accounted for 14 of the 16 missed wrong-pattern verdicts. The colorway errors missed most often were similar near-misses, such as "Black and white checkered phone case" on a beige and white one.
  • Ambiguous details. When a drawing is open to two readings, the model picks one and states it confidently.
  • Reasoning that contradicts the verdict. In a few runs the written explanation and the verdict disagreed. This is why our checker shows what the model saw for every image, so a person can overrule it at a glance.
  • Seeing it but not saying it. The model sometimes described an image correctly and still passed alt text that contradicted its own description.

A failure this benchmark could not see

The previous layout scored well on the 124 product photos, yet live deep audits of the demo store behind our sample report sometimes judged an alt text against the next picture on the page: alt="banner2" on a boot photo was judged against the image that followed it. When a line of alt text sits between two images, the model can read it as belonging to either one.

We tested this separately with the five and six images our live tools send for that page, 15 runs per layout. With the previous layout, at least one verdict landed on the wrong picture in 7 of 15 runs on the five-image page (21 of 75 verdicts). With the current layout, which puts each image's number before it and its alt text after it, that happened in 0 of 15 runs. Neither layout mispaired on the six-image page. We changed the production layout on October 11, 2026.

The lesson is broader than our tool: one test set, however carefully built, can miss a failure that a different page triggers. Measure on more than one kind of input.

The limits of this test

  • It is small. 124 alt texts on 32 images. The intervals in the tables show how much the percentages could move with a different sample.
  • It is narrow. 30 studio product photos (phone cases, a ring holder and a wrist strap) and 2 decorative gradients. The photos come from the catalog of Cocomii, a phone-accessory brand related to GotAlt's parent company, COCOMII LLC. No people, charts, screenshots, logos or images of text. Results on other kinds of image may differ.
  • The planted problems were written by us. Real-world errors can be subtler or stranger.
  • The answer key was drafted by an AI model. Each true description was written by Claude Opus 5.5 from the image and reviewed by a person before any run. That model comes from the same vendor as the model under test, which could favor it. The two true alt texts the check rejected (four of the seven wrong verdicts above) are arguably disputes with the answer key rather than model errors.
  • It is one point in time. Models and their settings change. These results describe October 10 and 11, 2026.
  • It is vendor-run. GotAlt built the test, picked the images and graded the results, about a product we sell. It has not been independently reproduced.
  • It measures judging, not writing. AI-generated alt text, including our own generator's, was not measured.

What this means if you use an AI alt text generator

We did not measure generators, so this section is reasoning from what we did measure, not a test result.

  • Treat every AI description as a draft. The model that got 98% of our verdicts right still confused leopard print with tortoiseshell and argued with a cartoon headband. A generator using a similar model can write those same confusions straight into your alt text.
  • Check the words that distinguish one product from another. Colors, patterns and materials are where our test's errors clustered, and they are exactly what tells a shopper which item this is. See alt text for product images.
  • Prefer short over exhaustive. Under the production setting, every false accusation landed on a long description; short ones were never accused. Long AI output has more chances to be wrong and costs screen reader users more time.
  • Add the context a model cannot see. A model looking at one image cannot know which variant the photo is attached to, whether the image is decorative on this page, or what the surrounding text already says. Our alt text examples by image type show how context changes the right answer.
  • Our own generator is in the same position. The GotAlt alt text generator runs on the same model family at the same medium setting, and it was not part of this test. Use its output as a first draft and read it against the image before you publish.

Why a check that compares the description to the image matters

Most automated alt text checks read the HTML. They can tell whether the alt attribute exists, whether it looks like a filename, and whether it is suspiciously short or long. Our free scan checks for missing and placeholder alt text that way, and in this test the vision check also flagged every filename (30 of 30). What a text-only check cannot do is notice that "Glossy navy blue phone case" sits on a photo of a caramel orange one. That alt text is well formed, specific and plausible. Only opening the image reveals it is false.

In this test, wrong-colorway alt texts were the largest group of planted problems, and the production check flagged all 23 of them in all three runs (69 of 69 verdicts). That kind of error is easy to introduce wherever catalogs are built by duplicating listings, as our Shopify guide describes. It is also the error an AI generator cannot catch for you after the fact, because the generator never sees the alt text that ends up on the page. A check that downloads the live image and compares it with the live alt text closes that gap. Our methodology page explains exactly what is checked, and why the deep audit opens up to six images per page rather than all of them. For how this differs from rule-only scanners, see why scanners miss issues.

Methodology

Images

32 images: 30 real product photos from the Cocomii phone-accessory catalog described above (square phone cases in colorways and patterns such as crocodile texture, abalone shell, leopard, tortoiseshell, marble, checkerboard and cartoon prints, plus a magnetic ring holder and a beaded wrist strap) and 2 plain decorative gradients. A script picked and resized the photos deterministically, so the set can be rebuilt exactly.

The answer key

For each product photo the key holds a full true description, a short true description and a near-miss: alt text copied from a different colorway or pattern of the same product line, so it is plausible but false. The descriptions were drafted by an AI model from each image and reviewed by a person before any run. Changing the key means rerunning every configuration, never one.

What the check was shown

The 124 alt texts in the test set
Kind Count Correct verdict
True full description30Accurate
True short description30Accurate
Empty alt on a decorative image2Accurate
Wrong color from another colorway23Inaccurate
Wrong pattern from another design7Inaccurate
Filename (for example "IMG_4021.jpg")10Inaccurate
Vague label10Vague
Empty alt on a product photo10Decorative mismatch
Description on a decorative image2Decorative mismatch

That is 62 alt texts that should be left alone and 62 that should be flagged. They were packed into 21 pages of up to six images, the most one of our audits sends at once, with true and planted alt texts mixed on each page so a model could not pass by flagging everything.

How it ran

Each page went through the production audit code unchanged: the same instructions, the same answer format and the same four verdicts (accurate, inaccurate, vague, decorative mismatch) a customer gets. Every page ran three times per configuration, 63 page requests in all, so a verdict that flips between runs shows up.

How it was graded

A true alt text counts as left alone only when judged accurate. A planted problem counts as flagged when given any verdict other than accurate; "exact verdict" also requires the right category. Results are pooled over every alt text and run, with 95% confidence intervals. The second vendor's model received the same instructions, image labels and answer format through a separate adapter.

What we changed because of it

  • The production check runs at medium effort, never low.
  • Each image is introduced by its number, and its alt text follows the image, so the model looks before it reads the claim and cannot attach it to a neighbor.
  • A stricter rule for alt text that omits information was tried and not shipped. On a separate set of six images whose message is their words or data (a sale banner, a chart, a coupon, a store-hours sign), the existing instructions already flagged 17 of 18 verdicts on alt text that dropped the message (six alt texts, three runs each), and the stricter rule cost accuracy on the main test.

How we decide what to publish is set out in our editorial policy.

Questions about AI alt text accuracy

Can AI write accurate alt text?

Often, but not reliably enough to publish unread. We measured AI checking alt text rather than writing it, and even the best configuration confused look-alike patterns and argued with ambiguous details. A generator built on a similar model can make the same mistakes. Use AI output as a first draft and read it against the image.

How accurate is AI at checking alt text?

In our test of 124 labeled alt texts, the production setting left 97.8% of correct alt text alone and flagged 98.4% of planted problems, over three runs each. That is one small, vendor-run test on product photos, so treat it as evidence, not a promise of how it will perform on your images.

Did you test the GotAlt alt text generator?

No. This test measured the verdict behind the alt text checker, the deep audit and the paid weekly audit. The generator runs on the same model family at the same setting, but its output was not graded here.

Why not just use the AI's description and skip checking?

Because the errors that matter most are plausible ones. "Brown leopard print phone case" on a tortoiseshell case reads perfectly well. Only comparing it with the picture shows it is wrong, and only someone who knows the page knows whether the image needs a description at all.

Is this an independent study?

No. GotAlt designed the test, chose the images and graded the results, and we sell the product it measures. We publish the composition of the test set, the grading rules, every miss and the limits so you can judge the evidence yourself.

Does accurate alt text make a site accessible?

It covers one success criterion, WCAG 1.1.1, and only for images. Forms, headings, keyboard access, contrast and much more also matter, and no automated tool can confirm a site meets WCAG. See our WCAG 2.2 checklist for the rest.

Find out whether your alt text is true

The checker downloads the first six described images on any page and compares each one with its alt text, using the production setting measured above. Free, no signup.

Check if my alt text is actually right

Find out what your images are missing

The free scan finds missing and filename alt text across up to 3 pages of your site, alongside the other rule-based WCAG checks. No signup.

Scan my site for free