A rule engine can ask if alt text exists.
It can't ask if it's true.

See it run on a real broken page first

The deep audit downloads your images and puts each one in front of a vision model, alongside the alt text your site currently serves for it. Then it tells you where the words and the picture disagree. A free audit covers up to 6 images on one page; paid plans work through a page at a time across your site.

Free tier: one deep audit a day, up to six images on the page. Paid plans run it across a set of pages, on a weekly schedule.

The question a rule engine is built to ask, and the one it isn't

Rule-based scanners, including the free 16-check scan GotAlt itself runs, work by reading your HTML. That's genuinely useful, and it's cheap enough to offer generously. Ours allows 5 free scans a day. It also has a hard ceiling: nothing in an HTML document tells a parser whether a sentence is accurate.

What a rule engine can ask

  • Is the alt attribute present on this image?
  • Is it empty on an image that looks decorative?
  • Is the text suspiciously short, or does it match a filename pattern like IMG_4021.jpg?
  • Does it contain placeholder junk like "image" or "photo"?

All of this is answered by reading characters. The image file itself is never opened.

What the deep audit can ask

  • Does this sentence describe this specific image?
  • Is it vague to the point of being useless, even though it's grammatically fine?
  • Is a purely decorative image being narrated to screen reader users who don't need it?
  • Does a link's accessible name make sense read on its own, out of surrounding context?

Answering these requires downloading the image and examining it with a vision model, which has a cost per image, as explained below.

Four ways alt text fails that no parser can catch

These are the failure modes the audit is built to separate. Each needs a look at the actual picture, not just the sentence next to it.

Inaccurate

The words describe something the image doesn't show. Common cause: alt text copy-pasted from a similar product, or a template default that was never updated when the photo changed.

Before: alt="Black leather ankle boot, side profile" on a photo that is actually a tan suede boot, front-on. After: alt="Tan suede ankle boot, front view, block heel". A rule engine sees a well-formed sentence in both cases and passes it. Only opening the file shows the mismatch.

Vague

Technically true, but empty of the detail that made the image worth including in the first place.

Before: alt="Product" on a photo of a specific item. After: alt="Cast-iron skillet, 26cm, pre-seasoned, with pour spout on each side". Nothing here is false, but it conveys nothing a shopper deciding whether to buy would need.

Filename or placeholder text

The one failure mode rule engines already catch reasonably well, and GotAlt's free 16-check scan flags it too. Pattern-matching on things like IMG_4021.jpg, banner-2, or the literal word "image" is cheap and doesn't need a vision model.

Before: alt="DSC00456". After: alt="Team of six standing outside the Bristol office entrance". Listed here for completeness. The deep audit's contribution is everything above and below this one, not this one.

Decorative mismatch

An image that carries no information (a divider, a background flourish, a repeated icon that's already explained by adjacent text) is being narrated anyway, forcing screen reader users to sit through noise for nothing.

Before: alt="Green decorative swoosh graphic" on a purely ornamental divider. After: alt="", so screen readers skip it entirely. Recognising "this image is decorative" requires judging what the image contributes, not just whether it has a caption.

What else the audit looks at

Images are the clearest example, but the same problem (a check that needs judgment, not just parsing) shows up in a couple of other places the deep audit also covers.

  • Link text out of context. A rule engine can confirm a link has some accessible name. It can't judge whether "click here" or "read more", heard alone in a screen reader's list of links with no surrounding sentence, tells you anything about where it goes.
  • Heading descriptiveness. A heading structure can be validated for skipped levels by counting tags. Whether "Overview" or "Details" actually describes the section under it, for someone navigating heading-to-heading, needs a reader, not a counter.

How it works

  1. Fetch the page

    We request the URL and read the HTML your server actually sends: the same thing a browser gets before any client-side JavaScript runs.

  2. Extract each image and its current alt text

    Each <img> is pulled out along with whatever alt attribute (or absence of one) it currently has.

  3. Download each image file

    The actual picture is fetched, not just its filename or surrounding markup. This is the step no rule-based scanner performs.

  4. The vision model compares words to picture

    The image and its current alt text are shown to a vision model together, and it judges whether the description is accurate, sufficiently specific, and appropriate to whether the image is decorative or meaningful.

  5. You get a verdict and a suggested replacement

    Each image comes back flagged as fine, or flagged with what's wrong and a concrete suggested alt text you can paste in directly.

What the audit cannot see

Where the deep audit stops

Opening the image file moves the line, but it doesn't erase it. The deep audit, like the free rule-based scan, cannot detect: computed colour contrast, keyboard traps, focus order or focus visibility, how a screen reader actually announces a page in practice, logical reading order, cognitive load, or anything a script injects into the page after it loads. Some of these need a real browser rendering the page with layout and CSS applied, which our fetch-based approach doesn't do. Some need a human using assistive technology. No automated tool, ours included, checks all of WCAG. Independent estimates put the machine-testable share at roughly a third of it (Karl Groves puts definitively-testable best practices at 25–29%, with about 40% untestable by tooling at all).

Why this isn't unlimited

The free rule-based scan can afford to run as many times as you like, because reading HTML is nearly free. The deep audit is different: every image is downloaded and sent to a vision model, so the work grows with the number and size of the pictures on a page. That is why the free tier allows one deep audit a day and paid plans are priced by monitored page.

That's why the free tier caps out at one deep audit and one alt text check a day, why each audit opens up to six images on a page rather than all of them, and why paid plans price by page count instead of promising unlimited everything. It's also the answer to "why isn't this free like the rule scan": opening image files really does cost us something each time, and unlimited-and-free would mean either losing money on every audit or quietly cutting corners on how carefully the model looks.

Paid access runs the deep audit on a set of pages you choose, weekly, rerunning it whenever a page's images or their alt text change, with regression alerts when a new upload breaks something that used to be fine, and a full evidence log for every issue. Not included: exporting the evidence log. Professional is $69/mo for one site and 50 AI-monitored pages; Agency is $229/mo for 100 pooled pages across unlimited client sites. Extra pages are $1.30–$1.40 each per month. See the full pricing breakdown. Paid plans start by email: write to hello@gotalt.com and we confirm the page count and start date. Billing then runs through Stripe automatically.

See it on your own site first

The free alt text checker runs this exact process (download, compare, verdict) on the first six images of one page. No signup, no card.

Check if my alt text is actually right