返回新闻
工业AI Understanding 简报

AudioEye 研究发现人工智能编码工具无法满足无障碍标准

AudioEye 的新研究表明,五种领先的人工智能生成代码工具生成的网站充满了可访问性违规,但 81% 的开发人员仍然相信输出能够满足 WCAG 2.2 AA 合规性。

4 min readRead the linked source
Source-provided image accompanying AudioEye study finds AI coding tools fail to meet accessibility standards
来源参考来源记录
出版商
prnewswire.com
来源链接
prnewswire.comhttps://www.prnewswire.com/news-releases/audioeye-study-finds-ai-coding-tools-do-not-write-accessible-code-302893917.html
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

大语言模型(LLM)
在海量文本语料库上训练来生成和分析文本的语言模型。
XAI(可解释的人工智能)
使人工智能预测更加透明和易于理解的技术和实践。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
测试一下自己人工智能道德测验

发生了什么

AudioEye, a digital‑accessibility firm, evaluated five popular AI coding assistants—OpenAI, Anthropic, Google, xAI and Lovable—by asking each to build three distinct websites (news, e‑commerce, and financial services) that should comply with WCAG 2.2 AA standards. The resulting 15 sites were scanned for accessibility defects. Across the sites, testers recorded 306 unique issues that appeared more than 59,000 times, with 91% classified as medium or high severity. The AI‑generated pages averaged 55 issues per page, only slightly better than the 62‑issue average of a typical 2026 website in AudioEye’s Digital Accessibility Index. In a parallel survey, 81% of developers said they were confident the AI‑written code met accessibility requirements, but 73% reported a rise in accessibility complaints after adopting the tools, and 46% had received a demand letter or lawsuit—71% of those cases involved AI‑generated pages.

AudioEye’s methodology involved giving each AI tool the same brief: create a news site, an online store, and a financial‑services site that meet WCAG 2.2 AA criteria. After the tools generated the code, AudioEye’s accessibility platform scanned the pages for known WCAG failure modes, including keyboard focus traps, missing ARIA labels, improper dialog handling, and inadequate form error announcements.

The scan uncovered 306 distinct accessibility defects, which collectively manifested over 59,000 times across the 15 sites. Severity ratings placed 91% of the defects in the medium‑to‑high range, meaning they could impede navigation for users with visual, motor, or cognitive disabilities and increase legal exposure.

When compared with AudioEye’s 2026 Digital Accessibility Index—a of typical web sites—the AI‑generated pages performed only marginally better (55 versus 62 issues per page). This suggests that simply prompting an LLM to “build an accessible site” does not substantially improve outcomes.

The accompanying developer survey revealed a trust gap: while a large majority (81%) believed the AI output complied with accessibility standards, a similarly large share (73%) observed a rise in accessibility complaints after adopting the tools. Nearly half (46%) reported receiving a formal demand letter or lawsuit, and 71% of those cases involved AI‑generated pages.

来源详情: prnewswire.com ↗

为什么这很重要

The study highlights a critical gap between developer confidence in AI‑generated code and the actual compliance of that code with accessibility law. Because WCAG violations can trigger legal action, organizations that rely on AI assistants without independent accessibility testing may expose themselves to lawsuits and regulatory penalties. The findings also suggest that current LLMs are trained on largely inaccessible web content and lack the ability to self‑audit for accessibility, underscoring the need for dedicated testing tools or human review before deployment. As AI‑driven development accelerates, the risk of widespread non‑compliant code could increase, prompting tighter industry standards and possibly new guidance from regulators such as the U.S. Department of Justice or the European Accessibility Act.

The discrepancy between perceived and actual accessibility compliance raises immediate legal and reputational risks for firms that deploy AI‑generated code at scale. WCAG violations are increasingly cited in lawsuits, and regulators are beginning to enforce accessibility standards more aggressively.

The study underscores a fundamental limitation of current LLMs: they inherit the biases and gaps of the training data, which includes a web that is historically non‑compliant. Without explicit corrective mechanisms, AI assistants cannot reliably self‑correct for accessibility, necessitating external validation.

For organizations, the findings suggest that AI coding tools should be paired with dedicated accessibility testing—either automated scanners like AudioEye’s platform or manual expert reviews—before production release. This adds a layer of cost and time but mitigates the risk of costly legal action.

Regulators may use this evidence to justify stricter guidance or enforcement actions, especially as AI‑generated content becomes more prevalent in public‑facing services.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
交互式概念检查+10 Points
AI Ethics Quiz

Why can ethical evaluation not be reduced to one model score?

接下来看什么

Future research that compares newer AI coding models or incorporates fine‑tuning for accessibility will be closely watched. Industry responses—such as the rollout of built‑in accessibility validators within AI IDEs, or partnerships between AI vendors and accessibility firms—could mitigate the risk. Legal developments, including any class‑action suits or regulatory enforcement actions targeting AI‑generated web content, will also shape how companies adopt these tools. Finally, adoption metrics for AI coding assistants in regulated sectors (finance, healthcare, government) will indicate whether organizations adjust their workflows in light of the study’s findings.

Updates from AI vendors that integrate accessibility checks directly into their code‑generation pipelines could shift the risk landscape. Monitoring announcements from OpenAI, Anthropic, Google, and emerging competitors will be essential.

Legal developments, such as new case law or enforcement actions targeting AI‑generated web content, will indicate how courts interpret responsibility for accessibility compliance when LLMs are involved.

Adoption trends in regulated industries (finance, healthcare, government) will reveal whether organizations adjust their AI coding strategies in response to the study’s findings.

Academic and industry research that explores fine‑tuning LLMs on accessibility‑focused datasets may produce models that generate more compliant code, offering a potential mitigation path.

相关指南和测验

AI 伦理人工智能模型解释AI 的未来人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 资金追踪器
觉得这有用吗?