Part 8 of the “UX × AI” series.

For seven articles, this series has been about how designers should think about and work with AI. How to reframe the fear. How to use prompting as a design skill. How to protect genuine empathy. How to build deliberate practice. How to evaluate tools with rigour.

This article turns the lens around.

Because there is a question the design community has not asked loudly enough, and it is overdue. We talk constantly about how UX professionals should adapt to AI. We talk far less about whether AI products themselves meet the standards of UX that this profession has spent decades establishing for every other category of digital product.

They do not. Not consistently. Not nearly enough.

The AI products being shipped right now—by some of the most well-resourced technology companies in the world—routinely violate usability heuristics that have been established wisdom since Jakob Nielsen first articulated them in 1994. They fail at transparency. They fail at error recovery. They fail at building accurate mental models. They fail, specifically and repeatedly, at the things that UX as a discipline exists to get right.

This is not a minor irony. It is a serious problem, with real consequences for real people, and it deserves the same critical attention from the design community that we apply to every other category of product we evaluate.

The specific failure: Confident wrongness

Let me start with the failure mode that I believe causes more harm than any other in AI product design, because it strikes at something more fundamental than usability—it strikes at trust.

AI is often too fluent, too confident, and too fast. When something sounds authoritative but is occasionally deeply wrong, users feel a specific and corrosive cognitive dissonance. The interface presentation gives no signal that distinguishes a high-confidence, well-grounded answer from a low-confidence, fabricated one. Both arrive in the same typeface, the same tone, the same visual authority.

This is a UX failure of the most basic kind—the failure to communicate system state honestly. Every usability heuristic that Nielsen established in 1994 begins with visibility of system status: the system should always keep users informed about what is going on, through appropriate feedback within reasonable time. An AI product that presents a fabricated answer with the identical visual confidence of a well-grounded one is violating this principle at the most foundational level.

And the research on how users respond to this failure should alarm every product team shipping AI features. Researcher Berkeley Dietvorst’s work on algorithm aversion found that people often prefer flawed human judgment over algorithmic decision-making—especially after witnessing even a single error. The shift is not gradual. It is a cliff edge. One visible mistake from an AI system, and the user’s trust in that system—even for tasks it performs well—collapses disproportionately. Compare this to how users respond to human error: with patience, often even sympathy. The AI system gets one strike. The asymmetry is well-documented, and AI product design has done very little to account for it.

This means the cost of a confidently wrong output is not just the immediate harm of that wrong answer. It is the abandonment of a feature that might have helped the user in every subsequent interaction. The product team that ships a fluent, confident AI feature without honest uncertainty signalling is not just risking individual errors. It is risking the complete collapse of user trust in the capability—which is a UX failure with a measurable and severe business cost.

“Trust failures often happen before users ever interact with the feature, because they arrive with the wrong mental model. Use onboarding to build the right mental model so the system’s behaviour feels expected, not surprising.”Google People + AI Research (PAIR) team

Case studies in how badly this has gone

I want to ground this argument in specific, documented failures—not hypothetical concerns, but real products that real people relied on, and the real harm that resulted when AI UX failed.

A government-facing public chatbot provided incorrect information about tax filing deadlines and eligibility requirements, because it had not been updated with the latest regulations and relied on outdated training data. The result was citizen confusion, missed deadlines, potential tax penalties for real people, and an erosion of public trust in the institution. This is not a minor bug. It is a failure of the most basic UX responsibility—ensuring that the system’s output is current, accurate, and appropriately caveated when its currency cannot be guaranteed.

A customer-facing airline chatbot confidently provided a customer with incorrect bereavement fare policy information. When the airline was challenged, it attempted to argue that the chatbot was effectively a separate legal entity, not bound by the same standards as a human representative. A tribunal ruled against the airline. The case is now a widely cited example of what happens when an organisation treats its AI interface as exempt from the basic accountability that the rest of its customer-facing product is held to. The interface spoke for the company. The company was responsible for what it said. This should have been obvious from a UX accountability standpoint before it became a legal precedent.

An internal enterprise chatbot exposed personally identifiable information from other customers’ accounts when prompted—the result of insufficient access controls and the absence of row-level security. This produced a GDPR violation investigation, a potential multi-million dollar fine, and an emergency shutdown. The interface had no design consideration for the boundary between what the underlying system could technically access and what the conversational surface should be permitted to retrieve and present. This is a UX failure as much as a security failure—the interface gave users no indication of where its access boundaries were, because the team that built it had not designed those boundaries with intention.

A retail chatbot applied promotional codes multiple times to the same order, producing negative prices—the company effectively paying customers to purchase products. The result was a direct financial loss exceeding $150,000 before the pattern was detected, across more than 2,400 fraudulent orders. There was no validation that the discount logic produced valid prices, and no limit on discount stacking. This is a failure of basic error prevention—the fifth of Nielsen’s ten usability heuristics, established three decades ago: prevent problems from occurring in the first place, by eliminating error-prone conditions or checking for them before the user commits to the action.

A travel booking chatbot quoted flight prices significantly lower than actual fares, generating prices from learned patterns rather than querying live pricing systems. The result was widespread booking failures, customer complaints, and a wave of reviews describing the experience as a bait-and-switch. The fix, as documented in the post-incident analysis, was straightforward and should have been obvious from the start: always query live data sources for information that changes in real time, and never allow a generative system to produce numbers that should come from an authoritative source.

These are not edge cases. They are a pattern. And the pattern reveals something specific: AI products are routinely shipped without the basic UX discipline—error prevention, system status visibility, appropriate use of authoritative data sources, clear accountability boundaries—that any other category of digital product would be expected to meet before launch.

The four properties of good AI interface design, and how routinely they are missed

There is an emerging and useful framework for evaluating AI conversational interfaces specifically, and it is worth examining closely because of how clearly it maps onto failures that are happening at scale right now.

A strong AI interface is judged on four properties: capability transparency, recovery patterns, confidence display, and accessibility. Capability transparency means users can tell what the system can and cannot do before they engage with it—not discovering its limitations through trial and error, frustration, and failed attempts. Recovery patterns determine what happens when generation fails—whether the system has a graceful path back to a useful interaction, or whether the user experience simply breaks at the first misunderstood query. Confidence display signals when the system is uncertain, giving users the information they need to apply appropriate scepticism. And accessibility ensures the interface serves users with the full range of abilities and contexts that any well-designed product must account for.

The most common failure, and the one I see most consistently across AI products in the market right now, is the absence of recovery patterns. The bot has no graceful path when generation fails. Three distinct things can go wrong when a system processes a request—a technical failure such as a timeout, a content policy refusal, or a fundamental misunderstanding of the request—and most AI products collapse all three into one generic error message. This trains users to interpret every failure as “the AI is broken,” when frequently the more accurate read is “I asked a question in a way the system cannot process—and a different phrasing would work.”

This collapsed error handling is a direct and severe UX failure. It violates the principle of help users recognise, diagnose, and recover from errors, another of Nielsen’s foundational heuristics. The fix is not technically complex—differentiated error states that match the actual failure type, paired with recovery affordances specific to that failure. The fact that most products in the market still have not implemented this, years into the mainstream deployment of conversational AI, reflects a genuine underinvestment in the UX discipline that this category of product most urgently needs.

“Trust isn’t a soft word here. It’s a quantifiable business problem. A 2023 KPMG global study found over half of respondents worldwide are unwilling to trust AI, citing data privacy and opaque decision-making. The Nielsen Norman Group’s State of UX 2026 put it plainly: trust is now a major design problem for AI experiences.”Medium / Design Studio UI-UX, 2026

The agentic frontier: A new category of failure

As AI products move from conversational interfaces toward agentic systems—AI that takes autonomous action on a user’s behalf rather than simply responding to queries—a new and more consequential category of UX failure is emerging, and the design community needs to understand it before agentic products become as widely deployed as chatbots are today.

AI agent interfaces are not chatbot skins with autonomy added. They are a distinct design discipline that requires specific patterns for transparency, control, status communication, and recovery. Traditional interface design assumes the user initiates every action and the system responds. Agentic systems reverse that relationship—the system initiates actions, makes decisions, and changes state without being explicitly asked at each step. An interface built on the traditional assumption has no established pattern for showing unsolicited actions, no way to explain why they happened, and no clear path for the user to intervene.

The teams that are getting this wrong skip the transparency layer—the part that makes the system’s reasoning visible and gives the user a clear path to intervene. That missing layer is consistently where agentic products lose user trust within the first two weeks of deployment. The user discovers that the system has taken an action they did not anticipate, cannot understand why it made that decision, and has no clear mechanism to undo or redirect it. This is the agentic equivalent of the confident wrongness problem—except the stakes are higher, because the system is not just providing potentially wrong information. It is taking potentially wrong action.

This is precisely the territory that we will explore more fully in Article 9 of this series, on agentic UX. I raise it here because it is the clearest evidence that the industry has not yet internalised the lesson that the chatbot failures of the past several years should have taught it: autonomous and semi-autonomous AI systems require more UX discipline, not less, than traditional interfaces—and the discipline required is not being applied at the pace the technology is being deployed.

Why this keeps happening: The structural explanation

It would be easy to attribute these failures to individual incompetence or carelessness on the part of specific product teams. I do not believe that is the accurate or useful explanation. The pattern is too consistent, across too many organisations with genuinely talented engineering and design teams, for individual explanations to be sufficient.

The structural explanation is this: AI product development has been led overwhelmingly by technical capability rather than by user experience discipline. The teams building these products are often organised around the question “what can this model do?” rather than “what does the user need to understand, control, and trust in order for this capability to be genuinely useful and safe?” The model’s capability is the starting point. The interface is frequently treated as a thin presentation layer applied after the underlying technical work is complete—rather than as the accountability layer between the system’s autonomous behaviour and the human being who bears the consequences of that behaviour.

This is precisely the dynamic I described in Article 2 of this series—the risk that prompting and AI product design get positioned as technical work, handled by technical teams, with UX brought in late to make the output presentable rather than involved from the start in defining what the system should do and how it should communicate about what it is doing. The chatbot failures, the agentic UX failures, the trust collapses documented in this article—these are the direct, measurable cost of that structural exclusion.

The organisations that are avoiding these failures are the ones that have brought UX expertise into AI product development at the architecture stage, not the polish stage. They are asking the questions that good UX practice has always asked—what does the user need to know, when do they need to know it, what happens when something goes wrong, how do we build an accurate mental model before the first interaction—and they are asking them before the model is deployed, not after the first trust-collapsing failure forces a redesign.

What good AI product UX actually requires

Drawing from the failures documented in this article and from the emerging frameworks for evaluating AI interfaces, here is what genuinely good AI product UX requires—a standard that the design community should be holding AI products to with the same rigour we apply to every other category of digital product.

Honest uncertainty signalling. The interface must communicate, visually and consistently, when the system’s output is well-grounded versus when it is uncertain, generated from pattern rather than retrieved from an authoritative source, or operating outside the bounds of its reliable capability. This is not a technical feature to be added later. It is a foundational design requirement, as fundamental as showing a loading state or an error message.

Differentiated, specific error recovery. Every category of failure—technical error, content policy boundary, genuine misunderstanding—requires its own recovery pattern, with messaging and next-action guidance specific to what actually went wrong. The generic error toast that treats every failure identically trains users to distrust the entire system rather than to understand the specific limitation they have encountered.

Mental model construction before first use. Users need to understand what the system can and cannot do before they rely on it, not through trial and frustration. This means onboarding that demonstrates real capability honestly, explicit statements of what the system will and will not do, and consent and permission requests written in plain language rather than buried in terms of service.

Transparency into autonomous action. For any system that takes action on a user’s behalf—even partially autonomous action—the interface must make that action visible, explain the reasoning behind it in terms the user can evaluate, and provide a clear, accessible mechanism to intervene, redirect, or undo. This is not optional polish for agentic systems. It is the core accountability mechanism that determines whether users can trust the system enough to use it for anything that matters.

Validation and error prevention at the point of consequence. Wherever an AI system’s output has financial, legal, medical, or otherwise consequential implications, the system must validate that output against authoritative sources and sanity-check it against basic constraints before presenting it as fact or executing it as action. The flight price that should come from a live pricing API must not be generated from a language pattern. The discount calculation that could produce a negative price must be checked before it is applied.

Accountability that does not evaporate at the interface boundary. The organisation deploying an AI interface is responsible for what that interface communicates to users, in exactly the same way it is responsible for what its human representatives communicate. The “the chatbot is a separate legal entity” defence has already failed in tribunal. It should never have been attempted. The interface speaks for the organisation. UX accountability and organisational accountability are the same thing.

Applying LucyUX to AI product design

The LucyUX framework—Listen, Understand, Conceptualize, Yield—is, in this context, a direct critique of how most AI products have been built, and a roadmap for how they should be built instead.

Listen—to users encountering AI products for the first time, and specifically to the moments of confusion, frustration, and broken trust that occur when the system’s behaviour does not match their expectations. Most AI product teams listen to model performance metrics. Far fewer listen, with the same rigour, to the human experience of encountering a confident wrong answer, an opaque autonomous action, or an error message that explains nothing. That listening—directed at the actual human experience of using the product, not at the technical performance of the underlying model—is where the failures documented in this article would have been caught before launch.

Understand—building an accurate model of the gap between what users believe the system can do and what it actually does reliably, and designing specifically to close that gap. This is the mental model work that Google’s PAIR team identifies as foundational, and it requires UX research with the same depth applied to traditional product research—not assumptions about what users will intuitively understand about an unfamiliar and rapidly evolving category of technology.

Conceptualize—designing the transparency layer, the error recovery patterns, the confidence signalling, and the intervention controls as core product architecture, not as features to be retrofitted after a trust-collapsing failure. This requires UX involvement from the earliest stages of AI product development—not brought in once the model’s capability has already determined the shape of the product.

Yield—measured in sustained user trust over time, not in initial demo impressiveness or technical benchmark performance. The yield that matters is whether users continue to rely on the AI feature after their first encounter with its limitations—and that yield is determined almost entirely by whether the UX of the product handled that encounter with honesty, clarity, and a genuine path to recovery.

Your action this week

Pick one AI product or feature that you or your organisation uses regularly. Evaluate it against the four properties: capability transparency, recovery patterns, confidence display, and accessibility.

Specifically: does the product tell you, before you rely on it, what it can and cannot reliably do? When it fails—and deliberately try to make it fail, by asking something it should reasonably struggle with—does it give you a specific, useful explanation, or a generic error? Does it ever signal uncertainty, or does every output arrive with identical confidence regardless of its reliability? And if the product takes any autonomous action on your behalf, can you see why it took that action, and can you easily intervene?

Write down what you find. If you are a UX professional working on or adjacent to an AI product, bring these findings to your team as a direct, specific UX audit—not a general conversation about AI ethics, but a concrete evaluation against established usability principles. This is work the design community is positioned to do, and it is work that is not being done with nearly enough rigour across the industry right now.

My perspective: What I actually believe

The design community has spent the better part of this series learning how to work with AI thoughtfully. I believe that work is necessary and I have no intention of walking it back. But I want to close this article with a direct statement, because I think it needs to be said clearly and the design community needs to hear it from within its own ranks.

We have allowed an entire category of products—products that millions of people now rely on daily, for tasks ranging from customer service to financial decisions to government services—to be shipped with a level of UX discipline that would never be acceptable in any other product category. We would not accept a banking app that gave users no indication of whether a balance was accurate or stale. We would not accept an e-commerce checkout that occasionally, silently, charged the wrong amount. We would not accept a government services portal that gave incorrect information about deadlines without any mechanism for correction or accountability.

We have accepted exactly these failures, repeatedly, in AI products—because the technology is new, because the capability is genuinely impressive, and because the design community has been too occupied with learning to use these tools to apply the same critical standard to the tools themselves.

That needs to change. The UX of AI products is not a side conversation. It is one of the most consequential UX challenges of this decade, because the products in question are being deployed at a scale and into a set of consequential contexts—government, healthcare, finance, customer service—faster than almost any technology category in the history of digital products.

The design community has the expertise to fix this. The question is whether we will apply it with the urgency the moment requires.

Twenty-five years in this field has taught me that the principles of good UX do not change because the underlying technology does. Visibility of system status, error prevention, recognition over recall, user control and freedom—these principles were true for desktop software in the 1990s, true for mobile apps in the 2010s, and they remain true for AI products in 2026. What has changed is the consequence of getting them wrong, because AI systems are being deployed into higher-stakes contexts, with greater autonomy, at greater speed, than any technology category before them. The discipline that UX as a profession has spent decades building is more necessary now than it has ever been. I would like to see us apply it—to the tools we use, and to the tools being built.


Up next in the “UX × AI” series: “Agentic UX: When the Interface Stops Waiting for You.” This article showed what happens when transparency and control are missing from AI products that respond to user requests. Article 9 goes further into the frontier where the challenge is most acute: AI agents that act on a user’s behalf, autonomously, without being asked at every step. What does UX design look like when the user is not always in the loop? We examine the emerging patterns, the early failures, and what the design community needs to get right before agentic products become as ubiquitous as the chatbots this article has spent its length examining.


References & further reading