Design Health AI to be Doubted

A responsible health-AI output isn’t one a clinician just accepts—it’s one they keep double-checking, and that starts long before the screen.

When a clinical AI tool isn’t quite right—when users overlook details, agree too quickly, or feel unsure—we often want to tweak the display. We might adjust the layout, polish empty screens, add confidence scores, use colors, or insert a disclaimer. Since the interface is what we see and can update quickly, it’s where we naturally turn.

But that’s usually not the best place to begin.

Justin Ranton, who designs clinical AI at SmarterDX, shared a clear thought during our chat for an upcoming episode of the Healthcare UX podcast:

"You could redesign that screen in theory forever. Better hierarchy, better typography, better empty states, and the reviewer is still going to do the same cognitive work. So when we think about designing above the interface, what that means is the decision. We move it upstream into what the structure is, what the model is asked to produce, what shape the output has to take, what counts as a complete explanation."

The choices that make a clinical AI output truly useful—and safe for someone with just seconds to act—happen before any screen exists. What the model produces, what makes an explanation complete, what’s included or left out—these are design choices. They just don’t look like design because we’ve come to think of design as only the surface.

Clearer explanations can raise the stakes

Once we realize that the output’s structure is where the real work happens, it’s tempting to make it as clear and convincing as possible—to create an explanation a clinician trusts at a glance. That’s exactly where health AI can become risky.

The research here is clear. In a 2012 review of automation bias in clinical tools, Goddard and colleagues noticed that over-reliance popped up in about 6 to 11 percent of cases—not just a rare slip by careless users, but a steady pattern. Experienced reviewers learned to double-check the system, while less seasoned ones tended to go along with it. And the problem doesn’t announce itself. A 2024 study of pathologists under time pressure found that stress didn’t make automation bias happen more often; it made the mistakes more severe.

The most surprising takeaway—and the one that really matters for anyone building explanations—is that adding explanations to AI outputs doesn’t reliably cut down on over-reliance. Sometimes it even boosts it. Buçinca and colleagues showed this in 2021: people given smooth, complete AI explanations engaged less, not more. They used the explanation as a reason to stop thinking instead of a nudge to keep going. Smoothness feels like credibility, even if the answer underneath is off.

So, the better your explanation design, the bigger the risk when the model misses something. A well-crafted rationale can act like a persuasion tool. When it matches a correct answer, it saves clinicians time. But when it backs a wrong one, you’ve handed them the most convincing case for an error, right when they have seconds and every reason to accept it.

The goal is checking, not acceptance

That means the aim for responsible health AI isn’t getting clinicians to accept the output. It’s closer to the opposite: keeping them doing what only they can do—looking for what’s missing on the screen, like a conflicting lab result, a note that didn’t carry over, or a detail that breaks the neat story the model just told.

Ranton frames the design goal not as getting the clinician to accept the recommendation but as getting them to work with it:

“I’m not answering the question for them. I’m giving them enough pieces so they can say, ‘Oh, I put this together.’ … They still need to confirm.”

We call this “cognitive forcing”—setting things up so people really engage instead of just going along. It’s not a new concept. Researchers in human-computer interaction have been testing these ideas for years, and they do work. The same 2021 study showed that designs the force analytical engagement cut down on over-reliance, while simple explanations did nothing.

Here’s the tough part: people don’t always like it. In that study, participants preferred smooth, easy designs—even though they did worse with them. The design that keeps users safe is often the one they like least.

That’s the real challenge in health-AI design, and there’s no perfect fix. You’re trying to create something a clinician can use in seconds while staying skeptical enough to catch mistakes. Speed and doubt tug in opposite directions. Balancing both is the puzzle. If we pretend we can give effortless results and careful review at the same time, people can get hurt.

You can’t just add this to the screen

None of this is solved just by tweaking the interface. Designing for double-checking is built into the structure. It’s about whether the output is a final answer with evidence or a path a clinician can explore and question. It’s about showing where evidence is weak instead of smoothing it over. It’s about keeping the model’s uncertainties—missing data or guesses—visible rather than hiding them. Those choices happen early, in how the model is built, long before anyone picks a button color.

I’m not arguing against clarity, and I don’t want to overwhelm clinicians with friction for its own sake. I’m saying we tend to aim for the wrong goal. A health AI meant to be followed is a design flaw, no matter how neat it looks. And you don’t build that on the screen. You build it upstream, in the output itself, well before the interface comes into play.

Design it to be doubted.

Sources