UX/UI design · Healthcare

A clinician opens your survey between two patients, and the design decides whether she finishes

The short answer

UX and UI design for healthcare research is judged on completed responses, not clicks. Your respondent is a clinician answering between patients on a phone, so save state, honest progress and question types that survive a small screen decide how many finish. The admin your research team lives in matters just as much.

Your respondent answers this on a phone, standing up, in whatever gap the day gives her. She didn't set aside time for it and she won't come back to it later. If the next screen doesn't fit in a hand, everything you spent getting her there is gone.

Completion is the number, and it gets decided after they click in

Every status report on a study in field leads with the numbers that are easy to move. Invitations sent, open rate, click-through. Those measure how well you got somebody to the front door, and the moment a physician taps the link they're all spent. What's left is the interface, and it owns the whole distance between started and finished.

The useful thing about abandonment is that it has an address. People don't drift away from a questionnaire gradually. They quit at a specific question, on a specific screen, and in most instruments a handful of screens account for most of the loss. That makes it findable, which makes it fixable.

It also makes it a research task rather than a discussion. Everyone on your team has seen the instrument dozens of times, on a large screen, with a mouse, in a quiet room, with nobody waiting. That's the least representative viewing condition available and it's the one every internal review happens in. Intuition formed there is confidently wrong about which screen is doing the damage.

Five moderated sessions with real panel members, on their own phones, running the real instrument, get you further than a hundred hours of that. Two conditions decide whether they're worth anything. Recruit from the panel you actually field to, because a general consumer respondent will patiently do things a physician simply won't. And book at the hours those people are free: early morning, the gap after clinic, a Sunday evening. A one o'clock Tuesday slot recruits the clinicians whose schedules look nothing like your panel's.

Design for the interruption, because that is the actual usage pattern

The controlling fact about this audience isn't that they're busy. Everybody claims to be busy. It's that their attention arrives in fragments they don't control. A clinician starts your survey between patients, gets pulled out of it by something that outranks you, and comes back somewhere between four minutes and four days later, often on a different device.

Design for that and a broken-up eleven minutes still produces a completed response. Ignore it and the same eleven minutes produces nothing, because the session died and the work went with it. This is the highest-return area on the whole surface, and it's nearly invisible in a design review, because nobody reviews a design by walking out of the room halfway through. Three things carry the weight: what gets saved, how somebody gets back in, and what the interface tells them about how much is left.

  1. 1. Save on the answer, not on the page

    Persist each response as it's given rather than when a page is submitted. The difference shows up in exactly the case that matters, where somebody answers four of six questions on a screen and never reaches the button. Page-level saving throws away all four and shows them a blank screen when they return, which is the moment most people decide they've already done this once.

  2. 2. Make the way back in survivable

    The return link has to work from a different device, from an email opened at home, and after a login session expired, without anyone having to remember which browser they started in. Where identity has to be re-established, ask for the thing they'll have on them rather than the thing your database prefers. The W3C's redundant entry criterion, WCAG 2.2 success criterion 3.3.7, is the general form of it.

  3. 3. Treat the session timeout as a design decision

    An inactivity timeout is usually inherited from a framework default and nobody ever chose it. For this audience it decides completions. The W3C's timeouts criterion, 2.2.6, asks that people be warned about an inactivity period that could cause data loss unless the data is preserved for more than twenty hours. Read twenty as a floor rather than a target. A clinician interrupted at eleven in the morning often comes back that evening.

  4. 4. Show progress you can actually stand behind

    Most tools compute a progress bar as questions answered over questions in the instrument, which becomes fiction the moment there's branching. The respondent watches it reach sixty percent and slide back to forty-five, and that's where the study stops being the length they agreed to. A bar that can move backwards is worse than no bar, because it turns a vague feeling of length into evidence you misled them. Show named sections with the current one marked, or a time estimate from real completion data, or nothing.

  5. 5. Say what happens if they leave, before they leave

    Most respondents assume closing the tab loses everything, because almost every form they've used taught them that. If the truth is that their answers are held and the same link picks up where they stopped, that sentence belongs beside the questions, not in a confirmation email nobody opens. It turns leaving and returning from a failure into a supported path, and returning is where a real share of your completions come from.

Question design is interface design, and the trade is usually made in the analyst's favour

Most of the drop-off you're hunting isn't caused by the interface around the question. It's caused by the question. And the question types that hurt most share an origin: they produce a tidy data structure at the other end, and the cost of producing it was quietly transferred to the person answering.

That isn't a criticism of the analysts. A grid genuinely does give you comparable ratings across attributes in one clean block, which beats six separate questions for analysis. The trade was just made without anyone pricing the respondent's side, and on a phone that side is expensive.

So here's the position worth holding. Every question type that's hard to answer on a small screen has to justify itself against the completions it costs, and the person who wants it has to be in that conversation. A grid producing a beautiful crosstab for a question nobody will put in the report is the worst trade in the instrument, and most instruments have one. This is what the common offenders turn into in a hand.

Question types that cost completions with healthcare professionals, and what to do about each one.
Question typeWhy the analyst wants itWhat it becomes on a phoneThe design that keeps people in
Matrix grid, twelve attributes by a five-point scaleOne comparable block and clean crosstabs, with attributes rated against each other rather than in isolationA table wider than the screen with the scale header scrolled out of view, so row nine gets rated against labels nobody can seeOne attribute per screen with the scale pinned, and a cut list holding only the attributes that will reach the report
Long option list, a couple of hundred conditions, brands or specialtiesComplete coverage and coded data with no free-text cleanup afterwardsA picker scrolled past a hundred entries with a thumb while standing up, with mis-taps nobody notices until analysisType-ahead search with the ten most-chosen answers surfaced first, and an other field taking free text instead of forcing a wrong pick
Forced ranking of eight or more itemsA full preference order, which supports the strongest analysis of the setDrag and drop on a touch screen with items that jump under the finger and a list too long to see at onceA top-three pick or a run of paired comparisons, accepting the loss of a ranking tail that was never going to be reported
Constant sum or slider, allocate a hundred points across six itemsInterval data and clean shares rather than ordinal ranksFine motor control on a moving target, with a running total obscured by the thumb adjusting itNumeric entry with the remaining balance shown above the control rather than under it, and fewer buckets than the analysis plan asked for
Repeating loop, the same block for each item selected earlierPer-item data for everything a respondent uses, without running a separate studyA clinician who honestly selected nine products and now faces nine identical blocks, having been promised eight minutesA capped loop with the number of rounds stated up front, sampling items where the analysis allows rather than asking about all of them
Open text carrying the main insightVerbatims, quotes and the material that makes the deck persuasiveA phone keyboard covering two thirds of the screen while somebody types with one thumb in a corridorFewer open questions, placed after the closed ones have earned the effort, with the count stated so nobody is ambushed
Question types that cost completions with healthcare professionals, and what to do about each one.

The people who live in the admin are the second audience, and they cost more

Two very different people use what you build. A clinician meets the respondent surface once, for eleven minutes, and never opens that study again. Your research and operations team lives in the admin every working day for years.

Design budget goes almost entirely to the first one. That's understandable, because the first one carries the completion rate and it's the surface a client gets shown. It's also the smaller cost. The respondent experience is a fixed design spend against a fixed number of minutes. The admin is a recurring staff cost that never stops, and how it's designed sets the size of it.

The signal is easy to read and nobody reads it. Ask whoever runs your fieldwork what they do in a spreadsheet, and count. Every item on that list is something the admin can't do being done anyway, by somebody paid to do other work, with no record of it in the system.

What an admin has to support here isn't administration. It's triage against a clock. A study is in field, one specialty quota is filling and another isn't, three participants have written in saying something broke, and someone has to decide what to do in the next hour. Screens organised around the objects in your database are no help with that. Screens organised around the decisions made under time pressure are.

  1. 1. One screen that answers whether this study is going to land

    Quota progress against the segments you genuinely sample on, specialty and region and whatever else defines the frame, next to time remaining and the current completion rate. Not a list of responses. The question asked at nine every morning is whether this closes on time, and if the answer takes three screens and an export to assemble, it gets assembled in a spreadsheet instead and stays there.

  2. 2. A drop-off view that sits where the study lives

    Completion by question, inside the admin, for the study currently in field rather than in an analytics tool somebody has to remember exists and has a login for. Fieldwork is the only window in which finding the bad screen is worth real money. Found after close it's a lesson for next time. Found on day two it's a fix that saves this study.

  3. 3. A participant view showing what actually happened to them

    When somebody writes in saying the survey broke, whoever answers needs their status, where they stopped, what was saved and whether the incentive is affected, on one screen. Reconstructing that across four screens is a several-minute job happening dozens of times per study, and it's the most common reason support answers slowly enough to lose the person.

  4. 4. A way to see what a respondent is seeing right now

    A preview reflecting the live instrument at phone width, reachable from the study, without publishing anything or creating a test response somebody has to clean out later. A good share of support conversations during fieldwork are one person describing a screen to another who can't see it, both of them guessing.

  5. 5. Somewhere to keep what the team already learned

    People running fieldwork accumulate real knowledge. That this specialty needs a longer window, that this phrasing confused everybody last time, that this segment answers on Sundays. With nowhere in the system to put it, that knowledge lives in one person's head and leaves with them. A notes field on the study is close to free and it's the piece most often missing.

Accessibility here decides completion, not compliance

A physician panel skews older than almost any consumer panel you'd design for, and the reason is structural. Becoming a physician takes a long time and physicians practise late. So a national panel carries a large population reading your interface with presbyopia, the age-related loss of close focus that sets in for most people during their forties, in a corridor, on a phone, in whatever light the building has.

Which makes accessibility here a completion measure rather than a checklist bolted on before delivery. The numbers below are from the W3C's WCAG 2.2, and they're worth designing to directly rather than auditing against later.

Target size. Criterion 2.5.8 sets a minimum of 24 by 24 CSS pixels for a pointer target at level AA, and 2.5.5 sets 44 by 44 at AAA. For something answered with a thumb while standing up, design to the larger one. Radio buttons in a rating scale are the usual failure, because the tappable area is often no bigger than the visible dot.

Contrast. 4.5 to 1 for body text under criterion 1.4.3, and 3 to 1 under 1.4.11 for the parts of a control that signal it is a control. Placeholder text standing in for a label and pale grey helper text fail both, and helper text is usually where the instruction that makes a question answerable is hiding.

Colour on its own. Criterion 1.4.1 says colour can't be the only carrier of information. In an instrument that shows up in the same two places every time: the required-field marker and the selected state of a rating scale. Both need a second signal, a mark or a label or a change of shape.

Reflow and resize. Content has to work at 320 CSS pixels without scrolling in two directions under criterion 1.4.10, and survive text enlarged to 200 percent under 1.4.4. A respondent who has turned their phone's text size up, which describes many of this panel, is who your grid breaks for first.

None of it is expensive when it's decided before the question components exist. Retrofitting means going back over every question type in the library at once, under a client deadline.

When you don't need a designer

Not every study finishing under target has a design problem, and it's better to say that here than in a proposal.

If your instrument runs under about five minutes, your panel is warm and answers regularly, and completion among the people who start is already high, the interface isn't your constraint. Redesigning it produces a nicer survey and roughly the same numbers. Something else is holding you, usually one of three things.

Reach in a narrow specialty. Where the total population of that specialty in the country is small and you can contact a fraction of it, no interface recovers the rest. That's a sampling question, and the honest answer is sometimes that a defensible read isn't available at the sample you can get.

The questions themselves. If a study asks clinicians to comment on something they hold no view about, or asks for recall nobody has, people will start it and stop and the design isn't the reason.

The relationship. A panel that hears from you only when you want something answers worse than one that gets told what the last study changed. That's an operations fix, not a screen.

Where design genuinely is the constraint it announces itself the same way every time: high starts, low finishes, loss concentrated on two or three identifiable screens. If your numbers don't look like that, spend the money elsewhere first.

What the research says about finishing, and what we don't design

The gap that matters is between starting and finishing, and it has been measured. In a factorial randomised experiment published in the Journal of Medical Internet Research, Cook and colleagues at Mayo Clinic invited 3,966 physicians to an internet survey. 11.4 percent answered at least one question and 8.5 percent completed it. Offering a book and mailing paper reminders moved neither number.

Read that from the design side. Everybody in the second number had already agreed to take part, so what separates answering one question from finishing is the instrument itself. The Cochrane methodology review points the same way, putting shorter questionnaires ahead of longer ones at an odds ratio of 1.58, which is the bluntest design lever there is.

So the argument on this page about two audiences and an underspent admin is the shape of the work we do. We own the mechanics: how a question is presented, what it costs to answer on a phone, and what the people running fieldwork need to see.

What we don't do, plainly. We don't design clinical software. Not electronic medical records, not patient care systems, not regulated medical devices, and we hold no clinical certifications. Our room is research, medical communications and the knowledge work around clinicians. If what you need sits inside care delivery, that's not us.

Worth knowing before you start

  • Open your longest matrix grid on your own phone in portrait and try to rate row nine. If the scale header has scrolled out of view by the time you get there, that screen is where your completion rate is being decided, and it's the same screen in most studies you run.

  • Redesign the screener before anything else. It's the shortest surface, every invited person sees it, and it's the only place you're asking for effort before you've given anything back. Small changes there move a larger number of people than anything deeper in the instrument.

  • Check what your progress bar does when somebody branches out of a section. If it can move backwards, switch it off until it can't. A bar that loses ground reads as proof the study is longer than the one they agreed to.

  • Start a survey, close everything, and come back tomorrow on a different device. If the answer is start again, that's the most expensive default in your setup and it is very often a single setting rather than a build.

  • Sit with whoever runs your fieldwork for one hour while a study is live and count the things they do in a spreadsheet instead of in the admin. That list is your admin backlog, and it arrives already sorted by how often it costs someone time.

Further reading

Common questions

Usually inside the one you have, and that's where we'd start. Most platforms give you far more control over theme, question layout, mobile behaviour and progress display than teams end up using, and reworking the worst three screens inside your existing setup is quick and cheap. Building your own starts to make sense when the asset is the panel relationship rather than the instrument: recurring participants, balances, consent that has to persist across studies. That's a separate conversation and it should follow the cheap one rather than replace it.

You pay for their time the way you pay for participation, you book at the hours they're genuinely free rather than the hours your team is, and you need far fewer people than you'd expect. Five sessions on the real instrument will surface the screens costing you completions. Your internal clinical advisors are a useful first pass and not a substitute, because they already know what every question means and your respondents won't.

The comparison isn't good phone data against good desktop data. It's phone data against no data, because the careful desktop response you're protecting is one a lot of clinicians were never going to sit down and give you. The real quality risk is somebody straightlining through a fourteen-row grid to get out of it, and that happens on a large screen too. Question design fixes that. Screen size isn't what causes it.

Whether it's contractually required depends on your client, and sponsors increasingly write it into the agreement. The better argument is that the criteria that matter here are the ones your panel's age range needs anyway: larger tap targets, real contrast, information not carried by colour alone, and layouts that survive enlarged text. Treating it as a completion measure rather than a compliance one also gets it done early, which is the only point at which it's cheap.

For the genuinely generic parts, yes, and you should. A template gives you tables, forms and a user list without paying anyone to design them. What it can't give you is the view of a study in field: quota progress against the segments you sample on, time remaining, where people are dropping and what to do about it in the next hour. That view is specific to how your team fields studies, it's the screen they'd open first every morning, and it's the reason an admin is worth designing rather than assembling.

We don't design the instrument, and the division matters. The therapeutic content, the questions, the wording and what the study is trying to learn stay with your research team and their clinical input. What we own is the mechanics: how a question is presented, what it costs to answer on a phone, what happens when somebody is interrupted, and what the people running fieldwork need to see. Where the two overlap, which is mostly around question types, our job is to price the respondent's side of the trade so your team can make the call with that cost visible.

Related

Want to talk through your situation?

A short call is usually enough to tell whether this is the right work for you. If it isn’t, we’ll say so.