76 Open Research Questions in AI Usability
- Jakob Nielsen
- Aug 5
- 26 min read
Summary: AI is the biggest change to UX design in 60 years, yet we know embarrassingly little about how people actually use it. Here are 76 open research questions, organized into 7 lists sized for everyone from high school students to the big AI labs. Students who complete a research challenge gain the two qualities employers now screen for: agency and demonstrated AI experience.
AI is the first new user-interface paradigm in 60 years. It’s also the biggest uncontrolled experiment in the history of computing: billions of people are reorganizing how they write, search, decide, and learn around tools that shipped before anybody studied how humans would use them. Nobody is taking notes. The academic pipeline churns out benchmark papers by the thousand, while solid studies of real users doing real tasks with AI remain scarce. Thus, a motivated graduate student can still claim genuinely unexplored territory.
That gap is your opportunity. This article lists 76 open research questions about the usability of AI, most of them answerable with qualitative methods, modest resources, and a few weeks to a few months of work. Watch 10 users, code what happens, report honestly, and you’ll know something nobody else knows. So will the rest of us, once you publish.
Your Thesis Is a Job Application
Let’s be blunt about why students in particular should pick these projects. Entry-level jobs in user research are dwindling. This is even true beyond UX: Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen at the Stanford Digital Economy Lab analyzed payroll records covering millions of American workers and found that employment for workers aged 22–25 in the most AI-exposed occupations dropped 16% relative to older colleagues in the same jobs. Junior knowledge work is exactly what current AI does well, and entry-level research roles sit squarely in the blast zone.
But I refuse to join the doom chorus. Entry-level UXR jobs aren’t gone; they’re merely scarcer. And scarce jobs go to candidates who differentiate themselves on the two qualities hiring managers now screen for.
The first is agency. A hiring manager can’t tell from a transcript whether a graduate has initiative, but he or she can tell from a finished field study. The standard academic thesis, dutifully theoretical and read by three people, proves you can follow instructions for 9 months. A self-chosen study of a question that matters to business (and the world) proves you can find a problem worth solving and finish it without being told how. Employers hire evidence, not adjectives.

A traditional thesis = worthless to demonstrate agency, because it just says you followed the rules of academia, which are useless in business. Showing agency = you’re hired!
The second is AI knowledge and experience, which nobody can claim credibly but anybody can demonstrate. Every project below forces you to use AI tools in earnest, observe how other people use them, and form defensible opinions about where they fail. And the finished study is a portfolio piece that conducts half your interview for you.
8 Types of Research Insight
So much for your career. Now for what the rest of us get out of your work. I present my 76 questions sorted by the resources a study demands of the researcher. But a design team should read the same questions through a different lens: what kind of knowledge does each answer buy, and which design decisions does that knowledge upgrade from folklore to evidence? Sorted this way, my 76 open research questions promise 7 categories of insight into user behavior, plus one supporting set about research methods. Every unanswered question about AI usability marks a design decision currently made by guesswork.

The 8 groups of insights the world would gain from answering my 76 open research questions about AI usability.
Mental models: what users believe the machine is (questions 4, 11, 19, 29, 34, and 59). Users act on what they think an AI is: a search engine, a database, an oracle, a person. Each wrong model predicts its own class of errors, from trusting a “stable” system that’s random by design to confiding in a memory that isn’t there. Map the models, and designers can anticipate error classes before shipping, build onboarding that installs a workable model, and settle the anthropomorphism debate with evidence instead of taste.
Trust & verification (questions 2, 3, 5, 13, 14, 22, 30, 33, 39, 41, 45, 47, 53, 62, 63, 64, and 66). The biggest cluster, with 17 questions, because the gap between fluent output and reliable output is where AI departs furthest from classic UI. How do users decide to believe an answer? When do they verify, and when do they wave errors through? How much work will they delegate to an agent, what happens to generated drafts, and does trust ever recover after the AI burns somebody? The answers convert into credibility cues that track true reliability rather than visual polish, verification affordances people will use, oversight controls calibrated to real error rates, and honest quality metrics such as edit distance and downstream cleanup cost.
The articulation loop: from intent to prompt and back (questions 6, 12, 18, 32, 48, and 60). How do people translate the intention in their head into words the machine acts on, where do they hit the articulation barrier, and how do they repair the damage after a misunderstanding? These answers tell designers what scaffolding to build at the input box and which repair affordances pay off. They also settle a budget fight: spend on educating users or on fixing the interface? (My money is on the interface, but the studies will tell.)
Discoverability: learning what the AI can do (questions 8, 15, 37, 46, 50, 56, 57, 58, 61, and 65). Chat hides nearly everything behind a blank box, and embedded AI hides behind sparkle icons that users may already have trained themselves to ignore. This cluster asks what users find, what they never find, and whether the standard remedies (suggestion chips, example prompts, feature tours, progressive disclosure) accomplish anything. The payoff: entry points and empty states designed from observation, and feature investments based on what the median user adopts rather than what an imagined power user might.
Task–interface fit: when AI, and in what form (questions 7, 17, 20, 21, 23, 24, 35, 40, 49, and 51). What do people actually use AI for? When do they pick it over search, when does chat beat a GUI or voice beat typing, how do they read the answers, and how long will they wait for a thinking machine? Teams currently choose interaction paradigms by fashion. These answers would let them choose by principle, set response length and format from observed reading behavior, and design the wait around what users believe is happening during the pause.
Trajectories: how AI use changes people over time (questions 10, 16, 27, 36, 42, and 54). A usability test sees an hour; these questions ask about the months and years. Does usage settle into habit or drop off the novelty cliff? Do skills compound or atrophy? What happens to children who grow up with AI as the default helper, and to all of us when the model changes underfoot? The findings would separate durable value from curiosity clicks, steer education and training policy, and force product decisions to account for consequences that no single session can reveal.
The user range (questions 9, 25, 26, 31, 44, and 55). AI products are built by young, technical, mostly American teams and used by everybody else. This cluster measures how usability differs for users past 70, for blind and keyboard-only users, and across languages and cultures, plus how far a product team’s own usage sits from its customers’. The payoff is concrete inclusive-design fixes, and something subtler: knowing which findings generalize and which fall apart the moment they leave the lab’s zip code.
Methodology (questions 1, 28, 38, 43, 52, and 67–76). These questions study the research itself instead of the users: do our heuristics, sample sizes, think-aloud protocols, success metrics, and synthetic participants still work when the system under study is nondeterministic and changes monthly? They produce no design guidance directly. But whoever answers them makes every other answer on this page trustworthy, which is why methodology gets the closing list.
The answer categories deliberately cut across the question lists. Take trust: a high schooler can measure sycophancy with a scripted test (question 5), a master’s student can watch users verify or fail to (question 14), a Ph.D. candidate can build a theory of calibrated reliance (question 30), a product team can trace draft survival in its own logs (question 41), and an AI lab can mine trust recovery at population scale (question 53). Same insight, 5 budgets. So whatever resources you command, every kind of missing knowledge is within reach. Pick the insight you care about, then find the question you can afford.
How to Use the Question Lists
My 76 research questions run in one numbered sequence, grouped into 7 lists by the resources they demand: starter studies for high schoolers and bachelor’s capstones, 16 master’s thesis projects (the core of this challenge), Ph.D. dissertations, studies for in-house product teams, questions only the big AI labs can answer, replications of classic usability findings, and methodology research. Each entry states the question and why the answer matters. Deliberately, I give no research protocols: designing the study is your first test as a researcher, and half your learning.
The boundaries between lists are soft. Most master’s questions could shrink into a capstone or grow into a dissertation. So treat the categories as calibration, not law. And check the recent literature before committing: this list is dated August 2026, the field moves fast, and a few of these questions will be answered within a year. Good. That’s what open questions are for.

The project scope varies by student level, but you can scale most of my research ideas to be either more ambitious or more contained, to suit a more lavish budget or a tighter time frame. I recommend reviewing the lists at least one level above and one level below your intended project.
Starter Studies: High School and Bachelor’s Capstones
Constraint breeds focus. These 10 projects need no budget, no lab, and no recruiting agency: friends, family, classmates, and free AI products suffice. Each fits inside 2–6 weeks: right-sized for an ambitious high school student or a bachelor’s capstone. Small is not trivial. A narrow question answered cleanly beats a grand question answered vaguely. And several of these would hold their own in a working professional’s portfolio.
Heuristic evaluation of an AI assistant. Evaluate a leading AI product against the 10 usability heuristics and document every violation with screenshots. Which heuristics does conversational AI break most often? An expert review needs no participants and no budget, teaches the heuristics through application, and tests whether 30-year-old evaluation tools still bite on brand-new interfaces.
The fine print nobody reads. Every AI chat interface carries a warning like “AI can make mistakes.” Have 10 people use the product for 15 minutes, then ask what the footer said. Do users notice, remember, or act on it? Disclaimers are the industry’s main safety device; if they’re invisible, everyone should know.

How do users react to AI’s warnings?
Reference check. Ask an AI for sources on 20 topics, then verify every citation: does it exist, and does it say what the AI claims? Score hallucination rates by domain. No participants needed, only detective work, and the results speak to the AI problem that burns students most: fabricated references.
Ask twice. Pose the same 10 questions to the same AI on 5 different days. How much does the advice change in substance, not just wording? Users treat the AI as a stable oracle; measuring answer variance tests whether that folk belief survives contact with a system that’s random by design.
The pushback test. Feed the AI questions you know it answers correctly, then insist it’s wrong. How often does it cave, and does polite pushback work differently from aggressive pushback? Sycophancy undermines AI as a checking tool, and a simple scripted protocol can measure it without recruiting a single participant.

An assistant that abandons correct answers under pressure can’t function as a checking tool. The study would measure when correction becomes useful responsiveness and when it becomes sycophancy.
Does politeness pay? Run the same 10 tasks with blunt, neutral, and courteous prompt phrasings, then have judges rate the outputs blind. Prompting folklore claims “please” improves answers; testing the superstition costs an afternoon and teaches experimental control, blind evaluation, and healthy skepticism toward advice that circulates without evidence.

Does it matter how users phrase their requests, or will AI deliver the same results regardless?
AI versus search, with a stopwatch. Take 10 everyday questions, from store hours to visa rules. Answer each with a search engine and with an AI, logging time, steps, and correctness. It’s a miniature version of the biggest shift in information behavior since the web, and the numbers will surprise somebody, possibly you.
Do starter prompts get clicked? Watch 5 novices open an AI app that displays suggested prompts. Do they click a suggestion, type their own, or freeze? The empty text box is the most intimidating screen in modern software, and this tiny study tests whether the standard remedy does anything at all.
Chatbots without a mouse. Attempt complete AI sessions using only the keyboard, then run an automated accessibility checker: focus order, streaming announcements, copy buttons, stop controls. An accessibility audit requires no recruiting and produces an immediately actionable bug list that most AI vendors apparently never compiled themselves.
AI around the dinner table. Interview 5 households about who uses AI, for what, who refuses, and what family rules exist (homework, especially). Domestic adoption is invisible to lab studies, and households expose the social side of AI usability: trust negotiations, hand-me-down prompting habits, and one relative doing everyone’s AI chores.
The Main Event: 16 Master’s Thesis Projects
This is the core of the challenge and the reason I wrote this article. A master’s thesis buys a few months of focused work, real participant recruiting, and mainly qualitative methods, with targeted quantitative measures where they pull their weight. These 16 questions are sized for that research budget. And every one addresses a live gap: answer it competently, and practitioners will apply your findings the week you publish.
Folk models of how AI works. What mental models do everyday users hold of what an LLM is (a search engine, a database, a librarian, an oracle, a person), and which usage errors does each model predict? Most prompting failures and trust failures likely trace back to wrong models, but we lack a taxonomy grounded in observation of mainstream users rather than tech workers.

Users behave according to what they believe the AI is. Treating it as a search engine, database, librarian, oracle, or person produces different expectations and different predictable errors.
The articulation barrier in the wild. How do non-expert users translate an intention in their head into a prompt, and where exactly does the translation break down? Specifying requirements is a skill most people were never taught; knowing the precise breakdown points would tell designers what scaffolding to build. (A natural extension of my articulation barrier article.)

We need to know much more about how users attempt to climb the articulation barrier and translate the idea in their heads into prompts for the AI.
What makes an AI answer look credible. Which surface features (confident tone, citations, formatting, length, response speed) cause users to believe an answer, and do those cues correlate with actual accuracy? If trust rides on invalid cues, that’s both a danger and a designable problem.
Do users catch AI errors? In domains where users have little or moderate knowledge, what share of errors do they detect, and what triggers verification behavior versus passive acceptance? Verification is the safety net of AI use, and I suspect it rarely happens, but we have little direct observation of it.
The first session. How do new users explore an AI tool in their first hour: what do they try first, which capabilities do they never discover, and what determines whether they return? Chat interfaces have near-zero discoverability, so the blank text box hides almost everything the product can do. (A closing window: study participants who have never touched AI get scarcer every month.)
Why people quit AI. What made lapsed users abandon AI tools after an initial trial? Adoption research follows users on the way in; the ones heading out go unstudied, and their reasons amount to a roadmap of the unsolved usability problems. Recruiting quitters is harder than recruiting fans, which is exactly why nobody has done it.
AI or search? How do users decide between asking an AI and using a search engine for a given information need, and how do they interleave the two? Habits formed over three decades of searching are dissolving in real time, and the decision heuristics people apply are undocumented.
Conversational repair. When the AI misunderstands, what repair strategies do users attempt (rephrase, correct, scold, start over, give up), and which ones work? Repair is where sessions are won or lost, and a taxonomy of effective repair would translate directly into teachable technique and better UI affordances.
Folk theories of AI memory. What do users believe the AI remembers within a session, across sessions, and about them personally, and how do those beliefs shape behavior? Memory features are proliferating while user understanding of them is near zero, causing both privacy risk and wasted effort.
How people read (and want) AI answers. Do users read streaming responses linearly, skim them in something like an F-pattern, or jump to the end, and when do they want a long structured answer versus a short direct one? Response design is currently guesswork, and over-long answers are among the most common user complaints. (A chance to test whether my 20-year-old reading findings transfer.)
Waiting for the thinking machine. Do the classic response-time limits (0.1 second, 1 second, 10 seconds) hold when users believe the AI is reasoning, and what do visible chains of thought and progress messages do to perceived waiting? Reasoning models take minutes, so either the old limits bend under a new mental model or current products violate them at massive scale.

This cartoon user has remained dutifully attentive for a very long time because the interface implied that important reasoning was occurring. What do real users do in similar situations?
Supervising an agent. When an AI executes multi-step tasks, how closely do users want to watch, when do they want to be asked for confirmation, and what makes them feel safe delegating? The industry is racing toward agents while the human-oversight experience is nearly unstudied; getting it wrong produces either exhausting babysitting or blind trust.

If you have 5 butlers help you at breakfast, you can’t supervise them in every little detail, or you’ll never eat. Same when supervising agents.
Chat versus hybrid UI. For a well-defined task, does a task-specific graphical interface layered on top of AI beat a pure chat interface on success and satisfaction? Chat became the default AI interface by accident rather than by evidence, and even one careful comparison would inform the debate about whether chat is the final form.
Speak or type? For which tasks and contexts do users prefer voice interaction with AI over text, and what breaks down in each modality? Voice AI finally works well enough to be a genuine alternative, but modality fit is currently assumed rather than observed, and context (car, kitchen, open office) may trump task.

Modality preference depends on context as much as task. Voice reduces typing effort while creating new problems involving privacy, noise, correction, social embarrassment, and reviewability.
Older adults and conversational AI. What barriers do users past 70 encounter with AI assistants, and do their problems differ in kind from younger users’ problems? This group could gain the most from natural-language interfaces and from AI ideation, compensating for declining fluid intelligence, yet is largely absent from AI usability research, and findings would likely puncture the assumption that chat is inherently easy.
AI through a screen reader. How usable are AI chat interfaces for blind users, from streaming text to reasoning displays to generated images? Chat seems text-friendly on the surface, but dynamic updates and lengthy unstructured output may be hostile to assistive technology, and concrete findings would translate straight into fixes.
Ph.D. Dissertations: Questions That Need Years
Some questions don’t fit in a semester. The 10 below need what a dissertation uniquely offers: years of runway, longitudinal designs, validated instruments, and theory-building patience. Fair warning: parts of any AI study will rot as the technology moves (question 69 tackles that meta-problem). So pick a question whose core outlasts any single model.
Deskilling or upskilling? Follow novice writers or programmers who lean on AI through 2–3 years of their careers. Does the underlying skill atrophy, plateau, or compound? Separating augmentation from dependency demands longitudinal data that nobody has collected, and the answer will shape education policy, hiring, and tool design for a generation.

Some skills are nurtured over time and grow slowly into professional competence. We need longitudinal research on how AI shapes (and ideally accelerates) years-long competency development.
Usability heuristics for AI, validated. Derive candidate heuristics for AI systems from field data, then test whether expert evaluations using them predict the problems real users actually hit. Iterate until inspection results are reliable. The field is drowning in proposed AI guidelines; what it lacks is one set with demonstrated predictive power.
A stage model of AI mental models. How do users’ conceptions of AI evolve from first exposure to expertise? Longitudinal interviews and drawing exercises could yield a developmental stage theory, comparable to classic models of how novices become programmers. Each stage predicts characteristic errors, telling educators and designers what to teach, and when.
Designing for calibrated trust. Overreliance ships errors; underreliance wastes the tool. Which interface elements (confidence displays, visible sources, friction before acceptance) move users toward trust that matches the system’s actual reliability? Answering across domains with different stakes requires an experiment series plus theory: exactly the shape of a dissertation.

Confidence indicators help only when they correspond to actual reliability and influence behavior appropriately. A confident interface isn’t necessarily a safe bridge.
AI usability across cultures. Do prompting styles, politeness norms, anthropomorphism, and trust differ across languages and cultures enough to change design conclusions? Multi-country comparative studies would establish whether AI UX is one field or several, and which findings from American samples embarrass themselves when they travel abroad.
A prompting proficiency test. Build and validate an instrument that measures a person’s skill at working with AI, then demonstrate that scores predict real task outcomes. Think TOEFL, but for talking to machines. Education, hiring, and research all need a validated measure, because self-assessed AI skill is worth roughly nothing.

Can we develop a good exam to measure a user’s AI proficiency?
Learning to delegate. As agents execute multi-step work, how does delegation develop over months of real use? When do users stop checking, and is stopping rational or reckless given the agent’s actual error rate? A longitudinal field study could ground the first real theory of appropriate reliance on autonomous software.
Does humanlike design help or harm? Names, faces, “I” language, simulated emotion: anthropomorphic touches are everywhere, tested nowhere. Vary humanlikeness systematically across task types and measure performance, trust calibration, and attachment. The field runs on intuition and outrage here; a dissertation could replace both with a theory.
When language beats graphics. Build and test a task taxonomy predicting when conversational interaction outperforms direct manipulation, and when the GUI still wins. After 40 years of graphical interfaces and 4 of chatbots, product teams choose between them by hunch. A principled map would outlive every current model.

Top: Language excels at loosely specified planning. Bottom: Direct manipulation excels at simple, concrete actions. The enduring design question is where the crossover between the two occurs.
Growing up with AI. Follow a cohort of children for several years as AI becomes their default helper. How do question-asking, research skills, and epistemic trust develop differently from prior generations? Ethically demanding, methodologically slow, and impossible to rush: precisely why it suits a dissertation rather than a product roadmap.

Children who grow up with instant assistance may develop different habits in question-asking, research, verification, and epistemic trust. Only a longitudinal study can reveal which abilities strengthen and which are displaced as the wise AI owl raises our kids.
Studies Any Product Team Can Run This Quarter
You don’t need a university for the next 10 questions; you need your own analytics, support logs, and customers. Academics have none of these, so if you work in industry, you can beat them across this entire list. Any company that has shipped an AI feature can answer them, mostly within weeks. And each answer converts directly into a product decision, which makes this list a fine menu for in-house researchers who want to prove their worth in the AI era.

There’s an immense imbalance between the resources the AI labs spend developing new models and the research efforts to discover what works for users. Luckily, some studies are cheap enough to conduct even with limited resources.
Where’s the AI button, again? Instrument every entry point to your embedded AI feature. Which paths get discovered, which are never touched, and what are the opt-out and disable rates over time? Most mainstream users now encounter AI through features bolted into familiar software; discovery data shows whether yours exists in practice.

Where’s that big red AI button again? An AI feature may exist technically while remaining absent from users’ practical experience. (In this case, the user knows every button except the AI one.)
A taxonomy of AI support tickets. Mine your support contacts that mention the AI feature and classify the failures: wrong output, confusing interaction, unmet expectations, broken trust. The top categories are your usability roadmap, already prioritized by customer pain and conveniently priced in support costs your CFO understands.
Where the time actually goes. Run a before-and-after study of one team adopting your AI feature: where in the workflow does time shift? Vendors promise savings at the generation step. But a workflow-level audit shows whether those gains survive the editing, verification, and rework that follow, or quietly evaporate.
The unmet-intent backlog. Compare what users actually ask your AI feature with what it can do. The gap produces two deliverables: a ranked list of missing capabilities, and a list of requests to decline gracefully instead of failing silently. Few analyses turn raw logs into roadmap decisions this directly.
Accept, edit, or discard? Track the fate of AI-generated drafts in your product: sent unchanged, lightly edited, rewritten, or abandoned. The edit distance between draft and final version is an honest quality metric that no satisfaction survey can match, and it’s sitting in your data already.

The fate of generated drafts is a direct quality measure. Unchanged acceptance, light editing, complete rewriting, and abandonment reveal more than a satisfaction rating.
The novelty cliff. Plot usage cohorts for the 90 days after users first touch your AI feature. Does usage decay toward zero, settle into a habit, or grow with skill? Retention curves separate genuine value from curiosity clicks, and they’re the number your next AI investment decision should rest on.

All the penguins line up for the thrill of gliding down the new ice slide. But after one try, they wander off for the next adventure, showing that the slide was merely a novelty, not a durable improvement for the penguin colony.
Is the thumbs widget measuring anything? Compute what share of AI responses receive thumbs-up/down feedback, who gives it, and whether it correlates with any behavioral outcome such as retention or task completion. Many AI quality dashboards rest entirely on this widget. Testing whether yours means anything takes a week.

An impressive quality dashboard isn’t just useless; it’s dangerous if the underlying data it renders so beautifully is unreliable.
Dogfood versus reality. Compare how your employees use your AI feature with how customers use it: tasks attempted, prompt sophistication, success rates. Teams inevitably design for themselves. Measuring the gap tells you how far your internal intuitions drift from your actual market, in numbers nobody can argue away.
The cleanup tax. Trace errors in AI-generated content downstream: corrections, rework, support contacts, confused customers. Attach a dollar figure to wrong output. Quality debates inside product teams run on anecdotes; a costed error trail converts them into budget arithmetic, which is the only language that reliably wins those meetings.

High-volume AI production relocates work into verification, correction, customer support, and rework. Following the complete error trail turns abstract quality problems into visible financial costs.
Three example prompts. Run one A/B test: does showing worked example prompts at first use increase adoption and week-4 retention of your AI feature? It’s cheap, it’s decisive, and it establishes a template for settling AI design arguments with evidence instead of the loudest voice in the room.
Questions Only the Big AI Labs Can Answer
Some questions only yield to millions of conversations or A/B infrastructure, which means only the AI labs can answer them (with privacy-preserving aggregation, of course). In fact, the labs sit on the largest usability dataset ever assembled, and they publish a thin trickle of it. Consider this list a formal request.

Big data can answer some UX questions we mortals can’t tackle. The millionaires and billionaires working in the AI labs have a responsibility to humanity to do this research for us.
The abandonment signature. What share of conversations end in mid-task abandonment, and what behavior precedes the exit? Rapid rephrasing may be the rage click of AI. Classifying failure at population scale would give conversational AI its first true usability metric. And every product team on Earth would borrow it.
Is the world learning to prompt? Run cohort analyses across years of logs: do individual users’ prompts measurably improve, plateau within weeks, or stagnate forever? Whether prompt skill grows in the wild decides how much the industry should invest in user education versus fixing the interface.
What AI is actually for. Classify the true task distribution across millions of conversations and publish how it compares with what surveys and journalists claim. Every design decision downstream depends on the real mix of use cases, and self-report has misled interface designers since the dawn of computing.
Do power features matter? Which affordances (edit, regenerate, branching, memory controls, projects) get used at all, by whom, and does using them predict retention? Usage data would reveal whether AI products are overbuilt for an imaginary power user while the median customer types one sentence and prays.
The length experiment. Does answer length or formatting causally affect task success and return rates? Over-long answers are the most common complaint about AI, yet nobody has run the definitive test. Only an A/B experiment across millions of sessions can settle it, and only a lab can run one.
Measuring success at scale. Can session outcomes (solved, partially solved, failed) be classified automatically against validated human-labeled ground truth? Without an outcome metric, AI usability can’t be tracked, compared, or improved systematically. With one, half the other questions on this page become routine dashboard queries instead of research projects.
After the first betrayal. Locate users who caught the AI in a serious error (they announce it in the chat), then track what follows: verification habits, usage drops, churn. Every log corpus contains thousands of these natural experiments; they would show how trust recovers after failure, or whether it ever does.

How do you deal with an oracle that hallucinates (as the Delphi oracle famously did)?
When the model changes underfoot. Each release quietly swaps the product under millions of users. What happens to prompt phrasing, retention, and complaints when the AI’s behavior shifts overnight? Nobody has quantified the usability cost of constant model churn, and only the labs hold the before-and-after data.
Who pays the prompting tax? Do prompt vocabulary and structure differ systematically by region, age, or expertise, and do certain phrasings fail predictably? Population-scale linguistics would reveal which user groups pay a hidden usability tax that lab studies, with their double-digit sample sizes, can never detect.
Does memory earn its keep? Do cross-session memory features measurably reduce repeated context-setting and improve outcomes, or mostly generate privacy anxiety? Compare matched users with memory on and off at scale. Vendors ship memory as an obvious win; the behavioral evidence for that confidence remains unpublished, if it exists.
Replication Studies: Regression Testing the Usability Canon
Human nature is stable; technology isn’t. That makes 30 years of usability findings a ready-made hypothesis list, and this section is usability regression testing: rerun the classic result against the new interface paradigm and see what breaks. And each project doubles as a built-in literature review, which thesis advisors love. (Questions 20 and 21 already smuggled two replications into the master’s list.)
The paradox of the active user, squared. Jack Carroll and Mary Beth Rosson showed in 1987 that users skip manuals and plunge straight into doing. Do people read AI onboarding, help pages, or prompt guides even less? If the paradox of the active user has strengthened, every documentation-based fix for AI usability is dead on arrival.

The magpie ignores the thick AI instruction manual and instead uses it as a ramp to peck several unfamiliar buttons. Are human AI users like that? Conversational AI may intensify the active-user paradox because the empty field invites immediate experimentation without preparation.
Recognition versus recall at the blank box. A core usability principle says show people their options instead of demanding they remember commands. The empty prompt field is pure recall. Do suggestion chips, command menus, or capability galleries measurably raise novice success? The oldest finding in the book meets the newest interface.

How can we address the excessive demands on recall placed by the blank-box chat UI?
Jakob’s Law among the AIs. Jakob’s Law states that users spend most of their time on other sites, so they expect your site to work like the sites they already know. Do habits formed in one chatbot transfer to its competitors, and do deviations cause real errors? The law has run the web for decades; check whether it runs chatbots too.
The verbal disagreement problem, 40 years on. My erstwhile Bellcore colleagues Furnas, Landauer, Gomez, and Dumais showed in 1987 that two people pick the same name for the same thing less than 20% of the time, dooming command languages. Do LLMs finally dissolve the problem, or does it resurface as prompts that founder on each user’s choice of words?
Banner blindness for AI buttons. Users trained themselves to ignore anything shaped like an ad. Do the sparkle icons and “Ask AI” panels now injected into every application suffer the same fate? Eye tracking or simple click-through observation would show whether forced AI features have achieved literal invisibility, which would explain a lot.
The aesthetic-usability effect, oracle edition. Humans judge attractive things as working better. Does interface polish (typography, highfalutin vocabulary, animation, branding) similarly inflate users’ ratings of identical AI answers? Swap the same content between a slick shell and a plain one, and measure how much unearned trust beauty buys. My prediction: plenty.

Aesthetic polish can increase perceived intelligence, credibility, and usability even when the underlying answer is unchanged. Beauty may buy AI systems a hefty dose of unearned trust.
Peak-end conversations. Daniel Kahneman showed that people remember experiences by their most intense moment and their ending, not their average. Do users judge an AI session the same way, forgiving mid-conversation failures if the finish lands? If so, recovery design deserves far more attention than error prevention currently gets.
Error messages without the message. Classic guidelines demand that error states say what went wrong and how to fix it. AI failures (refusals, hallucinations, half-answers) rarely announce themselves as failures at all. Audit AI failure states against the old guidelines, then test redesigns. The 1990s rules may need an update; find out which parts.
Progressive disclosure of AI power. Showing little at first and more on demand has 4 decades of evidence behind it, as I recently reviewed. Can staged revelation of AI capabilities (memory, tools, file handling) beat both the blank box and the overwhelming feature tour for novice success and retention?
The power of defaults, draft edition. Users keep defaults, from browser settings to the first search result. Is the AI’s first draft the new default: accepted not because it’s good but because it’s there? Compare how often first outputs survive against how often users request regeneration, and interview them about why.
Research About Research: Methodology for the AI Age
Every study above must survive nondeterministic output, monthly model churn, and conversations too private to read. Our methods playbook was written for stable, deterministic software, and it creaks. So the final 10 questions are research about research: whoever answers them enables everyone else on this page, which is how methodologists earn their citations.

The caliper measures the ruler, the stopwatch times the hourglass, and the evaluation apparatus evaluates another evaluation apparatus. Traditional UX research instruments must now be tested, given that the system under study is nondeterministic, continually changing, difficult to score, and often too private to inspect.
How many users for a random system? Testing with 5 users assumed the system behaves the same for everyone. When output varies with every session, how many participants, and how many trials per participant, does a usability test need to find the same problems reliably? The most-quoted number in UX needs recalculating.
Does think-aloud distort prompting? Talking while clicking changed behavior only modestly, which is why think-aloud became the workhorse of usability testing. But verbalizing may interfere far more with composing prompts, which is itself linguistic work. Compare concurrent, retrospective, and silent protocols on identical tasks before we trust another AI study.
Findings with expiration dates. Which AI usability findings survive model updates, and which rot within a quarter? Rerun a set of published studies across model versions, then propose reporting standards (model, date, settings) so future studies age gracefully. Research that spoils faster than milk needs a date stamp.

Many research findings from studies conducted with older AI models may be spoiled milk.
Synthetic users on trial. Where can AI-simulated participants validly stand in for humans in early usability work, and where do they mislead? Run identical studies with real and simulated users, compare the problem lists, and map the safe zone. Half the industry already believes the answer. But nobody has measured it.

Synthetic research participants may be useful for some early checks while systematically missing confusion, hesitation, emotion, context, and compensatory behavior. Their safe zone must be measured rather than assumed.
Scoring the unscorable. Open-ended AI output has no single correct answer, which wrecks the traditional definition of task success. Develop and test practical rubrics that independent evaluators apply consistently, and report the inter-rater reliability. The field can’t measure improvement until it agrees on what counts as success.
Studying conversations too private to read. People’s AI chats contain therapy, medical fears, and trade secrets. Develop and compare donation-based, diary-based, and redaction methods for studying real usage without reading what users would never show a researcher. Ethics boards will demand these methods soon; better to have them ready.

The hermit crab reasonably views its shell as too private for the turtle to study using traditional user research methods. Some of the most important AI conversations may contain therapy, medical fears, intimate relationships, or trade secrets. Useful research methods must capture behavioral evidence without demanding unrestricted access to private lives.
Do benchmarks predict usability? Models are ranked by benchmark scores; users experience something else entirely. Correlate benchmark deltas between model pairs with human task success and satisfaction on realistic work. If the correlation turns out weak, the entire industry is optimizing the wrong number, which would be worth knowing.

The over-decorated prize horse won all the benchmark awards, but the humble donkey is the one that actually opens the gate. Are AI benchmarks similarly bad at predicting usability?
Longitudinal research on quicksand. How do you run a 90-day diary study when the product updates weekly? Develop designs (rolling cohorts, version logging, event-triggered sampling) that separate user change from product change. Most longitudinal AI findings published so far confound the two, which should worry us more than it does.
What counts as a task? Conversations sprawl across topics; sessions blur into each other; one “question” can hide 5 goals. Define and validate units of analysis for conversational AI so that metrics like success rate and duration measure comparable things. Unsexy, foundational, and cited forever if done well.

One conversation may contain several tasks, while one user task may span several sessions. Until the field defines a defensible unit of analysis, success rates and task durations will measure different things under the same names.
Wizard of Oz, retired or rehired? Faking the AI with a human behind the curtain built early voice-interface research. Now the AI is real but unpredictable. When does Wizard-of-Oz prototyping still beat testing live models, and how do you fake nondeterminism honestly? The oldest trick in HCI deserves a formal re-evaluation.
Conclusion: Somebody Should Take Notes
The biggest interface change in computer history is running as an uncontrolled experiment on billions of people, and the experimenters are too busy shipping to take notes. Be the one who takes notes.

ChatGPT recently gained a billion weekly users, passing this milestone less than 4 years after the product's release. Unprecedented change requires unprecedented research efforts, and yet almost no user research is being done on the impact of AI and how to design it better. Somebody must rise to the challenge: you!
So pick one question sized to your resources, and get going.
A final word to students. The research gap and the job squeeze are the same fact seen from two sides: the field is so new that nobody has the answers, so anyone who produces one owns something scarce. A degree certifies knowledge; a completed study demonstrates agency. You’ll walk into interviews carrying evidence instead of adjectives.
And when you finish one of these studies, tell me what you found. I want to read it.

A reminder: these are the 8 types of insights we stand to gain from answering my 76 open research questions about AI usability. (All images in this article made with GPT Image 2)
