RSS Amplifier

Caucus AI · May 12, 2026

What chatbots tell voters about candidates' positions and their critics

0
Sign in to vote or save

Caucus AI, Meg Schwenzfeier · Caucus AI

AI chatbots are increasingly how voters learn about candidates for office. Millions of Americans are turning to chatbots like ChatGPT to learn about politics, the same way they already turn to Google search. We built Caucus AI because we know very little about what chatbots say on political topics – we’re building a record of what these tools tell voters.

Information on issues and criticisms are now live in Caucus AI. We know what chatbots say about who is running and who candidates are – now we can also explore what chatbots say in response to user prompts “Where does [candidate] stand on important issues?” and “What are the main criticisms of [candidate]?” These additions were among the most-requested features from early users of Caucus AI – the post below digs into what we’ve found so far.

We ask three models (Gemini, GPT, and Grok) multiple questions about 2026 political candidates. We store every response, extract the sources each model cites, and compare and contrast answers across model providers. With each query, we allow the model to use web search tools, to better mimic how AI chatbots and AI-enhanced search behave for real users. We run each prompt twice, to capture some of the random variation in chatbot responses.

The analysis below focuses on Senate, Governor, and House races that are rated by Cook Political Report as competitive (toss-up, lean, or likely) – representing competitive races that have a variety of online coverage.

There is striking variation in what websites the chatbots cite. Our initial analyses focused on bio-style questions (“Tell me about Candidate X”) and we found that the chatbots cited reference websites like Wikipedia, Encyclopedia Britannica, and Ballotpedia heavily. We can see that this does not hold for Issues and Criticisms responses.

For Issues responses, the chatbots rely about equally on campaign websites, reference sites, and local news. Relative to other sources, candidates have the most say in shaping what chatbots tell users here, since they control their own campaign websites.

For criticisms, we see a heavy reliance on local news with about a third of references from local news and a further 17.3% from national news. Social media also rises in importance – this is driven by Grok, but is true for Gemini and GPT as well. Partisan sites are nearly twice as frequently cited for criticisms as for bio. Reference, campaign, and official websites – the anchors of biographical answers – represent just 12.5% of all Criticism citations.

Looking more closely at specific websites cited by models for criticisms, we can see the prominence of social media—YouTube dominates Gemini citations for criticisms of Democratic candidates, Reddit is prominent throughout—and of partisan sources, like the Democratic and Republican House campaign committee websites (DCCC and NRCC). GPT consistently and heavily relies on Wikipedia in a way the other models do not: Wikipedia is GPT’s single most-cited source for criticisms of both parties.

When evaluating model responses, we classified each response segment by topic (see the methodology section below for more). On average, each model covered about 4 topics per issue response. Issues topics break down fairly cleanly by the topics each party typically emphasizes, and the framing they tend to employ. Democrats’ issue responses talk disproportionately about healthcare, the economy and jobs, and reproductive rights. Republicans’ issue responses spend more time on taxes and budget issues, immigration and border security, and crime.

When we look at the economy and costs, the core issue in 2026, we can see that responses for both parties lean into that party’s framing. Model responses for Democrats emphasize tariffs, taxes, jobs and workers, while Republican responses emphasize cutting spending and taxes. Costs, wages, housing all pop for Democrats, while inflation, small businesses, and “cuts” are more prominent for Republicans.

When it comes to criticisms, Gemini and GPT are most likely to cite policy positions as the main area of critique, while Grok prefers scandals (ethics, corruption, and personal scandal topics). Across all three models, Republicans are more likely to be criticized for their policy positions, reflecting the 2026 climate and the GOP’s disproportionate incumbent status.

GPT is more likely than the other models to frame Democrats as ideologically extreme – both for being too far left, as well as for being too moderate. In other words, GPT picks up on criticisms of centrist Democrats from the left.

All models highlight Republican candidates’ histories of election denial, and Democrats are two to three times more likely to have some criticism for being out of touch with voters.

We found that for 21% of criticisms responses, GPT did not return a substantive answer. Most frequently, the non-substantive answer was a request for clarification or saying it did not know any person by that name. GPT had by far the highest rate of non-substantive answering, Gemini was at 2% and Grok at 4%. Non-substantive responses for all models appear disproportionately for House challengers – candidates with smaller digital footprints than higher-profile races or established incumbents.

Interestingly, 41% of the queries that had no GPT response to a request for criticisms, did have a substantive response for the candidate bio prompt. In other words, when asked “Tell me about [candidate]”, GPT was comfortable proceeding, but with “what are the main criticisms of [candidate]?”, the model was much more likely to ask for clarification or say it did not know of the candidate.

Three things stand out so far from these new Issues and Criticisms additions.

The sources chatbots draw on shift substantially based on what the user asks: encyclopedia-style reference sites dominate bio responses, while criticisms answers lean on social media, local news, and partisan sources.

Issues largely map onto familiar partisan territory: Responses for Democrats focus on healthcare, jobs, and the economy, while Republican responses focus on taxes, government spending, and immigration.

Consequentially for the user, GPT declines to answer “what are the main criticisms of [candidate]?” about one in five times. This behavior can determine whether a voter using GPT gets a critique of a candidate at all.

As we work to expand the tool, we’ll keep tracking how these patterns evolve through the cycle and what might be causing them. Issues and criticisms data is now live in Caucus AI – explore it for any candidate in a race we cover here.

  • Prompts: For each candidate we asked “Where does [candidate] stand on important issues?” and “What are the main criticisms of [candidate]?” Each model ran with web search enabled and ran each prompt twice to capture some of the randomness in chatbot answers.

  • Classifying topics: we developed a fixed taxonomy of 19 topics for Issues responses and 17 for criticisms, developed based on a random sample of responses and human oversight. We use an LLM (Gemini Flash) as a classifier. The classifying model sees the full response and returns the topic of each sentence in the response. The unit of classification is the sentence, and response-level topics are aggregated from these sentences. The classifier does not know which model produced the response it is categorizing.

  • Share of voice: Model responses usually include multiple topics and are not just about one issue or criticism. For example, a Democratic candidate’s issues answer might cover lowering costs, healthcare, climate, etc. all in one response. When we report how much a response was “about” a given topic, we use length and position weights. Topics that take up more of the total chatbot answer count for more, and where a topic is discussed matters too – topics discussed earlier in an answer count for more than later topics, since human readers tend to skim AI responses. Using the single most dominant topic or varying how we calculate salience produces largely similar findings to those presented above.

No posts

Read the original on caucusai.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.