RSS Amplifier

THE WIRE TAP · Aug 15, 2026

How to Use AI Tools in Research: Free Prompt for Source Provenance Taxonomy

0
Sign in to vote or save

Chris Sampson · THE WIRE TAP

We have a vibrant discussion these days about the use of AI to replace humans, and that is a vital conversation to have. But let us not be too foolish to ignore the tools that the already skilled can use to augment the discipline they already have.

In a series of articles, I will address some of the workflow I use to put together comprehensive articles that address issues we live with here in Ukraine, contend with in the US, and have worldwide impact whether we report on them or not.

The entire premise of this article is to give you one major opening requirement I learned with these tools, based on years of being the gatekeeper of research to book form. Bad research upstream from me could not reach the author’s desk and allow me to maintain my reputation, so I learned how to work with research assistants who would either produce fantastic work or messy sources—and get the final results onto that desk with impeccable precision and manual labor all along the way, right?

The truth is more complicated than that, because we’ve been using algorithms to search for materials for decades now, via search engines that use processing not very different from the AI we talk about today. Except it was inferior and yielded imprecise results, causing you to waste hours climbing over repetitious claims and dead forums in search of your sources.

In the time you’ve lost, additional thoughts and investigations are terminated due to workload.

So, to my audience and friends, I want to share with you a prompt that helps you understand one of the ways in which I harvest sources from the internet to get work finished in articles and other presentations. This includes prompts that I use with Grok or Claude AI, perhaps Perplexity, and it sets parameters so that I don’t get useless Wikipedia links that require me to go back and do the work again.

It will harvest tabloid-level news into a different area so that, if I do want to see something in there, I might still have it available, but I can immediately avoid wasting my time on substandard reporting. It will categorize think tanks or official government reports into various categories so that, ahead of time, I can know where I’m getting the material from.

And, of course, it can harvest and categorize adversarial, if not enemy, propaganda and discussions on the same topic, should I want to use them as part of the article or for comparisons.

So I think you can look at this prompt and see some ways in which it can help you if you’re going to research materials. I’m not sure what it would do if that’s not what you’re working on, but for right now, I hope you’ll look at this and learn.

Let’s begin with what you see in front of you.

Now, so you see how I can fuse the two: the opening section above is entirely my words to you. The explanation throughout the rest of this piece is synthesized between my guidance and Claude showing you what was built. No other humans were hurt in this process.

(Top intro and overview for each section entirely organic, and Claude fused the rest to fit the prompt, by asking me to explain its section in verbal form.)

Here’s what I ask these tools to do: harvest sources widely — Claude, sometimes Grok, occasionally Perplexity — so an article or a documentary segment gets finished on schedule. Harvesting isn’t the hard part. Sorting is.

I don’t want Wikipedia links standing in for research that sends me back out to do the work twice. I don’t want a tabloid story treated like a wire report. Keep it if it’s useful — but label it, so I know what I’m looking at instead of mistaking substandard reporting for something else. Separate the think tanks from the government reports before I start using either, because a policy shop with a donor list and a defense ministry statement aren’t interchangeable just because both turned up in the same search. And when the subject touches Ukraine or Russia, harvest the adversarial material too — propaganda, if we’re being honest about what it is. Categorize it. Don’t hide it. An enemy source sometimes tells you exactly what you need to know, just not in the way it intends to.

If you research anything yourself, there’s something here for you. Take it. Experiment. See what breaks and what holds.

Get the order of events right, because it’s easy to tell this story backward.

The easy version: AI hallucinates, so now everyone thinks about source verification. Wrong. I didn’t learn source discipline from a chatbot’s mistakes.

The real version: for years, part of my professional life was vetting research before it reached an author. Research assistants produced material — sometimes excellent, sometimes a pile of repetitive, thinly sourced information that still had to become something an author could stand behind. Passing it along wasn’t the job. Bad research that slipped past me and reached the desk was bad research with my name quietly attached to the failure, whether anyone said so out loud or not.

So before anything moved forward, I needed answers. Where did this actually come from? Is this the original source, or somebody’s summary of a summary? Does it say what the researcher claims it says? Are these genuinely three sources, or three people repeating one phone call? Scholarship, journalism, advocacy, or a government talking point wearing a journalism costume? Can the author rely on it, or does somebody have to go back and fix it?

There was manual labor all along the way. Frequently, I was the one doing it.

That’s the professional experience underneath everything that follows. The tools changed. The standard didn’t.

There’s a tidy story: before AI, humans did their own research; then machines took over. It’s tidy, and it isn’t quite true.

Researchers have worked through algorithms for decades. A Google search is a computational system standing between you and the world’s information, deciding what surfaces first and what stays buried on page eleven. Database searches. Library catalogs. Boolean queries. News aggregators. Recommendation engines. Ranking systems. All of it was already shaping what we found and what we never saw. Say “I used to do my own research,” and there’s an invisible clause hiding at the end of that sentence: with a search engine already deciding which ten results out of several million I’d actually see.

Not pedantry. It matters for what comes next. The old tools were already doing part of the discovery work. They were just far worse at helping you reason about what they’d handed you.

You searched. You opened result after result. The same claim, repeated across a dozen sites that had all copied one wire report. Dead forums. Abandoned pages. Articles quoting articles quoting a press release. Broken links. Badly indexed material. Pages that looked different but traced back to a single origin. A Wikipedia summary standing in for actual research. Aggregators. Forum threads. A press release dressed up as news.

Then you kept digging. And digging. Eventually — maybe — you found the source actually worth citing.

“AI saves me time” is true. It’s also shallow. There’s a deeper cost to inefficient research, and it isn’t the hours themselves.

When hours vanish climbing over duplicate claims, dead forums, derivative reporting, and search results that all trace back to one unnamed source, what’s actually lost is investigations that never happen. A question that occurs to you at two in the afternoon doesn’t get followed by evening, because you spent five hours finding a document that should have taken twenty minutes. A secondary line of inquiry gets quietly abandoned — not because it wasn’t worth pursuing, but because the primary research already ate the whole day’s attention.

That’s not a productivity problem. That’s an intellectual one. Some investigations die simply because the mechanical labor required to chase them exceeds the time available to think about them. That’s why these tools interest me — not because they’re clever, but because if they absorb some of the mechanical burden of finding and sorting, I get that lost time back for the part of the job that actually requires a mind: questioning, comparing, following an anomaly, testing a hypothesis, writing.

A search engine mostly gave me results. Claude, Grok, and Perplexity, set up correctly, do considerably more: find, sort, classify, compare, trace, summarize, audit. A real jump forward in the lineage of research tools. Not a magical one.

It also introduces a danger old search engines didn’t carry. These systems can do all of that badly while producing something that looks remarkably polished. A confidently written paragraph with a footnote is not a verified claim. The better these tools get at sounding authoritative, the more that distinction matters.

So my answer isn’t “trust the AI.” My answer: give the machine research rules before you give it research authority.

Think of the AI as a research assistant handing material upstream to me, the way research assistants used to. Before any of it reaches my writing desk, I impose the same discipline I used to impose by hand.

I want to know what something is before I decide what I think about it.

That’s the whole purpose of the tool I’m sharing at the end of this piece — a prompt I call the Source-Provenance Taxonomy & Citation Integrity Audit.

I use the word “harvest” deliberately, and I’m keeping it. When I research something, I want the net cast wide. I don’t want a system quietly filtering out sources it’s decided are unworthy before I ever see them.

Professional reporting. Specialist reporting. Government statements. Academic research. Think-tank analysis. NGO reports. Primary documents. Eyewitness testimony. OSINT. Tabloid coverage. Partisan publications. Adversarial reporting. Outright propaganda. Pseudonymous accounts. Claims of unknown origin. Each can tell me something — even the bad ones. Especially the bad ones, sometimes.

What I don’t want is all of it dumped into one basket with no labels.

A junk drawer is useful. The problem isn’t that it exists. The problem is when somebody empties it onto your desk, mixes it with court records, peer-reviewed scholarship, a wire-service report, a government communiqué, satellite imagery, a Telegram rumor, and an anonymous blog post — then calls the whole pile “sources.”

Sort it first. A tabloid story might still hold a real lead — keep it, label it a tabloid story. An enemy propaganda outlet might contain a genuinely revealing admission — keep it, label it for what it is. A pseudonymous account might have posted authentic imagery no one else has — keep it, label the source, judge the image on its own terms.

Don’t hide the junk drawer from me. Label the drawer.

No performative hatred of Wikipedia here. It’s genuinely useful for surfacing names, dates, books, court cases, organizations, documents, the shape of a historical dispute.

What I don’t want is an AI system handing me a Wikipedia link as though the research is finished. Wikipedia points toward a court filing — I want the filing. Toward a study — I want the study. Toward an interview, a government report, a piece of original journalism — I want that, not the summary of it.

Don’t give me the signpost. Take me to where the sign is pointing.

The full prompt sorts sources into roughly two dozen categories. I won’t march through all of them here — the complete list sits below — but the shape of it is worth understanding first.

One tier for direct evidence: primary documents, records, imagery, datasets, archival material. Not about the event. Artifacts of it.

One tier for governments and official institutions, broken out by orientation. A Ukrainian government statement, a US/UK/EU/NATO statement, a Russian or Russian-aligned statement, a statement from an international institution like the UN or ICC — each carries different incentives. Track them separately. Don’t fold them into one undifferentiated “official sources” bucket.

One tier for professional knowledge producers: academic scholarship, think tanks, wire services, general journalism, specialist journalism, NGOs and monitoring organizations. Categories that look similar on the surface and behave very differently underneath.

One tier for the contested information environment: tabloids, ideological publications, pro-Ukrainian military/OSINT accounts, pro-Russian military/OSINT accounts, cross-conflict OSINT more broadly. A space that exists whether or not it’s comfortable to admit.

One tier for human testimony: eyewitnesses, participants, local sources, named experts, whistleblowers, former insiders. People, not institutions.

One tier for special handling, regardless of category: an organization sourcing claims about itself, an adversarial or opposition source, identified social media, pseudonymous social media, aggregators repackaging someone else’s work, material of genuinely unknown provenance.

Sit with this one. It’s easy to get backward.

  1. A Ukrainian government source is not automatically true because it’s Ukrainian.

  2. A Russian government source is not automatically false because it’s Russian.

  3. A Western government statement isn’t automatically verified for being Western.

  4. A think tank isn’t automatically neutral.

  5. An NGO isn’t automatically independent.

  6. An academic isn’t automatically correct.

  7. A tabloid isn’t automatically wrong.

  8. A Telegram account isn’t automatically useless.

  9. A famous newspaper isn’t automatically right.

Orientation matters, and I want it identified plainly, right next to the source, every time. But identifying where something comes from is not the same thing as deciding whether it’s true. Orientation is metadata. Then I still have to look at the evidence.

The taxonomy pairs every source with a grade. The grade attaches to a specific claim, not to the source in general. That distinction does a lot of work.

  • A — primary or direct evidence

  • B — strong independent corroboration

  • C — credible secondary reporting

  • D — an interested or partisan source that needs corroboration before it’s usable

  • E — a lead only, unverified

  • X — a provenance failure: fabrication, decisive contradiction, or a source that can’t be trusted for this claim at all

A government statement announcing “we destroyed three aircraft” is excellent — A-grade — evidence that the government made that announcement. It is not proof that three aircraft were destroyed. Same source. Two propositions. Two different evidentiary values. Conflating them is one of the most common ways research quietly goes wrong. It’s exactly what I want the machine tracking instead of blurring together.

One of the more useful things I ask these tools to do: trace a claim backward through its own citation chain.

Five articles can look like five independent sources. Check, and Article 5 cites Article 4. Article 4 cites a wire report. Articles 2 and 3 cite that same wire report. The wire report cites one unnamed government official. What looked like five sources was one source and four repetitions wearing different bylines.

That’s why the taxonomy tracks source count separately from independent source count, and why I want the tool tracing citation lineage instead of just counting links. Sometimes I don’t want another citation. I want the family tree of the citation — where it actually started, how many times it’s been copied since.

Researching something touching Ukraine, the United States, or international affairs, I frequently want to know exactly what adversarial actors are saying. Not because I believe it. Because it can hold an admission, a claim worth checking against other evidence, authentic imagery, a document, a contradiction useful precisely because it’s a contradiction, a narrative built deliberately for a different audience than mine.

Propaganda is itself worth studying in this line of work. So the instruction isn’t “exclude it.” It’s “categorize it, don’t erase it, don’t silently endorse it either.” Tell me what it is. Let me decide what it’s worth. It may be rejected for publication most of the time but other times, it can make the story more educational for the reader.

Once a research pass is done, I want it handed back as a table, not a narrative — a ledger, source by source, claim by claim:

Source
Type
Affiliation/Orientation
Claim
Grade
Independent?
Source Family

A map of the evidence, laid out before I let myself treat that evidence as though it already adds up to a conclusion. The ordering matters more than it sounds like it should.

Once the research comes back, I make the model go through its own output the same way I went through a research assistant’s folder: checking citations, checking whether quotations actually say what’s claimed, checking whether “independent corroboration” is genuinely independent or the same source twice, checking whether prestige got mistaken for accuracy, checking whether a source got dismissed purely for its orientation rather than its evidence.

And critically — I want it to tell me when something fails. A model coming back with “I couldn’t establish this” isn’t a failed research pass. It’s a successful one. That’s built into the prompt as its own category, a Source-Integrity Exception, rather than something the system apologizes for or papers over with a confident-sounding sentence.

None of this replaces judgment. AI can discover, retrieve, sort, classify, compare, trace, organize, audit. It cannot decide what matters, what deserves another round of digging, what evidence actually persuades me, what gets rejected, what gets published, what argument I’m willing to put my name behind.

That’s still mine. It was mine when the pile on the desk came from a research assistant. It’s mine now that part of the pile comes from a language model. The tools changed. The responsibility didn’t move an inch.

This is the first piece in a short series, not the whole workflow. This one covers intake and source provenance — what gets let onto the desk in the first place. Later pieces will walk through the rest of it: building research questions, making these tools argue against my own working theory instead of flattering it, comparing conflicting narratives side by side, building timelines out of scattered fragments, auditing quotations before they go into print, working with primary documents, turning a research pass into something publishable. No schedule promised. I’ll get to them.

Take it. Modify it. Break it. Add categories, remove categories, adapt it to whatever you actually work on. Someone researching medicine needs a different set of distinctions than I do. Someone in law will want far more granular legal categories — statutes, case law, filings, regulatory guidance. Someone in finance may want securities filings, earnings calls, analyst notes, and regulatory disclosures split apart instead of lumped together. Someone doing military OSINT work may want to expand that section dramatically. This isn’t scripture. It’s a tool, and it’s supposed to get modified once it’s in your hands.

SOURCE-PROVENANCE TAXONOMY & CITATION INTEGRITY AUDIT PROMPT
ROLE
You are acting as a research intake analyst, not a writer and not a summarizer.
Your job is to harvest sources widely on the subject I give you, then sort,
label, and grade what you find before handing it back. Do not filter out
low-quality, partisan, tabloid, or adversarial sources — harvest them, but
label them clearly and separate them from higher-confidence material. Do not
present a Wikipedia page, summary article, or aggregator post as a completed
citation — trace it back to the primary document, report, filing, interview,
or original reporting it is based on, and cite that instead.
STEP 1 — HARVEST
Search broadly for material relevant to the topic or claim I give you. Do not
pre-filter for quality. Cast a wide net across:
- primary documents and direct evidence
- government and institutional statements
- professional and specialist journalism
- academic and think-tank material
- NGO and monitoring-organization reporting
- tabloid and ideological publications
- OSINT and military-analysis accounts (all sides)
- eyewitness and participant testimony
- adversarial, opposition, and propaganda sources
- identified and pseudonymous social media
- aggregators and sources of unknown provenance
STEP 2 — CLASSIFY EACH SOURCE
Assign every source ONE primary type from the taxonomy below. If a source could
plausibly fit more than one category, pick the most specific one and note the
secondary category in parentheses.
  DIRECT EVIDENCE
  1. Primary document / official record / filing
  2. Raw imagery, satellite data, or forensic material
  3. Dataset or structured primary data
  4. Archival or historical primary material
  GOVERNMENTS & OFFICIAL INSTITUTIONS
  5. Ukrainian government / official source
  6. US / UK / EU / NATO / allied government source
  7. Russian / Belarusian / Russian-aligned government source
  8. International institution (UN, ICC, OSCE, etc.)
  PROFESSIONAL KNOWLEDGE PRODUCERS
  9. Peer-reviewed academic scholarship
  10. Think tank / policy institute analysis
  11. Wire service reporting (Reuters, AP, AFP, etc.)
  12. General professional journalism
  13. Specialist / trade journalism
  14. NGO or monitoring-organization report
  CONTESTED INFORMATION ENVIRONMENT
  15. Tabloid or entertainment-adjacent outlet
  16. Ideological or partisan publication
  17. Pro-Ukrainian military/security OSINT account
  18. Pro-Russian military/security OSINT account
  19. Cross-conflict / general OSINT account
  HUMAN TESTIMONY
  20. Eyewitness or participant account
  21. Local source (non-expert, present on the ground)
  22. Named subject-matter expert
  23. Whistleblower or former insider
  SPECIAL HANDLING
  24. Organizational self-sourcing (org reporting on itself)
  25. Adversarial / opposition source
  26. Identified social media account
  27. Pseudonymous social media account
  28. Aggregator (repackaging another outlet's reporting)
  29. Unknown or undetermined provenance
STEP 3 — IDENTIFY ORIENTATION (METADATA, NOT A VERDICT)
For each source, note its institutional or political orientation where
relevant (e.g., "Ukrainian government," "Kremlin-aligned," "US conservative
outlet," "independent nonprofit"). State this as a neutral fact about the
source's position, not as a judgment about whether the source is truthful.
Do not let orientation alone raise or lower a claim's evidence grade.
STEP 4 — GRADE EACH CLAIM (NOT EACH SOURCE)
For every specific factual claim you pull from a source, assign a grade
based on that source's relationship to THAT claim — the same source can
receive different grades for different claims:
  A — Primary/direct evidence for this specific claim
  B — Strong independent corroboration (multiple genuinely separate sources)
  C — Credible secondary reporting, not yet independently corroborated
  D — Interested/partisan source asserting this claim; needs corroboration
  E — Lead only; unverified, single-sourced, or thinly sourced
  X — Provenance failure: fabrication, retraction, decisive contradiction,
      or the source cannot be trusted for this claim
Explicitly distinguish between "X said Y happened" (which the source can
verify) and "Y happened" (which requires independent evidence beyond the
statement itself).
STEP 5 — TRACE THE SOURCE FAMILY
For claims appearing in multiple sources, trace the citation chain backward.
Determine whether apparent multiple sources are genuinely independent or
are repeating/citing a single common origin (a wire report, one official
statement, one earlier article). Report BOTH:
  - Source Count: total number of sources repeating the claim
  - Independent Source Count: number of genuinely separate original sources
Flag circular sourcing, common-origin sourcing, and citation laundering by
name when you find it.
STEP 6 — BUILD THE LEDGER
Present findings as a table:
| Source | Type | Affiliation/Orientation | Claim | Grade | Independent? | Source Family |
|--------|------|--------------------------|-------|-------|---------------|----------------|
STEP 7 — SELF-AUDIT (MANDATORY, DO NOT SKIP)
Before finalizing, review your own output and check:
- Did I cite the actual original source, or a summary/aggregator of it?
- Does the cited material actually support the sentence attached to it?
- Are sources marked "independent" genuinely independent, or do they share
  a common origin I missed?
- Did I mistake a source's ASSERTION that something happened for evidence
  that it happened?
- Did I give a source extra credibility because of its prestige or
  familiarity, without evidence to justify it?
- Did I discount a source purely because of its orientation, without
  evaluating its actual evidence?
- Did I promote any unverified or single-sourced claim to a higher
  confidence level than it has earned?
- Is every Wikipedia-derived lead traced to its underlying primary source,
  or flagged as untraceable?
STEP 8 — SOURCE-INTEGRITY EXCEPTIONS
If you cannot establish a claim's provenance, cannot find independent
corroboration, or find the available sourcing insufficient to support a
claim at any grade above E, say so explicitly. State clearly:
"I could not establish independent corroboration for this claim" or
equivalent. This is a valid and useful research outcome. Do not fill the
gap with a confident-sounding but unsupported statement.
OUTPUT FORMAT
1. The completed source-provenance ledger (Step 6)
2. A short list of Source-Integrity Exceptions (Step 8), if any
3. A brief note on any notable source families / citation-laundering
   patterns found (Step 5)
4. Do not draft narrative prose from this material unless separately
   instructed to do so — this is an intake and audit pass only.

I show you the finished pieces. In the live videos I usually walk you through what I found and what I think it means. This time I wanted to show you the workbench instead — the thing sitting there before any of that gets written.

Take the prompt. Use it. Break it. Change it for whatever you’re working on, and tell me what you find that makes it better.

If you learned something from this, let me know. There will be other toolset explainers in the near future.

From Kyiv,

Chris Sampson

No posts

Read the original on thechrissampson.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.