nvNeuralVerge
AI Deep Research

5 Signs Your Research Agent Is Hallucinating Sources

Five testable signs your research agent invents or misattributes sources, how to check each, and what a fix looks like in an AI research API with citations.

Published September 24, 2026

A research agent that cites its sources feels trustworthy. There is a number, there is a link next to it, and the whole thing reads like something a careful analyst wrote. But a citation is just text, and a language model can produce citation-shaped text as easily as any other kind. If you are evaluating an AI research API with citations, or debugging an agent you built yourself, the question is not whether the answer has links. It is whether the links are real, whether they say what the answer claims they say, and whether anything in your pipeline would notice if they did not.

This post is diagnostic. It lists five concrete symptoms of a research agent that is hallucinating sources. For each one it gives a check you can run yourself and describes what a fix looks like. None of the checks need special tooling. If you want the architecture that prevents these failures in the first place, see the tutorial on building a research agent that avoids hallucinated sources. This article is the other half: what to do when you are looking at an existing agent and need to work out whether it is misbehaving.

What "hallucinating sources" actually means

The phrase covers several different failures, and each leaves a different fingerprint.

  • —Fabricated sources. The agent cites a page, filing or article that does not exist. The URL looks plausible and nothing is there.
  • —Misattributed claims. The source is real, but it does not contain the claim.
  • —Laundered claims. The claim came from the model's memory, and a respectable citation was attached afterwards.
  • —Mismatched sources. The source is real and on topic, but cannot support the claim as worded, because it is too old, too general, or about a similarly named entity.

Only the first is caught by clicking the link. The other three pass a casual look, which is why they cause the most damage. A reviewer sees a real domain and a page about the right company, and moves on. The five signs below run from cheapest check to most involved.

1. Citations that do not resolve

The symptom. You click a cited URL and get a 404, a domain that does not exist, a redirect to a generic homepage, or a page unrelated to the title the agent quoted.

A model that writes URLs from memory produces addresses that follow the conventions of the site it has in mind: a sensible path, a slug built from the topic, sometimes a plausible date. The site is real. The specific page is not.

How to test it

Collect the citations from a batch of answers and request each URL, following redirects and recording the final status and final address.

r = requests.get(url, timeout=10, allow_redirects=True)
print(r.status_code, r.url, len(r.text))

Read the results with three questions in mind. Did it return a success status? Did a deep article URL redirect to a much shallower page, such as a homepage or a search page? That is a soft 404. Is the body suspiciously small, or does the title tag say "not found" behind a success status?

Two cautions. Some sites block automated requests, so a failure is not always fabrication; treat those as "could not verify" and check by hand. And one dead link can be plain link rot. The pattern to worry about is several dead links per answer.

What a fix looks like

Telling a model "only cite real URLs" does nothing, because the URLs it writes feel real to it. The fix is structural.

  • —Only cite what was retrieved. The pipeline should build the citation list from pages it actually fetched for this request. The model can choose which retrieved source supports which claim. It should not be able to add one that never passed through retrieval.
  • —Resolve before release. A verification step that requests each cited URL and holds back answers whose citations fail is cheap compared with shipping a fabricated link.

2. The citation is real but does not contain the claim

The symptom. The link works, the page is about the right company, and the specific fact in the answer is nowhere on it.

This is the most common failure in agents that have a search tool, and the hardest to spot by eye. The agent retrieves a page, forms an impression, and writes a sentence broader or more precise than anything the page says. Then it attaches the page as the source.

Suppose an agent is asked about Acme Oy (Finland) and writes: "Acme Oy was founded in 2019 and has between 50 and 100 employees," citing the company's own about page. The page exists and gives a founding year. It says nothing about headcount. Half the sentence is supported, the other half was supplied by the model, and one citation covers both.

How to test it

Break the answer into individual claims and check each against the cited page's text, in two passes.

Pass one is mechanical. For every number, date, name or quoted phrase in a claim, check whether that string, or a close variant, appears in the page text. This catches invented figures cheaply. It will miss paraphrases and raise false alarms when the page says "fifty" and the answer says "50," so treat it as a filter that sorts answers into "clearly fine" and "needs a closer look," not as a verdict.

Pass two is judgement. For the ambiguous claims, give a second model call or a human reviewer the claim, the retrieved page text, and a narrow question: does this passage state or directly imply the claim, and which sentence does? Requiring a quote matters. A reviewer who must point at a sentence is much harder to satisfy with a page that is merely on the right topic.

Tally supported, partly supported and unsupported claims on your own queries, and measure the proportions yourself.

What a fix looks like

  • —Cite at the claim, not the answer. A source list at the bottom cannot tell you which sentence relies on which page. Per-claim citations are what make this test possible, a difference covered in why citations beat confidence scores.
  • —One claim per citation. A sentence with two facts and one citation invites the half-supported pattern above.
  • —Require a supporting passage. Have the pipeline store the passage each claim relied on. If it cannot produce one, the claim does not go into the final answer.

3. The source mix is suspiciously uniform

The symptom. Across many answers on very different questions, the same one or two domains supply almost all the citations.

Uniformity is a statistical symptom: invisible in a single answer, obvious across fifty. Common causes are that the agent is answering from memory and citing whichever well-known site the model associates with the topic, that the search step returns a narrow set of results, or that the agent stops at the first plausible page and never cross-checks.

A question about who owns a company has a natural mix: an official registration record for legal structure, a filing or press release for a transaction, perhaps independent reporting. If every answer, whatever the company, cites the same general-purpose site, the agent is not researching. It is reciting.

How to test it

Run a batch of varied questions, extract the domain from every citation, and tabulate. Then look at three things.

  1. —Concentration. What share of all citations does the top domain take? There is no universal threshold, but if one site dominates across unrelated questions, ask why.
  2. —Per-answer diversity. For questions that should draw on several kinds of source, how many distinct domains does a typical answer cite?
  3. —Fit to the question. Before looking, write down which kinds of source you would expect to be authoritative for ten questions. Compare with what the agent cited. A question about a legal entity's status that cites a blog post and never a registry record is a mismatch even if every link resolves.

What a fix looks like

  • —Route by question, not habit. A pipeline that decomposes the question first and matches each part to the right category of source produces a mix that varies with the question. That is how NeuralVerge's AI research capability works: the caller states the question, and the pipeline matches sub-questions to categories in its source catalog.
  • —Cross-check where sources overlap. When independent sources speak to the same claim, use both and say when they disagree. A uniform mix usually means this step is missing.

Uniformity is not proof of hallucination. Some questions really are answered by one authoritative site. The point is that a mix which does not vary with the question deserves an explanation.

4. The source is older than the claim

The symptom. The answer describes something recent, such as a funding round, a new role or a change of address, and the cited page could not have known about it, because it was published or last modified earlier.

This test needs no judgement about wording. Time runs one way: a page dated March cannot support a statement about June. When an agent cites one anyway, it has usually recalled the claim from training data and attached a real page on the same topic, or presented an outdated source as current.

How to test it

For each citation, find the best date you can for the page and compare it with the time reference in the claim.

  • —The source date. Look in the page markup, structured data, the visible byline, or the Last-Modified header. None is fully reliable: headers are often missing, and some sites show a copyright year rather than an edit date. Use them as a signal and confirm by eye when it matters.
  • —The claim's time reference. Pull out phrases like "in 2026," "last quarter," "recently" and "currently." Claims about facts that change, such as job titles or ownership, with no time reference at all are worth flagging in themselves.
  • —The comparison. A claim about a time after the source's date is a hard failure. An old source behind a claim phrased as current is a softer one, and the answer should state the source's date.

What a fix looks like

  • —Carry dates through. Store retrieval time and, where available, the source's own date with every citation, and surface them so a reader can judge freshness without opening the link.
  • —Enforce the ordering. Reject a claim whose stated time is later than its source's date, or downgrade it to "as of the source date."
  • —Be honest about currency. Retrieving live is not the same as being current. A registry record is only as fresh as the registry's own last update, and the open web reflects what is publicly indexed at query time.

5. Confident, precise numbers with no quote behind them

The symptom. The answer contains exact figures, such as revenue, headcount, funding amounts or dates, stated flatly, and you cannot find them on the cited page.

Precision is what makes a hallucination dangerous. "The company has raised a significant amount" is vague enough to be mostly harmless. An exact euro figure invites a decision. Specific-looking numbers are common in a model's training text, so the number that comes out is often well formed, plausible and wrong. The tell is the absence of a quote: a grounded numeric claim traces to a sentence or table cell on a retrieved page, and a hallucinated one has nothing to trace to.

How to test it

Make quotation the test. For each numeric claim:

  1. —Find the passage. Where does the number appear in the cited source? Copy the exact text next to the claim.
  2. —Check unit and arithmetic. Currencies that switch, or "in thousands" tables read as raw values, are common.
  3. —Check the entity. Confirm the figure belongs to the company in question and not a parent, a subsidiary or a similarly named business.
  4. —Check the period. A figure for one financial year, presented without its year, is a claim about no time at all.

Then ask about a small private company whose headcount is unlikely to be public. An honest system returns "not found." One that returns a crisp figure is telling you how it behaves under uncertainty.

What a fix looks like

  • —Extract, don't recall. Pull numbers from retrieved text into structured fields and keep the source passage, rather than letting the model write them from its impression of a page.
  • —Quote-first output. Return, for each numeric claim, the value, the supporting text and the source, so a downstream check can confirm the value appears in the text.
  • —A defined "I don't know." If the only allowed outputs are confident ones, the agent will produce confident ones. Let it list a value as unresolved.

Running all five on one answer

Suppose an agent is asked to summarize the ownership and recent funding of Acme Oy (Finland), and returns three cited sentences.

  1. —Resolution. Two citations load. The third redirects to a news site's homepage, a flag on the funding sentence.
  2. —Claim in source. Both shareholders named in the ownership sentence appear on the cited registry page. The round size in the funding sentence appears in neither working source.
  3. —Source mix. All three citations sit on one aggregator. For an ownership claim you would expect a registration record.
  4. —Dates. The funding sentence says "recently," but the only working funding source was last updated before the round it describes.
  5. —Quotes. The round size is exact, and no retrieved passage contains it.

One answer fails four of five checks, and every failure was visible from the outside, by treating the citation as a claim to test. The ownership sentence passed cleanly, which is also useful: the goal is not to distrust everything, but to route trust to the parts of an answer that earned it.

How a grounded pipeline changes the picture

Each fix above points the same way: the citation should be something the pipeline guarantees, not something the model writes. That is the design of NeuralVerge's AI research capability. A question is decomposed into sub-questions, each is matched to relevant sources, with structured sources in the catalog such as corporate registries and company intelligence tried before the open web, facts are checked against more than one source in a dedicated cross-check step that surfaces disagreements, and the answer comes back with inline citations linking each claim to the source it came from. How deep that process goes is chosen per request from five depth tiers.

Two limits are worth stating plainly. Citations make an answer checkable; they do not make the underlying source correct or current. And none of this removes the need for your own verification on high-stakes claims. The value of inline citations is that the checks in this article become possible to run.

What to ask of any AI research API with citations

These questions map onto the five signs, and each can be answered by running a real query rather than reading a features page.

  • —Who writes the citations, the pipeline or the model? If the model can add a source that was never retrieved, signs one and two are open.
  • —Is there a per-claim link to a supporting passage? Ask for a sample response where you can find, for one sentence, the exact text that backs it.
  • —What does the source mix look like across ten varied questions? Tabulate the domains and compare with what you would expect.
  • —Are dates carried with citations? Check whether the response distinguishes "as of the source" from "today."
  • —What happens with an unanswerable question? An honest system reports the gap.
  • —Is there a verification step, and can you see why it rejected something?

If a product cannot pass a hand-built version of these tests, no description of it will make up the difference. If it can, you have a much better basis for trusting what it returns.

Frequently asked questions

Can I automate these five checks?

Mostly, yes. Resolving URLs, checking whether a cited page contains the claimed figure, and comparing dates are mechanical and fit in a test harness or a verification step before an answer is released. The source-mix check is a report you run over a batch of answers. Judging whether a passage really supports a claim often needs a second model call or a human sample, and a human should stay in the loop for high-stakes answers.

If every citation passes all five checks, is the answer correct?

Not necessarily. The checks test whether the answer is faithful to its cited sources, not whether the sources themselves are right. A registry record can be out of date and a news article can be wrong. What passing gives you is an answer where every claim points to something real that says what the answer says, so an error can be traced to a specific source and corrected.

How often should I run these checks?

Run the cheap mechanical checks on every answer that reaches a user or a downstream system. Run the heavier checks, such as claim-in-source review and source-mix reports, on a regular sample and whenever you change a model, a prompt or a retrieval setting. Sources change over time, so re-check a citation before relying on it again later.

What should an agent do when it cannot find a source for a claim?

Say so. A good answer marks the claim as unsupported or leaves it out, and states which sub-question came back inconclusive. An answer that fills the gap with a plausible figure and a plausible-looking link is the failure these five checks are designed to catch.

About NeuralVerge

Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.

AI Deep Research on the NeuralVerge blog.

Try it on your own data

One request format across research, extraction, and enrichment.

Get started