nvNeuralVerge
AI Deep Research

Building an M&A Due-Diligence Agent With a Research API

How to build an M&A due diligence API workflow: an agent that runs cited first-pass target research so a deal team spends time on judgment, not lookups.

Published September 16, 2026

Before a term sheet, before a data room opens, an acquirer or investor has to decide which companies deserve serious attention at all. That decision rests on a pile of public facts about each candidate: who owns it, what it has raised, what customers say about it, whether its filings tell the same story as its website. Collecting those facts is slow, repetitive work that lands on analysts and associates in the first days of a process. An M&A due diligence API workflow automates that first pass: an agent takes a target name, asks a research platform the right questions, and returns a cited briefing that a human deal team can check and build on.

This article is about building that agent. It is not about replacing the deal team, and it makes no claim about legal or financial conclusions. It covers what work is worth handing to an agent, how to structure the questions, how the research pipeline answers them, and how to wire the result into a process people can trust.

What an M&A due diligence agent should and should not do

Due diligence is a stack of very different activities. Some are mechanical: confirm the legal entity exists, list its officers, pull its funding history, summarize public reviews. Others are judgment: decide whether a customer concentration is acceptable, whether a contract clause is a problem, whether the price is right. The first kind is where an agent earns its keep. The second kind stays with people.

A useful way to draw the line is by source. An agent working from public information can do the following well:

  • —Confirm the target's legal identity, registration status and officers from official registers.
  • —Assemble funding rounds, investors and stated headcount from company-intelligence sources.
  • —Summarize customer and buyer sentiment from public review platforms.
  • —Read the target's own website and public documents into structured fields.
  • —Cross-check those results against each other and flag where they disagree.

An agent working from public information cannot see a virtual data room, audited financial statements, customer contracts, or anything a target has chosen not to publish. It also cannot give legal advice or a valuation. Building it as if it could is the fastest way to lose the deal team's trust, so the design has to keep that boundary visible in every output.

Framed correctly, the agent is a research assistant that produces the first draft of a public-facts briefing. Analysts then spend their time on the questions the briefing raises, not on the tabs they used to open to assemble it.

Why a single search call is the wrong tool

The instinctive first version of this agent wires a web search into an LLM and asks it to "research the company." It produces fluent paragraphs that are hard to verify, because search results are links, not facts. Someone still has to open each page, decide whether it is authoritative, and reconcile it with the others. For due diligence the missing pieces matter more than the fluent ones: a stale ownership record, a subsidiary listed under a different name, a funding round reported differently in two places.

What the agent needs is a research capability that plans a compound question, routes each part to the right kind of source, extracts the facts, checks them against each other, and cites every claim. That is what NeuralVerge's AI research is built to do. The agent supplies the workflow and the questions; the platform does the retrieval and verification. For a longer explanation of the mechanism, see the guide to deep research pipelines.

The shape of the workflow

A sensible agent has four stages, kept separate so each is easier to test.

  1. —Intake. The agent receives a target: a name, a country, ideally a website. It normalizes these into a structured target record and asks for anything missing before starting research.
  2. —Screen. A light, fast pass answers the gating questions: does this entity exist, is it active, is it roughly what the deal thesis assumes.
  3. —Investigate. For targets that pass the screen, a deeper pass covers ownership, funding, reputation and any red flags in the public record.
  4. —Brief. The agent assembles the results into a memo with sections, citations and an explicit list of open questions for the human team.

Splitting screen from investigate is the main lever on turnaround and effort, and it lines up with how research depth works. More on that below.

Stage 1: intake and identity resolution

Most due-diligence errors start with the wrong entity: names collide across countries and brands differ from legal names. The intake record should carry the trading name, likely legal name, country of registration and website; if only a name is supplied, the first research question is an identity question.

Take Acme Oy (Finland) as an illustrative target. The intake record might read: trading name Acme, legal name Acme Oy, country Finland, website supplied by the deal team. The first question the agent asks is narrow: "Is Acme Oy (Finland) an active registered company, and what is its registered legal form?" That question routes to a national corporate registry, and it is quick to answer. If it comes back ambiguous, say two entities with similar names, the agent stops and asks a human rather than guessing. Everything downstream depends on this step being right.

Stage 2: the screening pass

The screening pass exists to drop candidates quickly. Its questions usually have one authoritative source:

  • —Is the entity registered and in good standing?
  • —What is its legal form, and when was it registered?
  • —Who are the registered officers?
  • —Does the company's own website describe the same business the deal thesis assumes?

Because these are narrow questions with clear sources, a lighter research tier is the sensible default. NeuralVerge's AI research has five depth tiers — Lite, Base, Core, Pro and Ultima — and the lightest ones answer well-defined questions fast, with Lite typically returning in 30 to 90 seconds. A screening pass over a long list of candidates is exactly the case where running exhaustive research on every name would be waste.

Where the target has a website, the agent can also use AI extraction to read the company's own pages into structured fields: products, locations, leadership names, stated customers, contact details. That gives the screening memo the target's self-description, which later gets compared against what independent sources say.

A real call for the Acme Oy example above looks like this:

curl -X POST https://api.neuralverge.ai/functions/v1/run-extract \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://acme.example",
    "instructions": "Extract the company name, products, office locations, leadership names and contact details.",
    "settings": {
      "country_code": "fi",
      "extract_schema_json": "{ \"type\": \"object\" }"
    }
  }'

The response carries the extracted fields directly in machine, ready to diff against the registry record from the identity check:

{
  "session_id": "7a9c2e10-...",
  "kind": "extract",
  "url": "https://acme.example",
  "machine": {
    "company_name": "Acme Oy",
    "products": ["Example product line"],
    "locations": ["Helsinki, Finland"],
    "leadership": ["Jane Example, CEO"],
    "contact_email": "hello@acme.example"
  }
}

The output of stage two is a short pass, fail or needs-review status per candidate, with citations. Anything that fails on identity or status drops out. Anything ambiguous goes to a human. Only survivors move to the deeper stage.

Stage 3: the investigation pass

This is where the compound questions live. For a shortlisted target, the agent asks a set of questions that each need different kinds of sources and are more likely to produce conflicting information:

  • —Ownership and control. Who are the registered owners and officers, and has that changed recently?
  • —Funding and financing history. What rounds has the company raised, from whom, and when?
  • —Market position. How do public analyst and category pages describe the company and its product?
  • —Reputation. What do customers or buyers say publicly, and are there recurring complaints?
  • —Consistency. Do the company's own claims match what the registry and third-party sources show?

Each of these is a different question, but the agent does not have to decide which website answers which one. The research platform plans the request into sub-questions and routes each to the relevant part of the source catalog: registries for ownership, company-intelligence data for funding, review platforms for reputation, and the open web for anything the structured sources do not cover. The catalog page lists the categories and which countries have registry coverage.

Asking questions the pipeline can actually answer

Question quality decides output quality. A vague prompt like "do due diligence on Acme" yields a vague answer. The better pattern is a fixed question set, one entry per diligence topic, with each question naming the entity precisely and asking for facts rather than opinions.

Ask for public facts with a defined source type, for example "Summarize the registered ownership and officers of Acme Oy (Finland), and cite the register," rather than asking whether the company is a good target. Store questions as templates parameterized by target, so every candidate gets the same set.

Cross-checking is the point

Diligence is largely about finding where things do not line up. The research pipeline's cross-check step compares facts across sources where more than one source touches the same claim. If a registry lists one set of officers and the company's website lists another, that difference should reach the memo. If a funding database reports a round the company's own site never mentions, that is worth a line too.

The design rule for your agent is simple: never flatten a disagreement. When the research answer reports a conflict, the agent should carry it into the memo as an open question, with both citations, and mark it for human review. A tidy memo that hides conflicts is worse than a messy one that shows them.

A worked example: one target from intake to memo

Here is an illustrative run for Acme Oy (Finland), with a deal thesis that assumes a small, profitable software vendor with a couple of institutional investors.

Intake. The deal team submits "Acme, Finland, acme.example." The agent builds the target record and asks its identity question against the Finnish company register.

Screen. A lighter-tier research call returns that Acme Oy is a registered private limited company, currently active, with a named managing director and board. Citations point to the register. The agent also runs extraction on the company's site and gets structured fields: a product description, two office locations and a leadership page. The site's leadership list and the register's officer list overlap but are not identical. The agent records that as a note rather than a failure, because a site listing a non-statutory executive is normal.

Investigate. The agent moves Acme to the deeper pass and sends the fixed question set. The ownership question routes to registry sources and reports the registered shareholders' structure where the register publishes it. The funding question routes to company-intelligence data and returns a list of rounds and investors. The reputation question routes to review platforms and summarizes recurring themes. The consistency question compares the site's stated headcount and founding year to what the third-party sources show.

Suppose two sources disagree on the year of the last funding round: the pipeline prefers the more authoritative source with a stated reason, or reports the discrepancy, and the agent carries it to the memo.

Brief. The agent assembles a memo with a section per topic. Each factual sentence keeps its citation. A final section lists open questions: the officer mismatch, the funding-date discrepancy, and the fact that financial statements were not part of this pass and need to be requested. A human analyst opens the memo, clicks through the citations that matter, and decides whether Acme moves forward. The agent has saved the assembly time. It has not made the call.

Choosing the right depth for each stage

Depth is a per-request setting, so a due-diligence agent can match depth to stakes instead of applying one setting to everything. A reasonable allocation looks like this:

  • —Screening a long list. Use a lighter tier. The questions are narrow and the sources are authoritative.
  • —Investigating a shortlist. Use a middle tier for the standard question set.
  • —Deep dives on a final candidate. Use the heaviest tiers for ambiguous or high-stakes questions, such as reconstructing a complicated ownership chain or reconciling conflicting accounts of a company's history.

Typical turnaround grows with depth tier: roughly 30 to 90 seconds for Lite, 1 to 2 minutes for Base, 2 to 4 for Core, 3 to 7 for Pro and 5 to 12 for Ultima. A screening pass over many names at a light tier and an investigation of a few at a heavier tier is a very different workload from running everything at the top tier.

Wiring the agent: REST, MCP and orchestration

There are two sensible ways to connect the agent to the research capability, and they fit different designs.

Direct calls from a pipeline. If your process is a fixed sequence, a backend job that runs intake, screen, investigate and brief in order can call research and extraction as ordinary requests. This is predictable and easy to test. Each stage has a known input and output, and you control retries and limits in code.

Research runs as a session: you start it, then poll until it finishes. A screening-tier call for the Acme Oy identity question looks like this:

curl -X POST https://api.neuralverge.ai/functions/v1/run-research \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "instructions": "Is Acme Oy (Finland) an active registered company, and what is its registered legal form? Cite the register.",
    "settings": {
      "country_code": "fi",
      "search_enabled": true,
      "deepsearch_model": "base"
    }
  }'

That returns a session_id right away, which you poll:

curl "https://api.neuralverge.ai/functions/v1/get-session-status?session_id=7a9c2e10-..." \
  -H "Authorization: Bearer YOUR_API_KEY"
{
  "session_id": "7a9c2e10-...",
  "status": "complete",
  "results": {
    "kind": "deepsearch",
    "human": "# Research summary\n\n- Acme Oy is an active Finnish private limited company (osakeyhtiö), registered in 2019. Source: Finnish Trade Register.",
    "machine": {
      "entity_status": "active",
      "legal_form": "Private limited company (Oy)",
      "registered_year": 2019
    }
  }
}

For the deeper investigation pass on a shortlisted target, the same endpoint takes a heavier deepsearch_model ("core" rather than "base") and, where it matters, a finalizer_model to pick which model writes the final memo text — useful when a deal team wants a stronger model on the synthesis step even for a lighter-tier search.

Tools for an agent. To let a model decide which questions to ask next, give it extraction and the data sources as tools on NeuralVerge's hosted MCP server (run_extract and the data-source tools), and research through a function-calling tool that wraps the run-research REST endpoint, since research is not available over MCP. Bound it: a maximum number of calls per target, a maximum depth tier, and a rule that every claim in the memo must carry a citation returned by a tool.

Full parameter and response reference lives in the API documentation, since field-level details evolve faster than a blog post should try to track; the MCP overview covers the tool route for an agent that calls extraction and data sources mid-task instead of a backend job calling them directly.

Designing the memo so people trust it

The memo is where trust is won or lost. A few design choices matter more than the model behind it.

Keep citations attached to claims. Do not let the summarizing step strip them. Each sentence that states a fact should still link to its source in the final document, so a reviewer can verify it in one click.

Show disagreements, not resolutions. If sources conflict, show both. Add a short note on which the pipeline preferred and why, if it said.

State the freshness. Each source is read at request time, but a registry record is only as current as the register's own last update, and open-web results reflect what was indexed when the call ran. Put the run date in the memo and treat it as a snapshot.

Where teams use this pattern

Acquirer target screening

Corporate development teams keep long lists of possible targets. A screen stage gives every name a consistent, cited first look.

Investor deal sourcing

Investors screening many inbound opportunities can produce the same baseline briefing for each, so partners' time goes to companies that pass the basics. The counterparty-check workflow applies the same idea to onboarding.

What to check before you build

Before committing to this design, verify the following against your own deal flow:

  • —Coverage for your jurisdictions. Check the source catalog for the countries your targets are usually in, and note which have registry coverage and which fall back to the open web.
  • —Citation granularity. Confirm that every claim in a sample answer links to a specific source, not just a list at the end.
  • —Handling of conflicts. Run a target you know has messy public data and check whether disagreements appear in the answer.
  • —Scope boundaries in the memo. Make sure the output states clearly what a public-source pass cannot see.

A good test is to take a completed deal your team already knows well and run it through the workflow. If the memo matches what your analysts eventually found, and flags the things they had to dig for, the design is close.

Frequently asked questions

Can an agent replace a deal team's due diligence?

No. The agent handles the mechanical first pass: finding, extracting and cross-checking public facts about a target. Judgment calls, legal review, financial analysis and anything that depends on non-public information such as a data room stay with the humans and advisors running the deal.

Does this cover financial statements or a virtual data room?

No. The pipeline reads public sources such as registries, funding data, review platforms and the open web. It does not see a target's private financials or a virtual data room. Those documents are reviewed separately, and the agent's job is to make sure the public picture is assembled and checked before that work starts.

How do you keep the agent from stating things it cannot support?

Every claim in a research answer carries a citation to the source it came from, and the pipeline's cross-check step surfaces disagreements between sources instead of hiding them. Your agent should carry both forward: keep the citations in the memo, and flag open conflicts for a human reviewer rather than resolving them silently.

Can the agent monitor a target between signing and closing?

Yes. It is an ordinary call, so you can re-run the same questions on a schedule and compare answers over time. A change of control, a new registry filing or a shift in public reputation is worth a human look, and a scheduled re-check is a way to notice it without anyone remembering to look.

Which jurisdictions does it work for?

It works where the source catalog has coverage, which today includes several national corporate registries alongside funding, reputation and open-web sources. For a target in a country without a supported registry, the pipeline falls back to the open web, and the citations show which kind of source backs each claim. Check the source catalog for the countries your deals actually touch.

About NeuralVerge

Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.

AI Deep Research on the NeuralVerge blog.

Try it on your own data

One request format across research, extraction, and enrichment.

Get started