150+ Data Sources, One Schema: How NeuralVerge Normalizes Company Data
How an MCP server for data sources turns registries, profiles and directories into one predictable shape an agent can use, and where normalization stops.
Published September 27, 2026
An agent that can call ten different data tools has a problem before it ever reaches the data: every tool answers in its own dialect. One returns company_status, another returns status, a third buries the same fact in a paragraph of text. Connect an MCP server for data sources without fixing that, and you have moved the integration work from your codebase into the model's context window, where it is slower, more expensive and less reliable. This article is about the layer that fixes it: normalization. It covers what it means to return 150+ catalog sources through one predictable shape, and which problems the layer should not pretend to solve.
What "one schema" actually means
The phrase is useful shorthand and slightly dangerous, so it is worth being precise. The NeuralVerge source catalog lists 150+ individual sources across several families: company intelligence, marketplaces, product reviews, corporate registries, contact enrichment and professional networks. No one schema can describe a Finnish company register entry, a product listing and a contact record with the same fields, and a system that claims otherwise is either hiding information or inventing it.
What can be uniform is the contract around the data. Every data source in the catalog answers with the same envelope:
- —a session id you can trace,
- —the kind of lookup that ran,
- —an echo of the request parameters, which for a URL-based lookup includes the URL it read,
- —a readable
humansummary, - —and a structured
machineresult.
Inside the structured result, the fields are typed and named for what that source holds: a registry returns identifiers, status, legal form and addresses, while a funding profile returns a founding year, a size band and a funding total. The catalog counts lookups, not publishers, so the number of distinct underlying publishers is smaller than the number of entries.
So "one schema" here means one envelope, one way to make a request, one way to read a result and one way to find out where a value came from. It does not mean one flat table that every source is forced into.
Why an MCP server for data sources needs this layer
If you have read the guide to Model Context Protocol, the architecture is familiar: a client connects to a server, the server advertises tools, and the model decides which tool to call and with what arguments. This article does not repeat those basics. The relevant point is what happens when the server advertises many tools that all return company data.
The model reads each tool's name and description to decide what to call, then reads the result to decide what to do next. Both steps get harder when the tools are inconsistent:
- —Tool selection gets noisier. Many tools with overlapping descriptions and different argument styles give the model more ways to pick the wrong one. Consistent naming and inputs narrow the choice.
- —Result reading gets more expensive. A model can read a raw, source-specific payload, but every quirk it has to interpret costs tokens and adds a chance of misreading. Typed fields with stable names cost less to read and are harder to get wrong.
For a data tool, normalization is most of the value. The protocol handles discovery and transport; the normalization layer decides whether the results are usable once they arrive.
The problems normalization has to solve
Normalizing company data is a set of separate problems, each with its own failure mode. Treating them as one problem is how systems end up confident and wrong.
1. Entity resolution: which company is this?
Before you can normalize what a source says about a company, you have to know you are asking about the right one. Legal names differ from trading names, and the same brand can operate through entities in several countries.
The catalog mostly avoids guessing. Registries are read with AI extraction against the register's own page for a given identifier, and profile sources take a profile URL. An identifier removes ambiguity; a name does not. When all you have is a name, resolving it to an identifier is a separate step, and it is the step where research is the right tool, not a lookup. The AI research capability exists for the question "which entity is this," and only then do direct lookups make sense.
The practical rule for an agent: resolve once, keep the identifier, and pass identifiers, not names, to every later call.
2. Field mapping: same fact, different names
Once you have the right entity, the next problem is that sources describe the same fact in different ways. A registry says a company is "Active," a professional profile says nothing about status at all, and a funding profile publishes an employee count as a band rather than a number.
Field mapping means deciding, source by source, what each field means and giving it a stable name and type. Two rules keep this honest. Map only what genuinely means the same thing: a registered office address from a legal register and a location listed on a profile are related but not identical, and merging them would erase a distinction that matters for compliance work. And preserve what the source actually said. A funding total published as text stays text, "as published," rather than being parsed into a number with false precision. A size band stays a band. Normalization that converts everything to the most convenient type quietly discards information.
3. Provenance: where did this value come from?
A value without a source is a claim you cannot check. Provenance is the answer to "says who, and when."
In this catalog, provenance travels at the level of the lookup. Every response echoes the request it ran — for a URL-based lookup, the URL it was read from — and carries a session id that can be traced back to the request. Because each source is its own lookup and returns its own envelope, the fields inside a result all share one origin, and the origin is stamped on the result.
That is worth being clear about, because it is a design choice with a limit. Provenance here is per result, not per individual field. If you assemble a combined record from three lookups, the way to keep provenance at field level is to keep each field next to its source URL when you merge, rather than flattening them into one object. For research questions, citations serve the same purpose at the claim level: each statement in the answer links to the source that supports it.
4. Conflicting values: what to do when sources disagree
Most normalization pitches skip this one. Two sources that both look authoritative will sometimes disagree, and no schema fixes that, because the disagreement is real. A registry holds the address a company filed; a profile holds the address the company chose to show; they can differ for entirely legitimate reasons.
There are three defensible ways to handle it, and none is "pick one quietly":
- —Prefer by field. Decide in advance which source is authoritative for which fact. The legal register wins on legal status, legal form and registered address. A professional profile is more likely to be right about current headcount and how the company describes itself.
- —Prefer by recency. Where one source was updated more recently, the fresher one is usually the better bet, if you can tell which.
- —Surface the disagreement. When neither rule settles it, report both values with their sources and say they differ.
A single-source lookup does not do any of this, because it reports one source and nothing else. Cross-checking happens when you ask a question through research, where facts are compared across sources and disagreements are carried into the answer rather than hidden. If you build reconciliation yourself on top of individual lookups, the rule of thumb is to keep both values until you have a reason to drop one.
5. Freshness: how old is this answer?
Freshness is easy to overstate. Source lookups in the catalog are read at request time, so the answer reflects the source as it stands when you ask, rather than a copy stored weeks ago. That is a real advantage over a static dataset.
Real-time fetching is not the same as real-time facts. A registry record changes when the company files something, which for many fields is a handful of times a year. A profile changes when its maintainers update it. A live read of a stale record returns a stale value, quickly.
The useful habit is to phrase answers as "according to the register as of this lookup" rather than as timeless facts.
The difference from a stored dataset is the main thing that separates this catalog from B2B contact and company databases.
6. Missing values: absent is not zero
The last problem is the quietest. A source does not hold every field for every entity, and a normalization layer has to decide how to express that.
In this catalog, a field the source does not hold comes back empty rather than guessed, and a lookup that resolves nothing returns nothing rather than a near match. Both behaviors matter to an agent. An empty field means "unknown," not "none," "false" or "zero." An agent that reads a missing employee band as zero employees will write a wrong sentence with great confidence.
Stating in the tool description that some fields can be empty, and what empty means, is easy and prevents a whole class of errors.
A worked example: Acme Oy (Finland)
To see the six problems in one place, take an illustrative task: an agent is asked to prepare a short profile of Acme Oy (Finland) before a supplier review. Everything below is illustrative, not a claim about a real company.
Resolve the entity. The agent starts with a name. Names are ambiguous, so it does not go straight to a lookup. It asks a research question first, something like which legal entity in Finland trades as Acme, and gets back a business ID along with citations. From this point the agent carries the business ID, not the name.
Read the register. The agent runs AI extraction against the Finnish company register's page for that business ID. The structured result holds fields such as company name, company form, home municipality, main line of business and registered postal address. The response carries the register URL and a session id. These are legal facts, and the agent labels them that way.
Read the network profile. The agent then calls the professional-network company search source, which returns a name, an industry, a headquarters location, a follower count and a description. This is how the company presents itself, and the agent labels it that way.
Compare. The two results do not agree on everything. The register lists a registered postal address in one city, while the profile lists headquarters in a neighboring one. The register carries a legal form; the profile carries none. The profile gives a follower count; the register does not. The agent does not merge these into a single "address" field. It reports the registered address from the register as the legal address, reports the profile location as a presented location, and notes that they differ.
Handle the gaps and report. Some fields come back empty, so the agent writes "not stated in the sources checked" instead of filling them in. The final output lists each fact next to the URL it came from and notes that each value was read at request time.
None of that required the agent to understand either source's internal layout. A question this broad, spanning resolution, two source types and a comparison, is also a natural fit for the research capability, which runs the cross-checks and returns cited claims in one call instead of leaving the agent to orchestrate them.
How an MCP server exposes all this as tools
Each source becomes a tool with a name, a description of what it takes and returns, and a set of arguments. For data sources, the design choices that matter are these:
- —One tool per lookup, named for what it returns. A tool that resolves a registry record is described as taking an identifier and returning the record. A tool that searches takes a query and returns a list. Keeping search and fetch as separate tools stops the model from confusing "find candidates" with "get the record."
- —The same result envelope from every tool. The agent learns the shape once. The origin URL and session id are always in the same place.
- —Descriptions that state limits. Which country a register covers, whether a name or an identifier is expected, and which fields can come back empty all belong in the description, because that is what the model reads when it chooses.
NeuralVerge runs a hosted remote MCP server at https://api.neuralverge.ai/functions/v1/mcp-server, using the same bearer token as the REST API, so there is no server for you to build or host. Search, AI Extract (run_extract) and every data source are tools on it; AI research runs over REST only. Each tool is a thin wrapper of its REST endpoint with the same inputs and outputs, so the code that reads results does not change when the transport does. The connection details are documented at docs.neuralverge.ai.
If you want to see what a single source page looks like, the Finnish company register source shows the input it takes, the fields it returns and its known gaps, which is the level of detail an agent should be given about every tool.
Where normalization stops
A normalization layer is honest when it says what it does not do. It does not make sources agree; it makes disagreements visible and attributable, and which value wins is a policy your application owns. It does not create data a source does not hold, so missing stays missing. It does not equalize coverage: a registry covers its own country and nothing else, and a company with no public profile returns nothing from a profile lookup. And it does not guarantee accuracy, because a normalized wrong value is still wrong. For compliance and due-diligence work, the origin URL is there so a person can open the source and check.
Where teams use it
- —Supplier and counterparty checks. Confirm a company exists, what legal form it has and who is listed as responsible, then compare that with how the company presents itself.
- —Account research. Assemble a company profile from a legal record, a size and funding profile and a professional-network page, with each fact tied to its origin.
- —Agent tool use. Give an agent lookups it can chain, with predictable inputs and outputs, so it can decide mid-task what it still needs to check.
What to check before you rely on a normalized data layer
- —What is uniform, exactly? Ask for three responses from three different sources side by side. Is the envelope the same? Are field names and types stable within each source?
- —Does every value carry an origin? Look for a source URL and a request identifier on each result, not a general statement that data is "sourced."
- —What happens on a miss? Test with an identifier you know does not exist. It should return clearly empty rather than a near match.
- —Is missing distinguishable from zero? And does anything fill gaps with inferred values without labeling them?
- —Can you keep both values in a conflict? If you only get one value per field with no attribution, you cannot reconcile disagreements later.
- —Is freshness stated honestly? "Read at request time" and "the record is current" are different claims.
- —Are coverage limits stated per source? They should be visible on each source, not implied.
Running one real entity through two sources and comparing what comes back with what you expected will answer most of these faster than any feature list.
Frequently asked questions
Does one schema mean every source returns the same fields?
No. Every source returns the same outer envelope — a session id, the kind of lookup, an echo of the request parameters, a readable summary and a structured result. The fields inside the structured result are specific to each source, because a corporate registry and a funding profile simply hold different facts. What stays consistent is the shape around them and the way each field is typed and named within its source.
What happens when two sources disagree about the same company?
A single-source lookup returns what that source says and does not compare it with anything else. When you ask a question through AI research, the pipeline checks facts across more than one source where more than one is available and surfaces disagreements instead of silently picking one. If you build your own reconciliation on top of individual lookups, keep both values and their source URLs and decide by field which source you trust.
How fresh is the data an MCP tool returns?
Source lookups are read at request time, so the answer reflects the source as it stands when you ask. Fetching in real time does not make the underlying record new: a registry record is only as current as the registry's own last update, and a profile is only as current as whoever maintains it.
What does a normalized result look like when a field is missing?
Fields a source does not hold for a given entity come back empty rather than guessed. An empty value means the source did not have it, not that the value is zero or false, and an agent should treat it as unknown.
Is every data source available as an MCP tool today?
Yes. Search, AI Extract and every data source are tools on the hosted MCP server, and corporate registries are read through the AI Extract tool. AI research runs over REST only. Each tool is a thin wrapper of its REST endpoint, with the same inputs and the same response shape.
About NeuralVerge
Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.
AI Agents & MCP on the NeuralVerge blog.
Try it on your own data
One request format across research, extraction, and enrichment.