nvNeuralVerge
AI Extraction

Extracting Structured Data From News Articles at Scale

How to build a document to json api workflow for news: schema design for events and entities, batching, deduplication, paywalls, and provenance for every field.

Published September 21, 2026

News is one of the richest sources of company and market information, and one of the least convenient to consume. Every article is written for a human reader: a headline, a lede, some quotes, background, a byline. What a downstream system needs is a row: who was involved, what happened, when, and where the claim came from. A document to json api turns the first into the second. Doing it for one article is easy. Doing it for tens of thousands, from outlets with different layouts, covering the same events repeatedly, is where the design decisions start to matter.

This article walks through one workflow: monitoring news coverage of a set of companies and turning it into structured event records. It covers schema design, batching, deduplication, paywalled and script-heavy pages, and provenance, so every field in your database can be traced back to the article it came from.

What a document to json api actually does

A document to json api takes a document, usually a URL, and returns fields you defined, in a fixed structure, as JSON. For news, an article page goes in and a headline, publication date, organizations, event type and summary come out. Three adjacent tools get used for the same job, and each fails differently.

Versus a scraper with selectors

A selector-based scraper finds fields by their position in the markup. That works on one outlet's template and breaks the day the outlet redesigns. For news the problem multiplies, because you read from many publishers, each changing templates on its own schedule. Extraction that works from meaning asks a different question: what is the publication date of this article, whatever the page looks like. One schema then applies to every outlet.

Versus a full-text dump

Many teams convert every article to clean text and store it. That is a fine archive, but text is not queryable. You cannot ask a pile of articles for every funding announcement last month and get a table back. Extraction is the reading step, done once and stored in a shape you can filter and join.

Versus a research pipeline

Extraction reads a page you already have. It does not decide which pages to read, and it does not check whether the article is right. A research pipeline starts from a question and works out which sources to consult. The two are complementary: AI extraction is the repeatable per-document step, and AI research is for when you have a question rather than a URL.

Why news is harder than it looks

The same event appears many times. A funding round or leadership change is usually reported by several outlets, then followed up. One row per article over-counts events.

Dates are slippery. An article has a publication date, and the event it describes has its own. Phrases like "last week" only make sense relative to the first. A single date field will be wrong for one of them.

Entities are ambiguous. A company appears as its legal name, a short name, a brand or "the Finnish group". Two companies can share a name.

Not everything is the article's own claim. Stories quote spokespeople, cite anonymous sources and repeat rumours. A flat "acquirer" field cannot say whether the article states something as fact or reports that someone said it.

Each of these shapes the schema.

Designing the schema for news

The schema decides what you can query later, and changing it after you have stored many records means re-running extraction. A reasonable starting point, with Acme Oy (Finland) as the illustrative subject, sketched here as a field outline (the request below shows how it becomes a JSON Schema):

{
  "headline": "string",
  "publication_date": "string, ISO 8601",
  "outlet": "string",
  "is_partial": "boolean, true if only a teaser or truncated text was visible",
  "events": [
    {
      "event_type": "funding | acquisition | leadership_change | product_launch | layoff | partnership | legal | other",
      "event_date": "string, ISO 8601, or null if not stated",
      "event_date_precision": "day | month | quarter | year | unknown",
      "summary": "string, one sentence",
      "organizations": [{ "name": "string", "role": "subject | acquirer | target | investor | partner | other" }],
      "amount": { "value": "number or null", "currency": "string or null" },
      "attribution": "stated_as_fact | attributed_to_source | rumour",
      "evidence_quote": "string, the sentence that supports this event"
    }
  ]
}

Each choice prevents a specific problem.

  • —Events as a list. One release can announce a funding round and a new finance chief. A single event field forces the extractor to drop one.
  • —Two dates and a precision field. Separating publication date from event date resolves the ambiguity above. When an article says "in the third quarter", a fabricated exact date looks tidy and is wrong. A null with precision unknown is honest and easy to filter.
  • —Roles, not a flat list of names. Knowing which company is the acquirer and which the target tells you what happened.
  • —Attribution. This does not tell you whether a claim is true. It tells you how the article presented it, which the article really does contain.
  • —An evidence quote per event. It is the seed of your provenance trail. A reviewer can check it in seconds, and if the quote does not contain the entities and event type in the record, the record is suspect.
  • —Enums where you group, free text where you must. Fixed values make counting possible. Include other so unusual events are not forced into the wrong bucket.

Start with a subset, run it on a sample of real articles, read the output, and add fields only when a real question demands them. Every field is another place the extractor can be wrong.

The extraction call

With NeuralVerge's AI extraction, one article against a schema is a single run-extract request: the article URL, a plain-language instruction, and the schema passed as a JSON Schema string in settings.extract_schema_json. A trimmed version of the schema above:

import json, requests

SCHEMA = {
    "type": "object",
    "properties": {
        "headline": {"type": "string"},
        "publication_date": {"type": "string"},
        "events": {
            "type": "array",
            "items": {
                "type": "object",
                "properties": {
                    "event_type": {"type": "string"},
                    "event_date": {"type": ["string", "null"]},
                    "summary": {"type": "string"},
                    "organizations": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "properties": {
                                "name": {"type": "string"},
                                "role": {"type": "string"},
                            },
                        },
                    },
                    "evidence_quote": {"type": "string"},
                },
            },
        },
    },
    "required": ["headline"],
}

resp = requests.post(
    "https://api.neuralverge.ai/functions/v1/run-extract",
    headers={"Authorization": "Bearer <API_KEY>"},
    json={
        "url": "https://news.example.com/2026/09/acme-oy-raises-round",
        "instructions": "Extract the headline, publication date and every business event the article reports, with the organizations involved and a supporting quote.",
        "settings": {"country_code": "fi", "extract_schema_json": json.dumps(SCHEMA)},
    },
)
record = resp.json()["machine"]

The extracted fields come back under machine in that shape, whatever the outlet's layout. Exact request details live in the documentation. Pages are rendered before extraction, so script-driven articles are not a special case in the request, and a press release published as a PDF goes through the same call and comes back in the same shape.

Running it at scale: batching

Each article is an independent request, which is what lets you parallelize, retry and resume without coordination. What changes at scale is everything around the call.

  • —Use a queue, not a loop. Each job holds a URL, a schema version, a status and an attempt count. A page that times out fails one job, not the run.
  • —Make jobs idempotent. Key results by URL and schema version. Re-running replaces instead of duplicating, and a schema change becomes a clean operation: bump the version and re-queue.
  • —Bound concurrency deliberately. Let your own constraints set it: write speed of your store and documented rate limits. Spread requests across outlets rather than draining one publisher at a time.
  • —Retry with a budget. Retry transient failures a few times with a delay, then park the job with its reason. A cluster of failures from one outlet points to a systematic problem that retrying will not fix.
  • —Validate before storing. Valid JSON is not a good record. Check required fields, enum values, parseable dates, numeric amounts and a non-empty quote. Failures go to a review queue, not the main table.

The call itself is rarely the constraint on a news pipeline. Review time and duplicate noise usually are.

Deduplication: one event, many articles

Separate two concepts: the article, a fetched document with its own URL and provenance, and the event, a real-world occurrence many articles may describe. Deduplication links the first to the second, in layers from cheap and exact to fuzzy.

Layer 1: identical pages

Normalize URLs: strip tracking parameters and fragments, lower-case the host, resolve obvious redirects. Hash the result. This catches the same page arriving through different feeds.

Layer 2: near-identical text

Wire copy appears on many sites with only the headline changed. Hash the cleaned text to detect it. Do not delete the copies. Link them to one cluster and mark them syndicated, because independent reporting versus republishing is meaningful.

Layer 3: same event, different words

Two outlets write their own stories on one announcement. Here the extracted fields do the work: two records are candidates when they share event type, overlapping organizations in the same roles, and close event dates. Score the overlap, set a threshold and send borderline pairs to review. For example, three articles about Acme Oy each return funding, Acme Oy as subject and a month-level date. Matching groups them into one cluster, not three rounds.

Keep the cluster, not just the winner

Pick a canonical record, perhaps the earliest article or the most complete one, but keep links to the rest. Disagreements such as different amounts stay visible, and the number of independent outlets is a useful signal. Fuzzy matching will sometimes merge distinct events or split one, so bias towards splitting when merging is costly, and log every merge so it can be undone.

Resolving entities

Extraction gives you names as written. Analysis needs identities. Treat resolution as its own step: store the extracted name exactly, then match it against a reference list using normalization (legal suffixes, punctuation, case) plus country and industry context from the same record. Link confident matches automatically and queue the rest. Registry records are better keys than name strings, and the source catalog includes corporate registry sources for that lookup.

Paywalls, script-heavy pages and pages you cannot read

An honest pipeline plans for failed and partial fetches.

Paywalls. A public request generally sees a teaser. Extraction can only work with what is visible, so it can often recover headline, date and outlet, but not the body. Set is_partial when the text is clearly truncated, and make null an explicit allowed answer in your schema, so an amount that is not on the visible page comes back empty rather than guessed. If you hold a licence for the full text, supply it yourself through your own authorized access. Circumventing a paywall is not something to build into a pipeline. Partial records still tell you an outlet covered something, which may trigger a follow-up elsewhere.

Script-heavy pages. Rendering handles pages that load body text with JavaScript. It does not make every page readable: consent walls, interaction-gated content and region-specific versions can still get in the way.

Blocks and failures. Some publishers block automated requests and some pages time out. Record the reason against the job. If an outlet fails consistently, that is a coverage gap to report as unknown, not a quiet news day.

The fair summary: a news pipeline sees the public, readable part of the news. "No news about this company this month" only means no news among the pages you could read.

Provenance: every field traceable

For each stored record, keep the source URL and outlet, the publication date as read from the article, the fetch timestamp, the schema version, the evidence quote and attribution for each event, and status flags such as partial content or manual overrides.

Two habits pay off. Never overwrite extracted values in place: store human corrections as a separate layer with the person and time. And never merge away the source articles of a cluster, so the answer to where a number came from is a list of links.

Provenance also marks the limit of what extraction can claim. The quote proves the article said something, not that it was right. Where a claim matters, corroborate it: give it to a research step as a question and it checks other sources and returns a cited answer.

A worked example

Suppose your watchlist includes Acme Oy (Finland), and three articles appear on a Tuesday.

  1. —Discovery surfaces three URLs, which are normalized and hashed. None match stored items. Discovery is whatever mechanism you use to find candidates, outside extraction.
  2. —Extraction runs per URL. The business daily yields a funding event with an amount and a quote. The regional outlet yields the same type with no amount and a month-level date. The paywalled piece yields a headline and date, is_partial true, and no events.
  3. —Validation passes the first two and flags the third rather than rejecting it.
  4. —Deduplication groups the first two by event type, subject and date. The third stays a standalone article record.
  5. —Entity resolution links "Acme" and "Acme Oy" to the watchlist entry.
  6. —Single-source flag. Only one article states the amount, so the cluster records it with a note that one source supports it.
  7. —Follow-up. The amount is material, so a research question is raised about recent funding for Acme Oy, and a reviewer compares the cited answer with the extracted record.

The result: one event, two supporting articles, one partial article and one flagged amount, with a trail from every field to a URL and a quote.

Where teams use this

  • —Account monitoring for sales. Track funding, leadership changes and launches across a target list and route events to account owners.
  • —Due diligence and risk screening. Build a dated timeline of coverage on a counterparty, with the quote behind each entry.
  • —Competitive intelligence. Follow launches and partnerships, clustered so one launch counts once.
  • —Analyst research. Turn a stream of sector articles into a table of events to filter and chart.
  • —Feeding agents. Give agents event records with sources attached instead of raw article text.

What to check before you commit

  • —Does one schema hold across outlets? Test at least ten publishers with different layouts.
  • —Can it return null instead of guessing? Give it a teaser-only page and ask for the amount.
  • —Does it render dynamic pages? Include some in your sample.
  • —Can you get a supporting quote per event? Without it, review is slow.
  • —Can you tell failures apart? Blocks, timeouts and paywalls should look different in the response.
  • —Is the output shape stable between runs? Wording may vary; structure should not.
  • —Can you re-run cleanly after a schema change? Keyed by URL and version, it is a queue operation.

Read a few dozen outputs against their source pages yourself before trusting any aggregate. Spot checks catch systematic errors, such as one outlet's date format, that summary statistics hide.

Frequently asked questions

Should I extract one record per article or one record per event?

Both, in two layers. Extract per article first, because an article is the unit you fetch and the unit you can cite. Then group article-level event records into event clusters in a separate step. Extracting straight to events loses the link between a field and the article that supplied it, which is the link you need when someone asks where a fact came from.

How do I handle articles behind a paywall?

Extraction only works on what a public request can actually see. If a page returns a login wall or a truncated teaser, the extractor can only read the teaser, so the right move is to record that the article was partial and mark the affected fields as missing rather than let the model fill them in. If you have a licence for the full text, feed that text in yourself. Working around a paywall is not something to build into a pipeline.

How do I avoid storing the same story dozens of times?

Deduplicate in layers. Normalize and hash URLs to catch identical pages, compare extracted fields such as the entities, event type and event date to catch the same story on different sites, and keep syndicated copies linked to a single cluster instead of deleting them. Keeping the copies matters because the number of outlets carrying a story is a useful signal in its own right.

Can the extractor tell me when an article is wrong?

No. Extraction reports what the article says, not whether it is true. Store the publication date, the outlet and a quote for each field so a reviewer can judge the claim, and use cross-checking against other sources, for example with a research step, when a claim matters enough to verify.

How many articles can I run in one batch?

The limit that matters is usually your own: how many fetches you want in flight and how quickly your downstream store can absorb the results. Each article is an independent call, so you can size a batch to fit those constraints and retry failures individually. Check the documentation for current rate limits before you plan a large backfill.

Do I need to define a schema before I start?

You can let the extractor infer a structure while you explore, but for anything you plan to store and query you should define the schema. Inferred structures vary between runs and between articles, which makes comparison and deduplication harder. Start with a small schema, run it against a sample of real articles, and expand it from what you see.

About NeuralVerge

Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.

AI Extraction on the NeuralVerge blog.

Try it on your own data

One request format across research, extraction, and enrichment.

Get started