nvNeuralVerge
AI Extraction

From HTML to JSON: What Changes When You Add AI to Extraction

An html parsing alternative ai teams can adopt step by step: what changes in code and operations, where determinism goes, and what still needs a parser.

Published September 29, 2026

For most of the web's history, getting data out of a page meant writing code that described where the data sat: this div, that class name, the third td in the second row. That code worked until the page changed, and then someone spent an afternoon fixing it. Adding AI to extraction replaces "where is it" with "what is it," and that swap is bigger than it sounds. It changes the code you write, the way it fails, what you have to test, and what you pay for. This article is a practical look at an html parsing alternative ai can provide, aimed at people who already maintain parsers and want to know what they would be trading.

The old contract: you describe the page, the parser returns what it finds

A classic parser encodes a relationship between your code and one specific document structure. You inspect a page, find a stable-looking anchor, and write a rule that reaches from that anchor to the value you want. CSS selectors, XPath, regular expressions and libraries like BeautifulSoup differ in syntax, but they share the same contract: the rule describes location, and the parser returns whatever sits at that location.

The rule is deterministic: the same HTML through the same rule gives the same output on any machine, so a unit test with a saved page means exactly what it says. It is also brittle in a specific way, because the rule knows nothing about what the value means. If a redesign moves the price, the rule returns nothing or, worse, the neighbouring element that now sits in the old position. The parser cannot notice that "Contact us" is not a price.

Teams accept this trade because, for stable pages, determinism plus a one-off writing cost is hard to beat. The pain comes from unstable pages and from the number of distinct layouts you eventually cover, which we look at from the operations side in the web scraping alternative comparison.

An html parsing alternative with AI: you describe the data

AI extraction inverts the contract. Instead of telling the system where a value lives, you tell it what you want, and it reads the rendered page to work out where that information is. In NeuralVerge's AI extraction product, the page is loaded the way a browser would show it, boilerplate such as navigation and ads is stripped, and the remaining content is mapped to your fields and returned as structured JSON.

The consequences:

  • —The unit of work is a field with a meaning, not a path to a node. "Founding year" survives a redesign that moves the founding year, because nothing in your request mentioned the position.
  • —The same request works across different layouts, and the output is typed JSON with no HTML in your code.
  • —The behaviour is probabilistic at the edges. A model reads the page, and reading involves judgement, so the determinism you used to get for free becomes something you build on purpose.

Before and after: the same task in code

Take a simple job: pull a company's name, founding year and headcount range from a profile page. First the parser version.

# Illustrative parser: tied to one page's markup
from bs4 import BeautifulSoup

def parse_profile(html: str) -> dict:
    soup = BeautifulSoup(html, "html.parser")
    name = soup.select_one("div.company-header h1.title").get_text(strip=True)

    facts = {}
    for row in soup.select("table.facts tr"):
        label = row.select_one("th").get_text(strip=True).lower()
        value = row.select_one("td").get_text(strip=True)
        facts[label] = value

    return {
        "company_name": name,
        "founded_year": facts.get("founded"),
        "employee_range": facts.get("employees"),
    }

Every line of that function is a bet about markup. The heading class, the table class, the label text "founded" and the assumption that the label sits in a th are all load-bearing. If a second source calls the label "Established" or renders the facts as a definition list, you write a second function. It also assumes the HTML already contains the content, which fails once the site renders facts with JavaScript.

Now the schema-driven version, sent to the run-extract endpoint. The request carries the URL, a plain-language instruction and, optionally, a JSON Schema passed as a string in settings.extract_schema_json. The full reference is at docs.neuralverge.ai.

# Describe the data, not the markup
import json, requests

SCHEMA = {
    "type": "object",
    "properties": {
        "company_name":   {"type": "string"},
        "founded_year":   {"type": "string"},
        "employee_range": {"type": "string"},
    },
    "required": ["company_name"],
}

resp = requests.post(
    "https://api.neuralverge.ai/functions/v1/run-extract",
    headers={"Authorization": "Bearer <API_KEY>"},
    json={
        "url": "https://example.com/company/acme-oy",
        "instructions": "Extract the company name, founding year and employee range.",
        "settings": {"country_code": "fi", "extract_schema_json": json.dumps(SCHEMA)},
    },
)
result = resp.json()["machine"]

There is no selector, no label mapping and no HTML in the function at all. A layout change on the source does not touch this code. A second source with different markup uses the same call. And rendering is handled on the extraction side, so client-side content does not require you to bolt a headless browser onto your own stack.

Where determinism goes, and where it does not

What you lose:

  • —Bit-exact reproducibility of values. Two runs over the same page may word a free-text value slightly differently. A field like a company description can vary in phrasing. A field like a year is far more likely to be stable, but you should not assume that without checking.
  • —A fixed mapping you can read. With a parser you can point at line 14 and say why a value came out the way it did. With extraction, the mapping is inside the model. When a value looks wrong, your debugging starts from the page and the request, not from a rule.
  • —Failure that announces itself. A selector that matches nothing throws an error or returns None. An extractor that misreads a page can return a well-formed, plausible, wrong answer. This is the change that most deserves your attention.

What you keep: the output shape (the schema decides keys and types), the ability to say "not found" as an empty field (NeuralVerge returns a field that is not on the page empty rather than guessed), and testability, if you build the test set. The certainty moves from the mechanism to the checks around it.

Getting determinism back: three layers

Wrap the extraction step so that what leaves your pipeline is checked. Three layers, each catching a different kind of problem.

1. Schemas pin the shape

A schema fixes property names, types and which fields are required. It is the first thing to write and the cheapest way to remove a whole class of surprises: no renamed keys, no numbers that arrive as prose, no field that appears on some pages and vanishes on others. The schema replaces the dictionary your parser function returned, declared up front and applied to every source; the schema-based extraction piece covers the mechanics.

A schema does not make the values right. It makes the shape checkable, which is what the next layer needs.

2. Validation checks the values

Treat every extraction result as untrusted input, the same way you would treat a form submission. Shape validation is the easy part, and a schema library will do it. The more valuable checks are the ones that encode what you know about the domain:

  • —Format rules. A founding year should look like a year. A country code should belong to a known set. A URL should parse as a URL.
  • —Range and consistency rules. A founding year in the future is wrong; a headcount range should be ordered.
  • —Required-field failures. A required field that comes back empty is a routing decision: review, retry, or mark incomplete.

These are ordinary, deterministic, cheap checks that do not care how the value was produced. The extractor supplies flexibility; the validator supplies the guarantee.

Keep provenance with each record (URL, time, schema version) so you can reproduce a disputed extraction.

3. Evals catch drift over time

Validation checks one record; evals check the system. An eval is a fixed set of pages with hand-verified values, run on a schedule and whenever a schema or prompt changes. It answers the question your old unit tests answered: did anything change that should not have?

A workable eval set is smaller than people expect: a few dozen pages covering your templates, including missing fields, ambiguous fields and recent layout changes, with values you would accept. Track correctness per field and stability across repeated runs. If a field wanders between runs, tighten its description or treat it as advisory. For categories, use an enumerated list in the schema so there is one right spelling.

We deliberately quote no accuracy percentages: any number would be about somebody else's pages. The eval is how you find your own.

A worked example: one profile, two pipelines

Consider the illustrative company Acme Oy (Finland) and two ways of keeping its profile data current.

In the parser pipeline, the directory ships a redesign and the facts table becomes cards. The selector returns None, the employee field is written as null, and someone must reverse-engineer the new structure, rewrite the rule and re-run the backlog. In the extraction pipeline, the request is unchanged, the headcount is found in its new place, and the record passes validation.

Now suppose the redesign also removes the headcount entirely. Both pipelines return an empty value; what happens next is decided by your schema. Because employee_range is not required here, the record is accepted with a gap. If it were required, the validator would route it to review.

The subtler case is where extraction fails quietly: the page shows a group headcount and a Finnish-entity headcount, and the extractor picks the wrong one. No selector broke and no schema was violated. Only a cross-check or an eval built around that ambiguity catches it, which is why the evaluation layer is not optional for anything feeding a decision.

Cost and latency: what to expect, stated plainly

Per-page work goes up. A parser running on your own machine does almost nothing per page once it is written. An extraction call does substantial work: render the page, clean it, run a model over the content. The right way to compare is total cost of ownership. Count the engineering hours spent writing, testing and repairing selectors across all your templates, not just the per-request bill; the pricing page covers the extraction side.

Latency goes up. Rendering and model inference take longer than a selector over HTML you already hold. That is irrelevant for an overnight refresh and may matter where a user is waiting; fetch ahead, cache, or keep a fast path for pages you can parse cheaply.

Cost does not scale with your template count. This is the counterweight. A parser pipeline's cost grows with the number of distinct layouts, because each layout is code you own. An extraction pipeline's cost grows with the number of pages. Few templates and huge volume favour a parser on cost. Many templates, moderate volume and layouts you do not control favour extraction. A long tail of one-off sources favours extraction by default.

Where the line falls depends on your volumes and templates. Measure your current maintenance time and set it against a sample extraction run on your own pages.

What still needs a classic parser

Keep a parser, or skip HTML parsing altogether, in these cases:

  • —Structured feeds you already have. If a source publishes JSON, CSV, an API, a sitemap or an RSS feed, use that directly. Sending structured data through a model to get structured data back is a step backwards.
  • —Pages you own or control. If your own application generates the markup, read from the source of truth behind it, not from the rendered page.
  • —Fixed, documented formats and very high volumes of near-identical pages, where a tested parser wins on price and speed.
  • —Bit-exact requirements. If a downstream process needs the same bytes out for the same bytes in, for example for audit reproducibility, parse deterministically and keep the model out of that path.
  • —Pre-filtering and routing. Parsers are cheap at deciding which pages are worth an extraction call.

For a company profile, the typical mixed design is to pull whatever a registry or feed offers directly, and use extraction for the unstructured remainder. NeuralVerge's 150+ ready-made sources, such as corporate registries and company profiles, already come with documented output schemas, so for those you need neither a parser nor an extraction schema of your own. The AI research side of the platform is relevant when you do not yet know which page holds the answer; extraction is the right tool once you do.

A migration strategy that does not bet the pipeline

If you run parsers today, the low-risk way to move is gradual and evidence-based.

1. Inventory your parsers by pain

Note how often each broke in the last year and how long repairs took. Frequent breakers covering many templates are the candidates.

2. Write the schema from the parser's output

Your existing function already defines the fields and types you rely on. Translate that into a schema, mark required fields deliberately, and add the domain validators you were previously doing implicitly in code. This step is where you discover what your parser was quietly assuming.

3. Run in shadow mode

Send the same pages down both paths and keep only the parser's output in production. Log every disagreement. Disagreements fall into a few piles: the parser was wrong and nobody noticed, the extractor picked a defensible alternative, or the extractor missed something.

4. Turn disagreements into evals

The pages where the two paths differed belong in your eval set. Label them by hand once.

5. Cut over one page type at a time

Keep the parser as a fallback until the new path has survived a real change on the source side. Expect a mixed system at the end.

What to check before you commit

  • —Can you define required fields and get an explicit empty value? You need the extractor to say "not found" rather than fill a blank with something plausible.
  • —Is the output typed, and is the shape identical across sources? If you still need a per-source cleanup step, most of the gain is gone.
  • —What does a failure look like? Ask to see an example of a page that does not contain the field. The response tells you how much you can trust the ones that do.
  • —Can you run your own eval set against it before committing? Any provider that discourages this is telling you something.

Frequently asked questions

Will AI extraction give me exactly the same output every time?

Not byte for byte, and it is better to plan for that than to assume it. What you can hold constant is the shape: a schema fixes field names, types and required fields, and a validation layer rejects anything that does not conform. Values on the same page should be stable when the page is stable, but you should verify that against your own pages with a small test set rather than take it on trust.

When should I keep using a classic parser?

Keep it for pages you own or that follow a fixed, documented format, for feeds and sitemaps that are already structured, for very high volumes of near-identical pages where per-page cost dominates, and for anything where you need bit-exact, fully reproducible output. A parser is also the right tool for cheap pre-filtering, such as deciding which pages are worth sending for extraction.

How do I migrate an existing parser without a big rewrite?

Run the two side by side. Keep the parser in production, send the same pages through extraction with a schema that mirrors the parser's output, and compare the results field by field. Switch over one page type at a time once the disagreements are explained, and keep the parser as a fallback until the new path has run cleanly for a while.

What happens when the page does not contain the field I asked for?

A well-behaved extractor returns the field as empty rather than inventing a plausible value. If the field is marked required in your schema, that empty result becomes an explicit signal that the record needs review, instead of a quiet gap that surfaces later in a report.

About NeuralVerge

Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.

AI Extraction on the NeuralVerge blog.

Try it on your own data

One request format across research, extraction, and enrichment.

Get started