Perplexity Sonar vs NeuralVerge: Cost per Cited Answer Compared
A Perplexity Sonar API alternative compared on cost per cited answer, once retries, follow-ups and verification are counted, with a model to fill in.
Published September 20, 2026
If you are looking for a Perplexity Sonar API alternative, the list price of one call is the least useful number to start from. What you are buying is not a call. It is a cited answer that someone can act on, and getting one usually takes more than one call, plus some checking. This post compares Perplexity's Sonar and NeuralVerge's AI research on that unit: the cost per cited answer, counted honestly.
We have already written a feature-level comparison of the two, covering how a single-pass search-grounded call differs from a multi-step research pipeline. This post does not repeat it. It takes the cost question on its own and builds a cost model you can fill in with your own numbers.
One ground rule first. We are not going to quote prices for either product. Price lists change, and a number printed here would go stale or be wrong the day you read it. We describe the variables in general terms and leave the numbers to you, taken from each provider's own pricing and from the charges you actually see.
Why the price of a call is the wrong unit
Pricing pages list prices per request, per million tokens, per thousand searches or per plan. Those are prices for inputs. Nobody's product requirement is "make 10,000 API calls". The requirement is "give the analyst a sourced answer about this company" or "fill in this field for every row in the list, with a link to where it came from".
The gap between input price and output value comes from four things that a per-call price hides.
- —Attempts. The first response is sometimes not the one you keep. The question was ambiguous, the sources were thin, or the answer wandered off the point. Each additional attempt is another charge for the same answer.
- —Follow-ups. A single question often leads to a second: which of these is the parent company, what changed since last year, why do these two sources disagree. In a real workflow one answer is frequently a small conversation.
- —Verification. A citation is only worth what it costs to check. If a reviewer has to open six links to confirm one sentence, that time is part of the cost of the answer, and it is usually the largest part.
- —Rejections. Some answers are simply not usable. They are paid for and thrown away, and the money they cost has to be spread across the ones that were kept.
A comparison that ignores these four is a comparison of price lists. A comparison that includes them is a comparison of what you will spend.
Two approaches to effort per question
Per-request and token-based pricing
Search-grounded answer APIs of the Sonar kind are generally priced around what moves through the system on each call: the tokens sent to and generated by the model, and some measure of how much search context or how many searches the request used. Different model tiers carry different rates. That is the general shape, and it is a reasonable model. Cost follows volume, small requests cost less than large ones, and you can predict spend if your requests are all about the same size.
What this model does not do is tie price to how much work a question deserved. A two-word factual lookup and a rambling, multi-part question can both fit in a similar number of tokens, or a short question can trigger a long answer. The price tracks text moved, not effort spent. For workloads made of similar questions that is fine. For workloads with a wide spread of question difficulty, per-call cost becomes a distribution you have to measure rather than a number you can look up.
To use this model in the formula below, take the current rates from the provider's pricing page, run a representative sample of your questions, and record the actual charge per call. Do not estimate from token counts alone. Measure.
Selectable depth tiers
NeuralVerge runs a research task at a depth tier you choose for that task. These are the tiers and their listed turnaround:
| Tier | Listed turnaround | Described as |
|---|---|---|
| Lite | 30s to 90s | Lightweight and fast |
| Base | 1m to 2m | Efficient for many tasks |
| Core | 2m to 4m | Balanced and strong for many tasks |
| Pro | 3m to 7m | Exploratory deep search |
| Ultima | 5m to 12m | Extensive deep search |
The thing that makes this approach different is that you pick the depth. A narrow lookup can go to Lite. A broad or ambiguous question can go to Pro or Ultima. You are choosing how much of the research pipeline a question gets, rather than every question running the same process. For questions that only need ranked links, Search returns results in a single synchronous call without a research task at all.
The turnaround times above are the ranges listed in the documentation for each tier. We are not making any claim beyond that listing, and we have not measured them against another provider.
A lever beyond depth: picking the finalizer
Depth tier controls how much searching and reasoning a task gets, but NeuralVerge also lets you choose the model that writes the final answer, independently of that tier. The settings object on a research call accepts a finalizer_model parameter alongside the depth setting, so a team can run a lighter search pass and still have a stronger model synthesize the final memo, or vice versa, instead of the tier dictating both decisions at once.
That matters for cost per accepted answer specifically because the finalizer step is what a reviewer actually reads and checks. If the synthesis quality is the thing driving rejections and follow-ups, upgrading only the finalizer — without moving to a heavier search pass you didn't need — is a more targeted way to raise acceptance than moving the whole task up a tier.
NeuralVerge advantages in this cost comparison
Depth and the finalizer are two of the levers. The others act on the checking cost (V in the formula below): per-claim citations cut the work of confirming a sentence to one link, and a dedicated cross-check step surfaces disagreements between sources before a reviewer has to find them. Here is the full set:
| NeuralVerge advantage | What it means for you |
|---|---|
| Per-claim citations | Every statement in the answer links to the source behind it, so a reviewer can check one sentence without re-reading a list of links |
| Primary sources first | Company and person facts come from 150+ structured sources — corporate registries, company intelligence, enrichment, professional profiles — before the open web |
| A dedicated cross-check step | Facts are checked against more than one source, and disagreements are surfaced instead of merged into one confident sentence |
| Five depth tiers per request | From a quick Lite lookup (30s–90s) to an exhaustive Ultima investigation (5m–12m), chosen question by question |
| Selectable finalizer model | Choose the model that writes the final answer independently of how deep the search goes |
| One platform for the whole agent | Research, Search, AI Extract and source lookups on one account with one response shape — Search, AI Extract and every data source are also MCP tools |
The cost model: what one accepted answer costs
Here is a formula you can use for either provider, or for any other. All the variables are things you can measure in your own environment.
Cost per accepted answer = (spend on all calls + cost of checking) divided by the number of answers accepted
Written out with the pieces that matter:
- —N is the number of distinct questions you ask in a period.
- —A is the acceptance rate: the share of those questions that end with an answer someone actually uses, after all attempts.
- —C1 is the average charge for a first attempt.
- —r is the average number of extra attempts per question (retries, rephrased questions, a second run at a higher depth).
- —Cr is the average charge for each extra attempt.
- —f is the average number of follow-up questions per question.
- —Cf is the average charge for each follow-up.
- —V is the checking cost per accepted answer, as time multiplied by an hourly cost, or a flat allowance per answer.
Then:
Total spend = N × (C1 + r × Cr + f × Cf)
Cost per accepted answer = (Total spend + N × A × V) ÷ (N × A)
Which simplifies to:
(C1 + r × Cr + f × Cf) ÷ A + V
Two things stand out when you read it that way.
First, acceptance rate sits in the denominator of the whole spend term. If A drops from 0.9 to 0.6, the cost of every accepted answer goes up by half, whatever the per-call price is. A cheaper call with a lower acceptance rate can lose to a dearer call with a higher one.
Second, V is added on its own and is not divided by A, because you only check answers you intend to use. It is often larger than every other term put together. If a person needs two minutes to confirm the citations behind an answer, at almost any realistic hourly cost that is more than the machine spent producing it. This is why the format of the citations matters for cost as well as for trust. We come back to that below.
A worked example with Acme Oy (Finland)
The numbers in this section are illustrative. They are chosen to show how the formula behaves, not to describe how either product performs. Replace them with your own.
Say a team researches suppliers and needs an answer for each of 1,000 companies in a month. The typical question is: "Who owns Acme Oy (Finland), and has it raised any funding in the last two years? Cite each claim."
Scenario A: a per-call model with illustrative variables
Take a hypothetical per-call price. Call it P per call, and do not put a real number in it: take P from the provider's pricing page for the model tier you would use, and confirm it against a few real charges. Suppose:
- —Each question needs one call, and 20 percent need one retry: r = 0.2, Cr = P.
- —Half of the questions prompt one follow-up: f = 0.5, Cf = P.
- —Acceptance after all of that is 80 percent: A = 0.8.
- —Checking takes one person 3 minutes per accepted answer, and you value that at V.
Then cost per accepted answer is (P + 0.2P + 0.5P) ÷ 0.8 + V, which is 2.125 × P + V.
So in this illustration you pay for a little over two calls per accepted answer, plus the check. The point is not the 2.125. The point is that it is not 1.
Scenario B: a depth-tiered pipeline, with the same kind of variable
Now the same 1,000 questions on NeuralVerge, choosing a depth tier per question. Suppose the team routes questions like this: 700 of them are simple lookups that go to a lower tier such as Base, and 300 are compound ones that go to a higher tier such as Core. Call the average charge you measure for a lower-tier question Q1 and for a higher-tier question Q2, read from your own usage records for the tiers you would actually use — measured the same way as P in Scenario A, not printed here. To keep the comparison honest, assume the same friction as before: 20 percent need a second run at the same tier, and half prompt a follow-up at the same tier as the original, with acceptance at 80 percent.
- —First attempts: 700 × Q1 + 300 × Q2.
- —Retries (20 percent of questions, same tier): 0.2 × (700 × Q1 + 300 × Q2).
- —Follow-ups (half of the questions, same tier): 0.5 × (700 × Q1 + 300 × Q2).
- —Total: 1.7 × (700 × Q1 + 300 × Q2), i.e. 1,190 × Q1 + 510 × Q2.
Accepted answers are 1,000 × 0.8 = 800. So the machine cost per accepted answer is (1,190 × Q1 + 510 × Q2) ÷ 800, which is 1.4875 × Q1 + 0.6375 × Q2, before the check.
The comparison that matters is not a number printed in an article against another number printed in an article. It is (1.4875 × Q1 + 0.6375 × Q2) plus V against 2.125 × P plus V, with your own P, your own Q1 and Q2 measured for the tiers you'd route to, your own acceptance rates measured on your own questions, and your own V. Neither side of that comparison needs a number we supply — if you get the acceptance rate and the retry rate from a real sample, the answer takes an afternoon to compute.
Where the hidden cost actually is
Retries
A retry is triggered by a bad first answer, so the way to cut retry cost is to match effort to the question, which is what a depth tier is for. Specific questions with a named entity, a time frame and a request to cite also get fewer retries under any provider.
Follow-ups and rejections
Follow-ups are cheap to ask and easy to forget in a budget, so treat a question and its usual follow-up as one unit. Rejected answers are paid for and discarded, and their cost is spread across the accepted ones through the A in the formula. That is why cost should be stated per accepted answer, never per call.
Verification
This is the largest term for most teams and the one most often left out. What drives it is how easily a reader can go from a sentence to its evidence.
A single list of sources at the end of an answer tells the reader the answer is backed by these links, somewhere. To confirm one sentence, they open several. Citations attached to each claim tell the reader which source backs which sentence, which cuts the checking work to opening one link per claim. We describe the difference in more detail in the earlier comparison. For cost purposes the takeaway is narrow: measure V under each option using the same reviewer and the same questions. If per-claim citations bring the check from three minutes to one, that saving multiplies across every accepted answer and can outweigh the difference in machine spend entirely.
We are not claiming a specific reduction. Measure it.
Perplexity Sonar vs. NeuralVerge: where each approach tends to fit
Being clear about where a fast, single-pass, search-grounded call makes sense keeps this comparison honest.
- —Chat-facing, single-turn questions where the user waits for an answer in real time, and one good search result is enough.
- —High volumes of small, similar questions where per-call cost is stable and acceptance is high on the first attempt. Here the retry and follow-up terms are near zero and the formula collapses to the call price plus a light check.
A multi-step research pipeline with selectable depth tends to fit differently.
- —Compound questions such as ownership and funding together, where one search pass would have to cover meaningfully different ground.
- —Work where the check is the expensive part, such as due diligence, supplier vetting and background research, where per-claim citations are meant to shorten review.
- —Mixed workloads where question difficulty varies widely and choosing a tier per task lets you put more work only where a question warrants it.
Here is how Sonar and similar LLM research and search APIs compare with NeuralVerge, dimension by dimension, which helps when deciding which questions to route where:
| Dimension | Perplexity Sonar & other LLM search APIs | NeuralVerge |
|---|---|---|
| What you get back | Ranked web results, page content, or a generated answer over them — depending on the product and mode | A finished, synthesized answer with inline citations — or a structured record straight from a named source |
| Where facts come from | Mostly a general web index or the open web | Structured sources first — corporate registries, company intelligence, contact enrichment, professional profiles — with the open web as fallback |
| Company & person records | Whatever web pages happen to say about the company or person | Direct lookups across 150+ sources: register records, officers and filings, funding rounds and investors, work email, phone, professional profiles |
| Citations | Varies — check whether sources are attached per claim or as one list per response | Per claim — every statement in the answer links back to its source |
| Cross-checking | Usually left to your agent or your prompt | A dedicated step checks facts against more than one source and surfaces disagreements instead of smoothing them over |
| Depth control | Varies by product — modes, processors or model tiers | Five tiers chosen per request, from Lite (30s–90s) to Ultima (5m–12m) |
| Synthesis model | Varies by product | Selectable per task via finalizer_model, independently of depth tier |
| Agent access | REST; many also ship an MCP server | AI research over REST; Search, AI Extract and every data source also as tools on a hosted MCP server |
| Best fit | Fast single-fact questions, chat answers, broad web topics | Company and person questions that need primary sources and a citation trail — account research, KYB, due diligence |
Many systems use both, routing by question type. Nothing about the two is exclusive.
How to run this comparison yourself
A short procedure gets you real numbers in an afternoon.
- —Build a question set. Take 30 to 50 questions from your actual workload, including a few of the awkward ones. Do not curate it to flatter any provider.
- —Fix the rules for acceptance. Decide in advance what counts as a usable answer: correct entity, all the requested fields present, each claim traceable to a source. Have the person who would use the answers do the marking.
- —Run each option the way you would in production. That includes your normal retry policy and the follow-ups a user would realistically ask.
- —Record every charge. For each option, read the charge for each request or task from its usage records rather than computing it from a price sheet.
- —Time the checking. For every accepted answer, have the reviewer confirm its citations and record the minutes. Convert to money with an hourly cost that suits you.
- —Compute the formula. Work out A, r, f, C1, Cr, Cf and V from the records, then the cost per accepted answer for each option.
If two options come out close, the tiebreaker is usually how easy the citations are to verify, since that affects V for every answer you keep.
What to check before you commit
- —Can you choose depth per request? Without that, easy questions and hard ones get the same amount of work, so easy ones take longer than needed or hard ones get under-served.
- —What is the retry rate on your questions? Measure it. It multiplies your spend directly.
- —How are citations attached? Per claim or per response. It changes V.
Frequently asked questions
Why not just compare the list price of one call?
Because one call is rarely one answer. Retries, follow-up questions and the time a person spends checking citations all belong to the cost of an answer you can actually use. A cheap call that needs three attempts and a manual check can cost more than a pricier call that is accepted the first time.
How do I estimate acceptance rate without guessing?
Measure it. Take a sample of real questions, run them through each option, and have the person who would use the answer mark each one accepted or not. Acceptance rate is the accepted count divided by the total. A sample of a few dozen questions from your own workload tells you more than any published figure, including ours.
Can I use both in one system?
Yes. Nothing stops you from sending quick, single-fact questions to a fast search-grounded call and sending compound or higher-stakes questions to a deeper pipeline. Routing by question type, and measuring cost per accepted answer for each route, is a sensible way to run both.
About NeuralVerge
Give your agents structured, cited, real-world data from 150+ sources and the open web — through one API or MCP server.
AI Deep Research on the NeuralVerge blog.
Try it on your own data
One request format across research, extraction, and enrichment.