Product

Sep 29, 2026

AI Agent Search: Owned Index vs. SERP API (Sept 2026)

SERP vs. Owned Index for Agents | Seltz AI September 2026

Most teams pick a search API for their agents the same way they'd pick one for a web app: look at the docs, check the price, ship it. The problem is that agents aren't people. They can't click a link, they can't skim a snippet and decide to dig deeper, and they experience every slow call in a sequence, each one adding up. Whether a SERP API or an owned index search API fits your workload comes down to one question: what does your agent actually need to get back on the first call?

TLDR:

  • SERP APIs return ~30-word snippets scraped from Google; owned index APIs return full documents with no fetch required.

  • Agentic workflows consume 5-10x the tokens of a single LLM call, and thin results compound that cost quadratically across loop iterations.

  • Score providers on cost per verified answer, not price per request; a cheaper API that forces two extra hops lands more expensive at volume.

  • Scraper-based SERP APIs carry active legal risk: Bing Search APIs went offline August 2025, and Google's SerpApi litigation remains unresolved.

  • Seltz AI runs four owned indexes (news, people, companies, wikipedia) with a 1.5x p99-to-p50 ratio. Built for entity-depth retrieval after open web discovery.

What a SERP API Is and How It Works

A SERP API sits between your code and a search engine's results page. Send it a query, and it fires that query at Google, parses the HTML, and returns a structured JSON payload: titles, URLs, and short snippets, typically around thirty words each.

The provider doesn't own any index. It scrapes a results page programmatically, so what you get back reflects whatever the underlying engine decided to surface at that moment, filtered through a parsing layer.

The snippet is the key constraint. It's a passage from the page preview Google generates, not the page itself. It tells you a document is probably relevant. It rarely gives you enough to act on.

How an Owned Index Search API Works

An owned index means the provider runs its own crawlers, builds its own index, and owns every layer between the raw web and your query response. Nothing passes through Google or Bing. The retrieval pipeline is fully internal.

That vertical ownership changes what's configurable. Crawl frequency, coverage scope, retrieval models, and return format are all product decisions, not inherited constraints. The right approach to choosing a web search vendor starts with understanding which of those layers the provider actually owns. A SERP API can only return what the underlying engine surfaces; an owned index returns what its architecture is built to return.

The practical difference shows up in the response. Because the provider controls the full pipeline, it can return complete documents instead of page preview snippets. The agent gets the full record in one call, not a hint that the record might exist somewhere behind a URL.

Why Agent Consumers Need Different Infrastructure

An agent making six sequential lookups doesn't experience two seconds once. It experiences two seconds six times, in sequence, while your user waits. Unlike a person, it can't click through to read a page, so any document the index doesn't return in the response has to be fetched and parsed at runtime, adding another round trip and more tokens to carry forward.

Every gap in the response triggers more work. Agentic workflows can consume 5-10x the tokens of a single LLM call, and most teams have no meter on how many loop iterations are driving that number. Infrastructure designed for humans scanning a results page exports that problem directly to your token bill.

Search tools built for humans are optimized for a consumer that tolerates ambiguity, follows links, and brings judgment to the gaps. Agents don't have those fallbacks: they need complete information on the first call, low and consistent latency across the full sequence, and a return format they can reason over without fetching anything else.

Return Format: Links, Snippets, and Full Documents Compared

Three return formats exist, and they differ in what they hand off to the agent downstream.

A link requires a fetch every time. The agent receives a URL, retrieves the page, strips the markup, extracts what it needs, and carries that content forward as context. The retrieval step is just the beginning.

A snippet-based search result often requires a fetch too. When the answer sits inside the thirty-word preview, you're done. When it doesn't, the agent either guesses or loops back. That second case is common on queries needing structured facts or full role histories.

A full document requires no fetch. The model reasons directly over what came back, with no runtime HTML parsing, no second round trip, no markup to strip.

The downstream complexity compounds fast. As one analysis of scraper-based SERP APIs notes, five searches with a 3-second delay each produces a 15-second block before any reasoning begins, and that's before accounting for the fetch each link still requires.

The Search Loop and Its Compounding Cost

The mechanism behind ballooning agent costs isn't the number of searches: it's what re-billing looks like at each hop. As Augment Code's analysis of agent loop costs explains, context accumulation in naive agent loops follows a quadratic cost curve because the entire history is re-serialized and re-injected at every step. Message history grows linearly. Billed input tokens grow quadratically.

By the third pass, the agent is mostly paying to carry pages it already read.

A cheaper API that returns thin results and forces a second query ends up costing more than a pricier one that resolves in a single call. A basic search call that returns a full document eliminates that loop entirely. Price per request is the wrong unit. Cost per verified answer, counting every loop iteration and every reasoning token, is what determines which architecture is actually cheaper at production volume.

Latency at Scale: Why Tail Latency Beats Median

Median latency looks clean in a benchmark table and tells you almost nothing about agent performance. Agents run sequences, and a sequence never experiences the median. It experiences whatever the slowest call in the chain was.

The relevant metric is the ratio of p99 to p50. A provider with a 500ms median and a 4x p99 ratio delivers a 2,000ms worst case on every unlucky call in a sequential chain. Six lookups with one p99 hit can burn several seconds before any reasoning begins. On Exa's open benchmarks, Seltz measured a p99 to p50 ratio of 1.5x against Exa's 1.8x and Parallel's 1.9x. Full results are available on the Seltz benchmark comparisons page. That gap compounds across chained calls in a way the median never reveals.

To pressure-test a provider, run your actual query distribution at realistic concurrency. Measure p95 and p99 alongside the median, since a provider that looks fast in isolation may show tail behavior under your specific call shape that their published numbers don't surface.

Platform Risk: Legal and Structural Dependency on Third-Party Indexes

Building agent infrastructure on scraped search results introduces a dependency you don't control. On December 19, 2024, Google filed a complaint against SerpApi in the U.S. District Court for the Northern District of California, alleging unlawful circumvention of Google's technological barriers to scrape copyrighted search results at scale. A federal judge dismissed the core of the case on July 20, 2026, but Google refiled on narrower grounds on August 10, and SerpApi moved to dismiss again on August 25. The litigation is unresolved.

The Bing side is already settled. Per Microsoft's retirement notice, Bing Search APIs went offline on August 11, 2025, with no replacement. Any pipeline built on that endpoint stopped working.

An owned index carries no equivalent exposure. The provider controls its own crawl, so the retrieval layer doesn't inherit the legal or structural risk of a third party's terms of service.

Hybrid Retrieval: How Lexical and Dense Models Work Together

Owned index providers that perform well on agent workloads typically combine two retrieval legs instead of committing to one.

The lexical leg matches on terms, pinning surface facts precisely: an exact company name, a location, a job title. The dense leg matches on meaning, catching query variations and semantic intent that a keyword query would miss entirely. Without the lexical leg, named lookups get fuzzy; without the dense leg, intent-bearing queries fall apart. Both are load-bearing on real workloads.

The two legs produce separate scored candidate sets. Seltz combines them via learned score interpolation before the re-ranker stage, weighting each leg's contribution based on query type instead of applying a fixed ratio. The challenge is cost. A cross-encoder re-ranker scores candidates jointly and produces the best ordering, but running it across every document in a large index is too slow to serve. The solution is distillation: train a large, accurate teacher model to judge relevance, then transfer that judgment into smaller first-stage retrievers. Those retrievers inherit re-ranker-quality signal without the re-ranker latency on every query. Quality comes from the training; speed comes from keeping the served models small.

When to Use a SERP API vs. an Owned Index

The choice depends on what your agent is trying to do, not on which architecture is abstractly better.


SERP API

Owned Index API

Best for

SEO rank tracking, competitive SERP monitoring, Google-fidelity page ordering

Full-record retrieval, structured facts, career histories, company firmographics, full article text

Return format

~30-word snippet + URL

Full document, no fetch required

Agent loop cost

High: thin results force extra hops; costs compound quadratically

Low: complete record in one call eliminates follow-up fetches

Latency profile

Depends on third-party engine; tail latency inherited, not configurable

Controllable; Seltz p99-to-p50 ratio of 1.5x

Platform risk

Legal exposure (Google vs. SerpApi unresolved); Bing APIs went offline Aug 2025

Provider owns its crawl; no third-party ToS dependency

Date filtering

Not directly exposed

Exposed as a first-class parameter

Typical use pattern

Standalone for discovery or rank checks

Chained after open web search for entity depth

SERP APIs are the right call for SEO rank tracking, competitive SERP monitoring, and any workflow where Google-fidelity rankings are the actual deliverable. For time-sensitive workloads, filtering search results by publication date is a capability only an owned index exposes directly. If your agent needs to know what position a page holds on a results page, or how local map results are ordered, a scraper-based SERP API gives you exactly that.

Owned index APIs earn their place when you need the full record, beyond a bare pointer to it. The production pattern that holds up across real pipelines is a chain: open web search to find the entity, owned index to retrieve everything about it. Understanding search concepts like hybrid retrieval clarifies why each leg of that chain handles different query types. A recruiting agent confirming who holds a specific title at a Series B company uses open web search to confirm the name, then pulls the complete career record in one structured call. Neither step does the other's job well.

The edge cases worth planning for: single-fact lookups where a snippet resolves the question entirely, and general discovery queries with no clear target entity yet. On both, a snippet API is cheaper and fast enough. The ordering only inverts once the snippet fails to answer and the loop begins.

How to Test a Search API Against Your Actual Workload

Vendor benchmark tables tell you how a provider performs on someone else's queries, not yours.

Pull two to three hundred queries from your actual product: not synthetic ones, not representative samples from a benchmark suite. Use the queries your users generate, weighted toward the hard cases where the agent currently loops or returns thin results. Run every provider under consideration against that set under identical conditions: same tier configuration, same parameter settings, same concurrency. For news workloads, make sure to test filter by date behavior, since freshness requirements often drive the hardest latency cases.

Controlling for tier is where most comparisons go wrong. A fast low-quality mode against a slow high-quality one produces a meaningless number. Match tiers to their intended use and keep settings constant across the run.

Score on cost per verified result. Count every search call, every reasoning token, every loop iteration it took to reach a confirmed answer. A provider priced at $5 per 1,000 requests that forces two additional hops on a third of your queries will land more expensive than one priced higher that resolves in a single call.

On latency, measure p95 and p99 alongside the median, at the concurrency your production workload actually sees. A provider that looks fast in a sequential test can show pronounced tail behavior under load that the median never surfaces.

How Seltz AI Approaches Owned Index Search for Agents

Seltz built its retrieval stack from the ground up, without fine-tuning general-purpose models. The first stage runs a hybrid retriever: a learned-sparse lexical model handles surface-level precision on names, locations, and titles, while a dense bi-encoder captures semantic intent. A cross-encoder re-ranker scores the top candidates jointly for final ordering. The step that makes that quality servable at agent latencies is distillation, training a large teacher model to judge relevance, then compressing that judgment into the smaller models actually served in production.

On Exa's open 1,400-query people search benchmark, Seltz scored 0.91 HR@1 at a 363ms median. On the 605-query company search benchmark, Seltz tied for the best HR@10 at 0.86 and returned results at 327ms, roughly 3.4x faster than Exa. On the Live News Search Benchmark, Seltz answers at 89% accuracy in under 250ms, 5 points above Exa.

Coverage is clear and bounded. Seltz offers four dedicated indexes: news, people, companies, and wikipedia. There's no general web scope. The production pattern that works is a chain: open web search for discovery, Seltz for depth once you know the entity you're retrieving.

Free credits with no credit card required are at console.seltz.ai.

Final Thoughts on Picking the Right Web Index API for Agents

The benchmark tables and cost models in this post are only a starting point. The call that actually matters is the one your agent makes on your real query distribution, at your production concurrency, scored on total cost per verified answer. Run that test, measure p95 and p99 alongside the median, and the right architecture will be obvious. If you want to pressure-test the numbers against your specific workload, a short conversation can help.

FAQ

What's the best web search API for RAG pipelines that need real-time data?

The right answer depends on query shape. For structured lookups where you need a complete record in one call, an owned index like the Seltz Web Knowledge API returns full documents with no runtime fetch step, which keeps the search loop short and token costs bounded. For open-ended discovery where Google-fidelity ranking matters, a SERP API is the cheaper starting point; the loop only gets expensive once the snippet fails to answer and the agent starts fetching pages.

How do I give my AI agent access to live web data without building my own crawler?

Connect to a web knowledge API that operates its own index. Seltz, for example, crawls and maintains its own news, people, companies, and wikipedia indexes, so your agent gets full documents back in a single API call with no HTML parsing or page fetching at runtime. The alternative (wrapping a SERP API and fetching each link) works but re-bills a growing context window on every hop, and costs compound quadratically by the third loop iteration.

SERP API vs. owned index search API: which search architecture should I use for AI agents?

Use a SERP API when your workload is SEO rank tracking, competitive SERP monitoring, or any task where Google-fidelity page ordering is the actual deliverable. Use an owned index when your agent needs the full record and not a bare pointer to it: structured role histories, company firmographics, full article text. The production pattern that holds across real pipelines is a chain: open web or SERP search to find the entity, owned index to retrieve everything about it in one structured call.

What is the retrieval tax, and how does it affect AI agent costs?

The retrieval tax is the tokens an agent burns preparing content before it can reason: fetching pages, stripping markup, and extracting usable text from what a search call handed back. It compounds because each loop iteration re-bills the entire growing context window, so by the third pass the agent is mostly paying to carry pages it already read. An independent experiment measured this directly: a three-hop web search loop cost over 28,700 tokens against a 600-token baseline, while a single owned-index call returning the full document cost ~6,900 tokens. A cheaper API that forces a second query ends up more expensive than a pricier one that resolves in one call.

How do I measure a web search API's latency for AI agent workloads?

Median latency is the wrong metric for agents because a sequential chain never experiences the median; it experiences whatever the slowest call in the chain was. Measure the p99-to-p50 ratio instead: a provider with a 500ms median and a 4x p99 delivers a 2,000ms worst case on every unlucky call. On Exa's open benchmarks, Seltz recorded a p99-to-p50 ratio of 1.5x, against Exa's 1.8x and Parallel's 1.9x. Run your own query distribution at realistic concurrency and score p95 and p99 alongside the median; a provider that looks fast in a sequential test can show pronounced tail behavior under your specific call shape that published benchmark numbers never surface.

Fast, up-to-date web data, providing context-engineered web signals with sources for real-time AI reasoning.

Fast, up-to-date web data, providing context-engineered web signals with sources for real-time AI reasoning.

Fast, up-to-date web data, providing context-engineered web signals with sources for real-time AI reasoning.