Council Post: Benchmarks Are The New GEO And AI-SEO Strategy For B2B SaaS Companies

August 2026 · 5 minute read

Fenil Suchak is the Cofounder and CEO of OpenFunnel (YC F24).

getty

For 20 years, getting discovered meant winning the search results page. Companies built whole disciplines around it: keyword research, backlinks, content calendars. Now that many buyers ask ChatGPT and Claude instead of Google, a new playbook has emerged under names like GEO (generative engine optimization) and AEO (answer engine optimization). The promise is familiar. Optimize your content and the AI will mention you, the same way SEO once got you onto page one.

My team recently ran a study that made me question this playbook, at least for one kind of buyer: the AI agent that has been asked to research vendors, compare them and come back with a recommendation. In the sessions we observed, agents did not read marketing pages the way the GEO playbook assumes they do.

What We Watched Agents Do

Earlier this year, my team ran 200 controlled buyer sessions through Openbenchmarks For Agents, a public, independent benchmark hub, to observe how frontier reasoning models pick a vendor for a specific technical purchase, in our case, a lookalike-audience API. We gave the agents realistic buying prompts at different levels of specificity, then logged which sources they fetched, which ones made it into their reasoning and which vendor they recommended.

When an agent fetched our benchmark, it used it. On the most common query type, agents that retrieved the benchmark incorporated it into their final decision in 44 of 45 sessions, or 98% of the time. Vendors' own pages got fetched, too, but the agents often set them aside as claims they could not verify.

Traditional authority signals mattered less than I expected. The vendor that led the independent benchmark was chosen up to 86% of the time on specific, in-market queries, despite having almost no SEO authority behind it. Within our test, backlink authority was a poor predictor of which vendor got recommended.

The effect also scaled with model quality. Stronger reasoning models pulled the benchmark into their decisions far more than weaker ones. I can't say for certain that trend continues as models improve, but I have not seen a reason to expect it to reverse.

A Plausible Explanation

I can't see inside the models, so this is interpretation, but the transcripts point somewhere specific. A reasoning model given a purchasing task appears to look for comparable, structured, third-party numbers it can cite back to the person who asked. A vendor page that says "industry-leading match rates" gives it little to work with because the claim cannot be checked. A table showing match rates across eight vendors under identical test conditions is directly usable, and in our sessions, the agents used it.

There is precedent for buyers behaving this way once independent evaluation exists. Database selection shifted toward public performance comparisons like TPC. Enterprise software purchases lean on analyst evaluations. The difference with agents is speed and reach: the verification step that a human might skip on a small purchase, the agent ran on every session we observed.

What This Means If You Sell B2B Software Or Data

Our study covered one category, one purchase type and a handful of models, so I would treat it as an early signal rather than a settled rule. But if the pattern holds beyond our test, a discovery strategy built mainly on rewriting website copy for language models is optimizing the source agents weighted least at the decision step.

A few things I would do instead.

Get represented in the independent benchmarks and structured comparisons that exist in your category, and make sure the data on you is accurate and recent. In our sessions, when agents could not find structured comparisons, they pieced together whatever comparable data existed, and the vendor had no say in what that was.

Publish numbers someone could check. Latency figures, coverage rates, accuracy on a named methodology. In our sessions, verifiable specifics got used and superlatives got skipped.

Start measuring agent traffic on your own properties. Most analytics setups I have seen still treat AI crawlers and agent sessions as noise. Knowing what agents fetch from your domain and what they ignore is the only way to test any of this against your own funnel.

Finally, treat your measured performance as an input to growth, not just an engineering metric. The benchmark leader in our study won the recommendation in the large majority of high-intent sessions. Whether improving a benchmarked number beats another quarter of content production will vary by company, but it is now a comparison worth running.

The Bottom Line

One study should not end anyone's content or GEO work. The narrower claim is this: In 200 sessions across one B2B category, agents consistently preferred independent measurements over vendor-written pages when making a recommendation, and the preference got stronger as the buyer's intent got more specific and the model got more capable. If agents keep taking on more of the research step in B2B buying, the highest-leverage page about your company may be one you did not write. It is worth finding out what that page says.​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?