AI engines disagree 2× on which brands to recommend, Treyci finds
In one B2B software category, one engine named the tracked brands in 81% of answers; another named them in 43%.
Treyci, an AI-visibility measurement firm operating as a fictitious name of Helix Apps LLC, released data on September 3 showing that two AI engines answering the same buying-intent prompts in a single B2B software category named the tracked brands at wildly different rates: 81% versus 43%. That’s a 2× gap on identical questions, across more than 1,200 scored answers built from roughly 100 prompts per category, each run three times per engine each month.
The methodology matters because the finding does. Treyci tested ChatGPT, Perplexity, Gemini, and Grok. Two of them told buyers something close to opposite stories about the same market.
“Marketing teams are making decisions about AI search based on one screenshot of one answer from one engine. The data says that’s a coin flip wearing a suit. The engines disagree with each other, and they disagree with themselves from one asking to the next. Until a brand measures the distribution of answers — across engines, on repeat runs — it doesn’t know its own numbers.” So says Keith Schilling, Treyci’s founder and formerly an AEO/GEO practitioner at PayPal. The Agile Brand Guide’s September 5 rundown put the incentive problem more bluntly: “the vendor that chooses your queries chooses your score.”
The disagreement lands into a category already famous for its own bloat. Scott Brinker and Frans Riemersma’s State of Martech 2026 counts 15,505 commercial products; Gartner’s 2025 survey pegs actual capability use at 49%. AI engines are now the discovery layer sitting on top of that mess, and they’re pointing at different vendors.
Adjacent pieces are moving the same week. Comscore rolled out AI Intelligence to measure sponsored chat placement, with three data points already showing hotel-related sponsored-ad presence in ChatGPT climbing from 6% in March 2026 to 14% in April to 24% in May. Treyci also scanned 100 B2B SaaS sites and found 41 publishing an llms.txt file, the emerging convention for talking to AI crawlers that Cloudflare’s default block on agent crawlers has already made contentious. The parallel to Google’s new charge for missed Local Services Ads calls is the through-line: measurement regimes are being written in real time, and the vendors writing them set the score.
Sources
- https://www.globenewswire.com/news-release/2026/09/03/3356050/0/en/new-measurement-data-ai-engines-disagree-by-2-on-which-brands-to-recommend-treyci-analysis-finds.html
- https://agilebrandguide.com/yesterdays-marketing-technology-ai-news-september-5-2026/
- https://martechseries.com/analytics/b2b-data/new-measurement-data-ai-engines-disagree-by-2x-on-which-brands-to-recommend-treyci-analysis-finds/
- https://www.martech360.com/news/stack-platforms/activecampaign-launches-active-intelligence-wavelength-marketing-ai-tuned-to-your-business-not-the-industry-average
- https://martechedge.com/news/activecampaign-launches-wavelength-for-business-tuned-marketing-ai
