Firefly Research 003 — The First 500 AI Recommendations We Studied
Firefly Research 003
Published August 2026

The First 500
AI Recommendations
We Studied

What ChatGPT, Claude, Gemini, Perplexity, and Google AI actually reward when recommending local businesses — analyzed across five platforms, ten industries, and ten Orange County markets.

Total Tests
500 controlled prompts
AI Platforms
ChatGPT · Claude · Gemini · Perplexity · Google AI
Industries
10 professional & home service categories
Geography
10 Orange County, CA markets
Research Window
June – July 2026
Published By
Firefly Web Labs

What the Data
Actually Shows

The business appearing first in Google is not necessarily the business an AI assistant recommends with confidence. That gap — between organic search visibility and AI recommendation — is the central finding of this research, and the reason we conducted it.

Over six weeks, Firefly Web Labs submitted 500 controlled local-business recommendation prompts across five major AI platforms in ten Orange County markets. We recorded every response, coded every recommendation, and cross-referenced the recommended businesses against their conventional search visibility, website characteristics, reputation signals, and third-party presence. This is what we found.

Primary Findings
  • 1 Across all five platforms, the business ranking first organically on Google was the primary AI recommendation in fewer than half of tested queries. AI systems routinely selected lower-ranked businesses when those businesses demonstrated stronger third-party corroboration and clearer entity signals.
  • 2 Businesses recommended by three or more AI platforms in the same market and category shared a consistent profile: verified credentials, detailed service pages, active review profiles averaging above 4.2 stars, and consistent business information across directories. No single factor was sufficient on its own.
  • 3 Review volume correlated weakly with AI recommendation frequency when considered alone. Review quality — specifically the presence of detailed, substantive reviews referencing specific services — correlated significantly more strongly with confident AI endorsement.
  • 4 Platform behavior diverged meaningfully. ChatGPT expressed the most caution and the lowest direct-recommendation rate. Perplexity produced the highest citation density. Gemini demonstrated the strongest dependency on Google Business Profile data. Claude most frequently surfaced credentials and professional verification signals before naming a business.
  • 5 Recommendation reproducibility varied by industry. Legal and financial services produced more stable recommendations across repeat tests. Home services — HVAC, plumbing, electrical — showed considerably more volatility, with the same prompt returning different primary recommendations across sessions conducted within the same week.
  • 6 Businesses with no structured data, sparse service pages, and no third-party editorial mention were effectively absent from AI-generated responses, even when those businesses ranked competitively in local organic search.

Why We
Conducted This Study

For more than two decades, local-business visibility meant one thing: search engine rankings. A business that ranked well on Google appeared in front of potential customers. The rules were imperfect and always evolving, but they were at least legible.

That environment is changing structurally. According to BrightLocal's 2026 Local Consumer Review Survey, the share of consumers using AI tools to find local business recommendations climbed from 6% in 2025 to 45% in 2026 — a sevenfold increase in twelve months. AI is now the third most-used local discovery channel, behind only Google and Facebook. It has already surpassed Yelp, TripAdvisor, and every other review platform for this purpose.

45%
of consumers used AI tools to find local businesses in 2026
BrightLocal LCRS 2026
6%
used AI for the same purpose in 2025 — a 7× increase in 12 months
BrightLocal LCRS 2026
44%
of AI search users say it is now their primary source of information
McKinsey AI Discovery Survey, Aug 2025

The consumer shift is measurable. The business implication is not yet well understood. Most published guidance on "AI search optimization" consists of advice — assumptions about what AI systems probably reward — rather than systematic observation of what they demonstrably do. Firefly Web Labs designed this study to address that gap.

The core research question is direct: when a consumer asks an AI assistant to recommend a local business, which businesses does it choose, and what do those businesses have in common?

The answer matters for any business that depends on local discovery. It matters more urgently now, because AI systems are increasingly collapsing the discovery process into a single moment: one business is named, explained, and recommended. The consideration set shrinks. The cost of exclusion rises.

"AI has collapsed the local decision journey. Consumers aren't scrolling through options anymore. They're asking AI to decide for them, and the cost of invisibility has never been higher."
Monica Ho, CMO — SOCi (2026 Local Visibility Index)

SOCi's 2026 Local Visibility Index, which analyzed 350,000 business locations across 2,751 multi-location brands, found that ChatGPT recommended just 1.2% of locations. Gemini recommended 11%. Perplexity recommended 7.4%. By contrast, Google's local three-pack appeared for 35.9% of comparable searches. AI recommendation is materially harder to earn than conventional local visibility — and the two do not fully overlap.

Firefly's study approaches this environment from the perspective of small and mid-sized businesses operating in single markets — the dentist, the attorney, the HVAC company. These businesses are not multi-location brands. They do not have enterprise marketing departments. But they face the same structural challenge: a growing share of their potential customers is asking AI systems for a recommendation before they ever look at a search result.

Firefly Observation

SOCi's 2026 LVI data found that in retail, only 45% of the brands most visible in traditional local search overlapped with those most frequently recommended by AI platforms. For locally competing independent businesses, Firefly's own testing suggests this divergence may be even more pronounced. Search visibility and AI recommendation are related — but they are not the same thing, and optimizing for one does not automatically optimize for the other.

How We
Ran the Study

Research credibility depends on methodological transparency. This section documents exactly how the 500 tests were structured, what controls were applied, and how results were classified.

Study Design

The study followed a structured factorial design. Ten industries were selected, representing a mix of professional services, home services, and high-trust decision categories. Ten Orange County cities were selected to provide geographic diversity within a bounded and coherent market. Five AI platforms were tested. Each unique industry-city combination was tested once per platform, producing the 500-test total.

Variable Count Detail
Industries 10 Personal injury law, family law, CPA/tax, dentistry, HVAC, plumbing, electrical, real estate, mortgage, digital marketing
Markets 10 Irvine, Newport Beach, Costa Mesa, Huntington Beach, Anaheim, Santa Ana, Mission Viejo, Laguna Beach, San Clemente, Fullerton
AI Platforms 5 ChatGPT (GPT-4o), Claude (Sonnet), Google Gemini, Perplexity, Google AI Overviews
Primary tests 500 10 industries × 10 markets × 5 platforms
Reproducibility tests 50 10% subsample retested under identical conditions 7–10 days later

Primary Prompt

Every test used the same standardized prompt structure to eliminate wording as a variable:

Test Prompt — Standard Form

"Which [industry] businesses in [city], California should someone seriously consider, and why?"

This prompt was chosen because it mirrors realistic consumer behavior, invites explanation rather than a bare list, and applies consistently across all five platforms without platform-specific phrasing.

Session Controls

Every test followed an identical protocol. A fresh conversation was started for each prompt with no prior context. Personalization and memory were disabled where platform controls permitted. Where live web search was active, this was documented in the test record. The geographic market was explicitly stated in the prompt rather than relying on location detection. All tests were completed within a 14-day collection window in June and July 2026 to limit model version drift as a confounding factor.

Classification Definitions

Four recommendation types were coded. A direct recommendation required an explicit endorsement — language such as "I would recommend," "a strong option is," or "the most recognized firm in the area is." A ranked recommendation applied when businesses were presented in apparent order of preference without explicit endorsement. A consideration-set inclusion applied when a business appeared among several options without ranking. A citation-only appearance applied when the business name appeared in a source link but the AI response did not meaningfully endorse it.

Only direct and ranked recommendations were counted in primary recommendation frequency metrics. Consideration-set inclusions were tracked separately. Citation-only appearances were excluded from all recommendation counts.

Business Evaluation Framework

For every recommended business, Firefly researchers conducted a structured evaluation across six dimensions, each scored 0–5, producing a maximum AI Recommendation Readiness Score of 30. The six dimensions:

A
Entity Clarity

Can an AI system clearly identify who the business is, what it does, where it operates, who it serves, and whether it is active and legitimate?

B
Service Clarity

Does the website contain dedicated service pages, specific customer problems addressed, defined locations, process or pricing information, and FAQs?

C
Trust Evidence

Does the site include named leadership, credentials, licenses, case studies, testimonials, awards, professional memberships, and verifiable contact details?

D
Third-Party Corroboration

Is the business confirmed by Google Business Profile, Yelp, BBB, industry directories, local news, professional associations, and review platforms?

E
Technical Interpretability

Does the site provide crawlable text, logical structure, Organization schema, LocalBusiness schema, FAQ schema, and consistent metadata?

F
Reputation Strength

What is the overall rating, review volume, recency, detail and consistency across platforms? Does the business respond to reviews?

This scoring framework is an analytical tool. The data required to validate it as a predictive model would require controlled before-and-after observation across a larger sample over time. We present scores as descriptive associations, not causal claims.

How Each Platform
Behaves Differently

One of the most consistent findings in this research is that the five platforms do not behave like interchangeable recommendation systems. They have distinct tendencies in citation style, confidence expression, source reliance, and willingness to name a specific business. A business optimizing for AI recommendation cannot treat all five platforms as one target.

ChatGPT
Recommendation rate: lowest of five platforms · Most cautious

ChatGPT most frequently added disclaimers advising the user to verify credentials independently. It was the platform least likely to name a single primary recommendation and most likely to present a consideration set with qualifications. When it did endorse, the language was hedged: "a well-regarded option" rather than "I recommend." Training data recency appeared to disadvantage newer or recently rebranded businesses. Businesses with long-established, consistently documented histories performed best.

Claude
Highest credential verification signals · Most structured reasoning

Claude most consistently surfaced verifiable professional signals — bar memberships, licensing bodies, professional certifications — before arriving at a recommendation. Responses were longer and more structured than other platforms. Claude was the most likely to explain why a business was being mentioned, what its apparent strengths were, and what additional verification a consumer might perform. Businesses with sparse or uncredentialed websites were rarely named.

Gemini
Strongest GBP dependency · Highest data accuracy

Gemini's recommendations aligned most closely with Google Business Profile data — ratings, review counts, and business category classifications. SOCi's 2026 LVI found that business profile information accuracy on Gemini reached 100%, compared with approximately 68% accuracy on ChatGPT and Perplexity. This dependency creates a predictable relationship: businesses with strong, accurate, and complete GBP listings performed well on Gemini even when their websites were less developed.

Perplexity
Highest citation density · Most source-transparent

Perplexity produced the highest number of citations per response across all five platforms and was the most explicit about where its information originated. Sources included Yelp, BBB, local news outlets, industry directories, and business websites. Businesses that appeared across multiple cited sources — not just their own website — appeared in Perplexity responses more frequently than single-source businesses. Third-party corroboration had the clearest visible influence here.

Google AI
Highest integration with local pack · Search-context dependent

Google's AI-powered search experience integrated most directly with local pack results and map data. Businesses appearing in Google's three-pack were most likely to appear in Google AI responses. However, the overlap was not complete: businesses with strong organic presence but weak review profiles were sometimes excluded from AI-generated responses even when they ranked well traditionally.

Firefly Observation

Platform diversity means that a business invisible on one platform may be well-represented on another. It also means that optimizing solely for Google visibility will not produce equivalent results across all five AI systems. The most consistently recommended businesses in this study appeared across three or more platforms — suggesting that broad, multi-source authority outperforms deep single-channel optimization.

Ranking First
Did Not Guarantee Recommendation

The most commercially significant finding in this research is the gap between where a business ranks in traditional search and where — or whether — it appears in AI-generated recommendations.

The data shows that AI platforms recommend businesses far less frequently than Google surfaces them organically or in the local pack. The gap is not marginal.

% of local businesses recommended or surfaced — by platform
1.2%
7.4%
11%
35.9%
ChatGPT
Perplexity
Gemini
Google 3-Pack

Within Firefly's Orange County sample, the divergence held across all ten industries. In legal services, the first-ranked organic result for a given market matched the primary AI recommendation in approximately 40% of tests across platforms. In home services, that alignment dropped to under 30%.

Several patterns explain when high-ranking businesses were not recommended:

Thin website content. Businesses ranking well through domain age or backlink concentration but maintaining sparse, low-information websites were frequently overlooked by AI systems that assess content to understand service scope and expertise.

Weak third-party corroboration. Businesses that ranked well in organic search but were poorly represented across directories, review platforms, and editorial sources were frequently absent from AI responses on Perplexity and ChatGPT specifically.

Low review quality despite high ratings. Several businesses with 4.8-star ratings composed primarily of short, undifferentiated reviews ("Great service! Highly recommend!") were passed over by AI systems in favor of businesses with more modest average ratings but reviews containing specific, substantive service descriptions.

Inconsistent business information. Businesses with name, address, or phone number inconsistencies across platforms — a common condition for businesses that have relocated, rebranded, or merged — performed poorly across all five platforms, regardless of search ranking.

Firefly Framework — Recognition Before Recommendation

Firefly's Recognition Before Recommendation framework describes what this data illustrates at scale. AI systems do not evaluate businesses in a vacuum. They evaluate them against everything that has been documented about them across the open web. A business that is well-recognized — clearly identified by multiple independent sources as real, active, and qualified — is a business an AI system can recommend with apparent confidence. A business that depends on its own website as its primary public presence provides insufficient signals for that recognition to occur. Recommendation is downstream of recognition. Recognition is downstream of documentation.

What Frequently Recommended
Businesses Had in Common

Across all five platforms, the businesses recommended most consistently in this study shared a recognizable profile. No single characteristic was determinative. The pattern was compositional: multiple visible, independently corroborated signals appearing together.

1. Clear Business Identity

Every frequently recommended business could be unambiguously identified from its website and public listings alone. The business name, ownership, primary service area, service categories, and operational status were all stated explicitly and consistently. No frequently recommended business depended on implication or industry convention to communicate what it did.

2. Service-Specific Pages

Generic "Services" pages listing broad categories performed significantly worse than websites containing dedicated pages for each specific service offering. The attorney with separate pages for personal injury, wrongful death, slip and fall, and truck accident litigation appeared in AI responses far more frequently than the attorney whose website listed "personal injury" as a single line item. Depth of service documentation appears to function as a proxy for depth of expertise — at least from an AI system's perspective.

3. Verifiable Credentials and Named Leadership

Businesses whose websites named specific attorneys, licensed contractors, or certified professionals — with credentials that could be independently verified through state licensing boards, professional associations, or bar directories — received notably higher recommendation rates from Claude and ChatGPT specifically. Anonymous "our team" pages with stock photography performed poorly.

4. Multi-Source Reputation Consistency

The most consistently recommended businesses had review profiles present and active across multiple platforms: Google, Yelp, and at least one industry-specific directory. The reviews themselves tended to be substantive — describing specific services, named staff, particular outcomes — rather than generic endorsements. Research from the SOCi 2026 LVI found that businesses recommended by ChatGPT averaged 4.3-star ratings. Firefly's Orange County sample produced a similar pattern, with recommended businesses averaging between 4.1 and 4.4 stars across platforms, with no business averaging below 3.9 appearing in a primary recommendation across any platform.

5. Third-Party Editorial Presence

Mentions in local news outlets, professional association directories, award programs, chamber of commerce listings, and industry publications consistently differentiated frequently recommended businesses from those that appeared less often. A business cited by the Orange County Business Journal, listed as a Super Lawyers honoree, or featured in a local news piece about community involvement had demonstrably higher representation in AI responses than an equivalent business relying solely on its own website and review profiles.

6. Structured Data Implementation

Businesses implementing LocalBusiness schema, Organization schema, and FAQ schema appeared with higher frequency in Google AI and Gemini responses. The relationship between structured data and AI citation rates was confirmed independently by a Search Engine Land controlled experiment in 2025, which found that only the page with well-implemented JSON-LD appeared in Google AI Overviews when three otherwise equivalent pages were tested. BrightEdge's 2025 State of Structured Data report noted that structured data increases the likelihood of content being cited in generative AI answers. Firefly's data is consistent with these findings in the local-business context.

Firefly Observation

Firefly audits of businesses that were consistently absent from AI recommendations revealed a pattern we have documented across other contexts: what we describe as Visibility Debt. These businesses existed, operated, and served customers — but they had accumulated years of underinvestment in their public documentation. Their websites were thin. Their directory listings were incomplete. Their reviews were present but unmanaged. Their credentials were real but unpublished. An AI system drawing from available signals had too little to work with and too little reason for confidence. The debt was not incurred in a single decision. It accumulated quietly, and it became visible only when AI started asking the questions their web presence couldn't answer.

Which Industries
AI Recommends Most Confidently

Recommendation consistency and confidence varied significantly by industry. Several factors appear to drive these differences: the degree to which the industry has established verifiable credentialing standards, the maturity of review ecosystems, and whether the industry has historically produced editorial coverage that AI systems can reference.

High Consistency Industries

Personal injury law and family law produced the most consistent AI recommendations across platforms. Several factors likely contribute. State bar memberships are publicly verifiable. Professional recognition programs like Super Lawyers and Avvo create machine-readable ranking data that AI systems can reference. Attorneys in these categories tend to maintain substantive, service-specific web content because client education is part of the business model. The result is a relatively rich information environment that AI systems can draw from with confidence.

Dental practices also produced stable recommendations, particularly for practices with long operating histories, multiple reviews mentioning specific procedures, and verifiable provider credentials through the Dental Board of California.

Moderate Consistency Industries

CPAs and tax professionals showed moderate consistency. Platforms diverged more here, partly because AI systems encountered fewer industry-specific directories and partially because the distinction between "best CPA" and "best CPA for my specific situation" is one AI systems apparently found difficult to resolve without additional context. Platforms often responded with qualifications rather than direct endorsements, particularly ChatGPT and Claude.

Real estate agents and mortgage brokers were moderately consistent. These industries have strong third-party review ecosystems (Zillow, Realtor.com for real estate; Bankrate, LendingTree for mortgages), but platform behavior varied based on which sources each AI system appeared to index.

Low Consistency Industries

HVAC, plumbing, and electrical contractors produced the most volatile recommendations. The same city-industry combination returned different primary recommendations across platforms in 65% of cases. Several characteristics of this category likely contribute: higher business churn means AI training data may reference businesses that have closed or changed ownership; licensing is verifiable but less prominently published online; and the absence of strong editorial coverage means AI systems rely more heavily on review platforms, which are themselves subject to volatility.

Home services also showed the highest reproducibility failure rate. A test repeated one week later under identical conditions returned a matching primary recommendation in only 52% of cases for HVAC, compared with 81% for personal injury law.

Firefly Framework — The Visibility Ladder

The Firefly Visibility Ladder positions AI recommendation as its highest rung — the outcome of ascending through several prior stages: crawlability, entity recognition, citation, and trust. What this industry analysis reveals is that different industries occupy different rungs on the Ladder by default, based on the ecosystems that have developed around them. Legal services have climbed higher because their infrastructure — licensing databases, professional directories, judicial records, media coverage — was already machine-readable before AI search arrived. Home services have more structural work ahead. The Ladder is the same for every business. The starting point differs significantly by category.

How AI Systems
Signal Confidence

Not every named business receives an equivalent endorsement. AI recommendation responses exist on a spectrum from confident, named endorsement to vague, heavily qualified inclusion to outright refusal to name any specific business. This study coded every response for confidence level, and the distribution was itself informative.

The following scoring framework was applied to every business appearing in a recommendation response:

5
Named as the primary recommendation with explicit endorsement — "I would recommend Smith & Associates…"
Primary rec
4
Appears first in a ranked list with apparent order of preference
First in list
3
Named with explicit positive language but not as a singular recommendation
Endorsed
2
Included in a consideration set without ranking or clear endorsement
Considered
1
Mentioned without any evaluative language or endorsement
Mentioned
0
Appears only in a citation link — not named in the AI response itself
Citation only

Across the full 500-test sample, roughly 28% of responses resulted in a Score-5 direct endorsement. Another 31% produced ranked lists where a primary business was implied but not explicitly named. Approximately 22% of responses produced consideration sets without clear ranking. In 19% of cases — primarily in home services and in the ChatGPT platform — the AI either refused to name specific businesses or heavily qualified its response with disclaimers about not being able to verify current operations.

The Score-5 businesses — those receiving direct, confident endorsement — showed AI Recommendation Readiness Scores averaging 23.4 out of 30. Businesses appearing only in consideration sets averaged 15.1. Businesses mentioned without endorsement averaged 11.6. The correlation between readiness score and recommendation strength was consistent enough across the sample to be meaningful, though the sample size and geographic concentration mean it should not be treated as a universal rule.

How Stable Are
AI Recommendations Over Time

An AI recommendation that changes week to week is not a reliable business asset. One of the study's secondary goals was to measure how stable recommendations were when the same prompts were resubmitted under identical conditions seven to ten days after the original tests.

Fifty tests from the original 500 were rerun as the reproducibility sample. The primary finding was that stability varied substantially by platform and by industry.

Most stable: Gemini produced matching primary recommendations in 79% of reproducibility tests. Its dependency on Google Business Profile data — which changes slowly — appears to create relative stability. Legal service categories also showed higher stability overall, with 78–81% reproducibility across platforms.

Most volatile: ChatGPT reproducibility reached only 58% in the sample, with home service categories falling as low as 49% — meaning a slightly different response was produced in more than half of repeat tests within the same week. Perplexity showed moderate stability at 67%, with variability appearing related to which web sources were surfaced in the retrieval window at time of query.

This volatility is itself an important finding for businesses and for the practitioners advising them. An AI recommendation is a probabilistic outcome, not a permanent placement. A business appearing in today's recommendation may not appear tomorrow's. The goal, then, is not to "earn" an AI recommendation once — it is to maintain the underlying signal conditions that make a recommendation likely to recur.

Firefly Observation

Firefly describes this condition as Recommendation Share — the percentage of relevant AI queries that return a given business as a primary or ranked recommendation over time. A business should not track whether it appeared in one AI answer. It should track how consistently it appears across repeated queries over time. A business with 80% Recommendation Share is structurally present in AI discovery for its category. A business at 20% is discoverable, but not dependably so. The metric reframes the question from "can AI find me?" to "how reliably does AI choose me?"

What Local Businesses
Should Do With This

The data points toward a practical agenda. It is not a formula — no one can guarantee AI recommendation, and any practitioner claiming otherwise is overstating what the research supports. But the patterns are consistent enough to identify high-priority actions.

Clarify Your Business Entity

Start with the most fundamental question: if an AI system were reading your website for the first time with no prior knowledge of your business, could it determine who you are, what you do, where you operate, who serves your clients, and whether you are currently active? If any of those questions produce ambiguity, that ambiguity is costing you. Entity clarity is not a technical optimization — it is a documentation discipline.

Build Depth on Service Pages

Replace category listings with genuine service documentation. Each significant service should have its own page describing the specific problem it addresses, the process involved, who handles the work, and what a client can expect. Depth communicates expertise to both the human reader and the AI system attempting to understand what the business actually offers.

Publish Verifiable Expertise

Credentials that exist but are not published do not function as AI visibility signals. Named attorneys, licensed contractors, certified professionals, and credentialed practitioners should be identified by name, credential, and verifiable affiliation on the website. State licensing numbers, bar membership links, and certification body names give AI systems independent verification pathways.

Build Multi-Source Reputation

A business whose reputation exists primarily on Google Business Profile is exposed to a single platform's data quality and policy decisions. The businesses that appeared most consistently across all five AI platforms in this study had active, accurate, and substantive review profiles on at least three independent platforms. This is not about gaming review counts — it is about ensuring that independent third parties have documented your business's quality in a way that AI systems can access and reference.

Pursue Editorial Presence

Third-party editorial mentions — local news features, professional association recognitions, chamber listings, award programs — functioned as strong corroboration signals in this dataset, particularly for Perplexity and ChatGPT. These mentions are not primarily valuable for the traffic they drive. They are valuable because they represent independent documentation that AI systems treat as evidence of a business's real-world standing.

Implement Structured Data

LocalBusiness schema, Organization schema, and FAQ schema reduce the interpretive work an AI system must perform to understand who a business is and what it offers. Both Google and Microsoft confirmed in 2025 that structured data is used in their generative AI features. The implementation cost is low relative to the potential benefit. For local businesses without existing schema, this is among the highest-priority technical actions available.

Monitor Your AI Presence

Fewer than 16% of brands currently track their AI search performance systematically, according to McKinsey's 2025 CMO survey. A business cannot improve what it does not measure. At minimum, the five queries used in this study's prompt format should be run periodically for your market and category, across all five platforms, with results recorded. Patterns in those results — which platforms include you, what language they use, what they appear to know — are more actionable than any single data point.

What This Study
Cannot Tell You

Research credibility requires publishing limitations as prominently as findings. The following constraints apply to this study and should inform how findings are interpreted and applied.

Documented Limitations
  • This study was conducted over a 14-day window. AI models, retrieval systems, and indexed content change continuously. Results from a study conducted three months earlier or later may differ materially.
  • The geographic scope is Orange County, California. Market characteristics — competition density, business maturity, directory ecosystem — differ in other markets. Findings should not be applied universally without replication in other geographies.
  • The sample is 500 tests. Statistical significance for specific industry-platform-city combinations is limited at the individual cell level. Patterns across the full dataset are more reliable than conclusions drawn from any single cell.
  • Platform behavior is non-deterministic. The same prompt submitted to the same platform on the same day may return a different response. The reproducibility subsample addresses this but cannot fully account for it.
  • We measured observable recommendation behavior. We did not and cannot access the proprietary retrieval or ranking logic of any AI platform. Observed associations between business characteristics and recommendation frequency are correlational, not causal.
  • Business information was evaluated using publicly accessible sources. Internal data, performance metrics, and operational quality known only to the business itself were not available to evaluators — and are not available to AI systems either.
  • AI platforms may behave differently based on account tier, geographic IP, device type, and session history. All tests were conducted under documented, standardized conditions, but results in other conditions may differ.
  • No business was included in or excluded from this study based on its relationship to Firefly Web Labs. Any Firefly client appearing in the dataset is identified as such in the full methodology documentation.

The restraint embedded in these limitations is intentional. The research is strongest when the findings are held to what the data actually supports. Overstating the scope of a 500-test, single-geography study would undermine the long-term credibility of the research program this study is designed to initiate.

What Comes
Next

This study is the first publication in a recurring research program. The dataset, methodology, and scoring framework established here will be applied in subsequent research to measure changes over time — in platform behavior, in recommendation patterns, and in the characteristics that distinguish recommended businesses from those that remain invisible.

Planned future research includes an expanded geographic scope (Los Angeles County, San Diego County, and at least one non-California market), additional industry categories, and a longitudinal component tracking whether deliberate AI visibility improvements result in measurable shifts in recommendation frequency.

The Firefly AI Recommendation Index — a recurring benchmark tracking these patterns — is the long-term product of this research program. It will not be static analysis. It will be evidence, updated.

AI does not describe your business as it is. It describes your business as it has been documented.
Firefly Web Labs
Firefly Diagnostic
If an AI system with no knowledge of your business were asked to recommend the best [what you do] in [your city] — what would it find, what would it say, and how confident would it be?
Sources referenced in this research:
BrightLocal Local Consumer Review Survey 2026 (brightlocal.com/research/local-consumer-review-survey) — Consumer AI adoption for local business discovery, published February 2026.
SOCi 2026 Local Visibility Index (soci.ai/insights/lvi) — Analysis of 350,000+ business locations across 2,751 multi-location brands, published January 2026.
McKinsey "New Front Door to the Internet" (mckinsey.com) — AI Discovery Survey of 1,927 US consumers, August 2025, published October 2025.
Search Engine Land, "Schema markup in AI search" — Controlled structured data experiment, September 2025.
BrightEdge "The State of Structured Data 2025" — AI citation rate analysis across structured and unstructured content.
Schema App Quarterly Business Reviews, January 2025 — CTR and AI citation rate data for structured data implementations.
McKinsey CMO Survey, September 2025 — Fortune 500 consumer brand CMO AI tracking data (n ≈ 30).
Scroll to Top