Loading the benchmark…
Whole Whale · AI Brand Footprint 2026 Benchmark · data collected July–August 2026

AI Brand Footprint Charity & Foundation Study

What AI tells people asking who’s best, who to fund, or who to give to

We asked seven AI models 255 questions across 17 cause areas, then read all 1,785 answers to see which nonprofits get named, which get skipped, and why.

Our thesisAI is becoming the unelected arbiter of donor, volunteer, funder, and grantee relationships. Its biases have the power to reshape philanthropy, whether or not any nonprofit has opted in.

See a visual breakdown of our prompt methodology Click to expandClick to collapse
17cause areas
×
15questions each
=
255questions asked
10 per cause: donor, volunteer & reputational ("best," "most impactful") 5 per cause: funder & foundation
See the 17 cause areas we tested Click to expandClick to collapse
See the types of questions we asked Click to expandClick to collapse
How to read, understand & interpret this report Click to expandClick to collapse

Interpreting This Report

The numbers in this report should not be taken at face value or verbatim. This report is a snapshot in time, beholden to randomness and neither the "from memory" nor the "with web search" (which both use our APIs) necessarily reproduce what consumers see directly on chatgpt.com, for example. Differences in the number of mentions between organizations should not be considered a horse race, or be seen as a commentary on organizations' reach, influence, reputation, program strength, etc. (Aka, don't take the numbers here too literally!) Instead, this report should be an entry-point for discussions within philanthropy sector about what it means for AI to become an insurgent philanthropic advisor of sorts, and how its biases, tendencies, assumptions, and preferences are replicated at scale. This report is the beginning, not the end, of a conversation about what mass adoption of consumer AI means for giving, philanthropy, and social impact broadly.

Understanding This Report

There is a lot of data here, we know! A great way to begin to understand the findings in the report is to begin with a question. "Are all charitable sectors seeing similar trends when it comes to how nonprofits are surfaced in donor prompts?" "Is my [xyz] social cause space heavily consolidated or more diverse within AI outputs?" "What kinds of 3rd party validators are being referenced with regards to nonprofit trust?" There are dedicated sections to ask both these and many more questions. If you want to read raw outputs themselves, please refer to the bottom of this deck, where we have a database of the top organizations mentioned by cause space along with raw prompt outputs.

What's The Deal With the Different Methods?

We used two methods for pinging our API with a list of prompts. The first is "from memory," which is just a basic API call to the core model itself, without any plugins, etc. While these outputs are not what the consumer sees, they represent the inherent tendencies and essence of the models themselves. The second method is the "with search" method, which allows these models to reference the internet when providing an output. Note that this method is still not a 1:1 with what the consumer sees on the front end when logged into chatgpt.com for example, but it helps us, conceptually, to understand how internet results might begin to reshape outputs in the real world. (Not to bury the lead, but we suspect that it bolsters the case for traditional SEO and web optimization strategies.)

Are There Limitations?

Yes, many limitations. First, we are nonprofit technology geeks but not data science PhDs. We used AI heavily at every point in this process. We let the AI decide how to count, how to estimate, how to interpret, and how to de-duplicate records. (Every model refers to Doctors Without Borders / Médecins Sans Frontières (MSF) differently, it turns out!) Prompt outputs, as mentioned, are not a 1:1 with what a consumer necessarily encounters, and are subject to randomness and variability. Also not all prompts are identical across cause spaces (i.e., health organizations got health-specific prompts other sectors didn't) so be careful about reading too much into the absolute value of mentions between cause spaces broadly. Again, this report is more making an argument, with data, as a proof of concept about how we should consider AI to be shaping philanthropy and engaging in knowledge production. See the Technical Notes at the end of this report for more caveats.

Does This Study Apply Beyond The United States?

This is a great question. This study was U.S.-focused with prompts only written in English. We imagine that our findings become more pronounced (and more severe) when juxtaposed with results using the consumer products in other countries. However, this raises critical questions about access, equity, tech ownership, digital sovereignty, and more. How pronounced are the biases of these models towards U.S. or Western organizations (particularly massive INGOs) when used globally? Do outputs further disenfranchise smaller, local, organizations, grassroots groups, or community functions in countries beyond the United States? We simply do not have the ability to assess that in this study, but the ramifications of these questions require urgent consideration from philanthropy globally.

Who Are You Anyway?

Thank you for asking. We are Whole Whale, a digital agency that works with nonprofits and social impact organizations of all types. In addition to providing world class digital services to nonprofits and NGOs, we care passionately about the sector itself. With a bird's-eye-view of philanthropy, trends in donor behavior, nonprofit technology, as well as a foot in the tech world, we believe we are uniquely positioned to help the charitable sector navigate this onslaught of new technology and how it's changing the sector we love. Increasingly, we try to use our platforms to advocate for the sector and push conversations in the service verticals we maintain expertise.

FINDINGS

What AI Tells Donors, Funders, and the Public: The Main Findings

These are the top-level takeaways: which organizations get named, which sectors dominate, and what that means for a nonprofit trying to be found, whether the question is “who should I give to,” “who is doing this best,” or “who funds this work.” See the full question set for exactly what was asked. Most of this comes from the models answering with no web access, which is the clearest read on what AI has absorbed about your cause. A few findings look at what changes once a model can search the web, and are labeled as such.

BY CAUSE

Cause by Cause: Who Gets Named in Your Space

Pick your cause to see the top 20 organizations AI names, then scroll down to see what each individual model recommends. Use the toggles to switch between answers given from memory and answers given with web search.

Swipe to see the full chart →

What each model names

Each model's own top ten for this cause. Names in bold appear in more than one model's list; those are the consensus picks. Names in italics appear in only one column, meaning a donor would only hear about them from that one assistant.

Swipe to see all 7 models →

THE FIELD

Interactive Organization Mention Map

Every organization the models named, across all 255 questions, placed by who named it. The seven models sit at the points of the shape. An organization all seven recommend has no pull in any direction and settles in the middle. One that only Grok recommends is pulled out to Grok's corner. The farther a dot sits from the center, the more its visibility depends on which assistant someone opens, whether they were donating, volunteering, or researching funders. This map is built entirely from answers given from memory (no web search) - it is not affected by the memory/web-search toggles used elsewhere.

COMPETITION

Some Causes Are Open, Others Are Locked Up

This chart shows how diffuse or consolidated the causes we tested are. Each bar is one cause, read left to right in four bands: the cause leader's share of mentions, then ranks 2–3, then ranks 4–10, then everyone else. In Breast Cancer the single leading organization gets 13% of all mentions and the top ten get 60%. In Poverty the leader gets 4% and the field stays open. For a nonprofit, that difference decides whether the goal is displacing an incumbent or joining a crowd.

Swipe to see the full chart →

BY QUESTION TYPE

Mentions By Cause By Question

The chart above blends every donor or funder question asked in a cause into one number. This one does not: pick a single question type and see how open or locked up each cause is for it. Only the 10 question types asked in every one of the 17 causes are included. Most causes get the same 2 prompt variants per question type, but three causes (Cancer, Lung Cancer, Breast Cancer) get 4 variants for "Best-of" specifically, so that one question type has a larger sample in those three causes than elsewhere - the exact prompt count for whatever you have selected is shown below the dropdown. Both charts below use answers given from memory (no web access).

These charts measure visibility in the tested AI answers, not funding received, organizational effectiveness, or the full population of organizations working in a cause. A broader set of names does not necessarily mean more equitable representation.

Swipe to see the full chart →

↳ Drill in: Top 10 Organizations, for This Question and Cause

Organization names, back in - this is who actually wins a single specific question in a single specific cause, and by how many mentions. Answers given from memory only, same as the chart above.

Swipe to see the full chart →

GROUNDING

Answering From Memory vs. Searching the Web First

When a model answers from memory, it is telling you what it absorbed about your sector during training. That is a measure of brand recall. When it searches first, it is telling you what it can find and cite on the open web right now. That is a measure of findability. Donors mostly get the second kind, and the two sets of answers barely resemble each other.

SOURCES

Charity Navigator and GuideStar Lead a Shared Evaluator Layer

Models repeatedly invoke third-party charity evaluators when discussing nonprofits, giving those platforms an important role in AI-mediated discovery. Charity Navigator and GuideStar appear across most models, while evaluators such as GiveWell, Candid, and BBB Wise Giving Alliance carry disproportionate weight with particular assistants. For nonprofits, getting validated is therefore both a broad reputation strategy and, increasingly, a model-specific visibility strategy.

Swipe to see the full chart →

Share of a model's answers that name each evaluator. Read down a column to see an evaluator's reach; read across a row to see a model's habits.
One gap in this chart. Perplexity cites its sources with numbered footnotes in 252 of these 255 answers, but the links behind those footnotes were not captured in the from-memory data, so it looks source-light here when it is in fact the most search-driven model tested. The web-search data does capture them: every Perplexity answer there carries sources, 2,372 links in all.
CHARITIES vs FOUNDATIONS

Charity Questions and Foundation Questions Return Different Organizations

Eighty-five of the questions ask who funds a cause. The rest ask who to give to, volunteer with, or trust as best-in-class, a wider frame than donations alone. Comparing the top 100 organizations named in each group shows the two lists are almost completely separate. Being visible when someone asks about a charity does not make you visible to funders, and the reverse.

MEMORY vs SEARCH

The Same Question, Asked Two Ways, Returns Two Different Lists

Both sets of answers are scored the same way, so the two rankings can sit side by side. The organization at the top of a cause usually survives the switch. Almost nothing else does, and which organizations get swapped out follows a pattern clear enough to change what a nonprofit should do about it.

WHO AI KNOWS

Which Organizations AI Names Across Every Cause

Charities tend to be named in their own cause and nowhere else. Foundations show up across nearly every cause in the study. Ranked by how many of the 17 causes an organization appears in, the top of the list is almost entirely grantmakers.

Swipe to see the full chart →

X-axis: number of the 17 causes an organization is named in at least once. The label on each bar is the average number of mentions per cause it appears in there - a high average means it dominates the causes it shows up in, not just shows up everywhere.
ACTIONS

What Your Organization Should Do About It

EXPLORER

Leaderboard

Swipe to see all columns →

# Organization Type Share of answers Avg. place Named first Models agreeing Evidence

APPENDIX

Technical notes

How the study was run and where the results should not be over-read. Nothing in here changes a figure above.

Click to expandClick to collapse
A1

What we actually asked

Every number in this report comes from 1,785 answers to a fixed set of questions: 17 causes × 15 questions × 7 models, with no organization named in any question. The questions are the ones a donor, a volunteer, a program officer, or someone just asking who's best would type, not prompts engineered to produce lists. Ten per cause cover donor, volunteer, and reputational framing ("best," "most impactful," "how to help"); five ask specifically who funds the cause, aimed at foundations and other grantmakers rather than the charities they fund. Which question you ask changes the answer more than which model you ask, so the full set is below with the share of mentions each one produced.

A2

How organizations were counted

A3

How big is the field in each cause?

How many distinct organizations the models named in each cause. Most names come up more than once and can be counted directly. A long tail of names come up exactly once, and many of those are misspellings, programs, or non-organizations, so that tail is estimated from a hand-checked sample rather than counted whole.

Swipe to see the full chart →

Organizations named in each cause. The solid bar counts organizations named more than once. The lighter extension is the estimated once-only tail, scaled from a 400-item sample in which 24% of once-only names were real organizations.
The field size varies a lot by cause. Some causes have a wide field of organizations getting named; others are much narrower. That range, across all 17 causes, is — organizations.
A4

What this study cannot tell you

Whole Whale

Whole Whale · AI Brand Footprint Benchmark 2026.

1,785 answers given from memory, collected 28 July 2026, and — answers given with web search, collected 24 August 2026, both through OpenRouter. Every from-memory figure can be traced to the model answer behind it through the evidence links in the leaderboard. Web-search figures are computed from their own answers, which are kept with the study but not embedded in this page.

Read more: Measuring Your AI Brand Footprint: The Hidden Visibility Challenge