What AI assistants recommend in five software categories
We asked three AI models fifty blind buying questions across five software categories on one day. The models disagreed sharply. Here is the measured data.
Webflow scored 15 on ChatGPT and 60 on Claude. Same ten questions about website builders for marketing sites, same day, no brand names in any of the questions. ChatGPT named Webflow in 3 of its 10 answers. Claude named it in 9.
That is not a rounding difference or a wording quirk. It is two assistants holding materially different views of the same market, and a buyer who asked only one of them would walk away with a different shortlist depending on which one they happened to open.
On 17 September 2026 we ran ten blind buying questions in each of five software categories across three models and recorded every brand named. The rest of this page is what came back, including the parts that are inconvenient.
How this was measured
Ten questions per category, written the way a buyer would phrase them, with no brand names in the question itself. Each question went to three models through their official APIs: ChatGPT (gpt-4.1-mini), Claude (claude-opus-5) and Gemini (gemini-2.5-flash).
Spegla normally queries four models. Perplexity was not configured for this run and returned no data at all, so everything below is three models, not four. Nothing here has been adjusted to pretend otherwise.
Every model answered all ten questions in every category, which gives 30 answers per category. Where you see a competitor with a count of 26, that means 26 of those 30 answers named the brand. The maximum possible number is 30. These are not percentages and they are not market share.
Two numbers describe the subject brand of each category. Mention rate is how often it was named at all. The score weights where it appeared inside the answer, so a brand named last in every answer has a high mention rate and a low score. The subject brand is excluded from its own competitor list by design, which is why Asana does not appear in the project management list below.
One more caveat that applies to any measurement of this kind, including ours. Answers through the official APIs approximate what the consumer apps say without being identical to them, and the same question asked twice can come back differently. This is one run on one day in one language. Treat it as a snapshot, not a trend.
The five categories
Project management software
Asana scored 58 overall with a 93% mention rate. ChatGPT gave it 60, Claude 55, Gemini 60. Claude named Asana in all ten of its answers and still produced the lowest score of the three, which is exactly what a mention rate hides: being present every time is not the same as being recommended first.
Brands named across the 30 answers:
- Monday.com 27
- ClickUp 26
- Trello 23
- Wrike 20
- Jira 18
- Basecamp 14
- Notion 13
- Smartsheet 13
- Linear 9
- Zoho Projects 7
CRM software for small business
HubSpot scored 73 overall with a 93% mention rate, and the spread between models was the second widest in the dataset: 70 on ChatGPT, 60 on Claude, 90 on Gemini. Gemini named it in all ten answers.
Brands named across the 30 answers:
- Pipedrive 26
- Zoho CRM 25
- Freshsales 17
- Salesforce 11
- Insightly 10
- Less Annoying CRM 9
- Copper 6
- Monday CRM 6
- ActiveCampaign 6
- Jobber 6
Pipedrive and Zoho CRM were named more than twice as often as Salesforce, which says less about the size of those companies than about how these questions were framed. Ask for a small business CRM and the answers reflect what third party pages say about small business CRMs.
Email marketing platform
Mailchimp scored 62 overall with a 93% mention rate, and the models were unusually close here: 65 on ChatGPT, 60 on Claude, 60 on Gemini. Claude named it in all ten answers.
Brands named across the 30 answers:
- ActiveCampaign 26
- MailerLite 20
- Klaviyo 18
- Constant Contact 14
- ConvertKit 14
- Brevo 11
- HubSpot 11
- Omnisend 10
- Sendinblue 9
- Amazon SES 8
Brevo at 11 and Sendinblue at 9 are the same company under its old and new names. More on that in the limitations below, because it matters.
Website builder for marketing sites
This is where the models parted company. Webflow scored 38 overall with a 63% mention rate: 15 on ChatGPT, 60 on Claude, 40 on Gemini. Mention rates ran 30, 90 and 70 in the same order.
Brands named across the 30 answers:
- Wix 27
- Squarespace 26
- Shopify 16
- WordPress 12
- Framer 10
- Weebly 10
- WordPress.com 8
- Unbounce 7
- Elementor 6
- Carrd 6
Password manager for teams
1Password scored 68 overall with an 80% mention rate, and the models agreed more closely here than anywhere else: 60 on ChatGPT, 70 on Claude, 75 on Gemini.
Brands named across the 30 answers:
- Bitwarden 16
- LastPass 12
- Passbolt 11
- Dashlane 10
- Dashlane Business 9
- Keeper Security 8
- Keeper 7
- Vaultwarden 7
- Zoho Vault 6
- LastPass Business 6
Note the ceiling. The most named competitor here reached 16, while the most named competitor in each of the other four categories sat at 26 or 27. The answers spread themselves across more names rather than converging on two or three.
This was also the only category where any answers came back with cited domains attached: bitwarden.com, github.com, passbolt.com, teampass.net and lesspass.com, one citation each. The other four categories returned no cited domains in this run. That tells you the answers arrived without a source list, not that no source existed behind them.
What the disagreement means in practice
Compare 1Password and Mailchimp. Mailchimp has the higher mention rate at 93% against 80%, and the lower score at 62 against 68. Counting mentions would rank them the wrong way round. Where a brand lands inside the sentence is doing real work.
The wider point is the spread between models. Webflow at 15 and 60. HubSpot at 60 and 90. Asana, by contrast, sat at 60, 55 and 60, and Mailchimp at 65, 60 and 60. Some categories have a settled consensus that every model repeats, and some do not, and you cannot tell which kind you are in without asking more than one model.
So the practical conclusion is unglamorous. Checking one assistant tells you very little about your position. If you have run your brand through ChatGPT once and felt reassured, you have one reading from one model on one day. Webflow would have drawn the opposite conclusion from two equally valid checks.
Limitations of this run
Three models, not four. Perplexity returned nothing because it was not configured. Perplexity retrieves live on nearly every answer, so its absence likely removes the most web dependent view in the set.
One run, one day, one language. All questions were asked in English on 17 September 2026. Model answers vary between runs. Nothing here supports a claim about change over time, and we are not making one.
API answers are not consumer app answers. They approximate them closely enough to be useful and they are not identical.
Brand name variants were counted separately. The extraction treated "Dashlane" and "Dashlane Business" as two brands, likewise "Keeper Security" and "Keeper", "LastPass" and "LastPass Business", and "WordPress" and "WordPress.com". Worse, "Sendinblue" and "Brevo" are the same company before and after a rename. Consolidating those would raise several counts and change the order of the password manager list in particular. We are reporting the numbers as measured rather than quietly merging them, but you should read those lists knowing the totals are split.
Subject brands are excluded from their own lists. Asana is absent from the project management list, HubSpot from the CRM list, and so on. HubSpot does appear in the email marketing list at 11, because there it is a competitor rather than the subject.
Checking your own category
The method costs you an afternoon and no money. Write ten questions a buyer in your category would genuinely type, with no brand names in them. Ask each one on more than one assistant. Record which brands get named, in what order, and which sources the answers lean on. Then compare the models against each other rather than averaging them, because the disagreement is the most useful thing in the data.
Spegla runs that loop for you. The free check needs no account, and daily tracking is on the paid plan. Either way, the finding above is the one worth keeping: one assistant is one opinion, and your category may be far less settled than a single check suggests.
Track this every day
Check my AI visibility