We Excluded 441 Responses From Our AI Companion Survey. Here's Why.
We paid for 2,000 responses and got 2,150. One in five of them answered the same way, and the dramatic numbers went with them.
I expected the hardest part of running our first large survey to be interpreting the results.
It turned out to be deciding which responses I trusted enough to include.
We paid for a target of 2,000 completed responses through SurveyMonkey Audience. SurveyMonkey reported 2,149 completions; the individual-response export contained 2,150. During data cleaning I found a group of 441 that I eventually excluded from every result we published, leaving 1,709 in the analyzed dataset. That is 20.5% of our survey of 2,150 U.S. adults, and this post is the full account of why.
The sample: 2,150 completed, 441 excluded, 1,709 analyzed
SurveyMonkey reported 2,149 completions; the individual-response export contained 2,150. The 441 were excluded from analysis, not deleted.
The Numbers Looked Almost Too Interesting
What first caught my attention was not any individual response. It was the combined statistics, which looked unusually extreme for a first pass through a general-population sample.
Here is what the uncleaned dataset was telling us, next to what we eventually published.
| Finding | All 2,150 | Analyzed 1,709 |
|---|---|---|
| Currently use an AI girlfriend or boyfriend app | 30.3% (652) | 12.3% (211) |
| Would definitely consider using one | 29.4% (633) | 11.2% (192) |
| Could genuinely fall in love with an AI | 27.1% (583) | 8.3% (142) |
| Any romantic or sexual AI interaction by a partner is cheating | 47.0% (1,011) | 33.4% (570) |
| Could see themselves preferring an AI over a human relationship | 29.1% (625) | 10.8% (184) |
| Would choose dating apps if single | 52.6% (1,130) | 52.2% (892) |
Thirty percent of American adults currently using an AI girlfriend or boyfriend app. Twenty-seven percent able to fall in love with one.
Those are the kind of numbers that travel, and they are the numbers I would have been publishing if I had stopped at the top-line tables.
I want to be honest about my reaction, because it was not suspicion at first. It was excitement. The findings were more dramatic than anything I had seen from other sources, and the sample was large. It took a second look at the individual rows to turn that into unease.
Then I Noticed 441 People Answering the Same Way
Sorting the export by answer combinations, one profile kept repeating. 441 respondents had selected the first-listed answer choice on all five of the survey's attitude questions.
That produced a respondent who, at the same time:
- currently used an AI girlfriend or boyfriend app;
- would definitely consider using one in the future;
- thought they could genuinely fall in love with an AI;
- could see themselves preferring an AI companion over a human relationship;
- and considered any romantic or sexual AI interaction by a partner to be cheating.
It is not impossible for one person to hold all five positions.
I want to be careful here. Someone could genuinely believe every one of those things. If I had found ten such respondents, I would have shrugged.
The issue was that 441 people produced the identical first-option pattern, and that pattern happened to line up with the answer order on the screen. One disclosure matters a great deal for interpreting that: our answer choices were displayed in a fixed order and were not randomized. That was a design decision I did not think hard enough about before fieldwork, and it is the reason a first-position streak is legible in the data at all.
Look back at the table above and you can see the 441 directly. On every one of those five questions, the difference between the uncleaned count and the analyzed count is exactly 441. The dramatic version of each statistic was the ordinary version plus this one group.
The Answer Pattern Wasn't the Only Thing That Stood Out
If the repeated pattern had been the only unusual thing about these 441 responses, I am not sure I would have excluded them. It was not.
The answer pattern was the criterion. Four other signals travelled with it.
Base: the 441 excluded responses. Comparisons are given only where the draft reports one for the rest of the sample.
None of these is proof of anything on its own, and I do not want to pretend otherwise.
A legitimate respondent can use an Android phone. A legitimate respondent can finish a short survey quickly. A legitimate respondent can mistype their age or have an outdated panel profile. And a legitimate respondent can select the first answer five times.
What raised the quality concern was the combination: 441 responses sharing the same answer pattern, and also sharing these other characteristics, at rates the rest of the sample did not come close to.
At First, I Thought About Keeping Them
My first reaction was not "these are bad responses, remove them." I considered leaving them in.
The questions were simple. A respondent had to read a question and pick the answer closest to their view; nobody needed several minutes to solve anything. And the survey itself was short. SurveyMonkey reported a median completion time of one minute eight seconds across all completions; the export produced a median of roughly 70 seconds. Against that baseline, 39 seconds is fast, but it is not absurd.
So I want to be precise about what the exclusion criterion was not. Thirty-nine seconds was not the criterion. Android was not the criterion. Holding unusual opinions was not the criterion. I did not remove anyone because their answers seemed unbelievable, and I had no way of knowing who these respondents were or why they answered as they did.
A weakness, stated plainly. This rule was developed after collection, during data cleaning, not defined in advance. I return to it below.
Removing Them Made Our Findings Less Dramatic
The three biggest changes: current use of an AI girlfriend or boyfriend app went from 30.3% to 12.3%, from 652 of 2,150 to 211 of 1,709. Could genuinely fall in love with an AI went from 27.1% to 8.3%, from 583 to 142. Could see themselves preferring an AI over a human relationship went from 29.1% to 10.8%, from 625 to 184.
Before and after the screen: the findings that depended on first-option answers collapsed, the one that did not barely moved
Each pair shows the figure on all 2,150 completed responses and on the 1,709 analyzed. Single-choice questions; "cheating" here is the strictest answer only.
Every one of those would have been an extremely attractive headline.
Each became a much more modest number, and each is now one I can defend to anyone who asks how it was produced.
Now the counterexample, which I think is the most important row in the table.
Asked what they would rather use if single, 52.6% of the full sample chose dating apps. In the analyzed sample, 52.2% did. Almost nothing happened. In the final dataset, 52.2% chose dating apps, 26.8% were not sure, 13.2% chose both, and 7.8% chose an AI girlfriend or boyfriend.
That matters because it shows the cleaning did not simply push every number downward. Findings that depended on the first-option answers moved enormously. Findings that did not, including the dating-app question, where "dating apps" happened to be the first option too but was the majority answer regardless, barely moved. The cheating question sits in between, and it is the one whose interpretation changed most in character rather than size; I worked through that in our analysis of whether AI relationships count as cheating.
We Didn't Erase the 441 Responses
"Excluded" means excluded from analysis. The 441 records were not deleted.
The master report states it plainly: 441 responses were excluded, 1,709 remained, the excluded responses were preserved in full and are available for inspection alongside the analyzed dataset, and no other exclusions were made. I also reported the 441 to SurveyMonkey for their own review, which as of publication is still pending.
- 2,150completed responses
- 441excluded from analysis
- 1,709final analyzed sample
- 20.5%share excluded
- 0records deleted
- 0other exclusions made
Other Researchers Pointed Out What I Could Do Differently Next Time
After discussing the survey on Indie Hackers, people raised points about attention checks and about how hard it is to control respondent quality in paid online surveys. Nothing there changed the exclusion decision, which I had already made, but it sharpened the list of what I would change.
- Include at least one attention-check question.
- Randomize answer order where the question allows it, so a first-position streak cannot form in the first place.
- Define exclusion rules before fieldwork begins, rather than discovering them afterward.
- Inspect the completion-time distribution before looking at any headline result.
- Compare survey answers against whatever panel metadata is available.
- Check for repeated response patterns before analyzing anything else.
And one lesson from my own assumption going in: a large sample does not solve a response-quality problem. It scales it.
Margin of Error Doesn't Fix This Problem
SurveyMonkey reported a margin of error of plus or minus 2.157% for the full sample under its own assumptions, and our report cautions that the figure does not apply to smaller subgroups.
But margin of error is not protection against respondents clicking through a survey without engaging with it. It describes sampling uncertainty under particular assumptions about who answered and how. If 441 people answer the same way for reasons unrelated to their opinions, the margin of error does not shrink to accommodate them; it simply does not know they are there.
Sample size and response quality are different problems. I knew that in the abstract. This was the survey that made it concrete.
What I'd do differently next time
The original survey did not include a dedicated attention-check question. That is the first thing I would change, and it is worth saying without hedging.
There will always be respondents whose main motivation is finishing a paid survey rather than helping with the research, and we cannot necessarily identify every one of them afterward. The goal is not to pretend the remaining 1,709 responses are perfect. It is to build stronger safeguards before the next survey opens.
The remaining limitations are real. The responses are self-reported and unverified. The overall median completion time was about 70 seconds. And answer patterns still differed between Android and iOS/desktop respondents after screening, which is why the report checks each highlighted finding for consistency across device types and flags the ones that do not hold. Removing the 441 did not eliminate every data-quality question. It removed the one I could see clearly.
Why the Rest of Our Research Uses 1,709
Every figure we have published from this survey uses the 1,709 analyzed responses, and every article in the series says so. The overview of the whole study is The State of AI Romance in 2026: What 2,150 U.S. Adults Told Us. The pattern that survived the cleaning best, and the one I find most interesting, is in our analysis of the 422 people who had used an AI companion: they took AI relationships more seriously than people who never had, on every attitude we measured, and that held on both device types.
This page is the canonical explanation of the 1,709-person sample. Future articles from the survey will link back here rather than re-explaining it. The headline tables themselves, all on the 1,709, are on the data page.
The Most Interesting Number Isn't Always the One to Publish
I would have preferred to keep all 2,150 responses. It would have given us a larger sample, and some of the results would have made much better headlines. Thirty percent of Americans with an AI girlfriend is a story. Twelve percent of a screened American sample is a footnote by comparison.
But that is exactly why the 441 responses are worth talking about. The results I was most excited to publish were the ones I had the strongest reason to question, and the two facts were not a coincidence. The same group of responses that made the survey dramatic was the group that made it unreliable.
The same rule runs the reviews: the number I can defend beats the number that travels. Every app on my roster was paid for and tested the same way.
See the AI girlfriend apps I've tested
