Market researchers should never put personally identifiable respondent data (names, locations, phone numbers, email addresses, IP addresses) into a public AI tool unless the tool has a contractual data processing agreement and you've confirmed it's active. If you're using a free or pro account on ChatGPT, Claude, or similar tools, assume everything you type is visible to the provider and could surface elsewhere.

This question comes up in almost every workshop I teach, especially from researchers in regulated industries. The answer changes as tools update their policies and enterprise features, but the principles stay consistent.

What counts as personally identifiable in a research context?

The obvious things: names, email addresses, phone numbers, physical locations, IP addresses. Strip these before anything goes near an AI tool.

But also watch for verbatim responses that include identifying details. A quote like "I've been a nurse at St. Mary's in Akron for 12 years" is potentially identifying, especially in smaller sample sizes. If you're uploading open-ended responses, scan for these kinds of details first.

Demographic combinations can also identify people indirectly. A 62-year-old female VP of marketing at a specific company in a specific city might be unique in your dataset. That's identifying information, even without a name attached.

What about business-sensitive information?

This is where freelancers and independent consultants need to be especially careful.

If you're a solo researcher using a free or pro account on ChatGPT or Claude, don't put unpublished client findings, proprietary methodologies, or sensitive business strategy into those tools. Those accounts don't have data processing agreements. Your inputs could be used for training. That's a risk to your client relationships and potentially a breach of your contract.

Enterprise accounts are different. They typically include agreements that prevent your data from being used for model training. But you have to be on the enterprise plan and confirm the settings are correct. Being on an enterprise plan with training data collection still enabled happens more often than you'd think.

Which tools are actually safe?

Enterprise versions of ChatGPT, Claude, and Google's Gemini offer data processing agreements. Research-specific tools like Quillit and Flowres have built-in protections designed for sensitive respondent data. Always look for a vendor's Privacy Policy, read it thoroughly, and look through their Trust Center documents if they have one.

The rule of thumb: if you haven't read the Privacy Policy and Terms and Conditions yourself, don't assume you're protected.

What about researchers in regulated industries?

Pharma, healthcare, financial services, and legal research teams face additional requirements. GDPR, HIPAA, and industry-specific regulations add layers of compliance that general AI tools may not address.

If you're in a regulated industry, your legal or compliance team almost certainly has guidance on AI tools. Find it. Read it. If it doesn't exist yet, ask for it. Researchers in these industries have told me that the biggest risk isn't using AI. It's using AI without checking what's already been approved or prohibited internally.

The ICC/ESOMAR International Code on market, opinion, and social research and data analytics is the definitive external reference here. It covers data minimization, privacy, transparency, and human oversight in the context of AI. The MRS AI Guidance is another resource worth bookmarking, especially for UK and European practitioners.

What's the practical checklist before uploading anything?

Four questions to ask yourself:

  1. Does this contain PII (name, email, phone, location, IP address)? If yes, strip it first or don't upload.
  2. Does this contain unpublished client business information? If yes and you're on a free or pro account, don't upload.
  3. Does this contain information that's sensitive under industry regulations? If yes, check with compliance before uploading.
  4. Do I know whether this tool has a data processing agreement that covers my use case? If no, find out before uploading.

Go from reading to doing

Browse Classes → Free Newsletter