Production AI Institute · Public record
Public index · reviewed July 2026

Does this company use my data to train AI?

The AI Data Use Index reads the public record: what companies say about training reuse, opt-outs, retention, human review, and what ordinary people can actually tell from the disclosure in front of them.

63 products reviewed5 state training use20 user-controlled

This is a transparency index, not legal advice. It records what is publicly stated today and where the public answer is still incomplete.

Indexed products

What the public can tell from the public record.

OpenAI

ChatGPT

User-controlled

OpenAI says ChatGPT conversations may be used to improve models unless the user turns off model improvement in Data Controls.

Consumer serviceReviewed 2026-06-15
Google

Gemini Apps

Uses data when activity is on

Google says that when Gemini Apps Activity is on, Gemini data is used to improve Google AI with help from human reviewers.

Consumer serviceReviewed 2026-06-15
Microsoft

Microsoft 365 Copilot

Not used to train foundation models

Microsoft says files, communications, prompts, responses, and Microsoft Graph data used with Microsoft 365 Copilot are not used to train foundation models.

Commercial serviceReviewed 2026-06-15
Meta

Meta AI

Uses eligible public data

Meta says it uses public information from adult accounts and interactions with AI at Meta features to develop and improve generative AI models, with a right to object.

Consumer serviceReviewed 2026-06-15
X / xAI

Grok on X

Uses public data and interactions

X says public X data plus interactions, inputs, and results with Grok may be shared with xAI to train and fine-tune Grok and other generative AI models.

Consumer serviceReviewed 2026-06-15
Perplexity

Perplexity

Training on by default for consumers

Perplexity says AI data retention is enabled by default for Free, Pro, and Max users, and that users can turn it off in account settings.

Consumer serviceReviewed 2026-06-15
Anthropic

Claude

User-controlled

Anthropic describes a model-improvement setting for consumer chats and says incognito chats are not used to improve Claude.

Consumer serviceReviewed 2026-06-15
GitHub

GitHub Copilot

No training by default

GitHub says that by default it, its affiliates, and third parties do not use individual-subscriber Copilot data, including prompts, suggestions, and code snippets, for AI model training.

Developer serviceReviewed 2026-06-15
Anysphere

Cursor

Depends on Privacy Mode

Cursor says code is not used for training when Privacy Mode is enabled; when Privacy Mode is off, Cursor may use stored codebase data, prompts, editor actions, and snippets to improve AI features and train models.

Developer serviceReviewed 2026-06-15
Notion

Notion AI

No customer-data training

Notion says it does not use Customer Data, including user content under personal terms, or permit others to use it to train the machine-learning models used to provide Notion AI.

Commercial serviceReviewed 2026-06-15
Slack

AI in Slack

No generative-AI training

Slack says Customer Data is not used to train generative AI models unless the customer gives affirmative opt-in consent.

Commercial serviceReviewed 2026-06-15
Zoom

Zoom AI Companion

No customer-content training

Zoom says it does not use customer audio, video, chat, screen sharing, attachments, or other communications-like customer content to train Zoom or third-party AI models.

Commercial serviceReviewed 2026-06-15
Grammarly

Grammarly

Training control varies by account

Grammarly says individual-account Product Improvement and Training is on by default, while enterprise and certain sales-led accounts have it off by default.

Consumer and commercial serviceReviewed 2026-06-15
Canva

Canva

Depends on privacy settings

Canva says privacy settings control whether general usage data and User Content can improve AI-powered features, and that Canva Education User Content is not used for AI training.

Consumer and commercial serviceReviewed 2026-06-15
Adobe

Adobe Firefly

No customer-content training

Adobe says Firefly does not train on customer data and that Firefly uses commercially safe datasets such as licensed content and public-domain material.

Creative serviceReviewed 2026-06-15
OpenAI

ChatGPT Business & Enterprise

No training by default

OpenAI says inputs and outputs from ChatGPT Business and Enterprise are not used to train or improve its models by default.

Commercial serviceReviewed 2026-07-17
Anthropic

Claude for Work

No training by default

Anthropic says it does not use inputs or outputs from Claude for Work and other commercial products to train its models by default.

Commercial serviceReviewed 2026-07-17
Google

Gemini for Google Workspace

No outside-domain training without permission

Google says content used with qualifying Workspace editions is not human reviewed or used for generative-AI model training outside the customer domain without permission.

Commercial serviceReviewed 2026-07-17
Microsoft Azure

Microsoft Foundry direct models

No foundation-model training without permission

Microsoft says prompts, outputs, embeddings, and training data for Foundry direct models are not used to train generative-AI foundation models without customer permission or instruction.

AI development platformReviewed 2026-07-17
Google Cloud

Vertex AI managed models

No training without permission

Google Cloud says it will not use customer data to train or fine-tune managed AI models on Vertex AI without prior permission or instruction.

AI development platformReviewed 2026-07-17
Amazon Web Services

Amazon Bedrock

No base-model training

AWS says Amazon Bedrock does not use customer content to train models or share that content with model providers to improve base models.

AI development platformReviewed 2026-07-17
Amazon Web Services

Amazon Q Business

No service-improvement training

AWS says Amazon Q Business does not use customer data for service improvement or to improve its underlying large language models.

Commercial serviceReviewed 2026-07-17
Amazon Web Services

Amazon Q Developer

Depends on tier and opt-out

AWS says some Amazon Q Developer Free tier content may be used for service improvement, including model training; Q Developer Pro content is not used for service improvement.

Developer serviceReviewed 2026-07-17
GitLab

GitLab Duo

Feature and subprocessor handling varies

GitLab documents feature-specific model providers and retention paths. Optional Agent Platform usage collection is stated to support service improvement and debugging, not AI-model training.

Developer serviceReviewed 2026-07-17
JetBrains

JetBrains AI Assistant & Junie

No training without express agreement

JetBrains says it and its AI subcontractors will not use customer inputs, data, outputs, or suggestions to train generative models unless the customer expressly agrees.

Developer serviceReviewed 2026-07-17
Google

Gemini API & Google AI Studio

Depends on paid-service status

Google's Gemini API terms distinguish Paid Services from unpaid use. For Paid Services, Google says prompts and responses are not used to improve its products.

Developer serviceReviewed 2026-07-17
Microsoft

Microsoft Copilot Studio

No foundation-model training

Microsoft says Copilot Studio customer prompts, grounding data, and responses are not used to train or improve Azure OpenAI foundation models.

AI workflow platformReviewed 2026-07-17
OpenAI

OpenAI API Platform

No training by default

OpenAI says data sent to the API is not used to train or improve OpenAI models unless the organisation explicitly opts in to share data.

Developer serviceReviewed 2026-07-17
Microsoft

Microsoft Copilot (consumer)

User-controlled

Microsoft says consumer Copilot, Bing, and MSN interaction data may be used to train generative AI models unless the signed-in user is excluded or has opted out.

Consumer serviceReviewed 2026-07-17
Apple

Apple Intelligence

No private-user training

Apple says it does not use users' private personal data or user interactions when training the foundation models that power Apple Intelligence.

Consumer serviceReviewed 2026-07-17
Salesforce

Salesforce Agentforce

No generative-model training without opt-in

Salesforce says Agentforce runs through the Einstein Trust Layer with zero data retention agreements for third-party LLMs, and Salesforce does not currently use customer data to train generative AI models without an express opt-in.

Commercial serviceReviewed 2026-07-17
Atlassian

Atlassian Rovo & Intelligence

Depends on contribution settings

Atlassian says customer inputs and outputs are not retained or used by third-party LLM providers for training, while Atlassian may use contributed metadata to fine-tune Atlassian-hosted open-source models subject to data contribution settings.

Commercial serviceReviewed 2026-07-17
Box

Box AI

No training without approval

Box says neither Box nor its model providers train on customer data sent through Box AI, and Box will only train on customer content with explicit customer approval.

Commercial serviceReviewed 2026-07-17
Snowflake

Snowflake Cortex AI

No training for other customers

Snowflake says Usage and Customer Data for Snowflake AI Features, including inputs and outputs, are not used to train, re-train, or fine-tune models made available to other customers and remain inside the Snowflake security boundary.

AI development platformReviewed 2026-07-17
ServiceNow

ServiceNow Now Assist

Depends on data-sharing settings

ServiceNow says Now Assist inference data is processed transiently in regional compute hubs and is not commingled with other customers, while optional data sharing for Now LLM improvement can be opted out in admin settings.

Commercial serviceReviewed 2026-07-17
Databricks

Databricks AI & Foundation Model APIs

No foundation-model training

Databricks says it does not train generative foundation models on data, prompts, or responses submitted to Databricks AI assistive features, and partner-powered features use zero-retention endpoints.

AI development platformReviewed 2026-07-17
IBM

IBM watsonx.ai

Customer data stays private to the account

IBM says customer data on watsonx.ai is accessible only to the customer account, is used to train only that customer's models, and is never accessible or used by IBM or any other organisation.

AI development platformReviewed 2026-07-17
Mistral AI

Mistral AI API & Vibe

Depends on plan and controls

Mistral says API data is not used for model training under its current privacy controls, while free Vibe conversations can be used for improvement unless opted out; Pro, Team, and Enterprise Vibe conversations are not used for training by default.

Developer and workplace serviceReviewed 2026-07-17
Cohere

Cohere Platform

User-controlled

Cohere says enterprise SaaS customers can opt out of prompts and generations being used to train Cohere models in dashboard Data Controls; private and third-party cloud deployments do not send prompts or generations to Cohere.

Developer serviceReviewed 2026-07-17
Tabnine

Tabnine

No train, no retain

Tabnine says it never retains customer code beyond ephemeral inference and does not train its models on customer code, regardless of which Tabnine model is used.

Developer serviceReviewed 2026-07-17
Figma

Figma AI

Depends on Content Training setting

Figma says customer content may be used to train Figma AI models when the Content Training admin setting is on; third-party model providers are not permitted to train their own models on customer Figma content.

Commercial serviceReviewed 2026-07-17
HubSpot

HubSpot Breeze

No third-party model training

HubSpot says the AI service providers used for Breeze and other HubSpot AI features are not permitted to use customer data for model training, and HubSpot enforces zero data retention with those providers wherever possible.

Commercial serviceReviewed 2026-07-17
Intercom

Intercom Fin AI Agent

No third-party model training

Intercom says third-party AI providers used for Fin are contractually restricted from using customer data to train or improve their models, and Inputs and Outputs are treated as customer data.

Commercial serviceReviewed 2026-07-17
Cognition

Windsurf

No training without consent

Windsurf's commercial terms say customer data is not used to train or optimize AI models without prior written consent, and real-time customer data is deleted after output generation subject to listed exceptions.

Developer serviceReviewed 2026-07-17
DeepSeek

DeepSeek Chat

User-controlled

DeepSeek says it may use inputs and outputs, after encryption and de-identification, to develop or improve services and underlying technologies unless the user turns off model improvement.

Consumer serviceReviewed 2026-07-17
Midjourney

Midjourney

Training scope not explicit

Midjourney's privacy policy covers personal data collected through its services and through the process of training Midjourney machine-learning algorithms, but the reviewed text does not expressly state that every user prompt or upload is used to train models.

Consumer serviceReviewed 2026-07-17
xAI

xAI API

No training by default

xAI says it never trains on API inputs or outputs without explicit permission. By default it stores requests and responses encrypted at rest for 30 days for audit purposes, with team-level Zero Data Retention available.

Developer serviceReviewed 2026-07-17
Replicate

Replicate

Training when you initiate it

Replicate's terms license customer data as needed to provide outputs and train customer derivative models when customers create training workloads. API predictions remove inputs, outputs, files, and logs after one hour by default.

AI development platformReviewed 2026-07-17
Hugging Face

Hugging Face Inference Providers

No training by default

Hugging Face says it does not store user data for training purposes and does not store request bodies or responses when routing through Inference Providers.

AI development platformReviewed 2026-07-17
Groq

GroqCloud API

No training by default

Groq says it may not use inputs or outputs to train or fine-tune models unless the customer explicitly grants permission, and GroqCloud inference customer data is not retained by default subject to reliability, abuse, and persistence-feature exceptions.

Developer serviceReviewed 2026-07-17
ElevenLabs

ElevenLabs

Depends on plan and opt-out

ElevenLabs says certain submitted data may be used to improve audio models unless the account disables model improvement; Enterprise customer data is not used for that purpose by default.

Developer and commercial serviceReviewed 2026-07-17
Otter.ai

Otter.ai

User-controlled

Otter says it can train proprietary AI on de-identified recordings and transcripts, while an account feedback-and-training setting controls whether conversations are shared for training, product improvement, and human review.

Commercial serviceReviewed 2026-07-17
Dropbox

Dropbox Dash

Depends on improvement controls

Dropbox says it will not build generative AI models using customer content without consent, while Dash may use product-usage and interaction data to improve and fine-tune Dash subject to opt-out controls.

Commercial serviceReviewed 2026-07-17
Zendesk

Zendesk AI

Uses sanitized data for proprietary ML

Zendesk says its proprietary non-generative ML may use sanitized customer-service data, while generative AI features use third-party models that are not trained on Zendesk service data.

Commercial serviceReviewed 2026-07-17
Sourcegraph

Sourcegraph Cody

No training on customer code

Sourcegraph says it and its partner LLMs do not use customer code to train models, and partner LLMs operate under zero-retention terms for inputs, outputs, and code context subject to a narrow abuse-prevention exception.

Developer serviceReviewed 2026-07-17
Replit

Replit Agent

Depends on publishing and endpoints

Replit's commercial terms say private customer content and outputs are not used to train or improve AI features, while publicly published content may be used, and AI-integration endpoint privacy differs by plan.

Developer serviceReviewed 2026-07-17
Perplexity

Perplexity Enterprise

No training on enterprise content

Perplexity's Enterprise Terms prohibit Perplexity and authorized third parties from using Customer Content to train, retrain, fine-tune, or otherwise improve generative AI models.

Commercial serviceReviewed 2026-07-17
Fireworks

Fireworks AI

No training by default

Fireworks says it does not use prompts, training data, or API inputs to train or improve its AI models without explicit opt-in, and inference has zero data retention by default.

Developer serviceReviewed 2026-07-17
Anthropic

Anthropic Claude API

No training by default

Anthropic says inputs and outputs from commercial products, including the Anthropic API, are not used to train its models by default.

Developer serviceReviewed 2026-07-17
Oracle

Oracle OCI Generative AI

No inference retention or general-model training

Oracle says OCI Generative AI does not retain customer inference inputs or outputs, does not share them with third-party model providers, and does not use fine-tuning data to improve general OCI Generative AI use cases.

AI development platformReviewed 2026-07-17
Cisco

Cisco Webex AI Assistant

No customer-content model training

Cisco says Webex Assistant does not retain customer data for machine-learning training, and that Microsoft does not use Cisco customer content to improve Azure OpenAI models.

Commercial serviceReviewed 2026-07-17
Stability AI

Stability AI Platform API

User-controlled

Stability AI says Platform API customers can choose whether inputs and outputs are used to train models by setting the account preference 'Training: Improve the Model for Everyone' to No.

Developer serviceReviewed 2026-07-17
Asana

Asana AI

No customer-data AI training

Asana says neither Asana nor its AI partners use customer data to train AI models, and partners must delete customer data after completion of the individual AI query.

Commercial serviceReviewed 2026-07-17

If the answer is muddy

Ask for one clearer public disclosure.

Use the request builder to turn uncertainty into a competent email, public post, or procurement question set instead of leaving people to guess.

Open request builder →
AI Policy Change Watch

Join the change watch list.

Leave an email if you want notice when email alerts ship. Until then, the monthly editions below are the public record.

AI risk intelligence

Track what changed, read the weekly brief, and follow the public evidence record - the operator loop for production AI risk.

Live risk watch →AI risk brief →Save a watch →Public record →