ChatGPT
OpenAI says ChatGPT conversations may be used to improve models unless the user turns off model improvement in Data Controls.
The AI Data Use Index reads the public record: what companies say about training reuse, opt-outs, retention, human review, and what ordinary people can actually tell from the disclosure in front of them.
This is a transparency index, not legal advice. It records what is publicly stated today and where the public answer is still incomplete.
Indexed products
OpenAI says ChatGPT conversations may be used to improve models unless the user turns off model improvement in Data Controls.
Google says that when Gemini Apps Activity is on, Gemini data is used to improve Google AI with help from human reviewers.
Microsoft says files, communications, prompts, responses, and Microsoft Graph data used with Microsoft 365 Copilot are not used to train foundation models.
Meta says it uses public information from adult accounts and interactions with AI at Meta features to develop and improve generative AI models, with a right to object.
X says public X data plus interactions, inputs, and results with Grok may be shared with xAI to train and fine-tune Grok and other generative AI models.
Perplexity says AI data retention is enabled by default for Free, Pro, and Max users, and that users can turn it off in account settings.
Anthropic describes a model-improvement setting for consumer chats and says incognito chats are not used to improve Claude.
GitHub says that by default it, its affiliates, and third parties do not use individual-subscriber Copilot data, including prompts, suggestions, and code snippets, for AI model training.
Cursor says code is not used for training when Privacy Mode is enabled; when Privacy Mode is off, Cursor may use stored codebase data, prompts, editor actions, and snippets to improve AI features and train models.
Notion says it does not use Customer Data, including user content under personal terms, or permit others to use it to train the machine-learning models used to provide Notion AI.
Slack says Customer Data is not used to train generative AI models unless the customer gives affirmative opt-in consent.
Zoom says it does not use customer audio, video, chat, screen sharing, attachments, or other communications-like customer content to train Zoom or third-party AI models.
Grammarly says individual-account Product Improvement and Training is on by default, while enterprise and certain sales-led accounts have it off by default.
Canva says privacy settings control whether general usage data and User Content can improve AI-powered features, and that Canva Education User Content is not used for AI training.
Adobe says Firefly does not train on customer data and that Firefly uses commercially safe datasets such as licensed content and public-domain material.
OpenAI says inputs and outputs from ChatGPT Business and Enterprise are not used to train or improve its models by default.
Anthropic says it does not use inputs or outputs from Claude for Work and other commercial products to train its models by default.
Google says content used with qualifying Workspace editions is not human reviewed or used for generative-AI model training outside the customer domain without permission.
Microsoft says prompts, outputs, embeddings, and training data for Foundry direct models are not used to train generative-AI foundation models without customer permission or instruction.
Google Cloud says it will not use customer data to train or fine-tune managed AI models on Vertex AI without prior permission or instruction.
AWS says Amazon Bedrock does not use customer content to train models or share that content with model providers to improve base models.
AWS says Amazon Q Business does not use customer data for service improvement or to improve its underlying large language models.
AWS says some Amazon Q Developer Free tier content may be used for service improvement, including model training; Q Developer Pro content is not used for service improvement.
GitLab documents feature-specific model providers and retention paths. Optional Agent Platform usage collection is stated to support service improvement and debugging, not AI-model training.
JetBrains says it and its AI subcontractors will not use customer inputs, data, outputs, or suggestions to train generative models unless the customer expressly agrees.
Google's Gemini API terms distinguish Paid Services from unpaid use. For Paid Services, Google says prompts and responses are not used to improve its products.
Microsoft says Copilot Studio customer prompts, grounding data, and responses are not used to train or improve Azure OpenAI foundation models.
OpenAI says data sent to the API is not used to train or improve OpenAI models unless the organisation explicitly opts in to share data.
Microsoft says consumer Copilot, Bing, and MSN interaction data may be used to train generative AI models unless the signed-in user is excluded or has opted out.
Apple says it does not use users' private personal data or user interactions when training the foundation models that power Apple Intelligence.
Salesforce says Agentforce runs through the Einstein Trust Layer with zero data retention agreements for third-party LLMs, and Salesforce does not currently use customer data to train generative AI models without an express opt-in.
Atlassian says customer inputs and outputs are not retained or used by third-party LLM providers for training, while Atlassian may use contributed metadata to fine-tune Atlassian-hosted open-source models subject to data contribution settings.
Box says neither Box nor its model providers train on customer data sent through Box AI, and Box will only train on customer content with explicit customer approval.
Snowflake says Usage and Customer Data for Snowflake AI Features, including inputs and outputs, are not used to train, re-train, or fine-tune models made available to other customers and remain inside the Snowflake security boundary.
ServiceNow says Now Assist inference data is processed transiently in regional compute hubs and is not commingled with other customers, while optional data sharing for Now LLM improvement can be opted out in admin settings.
Databricks says it does not train generative foundation models on data, prompts, or responses submitted to Databricks AI assistive features, and partner-powered features use zero-retention endpoints.
IBM says customer data on watsonx.ai is accessible only to the customer account, is used to train only that customer's models, and is never accessible or used by IBM or any other organisation.
Mistral says API data is not used for model training under its current privacy controls, while free Vibe conversations can be used for improvement unless opted out; Pro, Team, and Enterprise Vibe conversations are not used for training by default.
Cohere says enterprise SaaS customers can opt out of prompts and generations being used to train Cohere models in dashboard Data Controls; private and third-party cloud deployments do not send prompts or generations to Cohere.
Tabnine says it never retains customer code beyond ephemeral inference and does not train its models on customer code, regardless of which Tabnine model is used.
Figma says customer content may be used to train Figma AI models when the Content Training admin setting is on; third-party model providers are not permitted to train their own models on customer Figma content.
HubSpot says the AI service providers used for Breeze and other HubSpot AI features are not permitted to use customer data for model training, and HubSpot enforces zero data retention with those providers wherever possible.
Intercom says third-party AI providers used for Fin are contractually restricted from using customer data to train or improve their models, and Inputs and Outputs are treated as customer data.
Windsurf's commercial terms say customer data is not used to train or optimize AI models without prior written consent, and real-time customer data is deleted after output generation subject to listed exceptions.
DeepSeek says it may use inputs and outputs, after encryption and de-identification, to develop or improve services and underlying technologies unless the user turns off model improvement.
Midjourney's privacy policy covers personal data collected through its services and through the process of training Midjourney machine-learning algorithms, but the reviewed text does not expressly state that every user prompt or upload is used to train models.
xAI says it never trains on API inputs or outputs without explicit permission. By default it stores requests and responses encrypted at rest for 30 days for audit purposes, with team-level Zero Data Retention available.
Replicate's terms license customer data as needed to provide outputs and train customer derivative models when customers create training workloads. API predictions remove inputs, outputs, files, and logs after one hour by default.
Hugging Face says it does not store user data for training purposes and does not store request bodies or responses when routing through Inference Providers.
Groq says it may not use inputs or outputs to train or fine-tune models unless the customer explicitly grants permission, and GroqCloud inference customer data is not retained by default subject to reliability, abuse, and persistence-feature exceptions.
ElevenLabs says certain submitted data may be used to improve audio models unless the account disables model improvement; Enterprise customer data is not used for that purpose by default.
Otter says it can train proprietary AI on de-identified recordings and transcripts, while an account feedback-and-training setting controls whether conversations are shared for training, product improvement, and human review.
Dropbox says it will not build generative AI models using customer content without consent, while Dash may use product-usage and interaction data to improve and fine-tune Dash subject to opt-out controls.
Zendesk says its proprietary non-generative ML may use sanitized customer-service data, while generative AI features use third-party models that are not trained on Zendesk service data.
Sourcegraph says it and its partner LLMs do not use customer code to train models, and partner LLMs operate under zero-retention terms for inputs, outputs, and code context subject to a narrow abuse-prevention exception.
Replit's commercial terms say private customer content and outputs are not used to train or improve AI features, while publicly published content may be used, and AI-integration endpoint privacy differs by plan.
Perplexity's Enterprise Terms prohibit Perplexity and authorized third parties from using Customer Content to train, retrain, fine-tune, or otherwise improve generative AI models.
Fireworks says it does not use prompts, training data, or API inputs to train or improve its AI models without explicit opt-in, and inference has zero data retention by default.
Anthropic says inputs and outputs from commercial products, including the Anthropic API, are not used to train its models by default.
Oracle says OCI Generative AI does not retain customer inference inputs or outputs, does not share them with third-party model providers, and does not use fine-tuning data to improve general OCI Generative AI use cases.
Cisco says Webex Assistant does not retain customer data for machine-learning training, and that Microsoft does not use Cisco customer content to improve Azure OpenAI models.
Stability AI says Platform API customers can choose whether inputs and outputs are used to train models by setting the account preference 'Training: Improve the Model for Everyone' to No.
Asana says neither Asana nor its AI partners use customer data to train AI models, and partners must delete customer data after completion of the individual AI query.
If the answer is muddy
Use the request builder to turn uncertainty into a competent email, public post, or procurement question set instead of leaving people to guess.
Open request builder →Leave an email if you want notice when email alerts ship. Until then, the monthly editions below are the public record.
Track what changed, read the weekly brief, and follow the public evidence record - the operator loop for production AI risk.