When people shop online or interact with chatbots, their behavior generates training data for artificial intelligence (AI). To regulate the use of such data, privacy regulations increasingly require organizations to disclose AI training and obtain explicit consent for this use. However, such disclosures may themselves alter user behavior and affect the collected data. In particular, three distinct mechanisms within the consent process could drive these changes. First, people who voluntarily consent to AI training may differ systematically from those who do not, potentially introducing selection bias into training data. Second, simply being aware that AI training is occurring may trigger behavioral changes. Third, behavioral changes may depend on users actively reading details about AI training, rather than merely knowing that such training is occurring. Across two experiments, we examined whether behavioral changes in training data arise from consent, mere awareness of AI training, or engagement with information about how their data would be used. To test these hypotheses, participants played the ultimatum game, deciding whether to accept monetary proposals. Some participants were informed about AI training, while others were not. In Experiment 1, we manipulated whether participants could opt into AI training and whether they had the option to read information about how their data would train AI for future participants. Although most participants opted into training, most chose not to read the information. Participants who opted in without reading behaved similarly to those unaware of AI training, while those who read rejected more unfair offers. In Experiment 2, we tested whether the amount of information affected engagement and found that, compared to brief disclosures, detailed disclosures produced similar reading times and behavioral changes. These findings show that informing users about AI training alters training data only when users engage with the information, regardless of how much detail is provided. This creates heterogeneous data in which informed and uninformed users behave systematically differently. This work highlights that disclosure practices intended to protect users may directly shape the data used to train AI, underscoring the need to carefully document and account for these behavioral shifts.
