In AI-assisted decision-making, explanations are increasingly personalized to match users’ decision criteria and improve usability and trust. However, aligning explanations with users’ decision criteria can also steer their decisions, especially when manipulations are inaccurate but still plausible. Although such manipulation is technically feasible, we lack quantitative evidence on how the degree of alignment affects decisions across tasks, how it interacts with individual differences, and how people evaluate the ethical acceptability and disclosure of aligned explanations. We study these questions in an online experiment with 167 participants and three decision-making tasks. For each participant, we vary how strongly explanations are aligned with that user’s inferred decision criteria and measure changes in decisions, subjective measures, and ethical acceptability judgments. The results show that as explanations more closely match users’ decision criteria, agreement with AI’s advice increases in the objective and subjective tasks, whereas in the high-stakes task alignment mainly raises subjective measures without improving decisions. Alignment also makes participants more uniformly susceptible, and participants judge inaccurate aligned explanations as much less acceptable in high-stakes than in subjective tasks, while still broadly expecting such manipulations to be disclosed. These findings indicate that making explanations intuitive and agreeable is not sufficient for responsible personalization and that alignment should be combined with mechanisms for verification and disclosure.
