Sycophancy (overly agreeable or flattering behavior) poses a fundamental challenge for human-AI collaboration, particularly in high-stakes decision-making domains such as health, law, and education. A central difficulty in studying sycophancy in large language models (LLMs) is disentangling sycophantic belief shifts from rational changes in behavior driven by new evidence or user-provided information. Existing approaches either measure descriptive behavior changes or apply normative evaluations that rely on objective ground truth, limiting their applicability to subjective or uncertain tasks.We introduce a Bayesian probabilistic framework, grounded in behavioral economics and rational decision theory, that explicitly separates sycophancy from rational belief updating. Within this framework, we propose two group-truth-independent metrics for studying sycophancy: (i) a descriptive metric that measures sycophancy while controlling for rational responses to evidence, and (ii) a normative metric that quantifies how sycophancy leads models astray from Bayesian-consistent belief updating. Applying our framework across multiple LLMs and three uncertainty-driven tasks, we find robust evidence of sycophantic belief shifts and show that their impact on rationality depends on whether models systematically over- or under-update their beliefs, with most baselines demonstrating significant increases in error due to sycophancy when the model over-updates. Finally, we propose a novel post-hoc calibration method and two fine-tuning strategies that reward Bayesian-rational updating (BayesSFT and BayesDPO). We find evidence that post-hoc calibration significantly reduces Bayesian error, and observe significant reductions in both sycophancy and Bayesian error associated with our novel fine-tuning methods.
