Health misinformation poses a significant public threat by eroding trust in scientific expertise and diminishing adherence to health guidelines, which collectively weaken community resilience to preventable diseases. For these reasons, detecting health misinformation is crucial to protect public health. However, manual detection requires substantial human effort and expertise, making it impractical at scale, particularly in low-resource settings where technological and linguistic resources are limited. Developing automated techniques for identifying false or misleading claims is therefore essential to ensure timely intervention. Advancing these automated detection methods depends on the development of robust datasets, as they enable more accurate modeling and adaptation for specific languages and contexts. To the best of the authors’ knowledge, no misinformation detection techniques or datasets have yet been developed specifically for the Estonian language within the health domain.Addressing this gap, the primary objective of this study is to develop a reliable system for generating ground truth labels for health misinformation in Estonian, thereby contributing to misinformation detection in low-resource settings. Leveraging pre-labeled datasets in English, the proposed Cross-lingual Alignment and Confident Prediction Sampling (CAPS) approach employs a hybrid two-phase methodology involving semantic similarity measurements, manual annotation, classification, and confidence sampling. This methodology enables the efficient generation of misinformation labels with minimal reliance on manual annotation, contributing a valuable resource for advancing misinformation detection in underrepresented languages. The resulting dataset of 8,795 annotated news articles represents a significant advancement in health misinformation detection for the Estonian language.
