Social Science Research Council Research AMP Just Tech
Citation

A framework for evaluating factual consistency in automated text summarization with large language models and prompting strategies

Author:
Islam, Md Moinul; Oussalah, Mourad
Publication:
Neural Networks
Year:
2026

The exponential growth of textual data has intensified the need for reliable automated text summarization (ATS) systems that can extract and synthesize knowledge while maintaining factual accuracy. Current evaluation frameworks for large language models (LLMs) in summarization tasks lack comprehensive assessment of factual consistency, particularly in knowledge engineering contexts where information integrity is paramount. This paper presents a comprehensive evaluation framework that systematically assesses factual consistency in LLM-generated summaries through advanced prompting strategies and multi-dimensional evaluation metrics. Our framework integrates five prompting methodologies, such as Zero-shot, Few-shot, Chain-of-Thought (CoT), Structured Chain-of-Thought (SCoT), and Chain-of-Verification (CoVe) with state-of-the-art (SOTA) factuality assessment approaches, such as FActScore, LongDocFACTScore (LDFActs) and AlignScore across eight LLMs and five diverse datasets spanning news, scientific literature, and conversational domains. Results demonstrate that Few-shot prompting achieves optimal performance across most domains except scientific literature, with LLMs consistently outperforming human-generated summaries. Our findings reveal trade-offs between completeness and precision, with models generating 2–10 times more atomic facts than human references while maintaining comparable or superior factual accuracy. The framework provides actionable insights for researchers developing reliable summarization systems, with open-source implementation available for reproducibility.