Skip to content
The Internet Compass

Topic Modeling

Benchmarking Large Language Models for News Summarization

Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, Tatsunori Hashimoto

Topic ModelingAdvanced Text Analysis TechniquesNatural Language Processing Techniques
Published January 1, 2024Read PDF ↗View on arXiv ↗

Abstract

Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LLM summaries are judged to be on par with human written summaries.

Sourced from arXiv · Updated September 2, 2026

Thank you to arXiv for use of its open access interoperability.

View original source ↗Spot an error on this page? Let us know →

FAQ

Common questions

What is "Benchmarking Large Language Models for News Summarization" about?

Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important o

Who wrote this paper?

Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, Tatsunori Hashimoto

Where can I read the full paper?

The full text is available as a PDF on arXiv (linked above), published January 1, 2024.

Does this paper have a DOI?

Yes: https://doi.org/10.1162/tacl_a_00632.