
Research Article
Comparison of Core Abilities of Four Large Language Models: GPT-3, GPT-4, LLaMA 2, and PaLM 2
@INPROCEEDINGS{10.4108/eai.22-5-2026.2365089, author={Guoyu Xu}, title={Comparison of Core Abilities of Four Large Language Models: GPT-3, GPT-4, LLaMA 2, and PaLM 2}, proceedings={Proceedings of the 4th International Conference on Image, Algorithms, and Artificial Intelligence, ICIAAI 2026, 22-24 May 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICIAAI}, year={2026}, month={8}, keywords={Large Language Models Transformer Natural Language Processing Model Evaluation}, doi={10.4108/eai.22-5-2026.2365089} }- Guoyu Xu
Year: 2026
Comparison of Core Abilities of Four Large Language Models: GPT-3, GPT-4, LLaMA 2, and PaLM 2
ICIAAI
EAI
DOI: 10.4108/eai.22-5-2026.2365089
Abstract
Different large language models vary dramatically in their resource requirements, with some costing millions to train while others can run on a laptop. This paper examines four models that represent different design philosophies: GPT-3, GPT-4, LLaMA 2, and PaLM 2. GPT-3, with its 175 billion parameters, first showed that scaling up could unlock few-shot learning. GPT-4 went further by adding image understanding, though OpenAI never published the architecture. Meta took the opposite approach with LLaMA 2, releasing weights publicly so anyone could experiment. Google's PaLM 2 focused on languages beyond English, training on over 100 languages. After reviewing how each model works and where it performs best, the paper discusses what problems remain unsolved, including that models still make up facts, training still costs too much, and no one fully understands why certain prompts fail.


