While LLMs appear to reason, their "thought" processes fundamentally differ from ours.

Human reasoning is a kaleidoscope of processes, integrating memory, context, culture, and more.

While LLMs can be improved, the underlying dynamics will always be different, according to some.

David Matos / Unsplash

A team of psychologists, computer scientists, and physicists from Italy, Slovenia, and South Korea has proposed seven “fault lines” between human and artificial intelligence (AI). The researchers question the extent to which AI is yielding intelligence, at least in the same form as human intelligence. The authors note the critical shift that occurred between statistical natural language processing (NLP), where AI retrieves and ranks existing information, with the (human) user then able to judge between them, and generative AI, which instead presents one fluent, “authoritative-seeming” answer. In short, while large language models (LLMs) produce content that seems cognitively informed and deliberated, the processes underlying how that content is produced are fundamentally different from human cognition.

One of the fundamental divergences between human and artificial intelligence highlighted by the authors relates to how they use language. AI does not understand language the way a human mind processes it and derives meaning; AI instead recognizes statistical patterns, garnered from human-produced text: “[LLMs] do not track truth conditions or causal structure; they track patterns of co-occurrence, association, and continuation in text.” The authors argue that by supplementing this ability with a number of additional mechanisms such as retrieval-augmented generation, they make LLM outputs look more reliable and more convincing without becoming more knowledgeable.

Examining the differences between human and artificial intelligence, the authors propose seven epistemic fault lines that highlight these differences. These correspond to seven sequential stages of human judgement:

Based on Quattrociocchi et al. (2025)

The grounding fault: The authors argue that humans arrive at judgments based on layers of information that are often not accessible to LLMs, such as facial expressions, tone of voice, and social cues. Accordingly, LLMs often underdetect expressions that are sarcastic or ironic, or other human-specific expressions. Despite that, LLMs appear to be getting better at detecting these, though that is related to more effective pretraining and conditioning rather than understanding.