Leibniz-Zentrum Allgemeine Sprachwissenschaft Leibniz-Gemeinschaft

Semantics circle: Mechanistic interpretability in AI as a method for probing the logical concepts of LLMs

Vortragende(r) Younchan Lee (UMass) & Guillermo Del Pinal (UMass)
Datum 10.07.2026, 14:00 - 15:30 Uhr
Uhrzeit 14:00 Uhr
Ort ZAS, Meierottostraße 8, Eberhard-Lämmert-Saal, Ground floor

Code of Conduct for ZAS events: The ZAS is committed to fair, respectful, and professional interaction at its events. Therefore, please observe the Code of Conduct for this event.

Abstract

One of the core research topics in current formal semantics and cognitive science concerns the logical primitives at the interface of language and cognition. Indeed, much of current research in semantics focuses on data patterns and puzzles that can be used to uncover the logical structure and components involved in functional terms, including quantifiers, modals, auxiliaries and connectives. In addition, important work at the interface between semantics and pragmatics focuses on phenomena which suggests that certain kinds of pragmatic enrichments previously thought to involve general cognition are instead triggered by covert logical operators at the syntax-semantics interface. Ultimately, these projects all help us make progress on a foundational project in cognitive sciences concerning the basic logical primitives of human thought and language, terms that are crucial to understand the full power and flexibility of human reasoning. There are good reasons to think that our access to such logical primitives is at least part of the reason why we are so efficient at language acquisition, at forming abstract and systematic thoughts, and more generally is involved in our distinctive capacity for robust generalization. These are all capacities that, according to many influential researchers, still distinguish humans from large language models (LLMs). Although a lot of research now focuses on this difference, little work has actually tried to test whether and to what extent the kinds of large language models which extensionally match our linguistic and logical inference patterns and competence actually use, in their internal representations, something like the kinds of logical primitives that, according to formal semanticists and cognitive scientists, are part of what explains our own relevant competence. Interestingly, AI researchers in a recent field called “mechanistic interpretability” have developed sophisticated methods for “mind reading” the internal representations learned and used by LLM models. These techniques have been applied with some success to uncover the abstract/concepts underlying open class words (e.g., ‘bridge’, ‘dog’), due to lack of the relevant interdisciplinary expertise, they haven’t yet been deployed on the patterns mentioned above to try to discover the logical primitives (or precursors) used by LLMs when they process functional words. In this exploratory talk, we will present preliminary evidence for the view that we can use mechanistic interpretability techniques to uncover the logical primitives used by LLMs to process functional terms, and will also discuss the prospects of this approach to help us settle theoretical disagreements about the specific logical decomposition of particular functional terms.