Preprint
Hypothesis

This version is not peer-reviewed.

The Evolutionary Origin of Predication and Syntax

Submitted:

07 August 2026

Posted:

10 August 2026

You are already at the latest version

Abstract
The hypothesis defended here is as follows. Before the emergence of predication, language consisted only of pre-words used exclusively to request something or call someone over, with only one pre-word occurring in each message and always with the same intonation. How did these meanings become detached from their invariant function and intonation and thereby turn into genuine words? The most plausible possibility is that a recipient was puzzled upon hearing a request for something that was no longer available or a call addressed to someone who was not in the area. The recipient would then repeat the pre-word, but without its conative function and intonation, using it instead as a thema, and would add another former pre-word that now functioned as a rhema or predicate. The transformation this entailed soon required a further innovation: vocal signifiers had to split into two mutually detachable planes, the articulatory-phonetic and the intonational. In this way, each signifier could be perceived as exactly the same both when it constituted a message used to call someone over or to ask for something and when it was integrated with one or more other signifiers into a single intonational pattern.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

This article addresses the origin of syntax and therefore presents a hypothesis about language evolution. As is well known, the only way to make progress in the study of language evolution is to formulate hypotheses and attempt to evaluate them using every resource available to us.
Predication and syntax are features of exclusively human communication. The question of how they emerged in evolution —or, more precisely, in gene–culture coevolution— may therefore also shed light on other uniquely human characteristics that can be linked in one way or another with fully developed language. These links cannot be explored here, although I will refer to some of them at various points.
The proposal advanced in this article focuses on vocal communication. Our present-day language, and very probably language at its origins as well, is certainly multimodal: we communicate not only through words but also through gestures. Nevertheless (in Section 2), I will explain my reasons for concentrating on the auditory modality.
Vocal communication is particularly well suited to the functions of requesting something or calling someone over. These functions, especially calling, often occur when the distance between producer and addressee is greater than is suitable for purely gestural communication. Thus, I propose (in Section 3) that, before the emergence of predication, language consisted only of pre-words used exclusively for requesting something or calling someone over, pre-words that were always used alone in each message, and with the same intonation every time. In this connection, I will examine the causal relationship between language and concepts and attempt to reverse the traditional view of that relationship.
According to the hypothesis, predication evolutionarily emerged because of a particular dialogical relation (Section 4). A chance hearer of a call addressed to someone who was absent, or the addressee of a request for something that was unavailable, might have been puzzled by what they had just heard. This would lead that person to reject, correct, or update the mental content they had just heard expressed by the other individual. Predication thus emerged, together with a syntax in an as-yet ungrammaticalized form.
The emergence of predication subsequently required—although, in evolutionary terms, not much later—another major change (Section 5). Pre-words did not possess the two differentiated and mutually detachable planes of phonetic articulation and intonation. A genuine word, by contrast, must be perceived as the same word both when it requests something or calls someone over and when it functions as a component of an intonational pattern encompassing at least two words. This required some other element of the sound to remain identical while intonation changed. In other words, articulatory-phonetic sequences had to be kept invariant by being learned and imitated with complete fidelity and precision.
Section 6 examines the relationship between Theory-of-Mind and the dialogical reply that constitutes the origin of predication.
In Section 7, the hypothesis turns to the predications used in the last communicative function of language to emerge: speaking to oneself.
In Section 8, we will return to the emergence of the two differentiated and mutually detachable planes of phonetic articulation and intonation. The need for this differentiation gave rise to the well-known fact that the two planes of word sound are processed in different cerebral hemispheres, although they must immediately be unified. Considering this fact within the framework of the hypothesis may prove informative. More specifically, with the aid of paleogeneticists, it might allow us to derive a tentative inference (or ‘hypothetical deduction’) concerning when syntax emerged.
Section 9 seeks to ground our hopes for the long-term outcomes of research on language evolution and, more broadly, evolutionary general anthropology.

2. Multimodality

We communicate not only through words but also through manual, facial, and whole-body gestures. When interacting with someone who does not know our language, we may even use pantomimes resembling children’s symbolic play. Human communication today therefore clearly operates in both the auditory and visual modalities, and it is highly plausible that this was also true at its origins. The hypothesis developed here will nevertheless focus on predominantly vocal communication.
I will begin, however, by acknowledging that highly distinguished scholars maintain that “gesture supports the emergence of communication systems”. The major figure Michael Corballis (2002) must be cited here, as must the group from Nicolaus Copernicus University (Poland) whose work focuses on pantomime 1.
What I find particularly important in this respect is that humans share ostensive and emotional gestures with the great apes. Indeed, among adult humans, purely gestural predicative communication most often takes the following form. First, the producer establishes eye contact with the addressee; next, the producer points to and looks at—or merely looks at—an object in the environment; finally, the producer looks back at the addressee and accompanies this gaze with a gesture expressing an emotion. We all know how to produce and understand purely gestural communication equivalent to “This speaker is a bore” or “This feast looks appetizing”. These emotional gestures are admittedly highly semiotized 2. More specifically, they are semiotized to a degree perhaps exceeding the maximum attainable by chimpanzees. Even if this difference is accepted, however, an evolutionary continuity with the signs of great apes remains, one that at first sight appears to favor a gestural origin of human language.
Since the 1990s, it has also been argued that gestural modality may provide further clues to the origin of language. This claim draws attention to emerging sign languages, which are progressively created when a group of deaf people live together over a period of years (see Senghas & Coppola 2001; Senghas et al. 2004). As Sandler et al. (2022) state, these languages “are the only human languages that can emerge de novo at any time. (…) Spoken creole languages are also young, but are different from emerging sign languages, in that the speakers of pidgins from which creoles are assumed to have descended already had native languages.”
Despite all this, I remain convinced that the vocal modality (or, more precisely, the predominantly vocal one) may be more fruitful for studying language evolution. Several considerations lead me to focus the hypothesis on the auditory-vocal modality:
  • Attention to both child development and pragmatics has led us to recognize the great importance of intonation. This is the most important of all these considerations.
  • The historical development of grammaticalization has proceeded only through the vocal channel.
  • Among people without hearing or vocal impairments, purely gestural communication is restricted to specific contexts in which the aim is for the message to reach only the addressee. I return to the example “This speaker is a bore”. Purely vocal communication, by contrast, is always entirely possible.
Finally, multimodality is now, and surely always has been, functionally differentiated. Certain communicative functions are much better suited to one modality than to the other. Messages used to ask for something and –even more so– to call someone over, typically involve a distance between producer and addressee that exceeds what is appropriate for purely gestural communication. Such communication would therefore have relied primarily on vocal modality.

3. The Type of Communication Immediately Preceding the Emergence of Predication and Syntax

At the outset, then, the vocal modality would have been used to ask for something and to call someone over, and it would also have relied on intonation. Even today, both as producers and as recipients, we continue to use intonation as the distinctive cue for differentiating calls, requests or commands, predications, and questions.
What the hypothesis proposed in this article emphasizes above all is that, unlike our words, these vocal signs
  • would have been used only for the communicative functions of requesting a particular object or calling a particular individual over,
  • would always have occurred one at a time—that is, always in holophrastic messages,
  • and would always have been produced with the same intonation.
  • In other words, these signs would not have been genuine words, but pre-words 3.
In human language, by contrast, the word ‘mama’ may of course be used as a call and the word ‘water’ as a request, but these words can also
  • fulfil many different syntactic roles,
  • be pronounced with different intonational patterns or as different parts of an intonational pattern,
  • and, most importantly, be combined with other words. Genuine words are therefore always ‘words for syntax’. Moreover, as syntax became increasingly grammaticalized, words came to be assigned to one or another “part of speech” or “word class,” and, in this way, they came to require one another reciprocally—especially nouns and verbs.
We have not yet considered a feature of pre-words that is particularly important to the hypothesis and therefore requires closer attention. Pre-words did not designate any referent; they served only to perform concrete acts of requesting an object or calling someone over. These signs certainly had what we might call ‘concrete worldly connections,’ enabling them to call a particular individual over or request a specific object. Those worldly connections, however, did not in any way turn them into referential signs to which a predicate could be added.
A further step is needed to characterize this phase of vocal pre-words. Words –especially words with meanings that can be evoked, such as nouns, verbs, and adjectives– have commonly been thought to derive from concepts formed by the mind independently of communication. The most explicit and elaborate version of this idea is ‘mentalese’ (Fodor 1975). Even without going that far, most scholars of language evolution regard concepts as the origin and model of content words.
I propose reversing this relationship. In Bejarano 2014 (a review of Hurford 2010) I criticized above all his acceptance of the conventional view. I argue that we never think of the multiple features contained in a perception as independent steps that jointly compose a higher-order unit. Perceptions simultaneously activate an enormous array of features. Size, distance, movement, action, direction, and so forth are just as constitutive of the perceptual unit as recognition of the ‘wolf’ or ‘apple’ pattern, and none of these features is apprehended in perception independently of the others. Indeed, progressively assembling perception would be not merely useless but counterproductive. Consider what would happen if the perception of a tiger had successively to add the features ‘large’, ‘coming toward me,’ and ‘rapidly.’ It therefore seems more likely that the concept, as a meaning that can be thought independently, emerged only after—and as a consequence of—the ‘semantics for syntax,’ or semantics of ‘parts of speech,’ created by human language.
This proposal –namely, language creates concepts that can be thought independently– closely resembles the ‘label-feedback hypothesis’. Lupyan 2012 defines it as follows: “Labels act as implicit category markers, highlighting commonalities among items that share a label”. Recently, he has demonstrated the power of his hypothesis even more convincingly. He did so primarily by focusing on the present: “This conception of language not only helps explain the effects of language on reasoning, categorization, and perception, but also the recent astonishing advances in artificial intelligence in which exposure to natural language transforms general-purpose neural networks into human-like minds”: Lupyan 2026. However, working within the well-established field of experimental semiotics, this author projects his hypothesis back onto the origins of language somewhat less explicitly than I would have preferred.
In summary, I propose that the original use of the vocal modality produced not genuine words, but instruments inextricably tied to a particular communicative goal and necessarily accompanied by their particular intonation. This is how they would have been learned by each generation. Detaching these signifiers from their function and intonation must therefore have been difficult.

4. How Could the Difficulties Surrounding the Original Emergence of Predication Be Overcome?

What has been said above about the enormous difference between pre-words and words seems to lead us into a dead end. Yet we know that access to vocal predication was not only possible but spectacularly successful. It should be recalled that, among people without vocal or hearing impairments, communication based solely on gesture is confined to very few occasions and highly specific contexts. We must therefore explain how, according to the hypothesis, genuine words and syntax emerged.
The most plausible possibility is that the addressee of a request, or the recipient of a call addressed to someone, was puzzled by what they had just heard. Today we might express this puzzlement by saying “How can this person be addressing a calling to someone who left several days ago?” or “How can this person ask me to bring something we no longer have?”. The recipient would then reply. Let us consider an example.
In Bejarano 2008, I described the following scene 4. “An adult and a child are playing with wooden blocks to build an increasingly tall tower. The adult then asks the child, who is holding the box of blocks, ‘Give me more blocks! ‘More!’ At that moment, the child, still holding the box, sees that it is empty and says in Spanish: ‘Más (= more) no.’”
This is, of course, very different from what the evolutionary origin would have been: the child had been surrounded by genuine words for more than two years. It can nevertheless serve as an example of the solution I am proposing. Whereas my “Más” was a message with conative intonation and function, the “Más” constituting the first part of the child’s reply was not conative at all; it served only to specify the point to which the second part applied. The second part, “no,” far from being a refusal to do what I had requested—that is, far from being the rejection-negation that the child had been using long before then—provided information about the object of the request. More specifically, it conveyed information still unknown to the addressee. A further feature must also be considered. ‘Más’ and ‘no’ can certainly often appear as the sole word in a message. The child’s reply, however, was a composition in which the two meanings formed a single intonational pattern.
In the preceding paragraph, I referred to the addressee’s ‘lack of knowledge’. This point, I propose, is crucial to the origin of predication and must be developed further. The thema (or, later, with the development of grammaticalization, the subject of the verb, as in ‘The blocks have run out’) does not correspond to the speaker’s own belief about the referent. Rather, it corresponds to the false, incomplete, or outdated belief that, in the speaker’s view, the addressee holds about that referent. The apprehension of false belief –the emblem of Theory of Mind– is central to my proposal concerning the origin of syntax; here, however (unlike in most Theory-of-Mind research 5) that belief has been heard. (Note, please, that it is not necessary to linguistically assert the predication corresponding to a belief for the hearer to become informed of that belief.) The rhema or predicate, for its part, must be understood as the feature that corrects, completes, or updates the addressee’s belief. I have defended this view of predication throughout my research career using various approaches. Chronologically, the first of these was a critical analysis of Frege’s “Sense and Reference”.
The hypothesis concerning the original emergence of predication has now been presented, but it can also be described in other terms. Repetition by another person of the conative, presyntactic message would have stripped that message—that pre-word—of its force and self-sufficiency. This would have constituted an extraordinary instance of ‘bleaching,’ a transformation of a magnitude never subsequently repeated in the languages for which the grammatical technical term ‘bleaching’ was coined 6. In this original bleaching, the genuine word emerged from the pre-word, and syntactic composition emerged with it.
I wish to emphasize that both ideas employed in this proposal—first, the clash between one individual’s mental content and that of another, and second, the emergence of predication in dialogical reply—lead us to regard dialogue as the cause that brought syntax into being. Proposals of this kind are now becoming increasingly widespread. Carpendale et al. (2024), for example, explain children’s achievements not in terms of an internal capacity but “as a distributed accomplishment emerging within shared practices.” My focus is primarily evolutionary rather than developmental, and the achievement that concerns me is the genesis of syntax (and, simultaneously, of concepts, as discussed above). Despite these differences in focus, I am pleased to observe that the general trend has now changed almost completely.

5. The New Requirement Generated by the Emergence of Predication

A genuine word must be perceived as the same word both when it serves as a call addressed to a particular individual or a request for something specific and when it functions as part of an intonational pattern encompassing at least two words. This required some other element of the sound to remain identical while intonation changed. The sound of genuine words therefore had to acquire two differentiated and mutually detachable planes: intonation and the articulatory-phonetic sequence. In other words, articulatory-phonetic sequences had to be kept invariant by being imitated with complete fidelity and precision.
This type of imitation was an evolutionary novelty. To appreciate its significance, we may begin by considering the research conducted at Stockholm University over approximately the past decade on the high neural cost of both representing and imitating motor sequences (see, e.g., Lind & Jon-And 2025). These types of imitation and learning were probably difficult to acquire. By contrast, earlier forms of cultural learning, although they could involve several steps, were assisted at each stage by the affordances presented by the object to which the task was applied. We can now turn to something—mentioned in Jon-And et al. 2025—that is more specific and directly relevant here, namely, imitation of the articulatory-phonetic sequence. In this type of imitation, it is clear that no such assistance provided by the object’s affordances is available; everything depends on the ability to represent the sequence.
This extremely demanding imitation of motor sequences has been undoubtedly applied to other tasks. But I propose that it first emerged in language. Note that, according to the hypothesis, the detachment of the articulatory-phonetic sequence from intonation —that is, the outcome of precise articulatory-phonetic imitation— provided the enormous adaptive advantage of making syntax less costly 7.
To explain more clearly how important, according to the hypothesis, the precise imitation of articulatory-phonetic sequences was, we can establish the following causal chain.
  • One, the earliest vocal signs lacked this precise imitation of articulatory-phonetic sequences.
  • Two, without such imitation, detaching intonation from the signs was virtually impossible.
  • Three, without this detachment, even after the evolutionary emergence of predication and syntax, understanding them remained difficult.
  • Four, following the evolutionary emergence of syntax, the precise imitation of articulatory-phonetic sequences must have evolved relatively quickly on an evolutionary timescale.
The distribution of present-day language across the two cerebral hemispheres should also be noted 8. In most individuals, the production and reception of articulatory-phonetic sequences are the exclusive responsibility of the left hemisphere. The right hemisphere, by contrast, is invariably involved in both emotional prosody and linguistic prosody, the latter serving to disambiguate speech acts and syntax. (Given this division of prosody, the term ‘intonation,’ formerly used to refer to linguistic prosody, could be dispensed with. I will nevertheless sometimes retain the older term here for convenient reference to the intonational pattern.)
The separation of the two major elements that constitute present-day linguistic signs clearly allows each to be processed with great efficiency. It is equally clear, however, that this new cerebral distribution must have represented a major change.
Before proceeding to the next section, I will comment on two points. Is there anything resembling a pre-word in present-day language? My answer is that there is nothing genuinely comparable. The only possible candidates that occur to me are purely expressive cries prompted by extreme emotion. In such cries—unlike linguistic interjections—the articulatory-phonetic sequence cannot be separated from intonation because there is no articulatory-phonetic sequence. In this respect, they bear some resemblance to pre-words. Yet there is a fundamental difference between these cries and ancient pre-words: cries are purely expressive and therefore lack communicative intentionality, whereas pre-words were genuine communicative instruments.
More generally, no pre-word could have survived as a “living fossil” 9. In light of the hypothesis proposed in this article, the following qualification can be made. Pre-words were used only to call someone over or to request something. By contrast, when we learn, for example, a person’s proper name, we assign it a referential meaning, whether in a narrated world or in the real world. Likewise, we learn words together with the other words with which they have been combined. In this way, we acquire them as what linguistics call “parts of speech” 10. Pre-words, however, could not have constituted ‘parts of speech’. In short, according to the hypothesis, living fossils can survive only from periods in which grammaticalization was still weak, not from the very origin of predication and syntax.
The other point concerns Gázquez et al. (2026), who propose that “hemispheric specialisation for language—especially production—can exert pressure on praxis pathways and manual skill asymmetry, more strongly than the reverse”. I find this issue interesting. However, rather than comparing two influences operating in opposite directions, we might suggest that development and evolution involve a spiral of steps caused alternately by hemispheric specialization and by manual asymmetry 11.

6. The Relationship Between Theory-of-Mind and the Origin of Predication

In Section 4, we saw that, according to the hypothesis, the apprehension of false belief (the emblem of Theory-of-Mind) is central to the original emergence of predication and syntax. More specifically, the dialogical reply that gave rise to the predicative communicative function resulted from reply-producer having heard a call addressed to someone who was absent or a request for an item that was no longer available. Certainly, Theory-of-Mind encompasses other highly diverse domains. The domain of interest here, however, may be the most neurally demanding.
After being puzzled by such messages, the producer of the impending dialogical reply simultaneously maintains two different mental contents concerning the same entity. Other Theory-of-Mind activities, by contrast, are based merely on imagining oneself in another’s circumstances. This may occur in the form practiced by animals—probably through vicarious expectations (Bejarano 2025, subsection 3.2)—or in the highly elaborate form of the 'addiction to narratives’ (see Sperber 1985 / Sterelny 2001). In these activities, the subject is momentarily identified with the individual to whom attention is directed; the costly duality concerning the same entity is therefore unnecessary.
How could the simultaneous duality of these contents –different from one another yet referring to one and the same entity– have originally emerged? I propose that this costly duality requires the subject engaged in Theory-of-Mind activity to be entirely unable, at that moment, to identify with the individual attended to and whose mental state the subject is apprehending. I further propose that this requirement is fulfilled in the hypothesis about the origin of predication presented above: Individual A cannot identify with Individual B while Individual B is speaking to Individual A.

6. The Communicative Function of Speaking to Oneself

Language is, of course, learned socially. Through continued use and through its influence on cognitive processes –an influence acquired by creating concepts as independent thoughts– it nevertheless becomes almost automatic. Speech, initially the paradigmatic instance of communicative intentionality, can thus sometimes become akin to a purely expressive emotional gesture. The only modification the subject can introduce into this mere automatism is to make it silent, especially when other people with whom the subject is not familiar are nearby. Unfortunately, the examples of inner speech provided by Vygotsky 1934 / 1986 are of this kind. No genuinely new function of language has yet emerged here.
At a later stage, however, the subject can address themself with deliberate communicative intent. This is a major innovation. Such self-address sometimes resembles a command and serves to remind oneself either of a command received or of the next step in a plan one has devised. Both cases reinforce a memory for the future; in other words, they serve as prospective memory.
At other times, intentionally self-directed inner speech takes the form of a predication, usually with an elliptical subject. These cases constitute the truly interesting aspect of inner speech (and not merely because they fall within the scope of this article). Unlike predication addressed to another person, such predication does not correct any lack of knowledge on the addressee’s part. By definition, the producer already knew what they say to themself. Moreover, because this is not an instruction to be carried out in the world but a description of what the subject knows about the world, it cannot constitute prospective memory. Could we say that the subject is attempting to reinforce, in the present, knowledge already possessed in the present? Would that be absurd? Perhaps not, if we assume that such knowledge may compel us to make decisions that we perceive as both reflecting reality and being very annoying. Through this type of inner speech, the subject would thus force themself to attend to that knowledge –or to look again at what has already been seen rather than passing it by.
This would constitute a new form of self-control. It may be compared with the control supplied by shame (see Baumard et al. 2013) and pride (Bejarano 2025, subsection 5.1). These self-conscious emotions are certainly highly adaptive and effective. The domains in which they can exert control, however, form a small closed set. Moreover, these emotions are not under the subject’s control; rather, they control the subject.
Self-directed speech, by contrast, has an extraordinary cost–benefit ratio, especially once full internalization greatly reduces its cost. It also enables the subject to ask anything of themself. It is therefore plausible that this type of inner speech constitutes a uniquely human characteristic which, although distinct from language, is nonetheless rooted in it.

7. Is There Any Way to Estimate When Predication and Syntax Originated?

The hypothesis presented above does not specify the evolutionary period in which predication and syntax emerged. It may nevertheless allow us to draw an inference that offers a tentative clue. For the moment, however, let us begin by focusing on the evolutionary emergence of an intonational pattern encompassing at least two words.
The function of such an intonational pattern corresponds perfectly to that of the appearance of the human eye. Hearing this intonational pattern and observing the producer’s eye movement serve to bind tightly together the different elements of syntactic predication within each modality. In gestural predications, such as the example equivalent to “This speaker is a bore”, the conspicuous horizontal movement of the iris across the broad white sclera leads the addressee to connect the producer’s gaze toward the object with the combined emotional gesture and gaze directed toward him –toward the addressee. This helps the addressee understand the emotional gesture as the predicate that the producer wishes to communicate about the indicated object 12. Analogously, in the vocal modality, the shared intonation turns the first meaning—the word “more” in the example described above—into the basis to which the predicate “no,” which corrects or updates it, is applied. Both resources help the recipient achieve syntactic integration.
We argued above, in Section 5, that the evolutionary emergence of vocal predication subsequently required —although, in evolutionary terms, not much later — an important change: the processing of intonation encompassing two words—that is, the intonation whose function we have just emphasized—and the processing of the articulatory-phonetic sequence constituting each word had to become separate, because each operated on different units. For both processes to remain efficient, they would have had to be located in different cerebral hemispheres; therefore, the immediate integration of their outputs would also have become indispensable. This probably required a major strengthening of the corpus callosum connecting the two hemispheres. Friederici et al. (2002) and Sammler et al. (2010) show that speech comprehension requires syntax and prosody to be reintegrated through interhemispheric interaction via the corpus callosum.
This provides the tentative clue—the inference derived from the hypothesis—that we were seeking. The gene or genes responsible for the larger size of the human corpus callosum relative to that of other hominins may perhaps be identified in present-day humans. In that case, our question about when predication emerged could then be reformulated as follows: assuming that this gene is identified in present-day humans, would paleogeneticists find it in extinct Homo species? And if so, in exactly which extinct species?

8. Concluding Remarks: A Reflection on Interdisciplinarity

Evolutionary general anthropology should study all characteristically human capacities together. These are numerous: language, reasoning, the capacity for moral choice in the most demanding sense, an attraction to fiction, imagination, curiosity capable of identifying problems to be solved, the sense of humor, etc. Such inquiry should also approach these capacities from every relevant scientific field. Put differently, in this domain we would do well to follow the maxim: “If you want to solve a puzzle, you must work with all the pieces”. Unfortunately, this conviction, which many of us ultimately share, is difficult to put into practice.
The difficulties are enormous. No one can acquire more than a superficial and ultimately useless familiarity with every field involved. Teams bringing together specialists from several related fields can certainly be highly productive. At present, however, a project aimed at assembling specialists from every relevant field would probably achieve little. Genuine interdisciplinarity culminates in a synthesis, not merely in an aggregate of studies on different topics.
These remarks may appear pessimistic, but they are not so in the long term. Indeed, although for now we must accept that the ideal strategy is not yet feasible, I am confident that this will change. Researchers will then have access to a far greater body of findings on which to draw. Moreover, future AI will provide an external memory of everything preserved from what has been said or writte; and, much more importantly, the capacity to identify, through associations across this immense body of human ‘semantic archives’, some element that might “make sense in relation to the very small fraction (of possible actions) that lead to successful discoveries” (Yaman et al. 2016), or, in other words, some element that might fit the still-empty profile of the solution we happen to be seeking at that moment. Embracing this long-term hope, however, requires us—despite our humble recognition of our limitations—to work with enthusiasm in the present.

Author Contributions

I am the sole author.

Funding

This research received no external funding.

Conflicts of Interest

The author declares no conflicts of interest.

Notes

1
I have decided to use notes so that I can present the hypothesis in a linear manner and thereby improve clarity. However, some of the discussions in the notes are important and may even prove more interesting than the main text.
2
“We define pantomime as a communication mode that is mimetic; non-conventional and motivated; multimodal (primarily visual); improvised; using the whole body rather than exclusively manual; holistic; communicatively complex and self-sufficient; semantically complex; displaced, open-ended and universal”: Żywiczyński et al. 2018. I also greatly value Wacewicz and Żywiczyński (2024). Pantomime, however, may be too costly and far too inefficient to be situated at the origin of language. (Note, please that, unlike pantomime, the example of purely gestural predication presented just below is closer to what these authors call a “conspiratorial whisper.” In addition, Placinski et al. 2023—an article co-authored by Żywiczyński and Wacewicz—acknowledges that pantomimes are communicatively inefficient: “Over repeated social interactions, the efficiency of gestures is due to a change from whole-body pantomimes to more efficient manual gestures”.) As noted above, I would locate the very origin of pantomimes in children’s symbolic play, not in communication.
3
Recall that “semiotics is in principle the discipline studying everything which can be used in order to lie” (Eco 1975), and consider the following account of what occurs here: “The intentional level would control and use the behavioral and even autonomic ones, i.e., those movements or expressions that originally were not intentionally communicative” (Lipschits and Geva 2024, a study that approaches the issue from a developmental perspective, unlike Graham et al. 2024, whose approach is comparative).
4
King & Janik (2013) state that “bottlenose dolphins develop their own unique identity signal, the signature whistle. This whistle encodes individual identity independently of voice features. The copying of signature whistles allows (as our experiments have shown) animals to address one another.” (my emphasis). The copying of signature whistles can thus be used to call a particular individual. We would therefore find here something resembling a proper name, but one restricted to the function of calling someone over—that is, precisely what is termed a pre-word in this article. If groups of social hunters such as dolphins can achieve this, groups of hunting hominins could have done so as well. Also relevant is Progovac’s (2026) claim that “enhanced face recognition in humans coevolved (probably in the fusiform gyrus area) with early naming strategies.” However, as the reader may already be beginning to see, my hypothesis leaves no room for the idea that “(flat) verb–noun compounds, which are typically used for (derogatory) naming and nicknaming” (ibid.) are “living fossils” from the truly ancestral stage. I will return to the relationship between living fossils and pre-words at Section 5.
5
I would now add that the scene involved my elder son (then 32 months old), who had been somewhat later than other children of his age in beginning to produce language, and me.
6
Likewise, Geurts (one of the figures in the field of language evolution who currently stands out for situating the field within the broader domain of anthropology) argues that “the mentalist view is unsuitable as a basis for an evolutionary theory of communication, because it presupposes that the ability to attribute mental states precedes the phylogenesis of human communication – which is controversial” (Geurts 2026). By contrast, as said above, I propose that still-prehuman messages, based on a single pre-word, can inform the hearer of the speaker’s false belief. In this way, false belief can be heard. Heard, yes. We all know that the hearer does not merely decode signs: the code-model fails because it demands too little. But this does not imply that the hearer in our example must elaborate their understanding to the point of turning it into: “He/she believes that So-and-so is nearby”. The hearer’s understanding need not be made explicit as a premise in a syllogism. Such explicit formulation, which arises only when the hearer or a third person recounts what happened, requires verbs such as “say” and “believe,” which, in my view, appeared thousands of years after humans first began to produce predications. (Apart from that, I completely agree with Geurts that “we use language to share commitments, which persist by default. This form of communication, which seems to be unique to our species, paved the way for cross-temporal cooperative activities, without which human culture and society would be unthinkable”, ibidem; see also Geurts 2022. Here we can discern his attention to general cognitive anthropology.)
7
‘Bleaching’ is commonly defined as the reduction or loss of specific lexical semantic content as a linguistic form develops more abstract, generalized, or grammatical functions. At least two books should be mentioned: Langacker (1987) and Hopper and Traugott (2003). In original bleaching, the ‘reduction or loss’ is maximal, but the gain—what Langacker calls “schematization”—is in this case literally incalculable, because the consequences of that bleaching continue to unfold.
8
All this –the learning of articulatory-phonetic sequences– also points, of course, to Hockett’s (1960) claim that “the duality of patterning (i.e., ‘the meaningful elements are made up of sequences of elements which are themselves wholly meaningless’) was the last property to be developed, because one can find little if any reason why a communicative system should have this property unless it is highly complicated”. Fleming, 2017 compares that “transition from monoplanar to dually patterned speech” with the transition known to have occurred many millennia later in writing systems. Initially, signs depicted things, and the parts of written signs were therefore not devoid of meaning. I have not studied any of those systems, but I wonder whether the first change in them might have occurred with factual negation (not prohibitive or rejection negation, which can be depicted as a human action, for example by two outstretched arms). Perhaps, factual negation—that is, the meaning ‘there is none’—could not appear in a wholly hieroglyphic system, because a crossing-out mark risked being interpreted as a fence or gate enclosing the object. As observed in the preceding section, negation may have been important in the origin of vocal syntax.
9
Regarding the respective specialization, so to speak, of each cerebral hemisphere, we can draw on Zatorre’s contributions. Across his work he proposed—and in one of his later studies (Albouy et al. 2020) ultimately confirmed experimentally—that the left hemisphere processes temporal information in sounds, which is crucial to articulatory-phonetic sequences, whereas the right hemisphere specializes in spectral information, including melody and pitch as well as recognition of vocal timbre. Gainotti (2026) may also be consulted. He proposes that, beyond the especially pronounced case of the amygdala and fear, “a general right hemisphere dominance for emotions exists in humans,” and assumes that hemispheric differentiation “was necessary to allow the emotional system and the cognitive-linguistic system to operate without interfering with one another”.
10
The idea of a living fossil of the earliest (/of earlier) stages of language goes back to Jackendoff 1999 and also appears in Dediu and Levinson (2013). We mentioned it in the preceding note 5 in connection with Progovac (2026) and her discussion of “(flat) verb–noun compounds, which are typically used for derogatory naming and nicknaming”. (See also Progovac, 2016; Benítez-Burraco & Progovac, 2024).
11
Jon-And & Michaud (2026) do not, of course, originate this 'use and generalization' based view of the grammatical categories of words and sentences, but they add two points to it. First, they explain it by applying the “uniquely human capacity for sequence representation” to syntax (and not only, as seen above in Section 5, to articulatory-phonetic sequences) while also introducing, as a second explanatory factor, the capacity for generalization that we share with non-human animals. Second, they have succeeded in getting a model to perform this learning task.
12
Some time ago, I was struck by a behavior often observed in children whose speech remained holophrastic. Having already learned to point with the hand—usually the right—these children reinforced increasingly angry repetitions of an unfulfilled request by moving the entire right arm up and down. An example is ‘¡Allí! (Spanish; in English ‘There!’) when the child asks not to be taken by his mother from the place where he wants to be. In this case, strong prosody that is simultaneously linguistic and emotional is complemented by shoulder movements; unlike the hand, the shoulder is controlled by the ipsilateral hemisphere. This may perhaps help intonation become still more clearly separated, somewhat later in the child’s development, from the incipient learning of cultural motor sequences.
13
The white, horizontally wide sclera of the human eye could serve that function of helping the recipient achieve syntactic integration in gestural predication. By contrast, proposals that such a sclera increases eye-gaze detectability (Kobayashi & Kohshima, 1997) or supports cooperative social interactions of any kind—as proposed by the Cooperative Eye Hypothesis (Tomasello et al., 2007)—have rightly been refuted (see Perea-García et al., 2026).

References

  1. Albouy, Philippe; Benjamin, Lucas; Morillon, Benjamin; Zatorre, Robert. Distinct sensitivity to spectrotemporal modulation supports brain asymmetry for speech and melody. Science 2020, 367, 1043–1047. [Google Scholar] [CrossRef] [PubMed]
  2. Baumard, Nicolas; André, Jean-Baptiste; Sperber, Dan. A mutualistic approach to morality. The evolution of fairness by partner choice. Behav. Brain Sci. 2013, 36, 59–78. [Google Scholar] [CrossRef] [PubMed]
  3. Bejarano, Teresa. Pragmatics and theory of mind: A problem exportable to the origins of language. In Proceedings of Proceedings of the 7th International Conference Evolang. World Scientific This is a conference paper, but, unfortunately, I did not attend., 2008; pp. 18–25. [Google Scholar]
  4. Bejarano, Teresa. Review of “Hurford, The origins of meaning”. Teorema 2010, 29, 157–164. [Google Scholar]
  5. Bejarano, Teresa. The Origin of Human Theory-of-Mind. Humans 2025, 5(1), 1–42. [Google Scholar] [CrossRef]
  6. Benítez-Burraco, Antonio; Progovac, Ljiljana. Syntax and the brain: language evolution as the missing link(ing theory)? Front. Psychol. 2024, 15. [Google Scholar] [CrossRef] [PubMed]
  7. Carpendale, Jeremy; Kettner, Viktoria; Guevara de Haro, Irene. A Genealogy of Gestures: On the Nature and Emergence of Forms of Gestural Communication within Shared Routines. Hum. Dev. 2024, 68(1), 30–43. [Google Scholar] [CrossRef]
  8. Corballis, Michael. From Hand to Mouth: The Origins of Language; Princeton University Press, 2002. [Google Scholar]
  9. Dediu, Dan; Levinson, Stephen. On the antiquity of language: the reinterpretation of Neandertal linguistic capacities and its consequences. Front. Psychol. 2013, 4. [Google Scholar] [CrossRef] [PubMed]
  10. Eco, Umberto. Trattato di semiotica generale; Bompiani: Milán, 1975. [Google Scholar]
  11. Fleming, Luke. Phoneme inventory size and the transition from monoplanar to dually patterned speech. J. Lang. Evol. 2017, 2(1), 52–56. [Google Scholar] [CrossRef]
  12. Fodor, Jerry. The language of thought; Harvard University Press, 1975. [Google Scholar]
  13. Beaney, M. (Ed.) Frege, Gottlob, 1997 (ed. original, 1892). On Sinn and Bedeutung. In The Frege reader; Black, M., Translator; Blackwell; pp. 151–171.
  14. Friederici, Angela; von Cramon, Yves; Kotz, Sonja. Role of the Corpus Callosum in Speech Comprehension: Interfacing Syntax and Prosody. Neuron 2007, 53(1), 135–145. [Google Scholar] [CrossRef] [PubMed]
  15. Gainotti, Guido. Emotions, the amygdala, and the right hemisphere. Brain Res. 2026. [Google Scholar] [CrossRef] [PubMed]
  16. Gázquez, José; Castelain, Thomas; Llorente, Miquel. Language as an evolutionary pressure of human handedness. Acta Psychol. 2026, 264. 106489. [Google Scholar] [CrossRef] [PubMed]
  17. Geurts, Bart. Evolutionary pragmatics: from chimp-style communication to human discourse. J. Pragmat. 2022, 200, 24–34. [Google Scholar] [CrossRef]
  18. Geurts, Bart. Evolutionary pragmatics from a normative point of view. Proc. Evolang 2026, 16, 141–143. [Google Scholar]
  19. Graham, Kirsty; Rossano, Federico; Moore, Richard. The origin of great ape gestural forms. Biol. Rev. Camb. Philos. Soc. 2024, 100, 190–204. [Google Scholar] [CrossRef] [PubMed]
  20. Jackendoff, Ray. Possible stages in the evolution of the language capacity. In Trends in Cognitive Sciences; 1999. [Google Scholar] [CrossRef] [PubMed]
  21. Hockett, Charles. The Origin of Speech. Sci. Am. 1960, 203(3), 88–97. [Google Scholar] [CrossRef]
  22. Hurford, James. The Origins of Meaning; Oxford University Press, 2007. [Google Scholar]
  23. Jon-And, Anna; Jonsson, Markus; Lind, Johan; Enquist, Magnus. The Evolutionary Costs of Sequence Representation: Why Only Humans Developed Language. Protolang conference, 2025. [Google Scholar]
  24. Jon-And, Anna; Michaud, Jérôme. Emergent categories in usage-based language learning. Proceedings of Evolang 16, 2026; pp. 196–203. [Google Scholar]
  25. King, Stephanie; Janik, Vincent. Bottlenose dolphins can use learned vocal labels to address each other. Proc. Natl. Acad. Sci. U S A 2013. [Google Scholar] [CrossRef] [PubMed]
  26. Kobayashi, Hiromi; Kohshima, Shiro. Unique morphology of the human eye and its adaptive meaning. J. Hum. Evol. 2001, 40, 419–435. [Google Scholar] [CrossRef] [PubMed]
  27. Lind, Johan; Jon-And, Anna. A sequence bottleneck for animal intelligence and language? Trends Cogn. Sci. 2025. [Google Scholar] [CrossRef] [PubMed]
  28. Lipschits, Or; Geva, Ronny. An integrative model of parent-infant communication development. Child Dev. Perspect. 2024, 18(3), 137–144. [Google Scholar] [CrossRef]
  29. Lupyan, Gary. From Chair to “Chair”: A Representational Shift Account of Object Labeling Effects on Memory. In Journal of Experimental Psychology: General; 2008. [Google Scholar] [CrossRef] [PubMed]
  30. Lupyan, Gary. The cognitive functions of language. Conference in Evolang 16, 2026. [Google Scholar]
  31. Hopper, Paul; Traugott, Elizabeth. Grammaticalization; Cambridge University Press, 2003. [Google Scholar]
  32. Perea-García, Juan Olvido; Kano, Fumihiro; Vojtech, Sibierska Marta & Fiala; Skrok, Marta K. & Danel Dariusz; Ratajczak, Ewa; Zywiczynski, Przemyslaw; Slawomir, Wacewicz. Context-Dependent Advantages in Primate Eye Morphology. Proc. Evolang 16 2026, 338–340. [Google Scholar]
  33. Placiński, Marek; Żywiczyński, Przemysław; Matzinger, Theresa; Sibierska, Marta; Boruta-Żywiczyńska, Monika; Szala, Anna; Wacewicz, Sławomir. Evolution of pantomime in dyadic interaction. A motion capture study. J. Lang. Evol. 2023, 8, 134–148. [Google Scholar] [CrossRef]
  34. Progovac, Ljiljana. A Gradualist scenario for language evolution: Precise linguistic reconstruction of early human (and Neandertal) grammars. Front. Psychol. 2016, 7, 1714. [Google Scholar] [CrossRef]
  35. Progovac, Ljiljana. Co-evolution of face-recognition and naming: Unifying various functions of the fusiform gyrus brain area. Proc. Evolang 16 2026, 379–387. [Google Scholar]
  36. Langacker, Ronald. Foundations of Cognitive Grammar. Volume I: Theoretical Prerequisites; Stanford University Press, 1987. [Google Scholar]
  37. Sammler, Daniela; Kotz, Sonja; Eckstein, Korinna; Ott, Michael; Friederici, Angela. Prosody meets syntax: The role of the corpus callosum. Brain 2010, 133, 2643–55. [Google Scholar] [CrossRef] [PubMed]
  38. Senghas, A.; Coppola, M. “Children Creating Language: How Nicaraguan Sign Language Acquired a Spatial Grammar”. Psychol. Sci. 2001, 12(4), 323–328. [Google Scholar] [CrossRef] [PubMed]
  39. Senghas, A.; Kita, S.; Özyürek, A. “Children Creating Core Properties of Language: Evidence from an Emerging Sign Language in Nicaragua”. Science 2004, 305(5691), 1779–1782. [Google Scholar] [CrossRef] [PubMed]
  40. Sperber, Dan (Ed.) Anthropology and Psychology: Towards an Epidemiology of Representations. In Sperber, Explaining culture; Blackwell, 1996. [Google Scholar]
  41. Sterelny, Kim. Review of “Explaining Culture: A Naturalistic Approach. Dan Sperber”. Mind 2001, 110, 845–854. [Google Scholar] [CrossRef]
  42. Tomasello, Michael; Hare, Brian; Lehmann, Hagen; Call, Josep. Reliance on head versus eyes in the gaze-following of great apes and human infants: The cooperative eye hypothesis. J. Hum. Evol. 2007, 52, 314–320. [Google Scholar] [CrossRef] [PubMed]
  43. Vygotsky, Lev (Ed.) Thought and Language; MIT Press: Cambridge, MA, 1986. [Google Scholar]
  44. Wacewicz, Slawomir; Żywiczyński, Przemysław. Two types of bodily-mimetic communication. In Perspectives on Pantomime; Żywiczyński, Przemysław, Blomberg, Johan, Boruta-Żywiczyńska, Monika, Eds.; Benjamins, 2024. [Google Scholar] [CrossRef]
  45. Yaman, Anil; Tian, Shen; Lindström, Björn. Semantic knowledge guides innovation and drives cultural evolution. Proc. Natl. Acad. Sci. 2026, 123. [Google Scholar] [CrossRef] [PubMed]
  46. Zatorre, Robert; Gandour, Jackson. “Neural specializations for speech and pitch: moving beyond the dichotomies”. Philos. Trans. R. Soc. B 2008, 363, 1087–1104. [Google Scholar] [CrossRef] [PubMed]
  47. Żywiczyński, Przemysław; Wacewicz, Slawomir; Sibierska, Marta. Defining Pantomime for Language Evolution Research. In Topoi; 2018. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings