Emergent Goal-Anticipatory Gaze in Infants via Event-Predictive Learning and Inference

Saved in:
Bibliographic Details
Title: Emergent Goal-Anticipatory Gaze in Infants via Event-Predictive Learning and Inference
Language: English
Authors: Gumbsch, Christian (ORCID 0000-0003-2741-6551), Adam, Maurits, Elsner, Birgit (ORCID 0000-0003-3441-2436), Butz, Martin V. (ORCID 0000-0002-8120-8537)
Source: Cognitive Science. Aug 2021 45(8).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 26
Publication Date: 2021
Document Type: Journal Articles
Reports - Research
Descriptors: Goal Orientation, Infants, Eye Movements, Cognitive Processes, Prediction, Psychomotor Skills, Familiarity, Learning Processes, Inferences, Comparative Analysis, Models, Infant Behavior
DOI: 10.1111/cogs.13016
ISSN: 1551-6709
Abstract: From about 7 months of age onward, infants start to reliably fixate the goal of an observed action, such as a grasp, before the action is complete. The available research has identified a variety of factors that influence such goal-anticipatory gaze shifts, including the experience with the shown action events and familiarity with the observed agents. However, the underlying cognitive processes are still heavily debated. We propose that our minds (i) tend to structure sensorimotor dynamics into probabilistic, generative event-predictive, and event boundary predictive models, and, meanwhile, (ii) choose actions with the objective to minimize predicted uncertainty. We implement this proposition by means of event-predictive learning and active inference. The implemented learning mechanism induces an inductive, event-predictive bias, thus developing schematic encodings of experienced events and event boundaries. The implemented active inference principle chooses actions by aiming at minimizing expected future uncertainty. We train our system on multiple object-manipulation events. As a result, the generation of goal-anticipatory gaze shifts emerges while learning about object manipulations: the model starts fixating the inferred goal already at the start of an observed event after having sampled some experience with possible events and when a familiar agent (i.e., a hand) is involved. Meanwhile, the model keeps reactively tracking an unfamiliar agent (i.e., a mechanical claw) that is performing the same movement. We qualitatively compare these modeling results to behavioral data of infants and conclude that event-predictive learning combined with active inference may be critical for eliciting goal-anticipatory gaze behavior in infants.
Abstractor: As Provided
Entry Date: 2021
Accession Number: EJ1310726
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGy0G8NcBWW9MEMEqSCrVbzAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDO2xQ1lSV_smj2Eq5AIBEICBm-edVQAJJYOshN5xbKNPKPIUq3r8YSRwaxYmJVGtO8F_yjvxbhHtvySukXOPoVqkTTmMx2fxAgnyaHIqDl9OlsKGUJttxaL-E-aqf41JVVxc3HTiM9_3GWSK-vtr5rJCpctvlMVtv-GYJ60_KbxvAmPLHbxYPmbblkeWMICPv4Hcp9b9mMo2gqK4tRmTAjlE1f46g9RStcUFn0RT
Text:
  Availability: 1
  Value: <anid>AN0152095508;cgn01aug.21;2021Aug28.05:04;v2.2.500</anid> <title id="AN0152095508-1">Emergent Goal‐Anticipatory Gaze in Infants via Event‐Predictive Learning and Inference </title> <p>From about 7 months of age onward, infants start to reliably fixate the goal of an observed action, such as a grasp, before the action is complete. The available research has identified a variety of factors that influence such goal‐anticipatory gaze shifts, including the experience with the shown action events and familiarity with the observed agents. However, the underlying cognitive processes are still heavily debated. We propose that our minds (i) tend to structure sensorimotor dynamics into probabilistic, generative event‐predictive, and event boundary predictive models, and, meanwhile, (ii) choose actions with the objective to minimize predicted uncertainty. We implement this proposition by means of event‐predictive learning and active inference. The implemented learning mechanism induces an inductive, event‐predictive bias, thus developing schematic encodings of experienced events and event boundaries. The implemented active inference principle chooses actions by aiming at minimizing expected future uncertainty. We train our system on multiple object‐manipulation events. As a result, the generation of goal‐anticipatory gaze shifts emerges while learning about object manipulations: the model starts fixating the inferred goal already at the start of an observed event after having sampled some experience with possible events and when a familiar agent (i.e., a hand) is involved. Meanwhile, the model keeps reactively tracking an unfamiliar agent (i.e., a mechanical claw) that is performing the same movement. We qualitatively compare these modeling results to behavioral data of infants and conclude that event‐predictive learning combined with active inference may be critical for eliciting goal‐anticipatory gaze behavior in infants.</p> <p>Keywords: Infancy; Goal‐anticipatory gaze; Computational model; Event cognition; Active inference</p> <hd id="AN0152095508-2">Introduction</hd> <p>Already during the first year of life, infants appear to develop a rudimentary understanding that human actions are directed toward goals. One associated paradigm investigates the development of goal‐anticipatory gaze shifts. In eye‐tracking studies, infants watch video sequences depicting action events, for example, a hand reaching for an object. If an infant looks at the goal of the shown event, such as a to‐be grasped object, before the movement, such as a reach, is completed, the infant successfully anticipated the goal of the event and thus apparently recognized the goal‐directedness of the action. The development of this ability seems to be supported by various factors such as familiarity with an observed event and the involved agent (Cannon & Woodward, 2012; Kanakogi & Itakura, 2011), the motor ability to perform the movement themselves, behavioral cues that indicate agency, and the saliency of the produced effect (Adam & Elsner, 2018; Adam et al., 2017; Kanakogi & Itakura, 2011). Despite the rather large conglomerate of findings, the involved internal representations and computational mechanisms are still mostly unknown and have been characterized only descriptively so far (see, e.g., Gredebäck & Falck‐Ytter, 2015, for a review).</p> <p>In this paper, we propose that goal‐anticipatory gaze shifts emerge in infants from two interplaying factors: (i) internally developing probabilistic generative models of action events and transitions between events, and (ii) the overall "objective" of the brain to minimize uncertainty in its currently activated generative models, that is, in its internal estimates about what is currently happening and what is about to happen in the outside environment.[<reflink idref="bib1" id="ref1">1</reflink>] In support of this proposal, we modeled the emergence of goal‐anticipatory gaze shifts in infants, merging recent insights from event‐predictive cognition and approximate free‐energy minimizing inference. The modeling system first learned about different object interactions in a virtual scenario, such as <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching for an object <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> or <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> transporting an object <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> . For each type of object interaction, it learned distinct event schematic, generative models, which encode the event boundary and the event dynamics in a probabilistic manner. Throughout training, we put our system in experimental conditions, similar to how goal‐anticipatory gaze shifts are tested in infants. The system was shown familiar action events (reaching) performed by agents that typically perform this kind of action (hand) or by unfamiliar agents (mechanical claw). We demonstrate that when the system used active inference, that is, it chose its gaze to minimize predicted uncertainty (Friston et al., 2015, 2016), the system showed similar gaze behavior as previously found in infants. Moreover, we analyze how experience with the events affected the gaze behavior, and we qualitatively compare these modeling results to behavioral data of different age groups. Seeing the closely, qualitatively fitting results, we hypothesize that event‐predictive, inductive learning biases combined with active inference principles appear to be highly important for enabling the inference of goals while observing others interacting with their environment.</p> <p>The remainder of this paper is structured as follows. We first give an overview on goal‐anticipatory gaze shifts and motivate our generative, event‐predictive modeling approach. Next we provide the algorithmic details of our learning and inference model. Section 5 evaluates the model. We summarize and conclude with a discussion of the results and its implications.</p> <hd id="AN0152095508-3">Goal‐anticipatory gaze behavior in infants</hd> <p>Looking behavior is one of the first behaviors in human development, and as such, has been widely used to investigate how infants form expectations about observed goal‐directed actions (e.g., Fantz, 1958; Gredebäck, Johnson, & Hofsten, 2010). A classical measure to capture these expectations is the assessment of infants' looking times to a still image of a goal state after an action has been completed (e.g., Woodward, 1998). However, because looking times are usually measured after the action has been completed, and are measured with low spatial and temporal resolution (Daum, Attig, Gunawan, Prinz, & Gredebäck, 2012), possible expectations that the infants might have formed can only be inferred post hoc. The focus of the present paper is on explaining whether and how infants of certain age form goal expectations during the observation of an action that is still unfolding. The production of goal‐anticipatory gaze shifts is cognitively demanding (Gredebäck & Falck‐Ytter, 2015) and requires measures with high spatial and temporal resolution, such as eye‐tracking technology. The research paradigm is mainly based on a seminal eye‐tracking study in which adult participants tended to shift their gaze to a to‐be‐attained goal before the goal was actually accomplished in a block‐stacking task (Flanagan & Johansson, 2003). Additionally, such anticipatory gaze behavior occurred when participants performed the task themselves and when they observed someone else performing it. This was taken as evidence for the so‐called direct‐matching hypothesis: observers visually track ongoing actions based on not only visual features but also on motor‐grounded predictive encodings of the perceived actions (e.g., Gallese, Fadiga, Fogassi, & Rizzolatti, 1996; Rizzolatti, Fadiga, Gallese, & Fogassi, 1996; Rizzolatti, Fogassi, & Gallese, 2001). The motor representations are thought to include instructions for the visual system to produce anticipatory gaze shifts, which result in similar gaze behavior during action execution and observation (e.g., Gredebäck & Falck‐Ytter, 2015).</p> <p>In a study supporting the idea of a direct‐matching process during infants' action observation, 6‐ and 12‐month‐old infants as well as adults observed how toys moved into a container, either being grasped and transported by a human or floating into the container without observable cause (Falck‐Ytter, Gredebäck, & Hofsten, 2006). The data revealed anticipatory gaze behavior in the human agent condition for the 12‐month‐olds and the adults, but not in the self‐propelled condition or for the 6‐month‐olds. Building on these results, several studies replicated and extended these findings. From about 7 months of age onward, infants show goal‐anticipatory gaze shifts when observing simple human grasping actions, while they keep reactively tracking unfamiliar back‐of‐hand actions or unfamiliar agents such as mechanical claws or self‐propelled spoons (Adam et al., 2016; Cannon & Woodward, 2012; Gredebäck & Melinder, 2010; Kanakogi & Itakura, 2011; Krogh‐Jespersen & Woodward, 2014). Moreover, infants' motor abilities and the amount of experience with certain actions are positively correlated with the infants' goal anticipations during observation of the respective actions (Cannon et al., 2012; Gredebäck & Melinder, 2010; Kanakogi & Itakura, 2011).</p> <p>As an interpretation for the aforementioned results, Elsner and Adam (2021) argue that actions can be perceived as events with a threefold structure. For example, in an unfolding grasping event, an observer will first see the initial phase of the action, including visual features of the agent and the goal object. Then, the grasping event will move into a dynamic phase, where the movement of the agent toward the goal can be perceived. Finally, the grasping event will move into its end phase or end state, where the agent typically arrives at the goal object and manipulates it. This then concludes the grasping event and starts the next observable event.</p> <p>According to Elsner and Adam (2021), an observer will draw on the available information from these different phases to reason about the goal‐directedness of the observed event. During the initial phase and the dynamic phase, only incomplete information about the unfolding event is available because the goal has not yet been achieved. Therefore, in order for the observer to correctly identify the goal ahead of time, they have to draw on stored top‐down event knowledge that is based on prior experience with similar actions or the same action as the one that is currently observed. If, however, no such top‐down knowledge is available, the observer can also draw on observable bottom‐up information from the action event itself. This entails, for example, the agent's behavior because the agent might act in ways or manipulate the goal object in ways that help the observer to reason about the action goal and to also store this new knowledge as top‐down information for future observations. This line of reasoning fits well to prior research showing that infants tend to anticipate goals more successfully when they have prior experience with the agent or the action, but are also able to predict the goal of unfamiliar agents or actions when bottom‐up information is provided during action observation (e.g., Adam & Elsner, 2018; Adam et al., 2017; Biro, 2013).</p> <hd id="AN0152095508-4">A developing event‐predictive inference model</hd> <p>Recent theories from different subfields of cognitive science suggest that humans tend to organize their sensorimotor experience by means of hierarchically structured, event encodings (Butz, 2016, 2017; Butz, Achimova, Bilkey, & Knott, 2021; Zacks, Speer, Swallow, Braver, & Reynolds, 2007; Zacks & Tversky, 2001) A crucial characteristic of an event is that it is perceived to have a beginning and an end, that is, an <emph>event boundary</emph> (Zacks & Tversky, 2001). In between two event boundaries, the event unfolds relatively uniformly and, in principle, predictably (Zacks, Speer, Swallow, Braver, & Reynolds, 2007). For instance, a <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> event typically starts with the initiation of an arm movement, while it ends when the hand has reached the object, grasping it. Between these boundaries, the hand typically follows a straight trajectory toward the object. Evidence for such event representations stems from different disciplines and can be found on different levels of processing, ranging from sensorimotor activations to semantic and linguistic representations (Baldwin & Kosie, 2021; Butz, Achimova, Bilkey, & Knott, 2021; Cooper, 2021; Franklin, Norman, Ranganath, Zacks, & Gershman, 2020; Kuperberg, 2021).</p> <hd id="AN0152095508-5">Generative, event‐predictive learning, and inference</hd> <p>From a behavioral perspective, the theory of event coding (TEC; Hommel, 2015; Hommel, Müsseler, Aschersleben, & Prinz, 2001; Hommel, 2009) suggests that actions and their effects are encoded in a common format, that is, <emph>event codes</emph>. Event codes develop from action‐effect learning, which initially starts with reflex‐like, behavioral exploration (Elsner & Hommel, 2001; Hommel, 2015). At a later stage of behavioral learning they can be used to infer anticipatory behavior: By considering a desired action effect, the agent activates the linked event code, which automatically also activates the associated motor activity (Elsner & Hommel, 2001). According to TEC, these action effects are not restricted to invoking own behavior, but can also be utilized for action understanding, in which case an action observation activates the own event production code.</p> <p>From an observational perspective, event segmentation theory (EST; Zacks, Speer, Swallow, Braver, & Reynolds, 2007; Zacks & Swallow, 2007) is based on the evidence that humans tend to consistently segment perceived streams of information, exhibiting good agreement among each other about when event boundaries occur (Zacks & Tversky, 2001). Event boundaries essentially correspond to significant changes in the unfolding interaction dynamics, which often correspond to subgoals in environmental interactions (Zacks, Speer, Swallow, Braver, & Reynolds, 2007). According to EST, this shared perception of distinct events is the result of internal <emph>event models</emph> that guide human perceptual processing (Radvansky & Zacks, 2014; Zacks, Speer, Swallow, Braver, & Reynolds, 2007). Event models are hierarchically organized, generative models that encode entities (such as agents and patients), their sequence of actions, and the resulting consequences in a spatiotemporal framework (Radvansky & Zacks, 2014; Stawarczyk, Bezdek, & Zacks, 2021). During the perception of an ongoing event, a subset of event models is active, predicting how the current events are going to unfold, and what will likely be perceived next (Franklin, Norman, Ranganath, Zacks, & Gershman, 2020; Zacks, Speer, Swallow, Braver, & Reynolds, 2007). At event boundaries, current event‐respective predictions will produce large transient errors, entailing significant changes of the activated event‐predictive models.</p> <p>In sum, humans appear to perceive temporal activity in terms of events. The outlined theories imply particular properties of the underlying event‐predictive encoding schemata:</p> <p></p> <ulist> <item> <emph>Event dynamics model</emph> : There is general agreement that event encodings are generative, that is, they model how an event typically unfolds, predicting perceptual information dynamics and correlating those with dynamics‐influencing motor activities.</item> <p></p> <item> <emph>End condition</emph> : EST additionally implies that event schemata predict the typical end‐effects, which may be equated with final behavioral action effects (e.g., holding an object), which fittingly can serve as desired goals or subgoals.</item> <p></p> <item> <emph>Start condition</emph> : Finally, event schemata are thought to specify particular preconditions necessary for an event to commence.</item> </ulist> <p>It has been shown that such event‐predictive encodings can be learned in a dedicated manner, when endowing the learning system with suitable inductive event‐predictive biases (Butz, Bilkey, Humaidan, Knott, & Otte, 2019; Franklin, Norman, Ranganath, Zacks, & Gershman, 2020; Gumbsch, Butz, & Martius, 2019). Here, we train separate event‐predictive models for individual interaction events as well as for transitions from one event to another.</p> <p>Our generative event‐predictive modeling perspective can be seamlessly merged with the free energy and active inference principles of cognition (Butz, 2016; Friston, 2010; Friston et al., 2015, 2016). It can also be closely linked to both Bayesian filtering (Knill & Pouget, 2004) and planning as inference (Botvinick & Toussaint, 2012). Accordingly, event‐predictive encodings are inherently probabilistic and generative. While processing sensorimotor information, free energy minimization causes event‐predictive activities to adapt, minimizing the deviance between predictions and encountered sensorimotor dynamics. In effect, bottom‐up and top‐down information are probabilistically integrated, dynamically maintaining, activating, and de‐activating currently applicable event encodings. Meanwhile, active inference generates behavior that attempts to minimize the uncertainty within and across current and subsequent events.</p> <hd id="AN0152095508-6">Inferring goal‐predictive gaze via developing event encodings</hd> <p>We propose a generative, event‐predictive learning and inference model, which may explain a variety of experimental findings on the development of goal‐anticipatory gaze behavior in infants, illustrated in Fig. 1. Anticipatory gaze behavior is modeled for a study scenario in which infants of different ages repeatedly observe a simple goal‐directed reach‐and‐grasp event. Our model assumes that infants younger than 6 months do not have a well enough learned event encoding for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> . As a result, they keep their gaze on the moving agent, thereby tracking the hand's movement to gain more information about the agent's future position. From about 7 months onward, when infants start to show goal‐anticipatory gaze, we hypothesize that they have developed a sufficiently well‐predicting generative event model. Visual cues, such as the appearance of the hand as well as movement‐based information, invoke the activation of the event schema for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> . The associated encoding of the end condition will predict that the hand will end up at the position of a target object. The 12‐month‐old infant will thus anticipatorily look at the target object in order to reduce uncertainty about when, where, and how exactly the observed <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> event will end. On the other hand, when 12‐month‐old infants observe a reaching movement by an unknown agent, for example, a mechanical claw, that does not exhibit any agency‐related cues, the infants tend to reactively track the claw (Adam et al., 2017). Our model assumes that, even though these 12‐month‐olds have learned an event‐generative model for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , this event schema will not be activated because some of the associated start conditions are not met.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0001.jpg" title="1 Illustrations of our model's hypotheses about experimental findings on goal prediction for (a) infants younger than 6 months, (b) for 12‐month‐old infants watching reaching motions done by hands, and (c) for 12‐month‐old infants watching reaching motions by mechanical claws. White eye symbols visualize gaze, colored circles visualize position predictions and their confidence (blue for the reaching hand, red for the target position). The screenshots are taken from Adam et al. (2016)." /> </p> <p></p> <p>In this paper, we investigate the validity of our theoretical considerations by implementing and testing the proposed computational model. Our implemented system, which we term cognitive action prediction model in infants (CAPRI), learns schematic, generative event encodings for different interaction events such as <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching for an object <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> transporting an object <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , etc. Each event is encoded by a simple generative model that predicts likelihood distributions over expected future observations. Meanwhile, the system continuously attempts to decrease uncertainty about the environment. It essentially attempts to infer which event is unfolding and which future events and event boundaries are likely to occur next. As a result, once particular events have been learned sufficiently well, the system, driven by its aim to reduce uncertainty, begins to direct its gaze in an anticipatory, information gain oriented manner much like the goal‐anticipatory gaze shifts generated by infants at different ages.</p> <hd id="AN0152095508-8">Cognitive action prediction model in infants</hd> <p>To evaluate our computational assumptions by means of a concrete free energy‐based inference formalization and an actual first implementation of the formalism, we assume that we have a system that "lives" in a virtual world. The system interacts with its world in every time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> by, first, receiving an observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> and, second, performing an action according to a policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> , which activates particular motor behavior, such as fixating a particular location in space or moving the hand toward an object. In our scenario, where the agent is an observer, policies correspond solely to gaze behavior.</p> <hd id="AN0152095508-9">Learning of event schemata</hd> <p>CAPRI learns event schemata that consist of three encodings: a <emph>start condition</emph>, the <emph>event dynamics model</emph>, and an <emph>end condition</emph>. All components are encoded as probabilistic models, that is, in the form of likelihood distributions.</p> <p>The <emph>start condition</emph><ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mtext>start</mtext></msubsup></math> </ephtml> models the likelihood <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>start</mi></msubsup><mrow><mo>(</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math> </ephtml> , with the observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"><mi>o</mi></math> </ephtml> , policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> , and time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> . For example, a <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> event might typically start with an observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> that contains a hand that just started to move toward a reachable object (cf. Fig. 2a). This encoded probability density of the corresponding observation will depend on the gaze location <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> .</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0002.jpg" title="2 Illustrations of two event schemata. An event schema is composed of three components: a start condition, an event dynamics model, and an end condition. The graphs exemplarily specifies potential encodings of two events: reaching (a) and falling (b)." /> </p> <p></p> <p>The <emph>event dynamics model</emph><ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> encodes the likelihood of an observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> given the last observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> and previously used policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> , that is, eye gaze location: <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mtext>event</mtext></msubsup><mrow><mo>(</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math> </ephtml> . For example, during a <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> event the position of the hand at a certain time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> , described by observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> , is expected to be closer to the object than during the preceding time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></math> </ephtml> , described by observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> (Fig. 2a).</p> <p>The <emph>end condition</emph><ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mtext>end</mtext></msubsup></math> </ephtml> encodes the likelihood <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>end</mi></msubsup><mrow><mo>(</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mi>κ</mi><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math> </ephtml> , with the observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"><mi>o</mi></math> </ephtml> , policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> , time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> , and a retrospective time horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>κ</mi><mo>≥</mo><mn>1</mn></mrow></math> </ephtml> . Thus, the end condition <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>end</mi></msubsup></math> </ephtml> models the likelihood that an observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> occurs at the end of event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> given the last policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> and some previous observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mi>κ</mi><mo>)</mo></mrow></math> </ephtml> , which lies <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"><mi>κ</mi></math> </ephtml> time steps in the past. For our reaching example (Fig. 2a), this means that at the end of a <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> reaching <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> movement at time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> the hand position, captured by the observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0057" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> , is at the same location as the reachable object, which can be predicted by some previous observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0058" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mi>κ</mi><mo>)</mo></mrow></math> </ephtml> .</p> <p>Thus, for every event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0059" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> CAPRI learns three separate likelihood distributions over sensory space. All distributions are modeled as multivariate Gaussians, which the system learns to encode by means of mixture density networks (MDNs) (MDNs; Bishop, 2006; details in Supporting Information Section 8.2). During training, which is composed of multiple episodes <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0060" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi></math> </ephtml> , the system learns MDNs in a supervised manner. Each episode <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0061" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi></math> </ephtml> can consist of one or multiple events, that is, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>E</mi><mo>=</mo><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mo>,</mo><msub><mi>e</mi><mi>j</mi></msub><mo>,</mo><mi>...</mi><mo>,</mo><msub><mi>e</mi><mi>z</mi></msub><mo>)</mo></mrow></math> </ephtml> . When an event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> starts at time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0064" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>t</mi><mn>0</mn></msub></math> </ephtml> , the starting condition <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>start</mi></msubsup></math> </ephtml> is updated using <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><msub><mi>t</mi><mn>0</mn></msub><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> as an input and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0067" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><msub><mi>t</mi><mn>0</mn></msub><mo>)</mo></mrow></math> </ephtml> as the nominal output. At every time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0068" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> during event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0069" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> , the event dynamics model <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0070" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>event</mi></msubsup></math> </ephtml> is updated using <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0071" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0072" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> as an input and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0073" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> as the nominal output. When an event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0074" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> ends at time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0075" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>t</mi><mi>τ</mi></msub></math> </ephtml> , the end condition <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0076" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>end</mi></msubsup></math> </ephtml> is updated <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0077" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>−</mo><mn>1</mn></mrow></math> </ephtml> times. <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0078" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>end</mi></msubsup></math> </ephtml> is updated using <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0079" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><msub><mi>t</mi><mi>τ</mi></msub><mo>)</mo></mrow></math> </ephtml> as nominal output and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0080" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><msub><mi>t</mi><mi>τ</mi></msub><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0081" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><msub><mi>t</mi><mi>κ</mi></msub><mo>)</mo></mrow></math> </ephtml> as inputs with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0082" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>t</mi><mi>κ</mi></msub><mo>∈</mo><mrow><mo>[</mo><msub><mi>t</mi><mn>0</mn></msub><mo>,</mo><mi>...</mi><mo>,</mo><msub><mi>t</mi><mi>τ</mi></msub><mo>−</mo><mn>1</mn><mo>]</mo></mrow></mrow></math> </ephtml> . These multiple model updates do not only increase the training data for the end condition, but also allow CAPRI to learn to generally encode how an event ends, given any observation that had previously occurred during this event.</p> <p>After successful training, CAPRI is able to predict the likelihood of an observation occurring at the beginning, end, or during an event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0083" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> . Let us assume that the system has two event models <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0084" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>fall</mtext></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0085" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>reach</mtext></msub></math> </ephtml> , encoding <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0086" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> An object is falling <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0087" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0088" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> A hand is reaching for an object <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0089" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , respectively (see Fig. 2). When the system perceives <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0090" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> a ball that just rolled over the edge of a table <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0091" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , the start condition <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0092" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>fall</mi></msub><mi>start</mi></msubsup></math> </ephtml> would predict a high likelihood for this particular observation. On the other hand, when the system observes <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0093" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> the ball hitting the floor <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0094" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> , the event end condition <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0095" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>fall</mi></msub><mi>end</mi></msubsup></math> </ephtml> would give a high likelihood for this observation. Thus, the likelihood estimates can be used to infer the probability of an event in the absence of explicit labels about the ongoing events.</p> <hd id="AN0152095508-11">Event‐predictive inference</hd> <p>During <emph>training</emph> CAPRI receives supervised information about which event is currently unfolding. In contrast, during <emph>testing</emph> CAPRI infers at each point in time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0096" xmlns="http://www.w3.org/1998/Math/MathML"><mi>i</mi></math> </ephtml> probabilities about which event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0097" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> currently unfolds given the available sensorimotor information, that is, the sequence of all previous observations <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0098" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>O</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>=</mo><mo>(</mo><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>,</mo><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo><mo>,</mo><mi>...</mi><mo>,</mo><mi>o</mi><mo>(</mo><mn>0</mn><mo>)</mo><mo>)</mo></mrow></math> </ephtml> and all performed policies <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0099" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi mathvariant="normal">Π</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>=</mo><mo>(</mo><mi>π</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>,</mo><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo><mo>,</mo><mi>...</mi><mo>,</mo><mi>π</mi><mo>(</mo><mn>0</mn><mo>)</mo><mo>)</mo></mrow></math> </ephtml> . That is, the system infers <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0100" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>)</mo></mrow></math> </ephtml> for every possible event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0101" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> . This event inference is performed iteratively every time step after executing an action based on policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0102" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></math> </ephtml> and receiving a new observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0103" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> via:</p> <olist> <item> <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0104" xmlns="http://www.w3.org/1998/Math/MathML">Pei(t)|O(t),Π(t)=∑ejP(ei(t)|o(t),o(t−1),π(t−1),ej(t−1))·P(ej(t−1)|O(t−1),Π(t−1)).</math> </ephtml> </item> </olist> <p>Hence, at every time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0105" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> , the system computes <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0106" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></math> </ephtml> for every combination of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0107" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0108" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>j</mi></msub></math> </ephtml> to update the event probabilities. We compute <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0109" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></math> </ephtml> as</p> <p>2 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0110" xmlns="http://www.w3.org/1998/Math/MathML">Pei(t)|o(t),o(t−1),π(t−1),ej(t−1)=Po(t)|o(t−1),π(t−1),ei(t),ej(t−1)·Pei(t)|ej(t−1)∑ehPo(t)|o(t−1),π(t−1),eh(t),ej(t−1)·Peh(t)|ej(t−1).</math> </ephtml></p> <p>The full derivation and the underlying assumptions can be found in Section 8.8. We set the event transition prior <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0111" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo><mo>=</mo><mn>0.9</mn></mrow></math> </ephtml> for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0112" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mo>=</mo><msub><mi>e</mi><mi>j</mi></msub></mrow></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0113" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mrow><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow><mo>=</mo><mfrac><mn>0.1</mn><mi>n</mi></mfrac></mrow></math> </ephtml> for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0114" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mo>≠</mo><msub><mi>e</mi><mi>j</mi></msub></mrow></math> </ephtml> , where <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0115" xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi></math> </ephtml> is the number of available events. This corresponds to the assumption that once the system is in one event, it typically stays in the same event with a probability of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0116" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>90</mn><mo>%</mo></mrow></math> </ephtml> .</p> <p>To compute the Bayesian posterior <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0117" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>)</mo></mrow></math> </ephtml> , we use the event schemata distributions, which were outlined in Section 4.1. There are two cases for computing this likelihood depending on <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0118" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow></mrow></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0119" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>j</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></mrow></math> </ephtml> : Either the world remains in the same event or an event transition has happened from time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0120" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></math> </ephtml> to time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0121" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> . When remaining in the same event, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0122" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mo>=</mo><msub><mi>e</mi><mi>j</mi></msub></mrow></math> </ephtml> and the likelihood of observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0123" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> can be computed by</p> <p>3 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0124" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mfenced separators="" open="(" close=")"><mrow><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow><mo>,</mo></mrow><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mrow></mfenced><mo>=</mo><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>event</mi></msubsup><mfenced separators="" open="(" close=")"><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>|</mo><mi>o</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo><mo>,</mo><mi>π</mi><mo>(</mo><mi>t</mi><mo>−</mo><mn>1</mn><mo>)</mo></mfenced><mo>,</mo></mrow></math> </ephtml></p> <p>with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0125" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>event</mi></msubsup></math> </ephtml> encoding the event‐respective dynamics distribution (see Section 4.1).</p> <p>On the other hand, when <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0126" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mo>≠</mo><msub><mi>e</mi><mi>j</mi></msub></mrow></math> </ephtml> , the likelihood needs to be computed by:</p> <p>4 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0127" xmlns="http://www.w3.org/1998/Math/MathML">Po(t)|o(t−1),π(t−1),ei(t),ej(t−1)=Peistarto(t)|π(t−1))·Pejend(o(t)|o(t−1),π(t−1),</math> </ephtml></p> <p>using the start and end conditions <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0128" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>P</mi><mi>start</mi></msup></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0129" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>P</mi><mi>end</mi></msup></math> </ephtml> (see Section 4.1). This corresponds to the idea, that the likelihood of the observation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0130" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> during an event transition from <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0131" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>j</mi></msub></math> </ephtml> to <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0132" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> is composed of the likelihood of the end condition of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0133" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>j</mi></msub></math> </ephtml> and the start condition of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0134" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> being satisfied.</p> <p>The process of event inference is initialized at time <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0135" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>=</mo><mn>0</mn></mrow></math> </ephtml> with</p> <p>5 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0136" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mfenced separators="" open="(" close=")"><msub><mi>e</mi><mi>i</mi></msub><mrow><mrow><mo>(</mo><mn>0</mn><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mn>0</mn><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mn>0</mn><mo>)</mo></mrow></mrow></mfenced><mo>=</mo><mfrac><mrow><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>start</mi></msubsup><mfenced separators="" open="(" close=")"><mi>o</mi><mo>(</mo><mn>0</mn><mo>)</mo><mo>|</mo><mi>π</mi><mo>(</mo><mn>0</mn><mo>)</mo></mfenced></mrow><mrow><msub><mo>∑</mo><msub><mi>e</mi><mi>h</mi></msub></msub><msubsup><mi>P</mi><msub><mi>e</mi><mi>h</mi></msub><mi>start</mi></msubsup><mfenced separators="" open="(" close=")"><mi>o</mi><mo>(</mo><mn>0</mn><mo>)</mo><mo>|</mo><mi>π</mi><mo>(</mo><mn>0</mn><mo>)</mo></mfenced></mrow></mfrac><mo>,</mo></mrow></math> </ephtml></p> <p>with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0137" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mi>i</mi></msub><mi>start</mi></msubsup></math> </ephtml> referring to the start condition of event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0138" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> (see Section 4.1).</p> <hd id="AN0152095508-12">Active inference</hd> <p>Event inference essentially infers probabilistically which event is currently unfolding using the available, previously learned event schemata and the available sensory information. However, sensory information is not some external, unchangeable quantity simply handed to the system. Rather, CAPRI can actively interact with the world by choosing policies that generate actions, which, in turn, affect the next sensory observation. We follow (Friston et al., 2015) to compute expected free energy for every policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0139" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> , as[<reflink idref="bib2" id="ref2">2</reflink>]</p> <p>6 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0140" xmlns="http://www.w3.org/1998/Math/MathML">FÊ(π,t)=Dτ,eiPo(τ)|ei(τ),πPei(τ)|O(τ),Π(τ)||Po(τ)|m(τ)︸predicteddivergencefromdesiredstates+Eτ,eiHPo(τ)|ei(τ),πPei(τ)|O(τ),Π(τ)︸predicteduncertainty,</math> </ephtml></p> <p>with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0141" xmlns="http://www.w3.org/1998/Math/MathML"><mi>D</mi></math> </ephtml> the Kullback–Leibler (KL) divergence, serving as a distance measure between two probability distributions, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0142" xmlns="http://www.w3.org/1998/Math/MathML"><mi>m</mi></math> </ephtml> a generative model of internal state preferences, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0143" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi></math> </ephtml> the expectation, and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0144" xmlns="http://www.w3.org/1998/Math/MathML"><mi>H</mi></math> </ephtml> entropy. The expected free energy is computed for a time horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0145" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> expanding from the present into the future. Because CAPRI perceives the unfolding sensory information in terms of discrete events, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0146" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> encodes the number of future event boundaries that are considered for estimating future free energy.</p> <p>Predicted KL divergence <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0147" xmlns="http://www.w3.org/1998/Math/MathML"><mi>D</mi></math> </ephtml> from desired states is computed based on a motivational generative model <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0148" xmlns="http://www.w3.org/1998/Math/MathML"><mi>m</mi></math> </ephtml> , which encodes internal state preferences. <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0149" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>|</mo><mi>m</mi><mo>(</mo><mi>τ</mi><mo>)</mo><mo>)</mo></mrow></math> </ephtml> specifies an according distribution over observations for particular, current state preferences <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0150" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>m</mi><mo>(</mo><mi>τ</mi><mo>)</mo></mrow></math> </ephtml> , which are often equated with states of internal homeostasis. This distribution is compared with the expected observations given the system's estimation of events. As a result, minimizing this term infers behavioral policies that are goal‐directed, attempting to generate observations that are in agreement with the desired ones over the specific temporal horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0151" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> . Let us, for example, assume that we have a system that can only interact with the world by looking at different positions. Hence, its policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0152" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> encodes different gaze positions. Let us further assume that the system currently would like to look at teddy bears. This desire is encoded in <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0153" xmlns="http://www.w3.org/1998/Math/MathML"><mi>m</mi></math> </ephtml> . Hence, the distribution of desired observations <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0154" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><mi>o</mi><mo>(</mo><mi>τ</mi><mo>)</mo><mo>|</mo><mi>m</mi><mo>(</mo><mi>τ</mi><mo>)</mo><mo>)</mo></mrow></math> </ephtml> describes visual images of teddy bears. To minimize the divergence between expected and desired observations, the system will choose a policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0155" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> via which it believes to see teddy bears, as shown in Fig. 3a.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0003.jpg" title="3 Illustration of active inference to minimize expected free energy for different policies. In this scene the system can only shift its gaze. Thus, policies π encode different gaze positions, which are illustrated by the white eye symbol. (a) Minimizing predicted divergence from desired states, when m desires visual images of teddy bears. (b) Minimizing predicted uncertainty for τ={0,1,2} when a ball unexpectedly starts to roll. Gaussian bell curves illustrate P(o(t)|ei(t),π) for o describing ball positions." /> </p> <p></p> <p>The second term of Eq. 6, partially competing with the first term, describes predicted uncertainty about future observations in the form of expected entropy, given that policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0161" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>τ</mi><mo>)</mo></mrow></math> </ephtml> is followed. Choosing a policy that minimizes the second term results in behavior that attempts to maximize certainty about future observations. Consider the scenario shown in Fig. 3b. Assuming again that we have a system whose policies correspond to different gaze position, let us assume that our system currently looks at its beloved teddy bear. Meanwhile, it assumes that all other objects and entities in its environment are in the <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0162" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> staying still and doing nothing <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0163" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> event, which we denote by <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0164" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>still</mtext></msub></math> </ephtml> . Suddenly a ball on a table in the periphery seems to move. This is highly unexpected since the system assumed that the ball was in event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0165" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>still</mtext></msub></math> </ephtml> . Since the system does not know in which state <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0166" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>e</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> the ball is in, the expected entropy over all possible events <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0167" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0168" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>E</mi><mrow><mi>τ</mi><mo>,</mo><msub><mi>e</mi><mi>i</mi></msub></mrow></msub><mi>P</mi><mrow><mo>(</mo><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi>π</mi><mo>)</mo></mrow><mi>P</mi><mrow><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>)</mo></mrow></mrow></math> </ephtml> is high. By looking at the ball, the system can perceive the most information about the current ball event. Hence, the system would look at the ball because it expects maximum information gain, that is, maximum entropy decrease.</p> <p>However, it is not enough for the system to know the current event of the ball, which may be denoted by <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0169" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>rolling</mtext></msub></math> </ephtml> . Eq. 6 implies that the system will infer actions from which it expects to minimize uncertainty about the ball behavior in the near future, considering a temporal horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0170" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> . The system can use <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0171" xmlns="http://www.w3.org/1998/Math/MathML"><msubsup><mi>P</mi><msub><mi>e</mi><mtext>rolling</mtext></msub><mi>end</mi></msubsup></math> </ephtml> to infer that this event will likely end at the edge of the table. Thus, the system will tend to look at the edge of the table next to minimize uncertainty about where and when the ball will go from <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0172" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> rolling <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0173" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> to <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0174" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> falling <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0175" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> . Given an even deeper temporal horizon and sufficiently accurate event model beliefs, the system may next look at the floor besides the table to know where and when the ball will go from <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0176" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> falling <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0177" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> to another event, such as <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0178" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> bouncing <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0179" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> . Hence, CAPRI's objective to minimize expected future uncertainty can cause anticipatory gaze shifts.</p> <p>Since we are interested in gaze and assume that attractiveness of the visual stimuli is well controlled and thus approximately uniform in the modeled experiments, CAPRI focuses only on optimizing its policy, that is, its gaze locations, for minimizing predicted uncertainty:</p> <p>7 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0180" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mrow><mi>F</mi><mi>E</mi></mrow><mo>̂</mo></mover><mrow><mo>(</mo><mi>π</mi><mo>,</mo><mi>t</mi><mo>)</mo></mrow><mo>=</mo><msub><mi>E</mi><mrow><mi>τ</mi><mo>,</mo><msub><mi>e</mi><mi>i</mi></msub></mrow></msub><mfenced separators="" open="[" close="]"><mi>H</mi><mfenced separators="" open="[" close="]"><mi>P</mi><mfenced separators="" open="(" close=")"><mrow><mi>o</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo></mrow><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi>π</mi></mfenced></mfenced><mi>P</mi><mfenced separators="" open="(" close=")"><msub><mi>e</mi><mi>i</mi></msub><mrow><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow></mrow></mfenced></mfenced><mo>.</mo></mrow></math> </ephtml></p> <p>CAPRI thus chooses the policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0181" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> for which it expects to minimize this free energy term, based on the active inference principle as follows:</p> <p>8 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0182" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>=</mo><mo form="prefix">arg</mo><munder><mi>min</mi><msub><mi>π</mi><mi>i</mi></msub></munder><mover accent="true"><mrow><mi>F</mi><mi>E</mi></mrow><mo>̂</mo></mover><mrow><mo>(</mo><msub><mi>π</mi><mi>i</mi></msub><mo>,</mo><mi>t</mi><mo>)</mo></mrow><mo>.</mo></mrow></math> </ephtml></p> <p>For <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0183" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>=</mo><mn>0</mn></mrow></math> </ephtml> this means that CAPRI does not consider future events and it only "wants" to minimize the entropy of the currently active event dynamics models:</p> <p>9 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0184" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mrow><mi>F</mi><mi>E</mi></mrow><mo>̂</mo></mover><mrow><mo>(</mo><mi>π</mi><mo>,</mo><mi>t</mi><mo>)</mo></mrow><mo>=</mo><munder><mo>∑</mo><msub><mi>e</mi><mi>j</mi></msub></munder><mi>H</mi><mfenced separators="" open="[" close="]"><msubsup><mi>P</mi><msub><mi>e</mi><mi>j</mi></msub><mi>event</mi></msubsup><mfenced separators="" open="(" close=")"><mi>o</mi><mo>(</mo><mi>t</mi><mo>+</mo><mn>1</mn><mo>)</mo><mo>|</mo><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo><mo>,</mo><mi>π</mi></mfenced></mfenced><mi>P</mi><mfenced separators="" open="(" close=")"><msub><mi>e</mi><mi>j</mi></msub><mrow><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow></mrow></mfenced><mo>.</mo></mrow></math> </ephtml></p> <p>For <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0185" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>=</mo><mn>1</mn></mrow></math> </ephtml> , CAPRI not only focuses on the present but also considers next possible event transitions. This means that the system additionally attempts to reduce the uncertainty about the next event boundary, encoded by the event end‐ and start conditions:</p> <p>10 <ephtml> <math display="block" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0186" xmlns="http://www.w3.org/1998/Math/MathML">FÊ(π,t)=∑ejHPejevento(t+1)|o(t),πPej(t)|O(t),Π(t)+∑ej∑ehHPehstarto(t′)|πPejendo(t′)|o(t),π·Pej(t)|O(t),Π(t).</math> </ephtml></p> <p>When minimizing this quantity, CAPRI attempts to minimize predicted entropy over the currently estimated event dynamics (as in Eq. 9) and over all possible next event boundaries. For <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0187" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>=</mo><mn>2</mn></mrow></math> </ephtml> another term can be added that additionally considers one more event boundary into the future, etc.</p> <hd id="AN0152095508-14">CAPRI model evaluation</hd> <p>We evaluate our CAPRI implementation in a simple three‐dimensional agent–patient interaction simulation, mimicking infants' probable action experiences. In the following, we first detail the simulation environment and the supervised learning procedure. We then detail the behavior of the system and show that goal‐anticipatory gaze behavior indeed selectively emerges over the course of learning.</p> <hd id="AN0152095508-15">Simulation setup and learning procedure</hd> <p>Our simulation always contains two entities—an agent, that is, the hand or another object such as a claw, and a patient, that is, a graspable object. The agent corresponds to the subject of an observed interaction while the patient corresponds to the object of the interaction. At every time step <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0188" xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math> </ephtml> , our system received an 18‐dimensional real‐valued observation vector <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0189" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>o</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> , which signals the three‐dimensional position of the agent <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0190" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>x</mi><mi>a</mi></msup></math> </ephtml> and the patient <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0191" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>x</mi><mi>p</mi></msup></math> </ephtml> , their respective three‐dimensional velocities <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0192" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>v</mi><mi>a</mi></msup></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0193" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>v</mi><mi>p</mi></msup></math> </ephtml> , the relative position of the agent relative to the patient <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0194" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>x</mi><mrow><mi>p</mi><mo>−</mo><mi>a</mi></mrow></msup><mo>=</mo><msup><mi>x</mi><mi>p</mi></msup><mo>−</mo><msup><mi>x</mi><mi>a</mi></msup></mrow></math> </ephtml> , and the Euclidean distance <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0195" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>d</mi><mrow><mi>a</mi><mo>,</mo><mi>p</mi></mrow></msup></math> </ephtml> between agent and patient. Furthermore, the observation contained a one‐dimensional "representation" of the agent's and patient's shapes ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0196" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>s</mi><mi>a</mi></msup></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0197" xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>s</mi><mi>p</mi></msup></math> </ephtml> , respectively). This shape signal serves as a simple cue to distinguish hands from other types of entities, such as claws, and, thus, can bias the system's inferred event probabilities. The shape values for hands, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0198" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>s</mi><mtext>hand</mtext></msub><mo>∈</mo><mrow><mo>[</mo><mn>0</mn><mo>,</mo><mn>0.5</mn><mo>)</mo></mrow></mrow></math> </ephtml> , and claws, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0199" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>s</mi><mtext>claw</mtext></msub><mo>∈</mo><mrow><mo>(</mo><mn>0.5</mn><mo>,</mo><mn>1.0</mn><mo>]</mo></mrow></mrow></math> </ephtml> , representing hands and mechanical claws were randomly determined for every simulation. The shape values for all other entities were randomly sampled from <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0200" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>[</mo><mn>0</mn><mo>,</mo><mn>1</mn><mo>]</mo></mrow></math> </ephtml> . Note that the sensory information was position‐encoded (starting with the agent followed by the patient and the relative encodings), thus sidestepping the challenge to assign agent and patient roles.</p> <p>Our system acted choosing different policies, that is, gaze positions. To do so, the system could choose to focus on one of three significant positions, influencing the standard deviation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0201" xmlns="http://www.w3.org/1998/Math/MathML"><mi>σ</mi></math> </ephtml> of the normally distributed noise levels on agent‐ and patient‐related sensory information. The system could look at the agent ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0202" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>a</mi></msub></math> </ephtml> ), yielding <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0203" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mo>=</mo><mn>0.1</mn></mrow></math> </ephtml> for all patient‐related components ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0204" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>x</mi><mi>p</mi></msup><mo>,</mo><msup><mi>v</mi><mi>p</mi></msup><mo>,</mo><msup><mi>s</mi><mi>p</mi></msup><mo>,</mo><msup><mi>x</mi><mrow><mi>p</mi><mo>−</mo><mi>a</mi></mrow></msup><mo>,</mo><msup><mi>d</mi><mrow><mi>a</mi><mo>,</mo><mi>p</mi></mrow></msup></mrow></math> </ephtml> ), but <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0205" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mo>=</mo><mn>0.01</mn></mrow></math> </ephtml> , for purely agent‐based information ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0206" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>x</mi><mi>a</mi></msup><mo>,</mo><msup><mi>v</mi><mi>a</mi></msup><mo>,</mo><msup><mi>s</mi><mi>a</mi></msup></mrow></math> </ephtml> ). When looking at the patient ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0207" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> ), the sensory noise scheme of all components was reversed, with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0208" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mo>=</mo><mn>0.1</mn></mrow></math> </ephtml> for the agent and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0209" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mo>=</mo><mn>0.01</mn></mrow></math> </ephtml> for the patient. When looking at neither of the entities ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0210" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>n</mi></msub></math> </ephtml> ), normally distributed noise with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0211" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>σ</mi><mo>=</mo><mn>0.1</mn></mrow></math> </ephtml> was added to all sensory components. Thus, we simulated a kind of object‐oriented attention, yielding clearer and noisier sensory information about the focused and unfocused entities, respectively, ignoring actual physical distance.</p> <p>Every event‐sequence simulation activated a new pair of agent <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0212" xmlns="http://www.w3.org/1998/Math/MathML"><mi>a</mi></math> </ephtml> and patient <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0213" xmlns="http://www.w3.org/1998/Math/MathML"><mi>p</mi></math> </ephtml> and positioned them randomly in the scenario. Four possible events were considered, mimicking events that may be encountered and produced by infants.</p> <p></p> <ulist> <item> During a <emph>standing still</emph> event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0214" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> ), agent and patient remained motionless for a fixed number of time steps (100 in our simulations), mimicking looking at stationary objects.</item> <p></p> <item> During a <emph>randomly directed motion</emph> event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0215" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>e</mi><msub><mrow /><mi>random</mi></msub></mrow></math> </ephtml> ), the agent moved constantly in one fixed, but randomly generated direction with decreasing velocity. The event ended when the agent was approximately motionless ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0216" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>v</mi><mi>a</mi></msup><mo><</mo><mn>0.0005</mn></mrow></math> </ephtml> ). This event mimics the observation of rolling or sliding objects, like a toy car.</item> <p></p> <item> During a <emph>reaching motion</emph> event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0217" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> ), a hand agent moved toward the patient with a randomly set, constant velocity. The event ended when the agent approximately reached the patient.</item> <p></p> <item> Hand agents could also perform a <emph>transportation</emph> event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0218" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> ), where agent and patient moved together to a randomly generated goal location with a randomly set, constant velocity.</item> </ulist> <p>By combining these events, three possible event sequences <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0219" xmlns="http://www.w3.org/1998/Math/MathML"><mi>E</mi></math> </ephtml> were trained (cf. Fig. 4): In <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0220" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>grasp</mtext></msub></math> </ephtml> , a hand agent reached for the patient ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0221" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> ), carried it to a certain location ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0222" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> ), and then let go of it and randomly moved away ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0223" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> ). <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0224" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>random</mtext></msub></math> </ephtml> started with a randomly directed motion event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0225" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> ) followed by both entities standing still ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0226" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> ). <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0227" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mi>still</mi></msub></math> </ephtml> showed both entities standing still ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0228" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> ). For testing, we considered a testing sequence <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0229" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> , which is generally identical to <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0230" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>grasp</mtext></msub></math> </ephtml> , but also allowed claw agents.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0004.jpg" title="4 Event sequences, their individual events, and corresponding exemplar visualizations, rendered from a bird's‐eye view (depth information can be deduced via entity size)." /> </p> <p></p> <p>During supervised training, the system was informed about the type of event that currently unfolded. Each training epoch was composed of 100 event sequences, each uniformly randomly chosen and lasting between 50 and 150 time steps. Additionally, the eye fixation policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0231" xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math> </ephtml> was randomly chosen and stayed fixed for every event sequence.</p> <p>During testing, we tracked the system's gaze while the system was shown grasping sequences, mimicking actual "experimental conditions" that investigated the occurrence of goal anticipation in infants (Adam et al., 2016; Adam et al., 2017; Cannon & Woodward, 2012; Kanakogi & Itakura, 2011). In contrast to training, the system had to infer which event was currently observed using its learned event‐predictive model components (cf. Section 4.2). Additionally, the system chose its gaze policy using active inference (cf. Section 4.3). Similar to Adam et al. (2016), we distinguished between two testing conditions: In the <emph>hand</emph> condition, the system was shown a grasping sequence <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0232" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> performed by a hand agent ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0233" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>s</mi><mi>a</mi></msup><mo>=</mo><msub><mi>s</mi><mtext>hand</mtext></msub></mrow></math> </ephtml> ), similar to the ones shown in training. In the <emph>claw</emph> condition, the same event sequence <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0234" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> was shown being performed by a claw agent, which differed in its shape sensory signal ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0235" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>s</mi><mi>a</mi></msup><mo>=</mo><msub><mi>s</mi><mtext>claw</mtext></msub></mrow></math> </ephtml> ). No model updates were performed during testing. Every test phase was composed of 10 <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0236" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> event sequences each for a hand and a claw agent. In order to assess whether goal‐anticipatory gaze behavior emerges over the course of learning, we alternated between training epochs and test phases.</p> <hd id="AN0152095508-17">System behavior</hd> <p>We focused on analyzing whether and under which conditions CAPRI generates goal‐anticipatory gaze after learning. We provide further experiments on learning progressions, parameter dependencies, and ablations in the Supporting Information Sections 8.5–8.7. We trained CAPRI for 30 training epochs each followed by a test phase. We used the temporal horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0237" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>=</mo><mn>1</mn></mrow></math> </ephtml> for policy inference. We ran this experiment 20 times with different, randomly chosen initial model weight and simulation parameter values. During the test phases, we mainly measured two quantities: (<reflink idref="bib1" id="ref3">1</reflink>) the internally estimated <emph>event‐predictive probabilities</emph><ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0238" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>i</mi></msub><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>|</mo><mi>O</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo>(</mo><mi>t</mi><mo>)</mo></mrow><mo>)</mo></mrow></math> </ephtml> , to quantify the system's beliefs about the unfolding events; (<reflink idref="bib2" id="ref4">2</reflink>) the chosen <emph>gaze policy</emph><ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0239" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>(</mo><mi>t</mi><mo>)</mo></mrow></math> </ephtml> , to test whether the system activated <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0240" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> (i.e., <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0241" xmlns="http://www.w3.org/1998/Math/MathML"><mo><</mo></math> </ephtml> looking at the patient <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0242" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> ) before the reaching event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0243" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> ends, that is, whether the system starts to exhibit anticipatory gaze shifts for the hand agent and/or for the claw agent after having accumulated sufficient training experiences.</p> <p>Fig. 5 shows the inferred event probabilities for two exemplary event sequences <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0244" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> for a fully trained system: (a) depicts the inferred event probabilities for a hand agent and (b) for a claw agent. During the beginning of a reaching movement of a hand, CAPRI inferred a high probability of standing still (i.e., <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0245" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>still</mi></msub><mo>)</mo></mrow></math> </ephtml> ). While accumulating more sensory observations, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0246" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>still</mi></msub><mo>)</mo></mrow></math> </ephtml> decreased while <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0247" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>reach</mi></msub><mo>)</mo></mrow></math> </ephtml> sharply increased. After approximately five time steps, the system correctly assumed that a reaching event was unfolding with more than 95% probability. As the hand got closer to the patient, the system started to anticipate that the reaching event will end soon, yielding a continuous decrease in <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0248" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mi>reach</mi></msub><mo>)</mo></mrow></math> </ephtml> and an increase in <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0249" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mtext>transport</mtext></msub><mo>)</mo></mrow></math> </ephtml> . Once the hand touched the patient and started transporting it, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0250" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mtext>transport</mtext></msub><mo>)</mo></mrow></math> </ephtml> sharply rose to approximately 100%. Later on, when the hand let go of the patient and started a random motion, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0251" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mtext>random</mtext></msub><mo>)</mo></mrow></math> </ephtml> quickly reached almost 100% probability.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0005.jpg" title="5 Exemplary event inference and policy inference after full training (30 epochs) over the course of one event sequence Etest for a hand agent (a) and a claw agent (b). The top row in each subfigure shows the active policy π(t) over time t. The bottom row shows the inferred event probability estimates for the four possible events over time t. Event boundaries are marked by dotted lines." /> </p> <p></p> <p>The inferred event probabilities drastically differ when a claw agent is performing the same movement (Fig. 5b). While the probability for a standing still event continuously decreased during the first half of the reaching event, <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0256" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mtext>random</mtext></msub><mo>)</mo></mrow></math> </ephtml> increased and reached approximately 100%. When the claw started transporting the object, CAPRI first incorrectly inferred that a standing still event was most likely, which then morphed into an approximately 50/50 guess in favor of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0257" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0258" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> , progressively favoring <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0259" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> slightly. Once the claw let go of the patient and started a random motion, the system correctly inferred <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0260" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>P</mi><mo>(</mo><msub><mi>e</mi><mtext>random</mtext></msub><mo>)</mo></mrow></math> </ephtml> , quickly reaching almost 100% probability.</p> <p>The differences in event inference also resulted in differences in gaze behavior for hand and claw agents. For hand agents, CAPRI started looking at the patient, that is, it activated <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0261" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> , once it was certain that it was observing a reaching event ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0262" xmlns="http://www.w3.org/1998/Math/MathML"><mo>></mo></math> </ephtml> 90 % probability). Thus, for a hand agent we observe goal‐anticipatory gaze shifts: the system looked at the reaching target before the target was reached by the agent. For a claw agent, the system kept tracking the agent, that is, it activated <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0263" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>a</mi></msub></math> </ephtml> , during the whole reaching movement. Thus, for a claw agent, the system did not show goal‐anticipatory gaze shifts.</p> <p>Fig. 6 shows the mean inferred event probability during the test phases for each of the four possible events <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0264" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> , <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0265" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> , <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0266" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> , and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0267" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> , as a function of progressively more learning experience for hand agents (left‐hand side) and claw agents (right‐hand side).[<reflink idref="bib3" id="ref5">3</reflink>] In all cases, CAPRI initially tended to infer <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0268" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0269" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> with high probability. Over the course of training, the system started to correctly infer progressively higher probabilities for the actual underlying events when the hand executed <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0270" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> , <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0271" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> , and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0272" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> . In contrast, when CAPRI observed a claw agent, it incorrectly inferred <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0273" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0274" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> with much higher probability, even when <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0275" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> or <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0276" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> were shown. Proper event inference failed because the executing claw agent did not match the start condition learned for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0277" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0278" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>transport</mtext></msub></math> </ephtml> .</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0006.jpg" title="6 Mean inferred event probabilities P¯(ei(t)|O(t),Π(t)) for ei∈{estill,erandom,ereach,etransport} during Etest over the course of learning for hand agents (a,c,e) and claw agents (b,d,f). (a and b) The event probabilities during the presentation of ereach, (c and d) during etransport, and (e and f) during e random." /> </p> <p></p> <p>Fig. 7a shows how the gaze behavior evolves with training experience. Here we measured at what time CAPRI on average activated the policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0284" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> during an event sequence <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0285" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> . Simulations where <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0286" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> was not activated counted as looking at the patient at the end of the sequence. We computed the mean time of activating <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0287" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> over simulations. When testing after very few training phases, the system did not look at the patient during the reaching event irrespective of hand or claw agent. From after the fourth training phase onward, however, the system began to systematically activate <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0288" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> in the beginning of the reaching event in <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0289" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> in the case of a hand agent. This goal‐anticipatory gaze behavior was not shown by CAPRI for claw agents.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01aug21/cogs13016-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13016-fig-0007.jpg" title="7 Analyzing anticipatory gaze for reaching movements. (a) The gaze behavior for CAPRI over the course of learning. We plot the mean time t during an event sequence Etest when the system first activated policy πp, which corresponds to looking at the patient. Event boundaries are marked by dotted horizontal black lines with t0 marking the agent's arrival at the patient. Shaded areas show the standard deviation across simulations. (b) The anticipatory gaze shifts in milliseconds for 12‐month‐old infants are taken from Adam et al. (2016). The dotted horizontal line marks the agent's arrival at the patient (t=0). Error bars show the standard error. In both plots a data point in the gray area, below the dotted line denoting the agent's arrival at the patient, marks an anticipatory gaze." /> </p> <p></p> <p>These results indicate that, once sufficiently trained on all four events and three event sequences, CAPRI generates gaze behavior that can be compared to that observed in 12‐month‐old infants (Adam et al., 2016). Fig. 7b shows the mean gaze arrival time for 12‐month‐olds watching movies of reaching movements performed by hand agents or claw agents (Adam et al., 2016). The movies started with a 500 ms still frame depicting an object, and then showed a hand or claw entering from the upper part of the screen (500–960 ms), pausing shortly (960–1800 ms), and then reaching for the target object (1800–2920 ms). To measure the gaze arrival times, two areas of interest (AOI) were created to cover the agent and target object for each movie. Mean gaze arrival times were calculated by subtracting the time when the agent entered the target AOI from the time of the first target AOI fixation. Thus, a negative gaze arrival time corresponds to goal‐anticipatory gaze. As shown in Fig. 7b, when the infants observed reaching movies done by hands, they tended to look at the target AOI before the hand arrived. For the claw, the infants tended to first look at the target AOI upon arrival of the agent. Similar behavior developed in CAPRI (compare results in Fig. 7a and b), indicating that infants' goal‐anticipatory gaze shifts may rely on processes and principles similar to those implemented in CAPRI.</p> <hd id="AN0152095508-21">Discussion of results</hd> <p>Over the course of training, CAPRI tended to improve its ability to correctly infer the event probabilities for a hand agent. For claw agents, the system produced incorrect inferences. This is mainly due to the system learning that reaching and transporting events are typically performed by hand agents, which can be recognized from their observed shape ( <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0295" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>s</mi><mi>a</mi></msub><mo>=</mo><msub><mi>s</mi><mtext>hand</mtext></msub></mrow></math> </ephtml> ). Agents with a different shape <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0296" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>s</mi><mi>a</mi></msub></math> </ephtml> , such as claw agents, are not known to execute reaching and transporting events. In conjunction with the active inference mechanism, CAPRI thus begins to show goal‐anticipatory gaze behavior over the course of training when a hand agent is observed, but not when a claw agent is observed. Similarly, 12‐month‐old infants, who already have some experience with self‐performed and observed human grasping, tend to shift their gaze to the goal of reaching before there is contact between hand and target, but they do not do so when a claw agent is observed (Adam et al., 2016).</p> <p>Note that CAPRI was not trained nor preprogramed to generate goal‐anticipatory gaze behavior. Instead, the behavior emerged purely from the system estimating the ongoing event via Eq. 1 and choosing its actions by means of active inference, which aims at decreasing expected uncertainty about the ongoing event and the upcoming event boundaries according to Eq. 10. During the beginning of training, CAPRI did not show goal‐anticipatory gaze shifts during <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0297" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>test</mtext></msub></math> </ephtml> because the system (i) did not recognize the underlying event, inferring a low probability for <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0298" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> , and (ii) had not yet learned which type of policy could be expected to result in a decrease in future uncertainty.[<reflink idref="bib4" id="ref6">4</reflink>] At this point <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0299" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mtext>random</mtext></msub></math> </ephtml> or <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0300" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>still</mi></msub></math> </ephtml> were inferred, during which CAPRI could minimize uncertainty best by looking at the agent, that is, choosing gaze policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0301" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>a</mi></msub></math> </ephtml> . Once the event‐predictive models were sufficiently precise and it became known that fixating the target object served to decrease uncertainty for both, the reaching event and the exact time when the next event boundary would occur, CAPRI generated goal‐anticipatory gaze shifts. In particular, for a reaching event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0302" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> , the system could minimize uncertainty best by fixating the patient, that is, choosing gaze policy <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0303" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>p</mi></msub></math> </ephtml> , because this resulted in less noisy information about the patient's position, which informed the system about where reaching would end. As a result, goal‐anticipatory gaze behavior emerged over the course of training when observing a hand reaching for an object.</p> <p>We propose that similar processes are involved when infants exhibit goal‐anticipatory gaze behavior (e.g., Elsner & Adam, 2021). For most events involving a moving agent, tracking the agent can give the most information about the future, such as the agent's future position and velocity. However, the end of a reaching event can be predicted best when the position of the target object is known, which can be estimated best when looking at the target before the reaching agent arrives at the target. Fittingly, goal‐anticipatory gaze behavior was found when adults performed, and when they observed another person perform, a block‐stacking task (Flanagan & Johansson, 2003). For observed reaching actions, infants' anticipatory gaze depends on their action experience and on the familiarity of the agent, starting from about 7 months for human hands, when infants typically achieve the developmental milestone of visually guided grasping (Adam et al., 2016; Falck‐Ytter, Gredebäck, & Hofsten, 2006; Kanakogi & Itakura, 2011). In CAPRI, goal‐anticipatory gaze depends on whether the reaching event is recognized. We assume that infants likewise associate a perceived start condition of visuospatial features of a hand together with a graspable object with the encoding of an event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0304" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> , and, as a result, only activate their generative model for reaching when sufficient (active or observed) grasping experience has been acquired and when sufficiently many relevant indicators are perceived. Based on their stored event schemas, infants can then create a forward model of the agent's future trajectory and the event's typical end condition (Elsner & Adam, 2021). When the start condition contains an unfamiliar agent, such as a mechanical claw, the event schemata necessary for eliciting goal‐anticipatory gaze behavior are typically not activated.</p> <p>Thus, according to CAPRI, 12‐month‐olds track the reaching motion of unfamiliar agents because they fail to recognize the start of a reaching event and, as a result, cannot tap on their stored event knowledge in order to anticipate how this event will end. An alternative explanation for the findings could be that the infants found it difficult to disengage their gaze from the interesting unfamiliar agent, which in turn made it more difficult for them to shift their gaze to the goal object ahead of time. However, recent research indicates that infants typically do not show longer looking times to unfamiliar claws compared to familiar hands in simple grasping events (e.g., Adam et al., 2016; Adam & Elsner, 2020). Hence, it is unlikely that this gaze behavior could solely be explained by general mechanisms such as a limited capacity for stimulus disengagement.</p> <p>While a comparable gaze behavior is shown by CAPRI and the 12‐month‐old infants, there are numerical differences worth discussing. For example, the gaze shifts toward the target object occurred later in infants than in CAPRI. Clearly though, the experimental results of CAPRI and the 12‐month‐olds are not quantitatively comparable, because Adam et al. (2016) used different stimuli and an AOI‐based measure for gaze recording. Moreover, our scenario is strongly simplified compared to the learning and testing of infants. CAPRI sees reaching movements exactly how they occur during training, while infants see the particular reaching movement for the first time at the start of the experiment, and can further process the movement (and relate it to their previously gained action experience) across about 12 repeated presentations. Besides this, CAPRI always observes the shown simulation, while infants can also avert their gaze from the screen or close their eyes. However, the aim of our study was not to exactly replicate the infant data, but to analyze the functional foundations of anticipatory, event‐predictive gaze behavior. Generally, CAPRI should achieve quantitatively similar results in richer simulation environments. Additional model evaluations provided in the Supporting Information largely support this claim: Decreasing the percentage of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0305" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>reach</mi></msub></math> </ephtml> events encountered during training (Section 8.5) or increasing the difficulty of recognizing the hand (Section 8.7) indeed delay the development of anticipatory gaze behavior. Along similar lines, we assume that a richer simulation environment would increase event inference difficulty and, thus, should likely delay the anticipatory gaze behavior of CAPRI, possibly matching the infants' data even more exactly.</p> <hd id="AN0152095508-22">Discussion and future work</hd> <p>In this paper, we have proposed how goal‐anticipatory gaze shifts in infants can emerge. We introduced and implemented CAPRI, a computational and algorithmic model that learns generative, schematic event encodings, which predict the dynamics of an event in a probabilistic manner and also predicts distributions modeling the typical conditions at the start and end of an event. CAPRI tries to constantly infer which one of the previously learned events is currently unfolding by iteratively inferring event likelihood distributions dependent on its learned event‐predictive schemata and the incoming sensory observations. Besides inferring events, CAPRI inferred its action policy aiming solely at minimizing expected future uncertainty, which was equated with the entropy of the current event distribution and the expected next possible event boundary distributions. As a result, CAPRI can be understood as a modified hidden Markov model, which processes dynamic Bayesian likelihood estimates as done recently elsewhere (e.g., Franklin, Norman, Ranganath, Zacks, & Gershman, 2020), but which additionally infers actions by means of event‐predictive active inference.</p> <p>We tested CAPRI in a simple agent–patient interaction scenario where policies corresponded to different gaze strategies. The system showed anticipatory gaze behavior similar to the goal‐anticipatory gaze shifts previously found in eye‐tracking studies in infants (Adam et al., 2016; Cannon & Woodward, 2012; Kanakogi & Itakura, 2011). When a hand agent was performing a reaching action that had been observed during training, CAPRI looked at the target before reaching was complete. These goal‐anticipatory gaze shifts could not be observed when CAPRI only had little experience in reaching. Similarly, 12‐month‐old infants look at the target of a hand‐reaching movement before the reaching is concluded (e.g., Adam et al., 2016), while infants younger than 6 months, who are believed to have little to no experience in reaching, do not. Moreover, when CAPRI observed a reaching event with an unfamiliar claw agent, which never performed reaching actions during training, our system did not identify reaching and preferred to track the agent over the course of the event. Similarly, when 12‐month‐olds observe a mechanical claw reaching for an object, they tend to perform tracking gaze (Adam et al., 2016).</p> <p>Note, that none of these effects were explicitly programed into the system. Both the structure of event‐representations and the minimization of predicted uncertainty have been proposed as key mechanisms of cognition over the last decades by various theories, such as the TEC (Hommel, Müsseler, Aschersleben, & Prinz, 2001), EST (Zacks, Speer, Swallow, Braver, & Reynolds, 2007), Predictive Coding (Rao & Ballard, 1999), and the Free Energy Principle (Friston, 2010). Showing how behavioral effects can emerge based on these theories further supports the proposition to merge these theories for developing a unifying theory of cognition (Butz, 2016) and allows for an integrative theoretical perspective on infants' anticipatory gaze behavior (Elsner & Adam, 2021). Goal‐anticipatory gaze behavior essentially emerged from the structure in which events were encoded during the training and from the objective to invoke behavioral policies, that is, eye fixation targets, that are expected to minimize predicted uncertainty in the near future.</p> <p>Nonetheless, various aspects and design considerations demand future modeling work. First, in the current implementation, we provided supervised information about the event that is currently observed during training. Additionally, we specified for each event which entity takes the role of the agent and the patient. Clearly this is unrealistic. Because we aimed at investigating whether goal‐anticipatory gaze behavior can emerge at all from learned event schemata in combination with active inference processes, we omitted the challenge of automatically individualizing event schemata models and the roles of the involved entities from the continuous stream of sensorimotor information. Nevertheless, in future work we want to investigate whether we can find the same effects when event schemata are learned in a fully self‐supervised manner. There are different methods for the self‐supervised extraction of event models from the continuous stream of sensorimotor information, such as surprise‐based segmentation (Gumbsch, Butz, & Martius, 2019; Reynolds, Zacks, & Braver, 2007) or retrospective gradient‐based hidden state inference (Butz, Bilkey, Humaidan, Knott, & Otte, 2019). Furthermore, the role of an entity with a specific shape could be inferred from the sensorimotor data, for example, via the detection of "mover" events, that is, events in which the movement of one entity leads to the movement of another entity (Ullman, Harari, & Dorfman, 2012). Further investigations are required to identify which type of segmentation methods, or combinations of such methods, are most suitable.</p> <p>Second, CAPRI chooses actions to minimize predicted uncertainty over a fixed temporal horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0306" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> into the future. This is sufficient for the investigated simple scenario. However, the effects of goal anticipation might emerge even more strongly when the system attempts to predict the future only when it is certain about the present. Therefore, our further research will consider a more flexible temporal horizon of predictions based on estimated uncertainty. This may strengthen the effects of goal anticipations found in the current study and may enable the emergence of goal‐anticipatory gaze shifts in more complicated simulations and without an explicit limit on the temporal horizon <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0307" xmlns="http://www.w3.org/1998/Math/MathML"><mi>τ</mi></math> </ephtml> .</p> <p>Third, CAPRI learns from own action experiences, utilizing those experiences for inferring subsequently observed action events. By including relative encodings between agent and patient, we somewhat sidestepped the perspective‐taking challenge (Moll & Meltzoff, 2011; Tversky & Hard, 2009), as has been done elsewhere (Schrodt & Butz, 2016). However, infant studies obtained comparable results for stimuli presented from various spatial perspectives (Adam et al., 2016, 2017), calling the relevance of this factor for anticipatory gaze behavior into question. Moreover, we somewhat ignored the fact that motor signals are available when executing particular actions, which probably support goal‐anticipatory gaze behavior via direct matching when observing others performing a similar action (e.g., Gredebäck & Falck‐Ytter, 2015). In future modeling work, we intend to enhance CAPRI to address these challenges in further detail.</p> <p>Besides such system expansions and modifications, we intend to model other effects of goal anticipation in infants. For example, it has been shown that 11‐month‐old infants perform a goal‐anticipatory gaze shift for mechanical claws when the claw shows cues of agency, such as self‐propelled motion and a salient action effect (Adam et al., 2017). Additionally, EEG studies found that from 9 months onward, infants show predictive sensorimotor‐cortex activity for actions performed by various agents such as a hand, a mechanical claw, or a self‐propelled toy (Southgate & Begus, 2013). After the infants had been familiarized with videos in which the agent grasped and transported the object to a different location, or in which the object moved to a new location by itself, predictive motor activity was found in the infants' brains even when only a still frame of the start of the action event was presented. Thus, instead of directly conditioning on the agent's appearance, more general agency cues, such as self‐propelled movement, may be encoded in the likelihood distributions of our event schemata. The presence of such agency cues adds visuospatial features that may help to identify goal‐directed events, such as reaching, thereby enabling the generation of goal‐anticipatory gaze shifts.</p> <p>Moreover, experimental data with adults demands further model‐based, computational explanations. Anticipatory gaze shifts have been investigated in various scenarios, including complex object interactions (Belardinelli, Barabas, Himmelbach, & Butz, 2016; Belardinelli, Stepper, & Butz, 2016; Hayhoe, Shrivastava, Mruczek, & Pelz, 2003). Moreover, in anticipatory cross‐modal interactions, the future temporal horizon appears to be event‐oriented and was shown to depend on predictable forthcoming event uncertainties (Belardinelli, Lohmann, Farnè, & Butz, 2018; Lohmann, Belardinelli, & Butz, 2019). We expect that our event‐generative inference model is applicable to the observed data patterns, because similar event‐generative encodings and inference processes should be at play.</p> <p>Lastly, it needs to be noted that for now, CAPRI only focuses on goal anticipation measured via anticipatory gaze shifts. Therefore, our results might not necessarily be applicable to post hoc measures such as looking times (e.g., Woodward, 1998) or to other predictive processes such as EEG activity (e.g., Southgate & Begus, 2013). As of now, the relations between these different measures are still widely unknown, and more research needs to be conducted to fully understand how these measures do or do not capture the same processes regarding infants' processing of goal‐directed actions.</p> <p>To conclude, let us consider the question which type of representations and inference processes may give rise to goal‐anticipatory gaze behavior in infants. It has been proposed that infants learn flexible goal representations for each action event (Cannon & Woodward, 2012). An alternative explanation suggests that infants rely on trajectory‐based information when estimating how an observed movement will end (Ganglmayer, Attig, Daum, & Paulus, 2019). Along similar lines, goal anticipation in infants was previously modeled in a complex robotic reaching scenario using closed‐loop sensorimotor forward simulations (Copete Copete, Nagai, & Asada, 2016). There seems to be experimental evidence that supports both flexible goal representations and trajectory‐based models (Cannon & Woodward, 2012; Ganglmayer, Attig, Daum, & Paulus, 2019). Our modeling results suggest that both types of representation are involved. Event dynamics models encode how the observation will change during one event over a short period of time, essentially predicting the exact trajectory of an ongoing motion. However, inferring the goal of an event using a closed‐loop simulation is costly and prone to accumulate error. The event‐generative encoding of end conditions enables their direct anticipation, but this requires previous (active or passive) experience with the observed action (Elsner & Adam, 2021). Similarly, start‐condition encodings enable hierarchical, closed‐loop simulations of event successions, without the simulation of detailed event dynamics. As a result, these types of representation can be directly accessed to infer goal‐anticipatory gaze behavior and thus to plan and reason on deeper, event‐predictive levels. The available developmental‐psychology research and the CAPRI model we have introduced in this paper suggest that we learn the compositional basis of such deeper cognitive abilities via the observation of action events during the first year of life.</p> <hd id="AN0152095508-23">Acknowledgments</hd> <p>This research was funded by the German Research Foundation (DFG) within Priority‐Program SPP 2134—project "Development of the agentive self" (BU 1335/11‐1, EL 253/8‐1). The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS‐IS) for supporting Christian Gumbsch. Martin Butz is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1—project number 390727645.</p> <p>Open access funding enabled and organized by Projekt DEAL.</p> <p>GRAPH: Supporting Information</p> <ref id="AN0152095508-24"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> To be more precise, the activated internal estimates usually consider only those aspects of the outside environment that are deemed relevant. Moreover, note that the objective to minimize uncertainty causes the system to attempt to maximize its precision estimates of its internal currently activated event‐predictive encodings.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> Since Friston et al. ([29]) use a different notation and do not consider internal event estimations <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0308" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> we provide a derivation of Eq. 6 in Section 8.3.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref5" type="bt">3</bibl> <bibtext> We compute for every event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0309" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>i</mi></msub><mo>∈</mo><msub><mi>E</mi><mtext>test</mtext></msub></mrow></math> </ephtml> and every simulation <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0310" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>s</mi><mi>k</mi></msub></math> </ephtml> the mean probability as <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0311" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>P</mi><mo>¯</mo></mover><msub><mi>s</mi><mi>k</mi></msub></msub><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>j</mi></msub><mo stretchy="false">|</mo><mi>O</mi><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mo stretchy="false">)</mo></mrow><mo>=</mo><msubsup><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><msub><mi>t</mi><mi>s</mi></msub></mrow><msub><mi>t</mi><mi>e</mi></msub></msubsup><mi>P</mi><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>j</mi></msub><mrow><mo stretchy="false">(</mo><mi>t</mi><mo stretchy="false">)</mo></mrow><mo stretchy="false">|</mo><mi>O</mi><mrow><mo stretchy="false">(</mo><mi>t</mi><mo stretchy="false">)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo stretchy="false">(</mo><mi>t</mi><mo stretchy="false">)</mo></mrow><mo stretchy="false">)</mo></mrow><mfrac><mn>1</mn><mrow><msub><mi>t</mi><mi>e</mi></msub><mo>−</mo><msub><mi>t</mi><mi>s</mi></msub></mrow></mfrac></mrow></math> </ephtml> , with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0312" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>t</mi><mi>s</mi></msub></math> </ephtml> marking the start of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0313" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0314" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>t</mi><mi>e</mi></msub></math> </ephtml> marking its end and for all <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0315" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>e</mi><mi>j</mi></msub><mo>∈</mo><mrow><mo stretchy="false">{</mo><msub><mi>e</mi><mi>still</mi></msub><mo>,</mo><msub><mi>e</mi><mtext>random</mtext></msub><mo>,</mo><msub><mi>e</mi><mi>reach</mi></msub><mo>,</mo><msub><mi>e</mi><mtext>transport</mtext></msub><mo stretchy="false">}</mo></mrow></mrow></math> </ephtml> . <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0316" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>P</mi><mo>¯</mo></mover><msub><mi>s</mi><mi>k</mi></msub></msub><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>j</mi></msub><mo stretchy="false">|</mo><mi>O</mi><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mo>,</mo><mi mathvariant="normal">Π</mi><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mo stretchy="false">)</mo></mrow></mrow></math> </ephtml> can be visualized when looking at Fig. 5. If one takes the mean probabilities during one event <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0317" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>e</mi><mi>i</mi></msub></math> </ephtml> , bounded by dotted lines, one gets the estimate of <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0318" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mover accent="true"><mi>P</mi><mo>¯</mo></mover><msub><mi>s</mi><mi>k</mi></msub></msub><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>j</mi></msub><mo stretchy="false">|</mo><msub><mi>e</mi><mi>i</mi></msub><mo>,</mo><mi>O</mi><mo>,</mo><mi mathvariant="normal">Π</mi><mo stretchy="false">)</mo></mrow></mrow></math> </ephtml> . This probability is computed for every test phase of every simulation and Fig. 6 shows <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0319" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mover accent="true"><mi>P</mi><mo>¯</mo></mover><mrow><mo stretchy="false">(</mo><msub><mi>e</mi><mi>j</mi></msub><mo stretchy="false">|</mo><msub><mi>e</mi><mi>i</mi></msub><mo>,</mo><mi>O</mi><mo>,</mo><mi mathvariant="normal">Π</mi><mo stretchy="false">)</mo></mrow></mrow></math> </ephtml> averaged over all simulations <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0320" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>s</mi><mi>k</mi></msub></math> </ephtml> .</bibtext> </blist> <blist> <bibl id="bib4" idref="ref6" type="bt">4</bibl> <bibtext> We further investigate how the experience with <ephtml> <math display="inline" altimg="urn:x-wiley:03640213:media:cogs13016:cogs13016-math-0321" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>E</mi><mtext>grasp</mtext></msub></math> </ephtml> in particular affects the goal‐anticipatory gaze in an additional experiment in Section 8.5 of the Supporting Information.</bibtext> </blist> </ref> <ref id="AN0152095508-25"> <title> References </title> <blist> <bibtext> Adam, M., & Elsner, B. (2018). Action effects foster 11‐month‐olds' prediction of action goals for a non‐human agent. Infant Behavior and Development, 53, 49 – 55. https://doi.org/10.1016/j.infbeh.2018.09.002</bibtext> </blist> <blist> <bibtext> Adam, M., & Elsner, B. (2020). The impact of salient action effects on 6‐, 7‐, and 11‐month‐olds' goal‐predictive gaze shifts for a human grasping action. PLoS One, 15 (10), 1 – 18. https://doi.org/10.1371/journal.pone.0240165</bibtext> </blist> <blist> <bibtext> Adam, M., Reitenbach, I., & Elsner, B. (2017). Agency cues and 11‐month‐olds' and adults' anticipation of action goals. Cognitive Development, 43, 37 – 48. https://doi.org/10.1016/j.cogdev.2017.02.008</bibtext> </blist> <blist> <bibtext> Adam, M., Reitenbach, I., Papenmeier, F., Gredebäck, G., Elsner, C., & Elsner, B. (2016). Goal saliency boosts infants' action prediction for human manual actions, but not for mechanical claws. Infant Behavior and Development, 44, 29 – 37. https://doi.org/10.1016/j.infbeh.2016.05.001</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Baldwin, D. A., & Kosie, J. E. (2021). How does the mind render streaming experience as events? Topics in Cognitive Science, 13 (1), 79 – 105. https://doi.org/10.1111/tops.12502</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> Belardinelli, A., Barabas, M., Himmelbach, M., & Butz, M. V. (2016). Anticipatory eye fixations reveal tool knowledge for tool interaction. Experimental Brain Research, 234 (8), 2415 – 2431. https://doi.org/10.1007/s00221‐016‐4646‐0</bibtext> </blist> <blist> <bibl id="bib7" type="bt">7</bibl> <bibtext> Belardinelli, A., Lohmann, J., Farnè, A., & Butz, M. V. (2018). Mental space maps into the future. Cognition, 176, 65 – 73. https://doi.org/10.1016/j.cognition.2018.03.007</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Belardinelli, A., Stepper, M. Y., & Butz, M. V. (2016). It's in the eyes: Planning precise manual actions before execution. Journal of Vision, 16 (1), 18 – 18. https://doi.org/10.1167/16.1.18</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Biro, S. (2013). The role of the efficiency of novel actions in infants' goal anticipation. Journal of Experimental Child Psychology, 116 (2), 415 – 427. https://doi.org/10.1016/j.jecp.2012.09.011</bibtext> </blist> <blist> <bibtext> Bishop, C. M. (2006). Pattern recognition and machine learning (information science and statistics). Berlin : Springer.</bibtext> </blist> <blist> <bibtext> Botvinick, M., & Toussaint, M. (2012). Planning as inference. Trends in Cognitive Sciences, 16 (10), 485 – 488. https://doi.org/10.1016/j.tics.2012.08.006</bibtext> </blist> <blist> <bibtext> Butz, M. V. (2016). Towards a unified sub‐symbolic computational theory of cognition. Frontiers in Psychology, 7, 925. https://doi.org/10.3389/fpsyg.2016.00925</bibtext> </blist> <blist> <bibtext> Butz, M. V. (2017). Which structures are out there. In T. K. Metzinger & W. Wiese (Eds.), Philosophy and predictive processing. Frankfurt am Main : MIND Group. https://doi.org/10.25358/openscience‐631</bibtext> </blist> <blist> <bibtext> Butz, M. V., Achimova, A., Bilkey, D., & Knott, A. (2021). Event‐predictive cognition: A root for conceptual human thought. Topics in Cognitive Science, 13 (1), 10 – 24. https://doi.org/10.1111/tops.12522</bibtext> </blist> <blist> <bibtext> Butz, M. V., Bilkey, D., Humaidan, D., Knott, A., & Otte, S. (2019). Learning, planning, and control in a monolithic neural event inference architecture. Neural Networks, 117, 135 – 144. https://doi.org/10.1016/j.neunet.2019.05.001</bibtext> </blist> <blist> <bibtext> Cannon, E. N., & Woodward, A. L. (2012). Infants generate goal‐based action predictions. Developmental Science, 15 (2), 292 – 298. https://doi.org/10.1111/j.1467‐7687.2011.01127.x</bibtext> </blist> <blist> <bibtext> Cannon, E. N., Woodward, A. L., Gredebäck, G., Hofsten, C. Von., & Turek, C. (2012). Action production influences 12‐month‐old infants' attention to others' actions. Developmental Science, 15 (1), 35 – 42. https://doi.org/10.1111/j.1467‐7687.2011.01095.x</bibtext> </blist> <blist> <bibtext> Cooper, R. P. (2021). Action production and event perception as routine sequential behaviors. Topics in Cognitive Science, 13 (1), 63 – 78. https://doi.org/10.1111/tops.12462</bibtext> </blist> <blist> <bibtext> Copete, J. L., Nagai, Y., & Asada, M. (2016). Motor development facilitates the prediction of others' actions through sensorimotor predictive learning. In 2016 joint IEEE international conference on development and learning and epigenetic robotics (ICDL‐EpiRob) (pp. 223 – 229). Piscataway, NJ : IEEE. https://doi.org/10.1109/DEVLRN.2016.7846823</bibtext> </blist> <blist> <bibtext> Daum, M., Attig, M., Gunawan, R., Prinz, W., & Gredebäck, G. (2012). Actions seen through babies' eyes: A dissociation between looking time and predictive gaze. Frontiers in Psychology, 3, 370. https://doi.org/10.3389/fpsyg.2012.00370</bibtext> </blist> <blist> <bibtext> Elsner, B., & Adam, M. (2021). Infants' goal prediction for simple action events: The role of experience and agency cues. Topics in Cognitive Science, 13 (1), 45 – 62. https://doi.org/10.1111/tops.12494</bibtext> </blist> <blist> <bibtext> Elsner, B., & Hommel, B. (2001). Effect anticipation and action control. Journal of Experimental Psychology: Human Perception and Performance, 27 (1), 229 – 240. https://doi.org/10.1037//0096‐1523.27.1.229</bibtext> </blist> <blist> <bibtext> Falck‐Ytter, T., Gredebäck, G., & Hofsten, C. von. (2006). Infants predict other people's action goals. Nature Neuroscience, 9 (7), 878 – 879. https://doi.org/10.1038/nn1729</bibtext> </blist> <blist> <bibtext> Fantz, R. L. (1958). Pattern vision in young infants. Psychological Record, 8, 43 – 47. https://doi.org/10.1007/BF03393306</bibtext> </blist> <blist> <bibtext> Flanagan, J. R., & Johansson, R. S. (2003). Action plans used in action observation. Nature, 424 (6950), 769 – 771. https://doi.org/10.1038/nature01861</bibtext> </blist> <blist> <bibtext> Franklin, N. T., Norman, K. A., Ranganath, C., Zacks, J. M., & Gershman, S. J. (2020). Structured event memory: A neuro‐symbolic model of event cognition. Psychological Review, 127 (3), 327 – 361. https://doi.org/10.1037/rev0000177</bibtext> </blist> <blist> <bibtext> Friston, K. (2010). The free‐energy principle: A unified brain theory? Nature Reviews Neuroscience, 11 (2), 127 – 138. https://doi.org/10.1038/nrn2787</bibtext> </blist> <blist> <bibtext> Friston, K., FitzGerald, T., Rigoli, F., Schwartenbeck, P., O'Doherty, J., & Pezzulo, G. (2016). Active inference and learning. Neuroscience & Biobehavioral Reviews, 68, 862 – 879. https://doi.org/10.1016/j.neubiorev.2016.06.022</bibtext> </blist> <blist> <bibtext> Friston, K., Rigoli, F., Ognibene, D., Mathys, C., FitzGerald, T., & Pezzulo, G. (2015). Active inference and epistemic value. Cognitive Neuroscience, 6, 187 – 214. https://doi.org/10.1080/17588928.2015.1020053</bibtext> </blist> <blist> <bibtext> Gallese, V., Fadiga, L., Fogassi, L., & Rizzolatti, G. (1996). Action recognition in the premotor cortex. Brain, 119 (2), 593 – 609. https://doi.org/10.1093/brain/119.2.593</bibtext> </blist> <blist> <bibtext> Ganglmayer, K., Attig, M., Daum, M. M., & Paulus, M. (2019). Infants' perception of goal‐directed actions: A multi‐lab replication reveals that infants anticipate paths and not goals. Infant Behavior and Development, 57, 101340. https://doi.org/10.1016/j.infbeh.2019.101340</bibtext> </blist> <blist> <bibtext> Gredebäck, G., & Falck‐Ytter, T. (2015). Eye movements during action observation. Perspectives on Psychological Science, 10 (5), 591 – 598. https://doi.org/10.1177/1745691615589103</bibtext> </blist> <blist> <bibtext> Gredebäck, G., Johnson, S., & Hofsten, C. von. (2010). Eye tracking in infancy research. Developmental Neuropsychology, 35 (1), 1 – 19. https://doi.org/10.1080/87565640903325758</bibtext> </blist> <blist> <bibtext> Gredebäck, G., & Melinder, A. (2010). Infants' understanding of everyday social interactions: A dual process account. Cognition, 114 (2), 197 – 206. https://doi.org/10.1016/j.cognition.2009.09.004</bibtext> </blist> <blist> <bibtext> Gredebäck, G., Stasiewicz, D., Falck‐Ytter, T., Rosander, K., & von Hofsten, C. (2009). Action type and goal type modulate goal‐directed gaze shifts in 14‐month‐old infants. Developmental Psychology, 45 (4), 1190 – 1194. https://doi.org/10.1037/a0015667</bibtext> </blist> <blist> <bibtext> Gumbsch, C., Butz, M. V., & Martius, G. (2019). Autonomous identification and goal‐directed invocation of event‐predictive behavioral primitives. IEEE Transactions on Cognitive and Developmental Systems, 13, 298 – 311. https://doi.org/10.1109/TCDS.2019.2925890</bibtext> </blist> <blist> <bibtext> Hayhoe, M. M., Shrivastava, A., Mruczek, R., & Pelz, J. B. (2003). Visual memory and motor planning in a natural task. Journal of Vision, 3 (1), 49 – 63. https://doi.org/10.1167/3.1.6</bibtext> </blist> <blist> <bibtext> Hommel, B. (2009). Action control according to TEC (theory of event coding). Psychological Research PRPF, 73 (4), 512 – 526. https://doi.org/10.1007/s00426‐009‐0234‐2</bibtext> </blist> <blist> <bibtext> Hommel, B. (2015). The theory of event coding (TEC) as embodied‐cognition framework. Frontiers in Psychology, 6, 1318. https://doi.org/10.3389/fpsyg.2015.01318</bibtext> </blist> <blist> <bibtext> Hommel, B., Müsseler, J., Aschersleben, G., & Prinz, W. (2001). The theory of event coding (TEC): A framework for perception and action planning. Behavioral and Brain Sciences, 24 (5), 849 – 878. https://doi.org/10.1017/s0140525x01000103</bibtext> </blist> <blist> <bibtext> Kanakogi, Y., & Itakura, S. (2011). Developmental correspondence between action prediction and motor ability in early infancy. Nature Communications, 2, 341. https://doi.org/10.1038/ncomms1342</bibtext> </blist> <blist> <bibtext> Knill, D. C., & Pouget, A. (2004). The Bayesian brain: The role of uncertainty in neural coding and computation. Trends in Neurosciences, 27 (12), 712 – 719. https://doi.org/10.1016/j.tins.2004.10.007</bibtext> </blist> <blist> <bibtext> Krogh‐Jespersen, S., & Woodward, A. L. (2014). Making smart social judgments takes time: Infants' recruitment of goal information when generating action predictions. PloS One, 9 (5), e98085. https://doi.org/10.1371/journal.pone.0098085</bibtext> </blist> <blist> <bibtext> Kuperberg, G. R. (2021). Tea with milk? A hierarchical generative framework of sequential event comprehension. Topics in Cognitive Science, 13 (1), 256 – 298. https://doi.org/10.1111/tops.12518</bibtext> </blist> <blist> <bibtext> Lohmann, J., Belardinelli, A., & Butz, M. V. (2019). Hands ahead in mind and motion: Active inference in peripersonal hand space. Vision, 3 (2), 15. https://doi.org/10.3390/vision3020015</bibtext> </blist> <blist> <bibtext> Moll, H., & Meltzoff, A. (2011). Perspective‐taking and its foundation in joint attention. In J. Roessler, H. Lerman, & N. Eilan (Eds.), Joint attention: New developments in psychology, philosophy of mind, and social neuroscience (pp. 286 – 304). Oxford, England : Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199692040.003.0016</bibtext> </blist> <blist> <bibtext> Radvansky, G. A., & Zacks, J. M. (2014). Event cognition. Oxford, England : Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199898138.001.0001</bibtext> </blist> <blist> <bibtext> Rao, R. P., & Ballard, D. H. (1999). Predictive coding in the visual cortex: A functional interpretation of some extra‐classical receptive‐field effects. Nature Neuroscience, 2 (1), 79. https://doi.org/10.1038/4580</bibtext> </blist> <blist> <bibtext> Reynolds, J. R., Zacks, J. M., & Braver, T. S. (2007). A computational model of event segmentation from perceptual prediction. Cognitive Science, 31 (4), 613 – 643. https://doi.org/10.1080/15326900701399913</bibtext> </blist> <blist> <bibtext> Rizzolatti, G., Fadiga, L., Gallese, V., & Fogassi, L. (1996). Premotor cortex and the recognition of motor actions. Cognitive Brain Research, 3 (2), 131 – 141. https://doi.org/10.1016/0926‐6410(95)00038‐0</bibtext> </blist> <blist> <bibtext> Rizzolatti, G., Fogassi, L., & Gallese, V. (2001). Neurophysiological mechanisms underlying the understanding and imitation of action. Nature reviews Neuroscience, 2 (9), 661 – 670. https://doi.org/10.1038/35090060</bibtext> </blist> <blist> <bibtext> Schrodt, F., & Butz, M. V. (2016). Just imagine! learning to emulate and infer actions with a stochastic generative architecture. Frontiers in Robotics and AI, 3, 5. https://doi.org/10.3389/frobt.2016.00005</bibtext> </blist> <blist> <bibtext> Southgate, V., & Begus, K. (2013). Motor activation during the prediction of nonexecutable actions in infants. Psychological Science, 24 (6), 828 – 835. https://doi.org/10.1177/0956797612459766</bibtext> </blist> <blist> <bibtext> Stawarczyk, D., Bezdek, M. A., & Zacks, J. M. (2021). Event representations and predictive processing: The role of the midline default network core. Topics in Cognitive Science, 13 (1), 164 – 186. https://doi.org/10.1111/tops.12450</bibtext> </blist> <blist> <bibtext> Tversky, B., & Hard, B. M. (2009). Embodied and disembodied cognition: Spatial perspective‐taking. Cognition, 110 (1), 124 – 129. https://doi.org/10.1016/j.cognition.2008.10.008</bibtext> </blist> <blist> <bibtext> Ullman, S., Harari, D., & Dorfman, N. (2012). From simple innate biases to complex visual concepts. Proceedings of the National Academy of Sciences, 109 (44), 18215 – 18220. https://doi.org/10.1073/pnas.1207690109</bibtext> </blist> <blist> <bibtext> Woodward, A. L. (1998). Infants selectively encode the goal object of an actor's reach. Cognition, 69 (1), 1 – 34. https://doi.org/10.1016/S0010‐0277(98)00058‐4</bibtext> </blist> <blist> <bibtext> Zacks, J. M., Speer, N. K., Swallow, K. M., Braver, T. S., & Reynolds, J. R. (2007). Event perception: A mind‐brain perspective. Psychological Bulletin, 133 (2), 273 – 293. https://doi.org/10.1037/0033‐2909.133.2.273</bibtext> </blist> <blist> <bibtext> Zacks, J. M., & Swallow, K. M. (2007). Event segmentation. Current Directions in Psychological Science, 16 (2), 80 – 84. https://doi.org/10.1111/j.1467‐8721.2007.00480.x</bibtext> </blist> <blist> <bibtext> Zacks, J. M., & Tversky, B. (2001). Event structure in perception and conception. Psychological Bulletin, 127 (1), 3 – 21. https://doi.org/10.1037/0033‐2909.127.1.3</bibtext> </blist> </ref> <aug> <p>By Christian Gumbsch; Maurits Adam; Birgit Elsner and Martin V. Butz</p> <p>Reported by Author; Author; Author; Author</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1310726
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Emergent Goal-Anticipatory Gaze in Infants via Event-Predictive Learning and Inference
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Gumbsch%2C+Christian%22">Gumbsch, Christian</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2741-6551">0000-0003-2741-6551</externalLink>)<br /><searchLink fieldCode="AR" term="%22Adam%2C+Maurits%22">Adam, Maurits</searchLink><br /><searchLink fieldCode="AR" term="%22Elsner%2C+Birgit%22">Elsner, Birgit</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-3441-2436">0000-0003-3441-2436</externalLink>)<br /><searchLink fieldCode="AR" term="%22Butz%2C+Martin+V%2E%22">Butz, Martin V.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8120-8537">0000-0002-8120-8537</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Cognitive+Science%22"><i>Cognitive Science</i></searchLink>. Aug 2021 45(8).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 26
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2021
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Goal+Orientation%22">Goal Orientation</searchLink><br /><searchLink fieldCode="DE" term="%22Infants%22">Infants</searchLink><br /><searchLink fieldCode="DE" term="%22Eye+Movements%22">Eye Movements</searchLink><br /><searchLink fieldCode="DE" term="%22Cognitive+Processes%22">Cognitive Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Psychomotor+Skills%22">Psychomotor Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Familiarity%22">Familiarity</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Inferences%22">Inferences</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Infant+Behavior%22">Infant Behavior</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/cogs.13016
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1551-6709
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: From about 7 months of age onward, infants start to reliably fixate the goal of an observed action, such as a grasp, before the action is complete. The available research has identified a variety of factors that influence such goal-anticipatory gaze shifts, including the experience with the shown action events and familiarity with the observed agents. However, the underlying cognitive processes are still heavily debated. We propose that our minds (i) tend to structure sensorimotor dynamics into probabilistic, generative event-predictive, and event boundary predictive models, and, meanwhile, (ii) choose actions with the objective to minimize predicted uncertainty. We implement this proposition by means of event-predictive learning and active inference. The implemented learning mechanism induces an inductive, event-predictive bias, thus developing schematic encodings of experienced events and event boundaries. The implemented active inference principle chooses actions by aiming at minimizing expected future uncertainty. We train our system on multiple object-manipulation events. As a result, the generation of goal-anticipatory gaze shifts emerges while learning about object manipulations: the model starts fixating the inferred goal already at the start of an observed event after having sampled some experience with possible events and when a familiar agent (i.e., a hand) is involved. Meanwhile, the model keeps reactively tracking an unfamiliar agent (i.e., a mechanical claw) that is performing the same movement. We qualitatively compare these modeling results to behavioral data of infants and conclude that event-predictive learning combined with active inference may be critical for eliciting goal-anticipatory gaze behavior in infants.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2021
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1310726
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1310726
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/cogs.13016
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 26
    Subjects:
      – SubjectFull: Goal Orientation
        Type: general
      – SubjectFull: Infants
        Type: general
      – SubjectFull: Eye Movements
        Type: general
      – SubjectFull: Cognitive Processes
        Type: general
      – SubjectFull: Prediction
        Type: general
      – SubjectFull: Psychomotor Skills
        Type: general
      – SubjectFull: Familiarity
        Type: general
      – SubjectFull: Learning Processes
        Type: general
      – SubjectFull: Inferences
        Type: general
      – SubjectFull: Comparative Analysis
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Infant Behavior
        Type: general
    Titles:
      – TitleFull: Emergent Goal-Anticipatory Gaze in Infants via Event-Predictive Learning and Inference
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Gumbsch, Christian
      – PersonEntity:
          Name:
            NameFull: Adam, Maurits
      – PersonEntity:
          Name:
            NameFull: Elsner, Birgit
      – PersonEntity:
          Name:
            NameFull: Butz, Martin V.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 08
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-electronic
              Value: 1551-6709
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 8
          Titles:
            – TitleFull: Cognitive Science
              Type: main
ResultId 1