File 012899
Engineering General Intelligence, Part 1: A Path to Advanced AGI via Embodied Learning and Cognitive Synergy (File 012899)
Academic book by Ben Goertzel and colleagues outlining a practical approach to artificial general intelligence (AGI) through the CogPrime architecture and OpenCog framework, published in September 2013.
Summary
This is Part 1 of a comprehensive two-volume academic work on engineering artificial general intelligence at human level and beyond. Written by Ben Goertzel, Cassio Pennachin, and Nil Geisweiller, the book outlines the CogPrime architecture and the OpenCog open-source software project as a practical pathway to AGI. The volume covers conceptual foundations of intelligence and mind, the novel CogPrime integrative architecture, and approaches for developing AGI systems through appropriate experience and learning. The authors present both theoretical frameworks and engineering blueprints for creating machines with flexible problem-solving abilities, creativity, and potential for self-modification beyond human cognitive capabilities.
Ben Goertzel with Cassio Pennachin & Nil Geisweiller &the OpenCog TeamEngineering General Intelligence, Part 1:A Path to Advanced AGI via Embodied Learning andCognitive SynergySeptember 19, 2013
This book is dedicated by Ben Goertzel to his beloved,departed grandfather, Leo Zwell – an amazinglywarm-hearted, giving human being who was also a deepthinker and excellent scientist, who got Ben started on thepath of science. As a careful experimentalist, Leo wouldhave been properly skeptical of the big hypotheses madehere – but he would have been eager to see them put to thetest!
PrefaceThis is a large, two-part book with an even larger goal: To outline a practical approach toengineering software systems with general intelligence at the human level and ultimately beyond.Machines with flexible problem-solving ability, open-ended learning capability, creativity andeventually, their own kind of genius.Part 1, this volume, reviews various critical conceptual issues related to the nature of intelligenceand mind. It then sketches the broad outlines of a novel, integrative architecture forArtificial General Intelligence (AGI) called CogPrime ... and describes an approach for giving ayoung AGI system (CogPrime or otherwise) appropriate experience, so that it can develop itsown smarts, creativity and wisdom through its own experience. Along the way a formal theoryof general intelligence is sketched, and a broad roadmap leading from here to human-level artificialintelligence. Hints are also given regarding how to eventually, potentially create machinesadvancing beyond human level – including some frankly futuristic speculations about stronglyself-modifying AGI architectures with flexibility far exceeding that of the human brain.Part 2 then digs far deeper into the details of CogPrime’s multiple structures, processes andfunctions, culminating in a general argument as to why we believe CogPrime will be able toachieve general intelligence at the level of the smartest humans (and potentially greater), anda detailed discussion of how a CogPrime-powered virtual agent or robot would handle somesimple practical tasks such as social play with blocks in a preschool context. It first describesthe CogPrime software architecture and knowledge representation in detail; then reviews thecognitive cycle via which CogPrime perceives and acts in the world and reflects on itself; andnext turns to various forms of learning: procedural, declarative (e.g. inference), simulative andintegrative. Methods of enabling natural language functionality in CogPrime are then discussed;and then the volume concludes with a chapter summarizing the argument that CogPrime canlead to human-level (and eventually perhaps greater) AGI, and a chapter giving a thoughtexperiment describing the internal dynamics via which a completed CogPrime system mightsolve the problem of obeying the request “Build me something with blocks that I haven’t seenbefore.”The chapters here are written to be read in linear order – and if consumed thus, they tella coherent story about how to get from here to advanced AGI. However, the impatient readermay be forgiven for proceeding a bit nonlinearly. An alternate reading path for the impatientreader would be to start with the first few chapters of Part 1, then skim the final two chapters ofPart 2, and then return to reading in linear order. The final two chapters of Part 2 give a broadoverview of why we think the CogPrime design will work, in a way that depends on the technicalviiviiidetails of the previous chapters, but (we believe) not so sensitively as to be incomprehensiblewithout them.This is admittedly an unusual sort of book, mixing demonstrated conclusions with unprovedconjectures in a complex way, all oriented toward an extraordinarily ambitious goal. Further,the chapters are somewhat variant in their levels of detail – some very nitty-gritty, some morehigh level, with much of the variation due to how much concrete work has been done on thetopic of the chapter at time of writing. However, it is important to understand that the ideaspresented here are not mere armchair speculation – they are currently being used as the basisfor an open-source software project called OpenCog, which is being worked on by softwaredevelopers around the world. Right now OpenCog embodies only a percentage of the overallCogPrime design as described here. But if OpenCog continues to attract sufficient fundingor volunteer interest, then the ideas presented in these volumes will be validated or refutedvia practice. (As a related note: here and there in this book, we will refer to the "current"CogPrime implementation (in the OpenCog framework); in all cases this refers to OpenCog asof late 2013.)To state one believes one knows a workable path to creating a human-level (and potentiallygreater) general intelligence is to make a dramatic statement, given the conventional way ofthinking about the topic in the contemporary scientific community. However, we feel that oncea little more time has passed, the topic will lose its drama (if not its interest and importance),and it will be widely accepted that there are many ways to create intelligent machines – somesimpler and some more complicated; some more brain-like or human-like and some less so; somemore efficient and some more wasteful of resources; etc. We have little doubt that, from theperspective of AGI science 50 or 100 years hence (and probably even 10-20 years hence), thespecific designs presented here will seem awkward, messy, inefficient and circuitous in variousrespects. But that is how science and engineering progress. Given the current state of knowledgeand understanding, having any concrete, comprehensive design and plan for creating AGI isa significant step forward; and it is in this spirit that we present here our thinking about theCogPrime architecture and the nature of general intelligence.In the words of Sir Edmund Hillary, the first to scale Everest: “Nothing Venture, NothingWin.”Prehistory of the BookThe writing of this book began in earnest in 2001, at which point it was informally referred toas “The Novamente Book.” The original “Novamente Book” manuscript ultimately got too bigfor its own britches, and subdivided into a number of different works – The Hidden Pattern[Goe06a], a philosophy of mind book published in 2006; Probabilistic Logic Networks [GIGH08],a more technical work published in 2008; Real World Reasoning [GGC + 11], a sequel to ProbabilisticLogic Networks published in 2011; and the two parts of this book.The ideas described in this book have been the collaborative creation of multiple overlappingcommunities of people over a long period of time. The vast bulk of the writing here was done byBen Goertzel; but Cassio Pennachin and Nil Geisweiller made sufficient writing, thinking andediting contributions over the years to more than merit their inclusion of co-authors. Further,many of the chapters here have co-authors beyond the three main co-authors of the book; andthe set of chapter co-authors does not exhaust the set of significant contributors to the ideaspresented.The core concepts of the CogPrime design and the underlying theory were conceived by BenGoertzel in the period 1995-1996 when he was a Research Fellow at the University of WesternAustralia; but those early ideas have been elaborated and improved by many more people thancan be listed here (as well as by Ben’s ongoing thinking and research). The collaborative designprocess ultimately resulting in CogPrime started in 1997 when Intelligenesis Corp. was formed– the Webmind AI Engine created in Intelligenesis’s research group during 1997-2001 was thepredecessor to the Novamente Cognition Engine created at Novamente LLC during 2001-2008,which was the predecessor to CogPrime.ixAcknowledgementsFor sake of simplicity, this acknowledgements section is presented from the perspective of theprimary author, Ben Goertzel. Ben will thus begin by expressing his thanks to his primaryco-authors, Cassio Pennachin (collaborator since 1998) and Nil Geisweiller (collaborator since2005). Without outstandingly insightful, deep-thinking colleagues like you, the ideas presentedhere – let alone the book itself– would not have developed nearly as effectively as what hashappened. Similar thanks also go to the other OpenCog collaborators who have co-authoredvarious chapters of the book.Beyond the co-authors, huge gratitude must also be extended to everyone who has beeninvolved with the OpenCog project, and/or was involved in Novamente LLC and Webmind Inc.before that. We are grateful to all of you for your collaboration and intellectual companionship!Building a thinking machine is a huge project, too big for any one human; it will take a teamand I’m happy to be part of a great one. It is through the genius of human collectives, goingbeyond any individual human mind, that genius machines are going to be created.A tiny, incomplete sample from the long list of those others deserving thanks is:• Ken Silverman and Gwendalin Qi Aranya (formerly Gwen Goertzel), both of whom listenedto me talk at inordinate length about many of the ideas presented here a long, long timebefore anyone else was interested in listening. Ken and I schemed some AGI designs atSimon’s Rock College in 1983, years before we worked together on the Webmind AI Engine.• Allan Combs, who got me thinking about consciousness in various different ways, at a veryearly point in my career. I’m very pleased to still count Allan as a friend and sometimecollaborator! Fred Abraham as well, for introducing me to the intersection of chaos theoryand cognition, with a wonderful flair. George Christos, a deep AI/math/physics thinker fromPerth, for re-awakening my interest in attractor neural nets and their cognitive implications,in the mid-1990s.• All of the 130 staff of Webmind Inc. during 1998-2001 while that remarkable, ambitious,peculiar AGI-oriented firm existed. Special shout-outs to the "Voice of Reason" Pei Wangand the "Siberian Madmind" Anton Kolonin, Mike Ross, Cate Hartley, Karin Verspoor andthe tragically prematurely deceased Jeff Pressing (compared to whom we are all mentalmidgets), who all made serious conceptual contributions to my thinking about AGI. LisaPazer and Andy Siciliano who made Webmind happen on the business side. And of courseCassio Pennachin, a co-author of this book; and Ken Silverman, who co-architected thewhole Webmind system and vision with me from the start.x• The Webmind Diehards, who helped begin the Novamente project that succeeded Webmindbeginning in 2001: Cassio Pennachin, Stephan Vladimir Bugaj, Takuo Henmi, MatthewIkle’, Thiago Maia, Andre Senna, Guilherme Lamacie and Saulo Pinto• Those who helped get the Novamente project off the ground and keep it progressing over theyears, including some of the Webmind Diehards and also Moshe Looks, Bruce Klein, IzabelaLyon Freire, Chris Poulin, Murilo Queiroz, Predrag Janicic, David Hart, Ari Heljakka, HugoPinto, Deborah Duong, Paul Prueitt, Glenn Tarbox, Nil Geisweiller and Cassio Pennachin(the co-authors of this book), Sibley Verbeck, Jeff Reed, Pejman Makhfi, Welter Silva,Lukasz Kaiser and more• All those who have helped with the OpenCog system, including Linas Vepstas, Joel Pitt,Jared Wigmore / Jade O’Neill, Zhenhua Cai, Deheng Huang, Shujing Ke, Lake Watkins,Alex van der Peet, Samir Araujo, Fabricio Silva, Yang Ye, Shuo Chen, Michel Drenthe, TedSanders, Gustavo Gama and of course Nil and Cassio again. Tyler Emerson and EliezerYudkowsky, for choosing to have the Singularity Institute for AI (now MIRI) provide seedfunding for OpenCog.• The numerous members of the AGI community who have tossed around AGI ideas with mesince the first AGI conference in 2006, including but definitely not limited to: Stan Franklin,Juergen Schmidhuber, Marcus Hutter, Kai-Uwe Kuehnberger, Stephen Reed, Blerim Enruli,Kristinn Thorisson, Joscha Bach, Abram Demski, Itamar Arel, Mark Waser, Randal Koene,Paul Rosenbloom, Zhongzhi Shi, Steve Omohundro, Bill Hibbard, Eray Ozkural, BrandonRohrer, Ben Johnston, John Laird, Shane Legg, Selmer Bringsjord, Anders Sandberg, AlexeiSamsonovich, Wlodek Duch, and more• The inimitable "Artilect Warrior" Hugo de Garis, who (when he was working at XiamenUniversity) got me started working on AGI in the Orient (and introduced me to my wifeRuiting in the process). And Changle Zhou, who brought Hugo to Xiamen and generouslyshared his brilliant research students with Hugo and me. And Min Jiang, collaborator ofHugo and Changle, a deep AGI thinker who is helping with OpenCog theory and practiceat time of writing.• Gino Yu, who got me started working on AGI here in Hong Kong, where I am living at timeof writing. As of 2013 the bulk of OpenCog work is occurring in Hong Kong via a researchgrant that Gino and I obtained together• Dan Stoicescu, whose funding helped Novamente through some tough times.• Jeffrey Epstein, whose visionary funding of my AGI research has helped me through anumber of tight spots over the years. At time of writing, Jeffrey is helping support theOpenCog Hong Kong project.• Zeger Karssen, founder of Atlantis Press, who conceived the Thinking Machines book seriesin which this book appears, and who has been a strong supporter of the AGI conferenceseries from the beginning• My wonderful wife Ruiting Lian, a source of fantastic amounts of positive energy for mesince we became involved four years ago. Ruiting has listened to me discuss the ideascontained here time and time again, often with judicious and insightful feedback (as sheis an excellent AI researcher in her own right); and has been wonderfully tolerant of mediverting numerous evenings and weekends to getting this book finished (as well as to otherAGI-related pursuits). And my parents Ted and Carol and kids Zar, Zeb and Zade, whohave also indulged me in discussions on many of the themes discussed here on countlessoccasions! And my dear, departed grandfather Leo Zwell, for getting me started in science.• Crunchkin and Pumpkin, for regularly getting me away from the desk to stroll around thevillage where we live; many of my best ideas about AGI and other topics have emergedwhile walking with my furry four-legged family membersxiSeptember 2013Ben Goertzel
Contents1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11.1 AI Returns to Its Roots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11.2 AGI versus Narrow AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21.3 CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31.4 The Secret Sauce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31.5 Extraordinary Proof? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41.6 Potential Approaches to AGI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61.6.1 Build AGI from Narrow AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61.6.2 Enhancing Chatbots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61.6.3 Emulating the Brain . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61.6.4 Evolve an AGI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71.6.5 Derive an AGI design mathematically . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71.6.6 Use heuristic computer science methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81.6.7 Integrative Cognitive Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81.6.8 Can Digital Computers Really Be Intelligent? . . . . . . . . . . . . . . . . . . . . . . . . 81.7 Five Key Words . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91.7.1 Memory and Cognition in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101.8 Virtually and Robotically Embodied AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111.9 Language Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121.10 AGI Ethics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121.11 Structure of the Book . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131.12 Key Claims of the Book . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13Section I Artificial and Natural General Intelligence2 What Is Human-Like General Intelligence? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192.1.1 What Is General Intelligence? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192.1.2 What Is Human-like General Intelligence? . . . . . . . . . . . . . . . . . . . . . . . . . . . 202.2 Commonly Recognized Aspects of Human-like Intelligence . . . . . . . . . . . . . . . . . . . 202.3 Further Characterizations of Humanlike Intelligence . . . . . . . . . . . . . . . . . . . . . . . . 242.3.1 Competencies Characterizing Human-like Intelligence . . . . . . . . . . . . . . . . . 242.3.2 Gardner’s Theory of Multiple Intelligences . . . . . . . . . . . . . . . . . . . . . . . . . . . 25xiiixivContents2.3.3 Newell’s Criteria for a Human Cognitive Architecture . . . . . . . . . . . . . . . . . 262.3.4 intelligence and Creativity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262.4 Preschool as a View into Human-like General Intelligence . . . . . . . . . . . . . . . . . . . . 272.4.1 Design for an AGI Preschool . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 282.5 Integrative and Synergetic Approaches to Artificial General Intelligence . . . . . . . 292.5.1 Achieving Humanlike Intelligence via Cognitive Synergy . . . . . . . . . . . . . . . 303 A Patternist Philosophy of Mind . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353.2 Some Patternist Principles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353.3 Cognitive Synergy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 403.4 The General Structure of Cognitive Dynamics: Analysis and Synthesis . . . . . . . . 423.4.1 Component-Systems and Self-Generating Systems . . . . . . . . . . . . . . . . . . . . 423.4.2 Analysis and Synthesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 433.4.3 The Dynamic of Iterative Analysis and Synthesis . . . . . . . . . . . . . . . . . . . . 463.4.4 Self and Focused Attention as Approximate Attractors of the Dynamicof Iterated Forward-Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 473.4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503.5 Perspectives on Machine Consciousness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 513.6 Postscript: Formalizing Pattern . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 534 Brief Survey of Cognitive Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 574.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 574.2 Symbolic Cognitive Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 584.2.1 SOAR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 604.2.2 ACT-R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 614.2.3 Cyc and Texai . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 624.2.4 NARS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 634.2.5 GLAIR and SNePS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 644.3 Emergentist Cognitive Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 654.3.1 DeSTIN: A Deep Reinforcement Learning Approach to AGI . . . . . . . . . . . 664.3.2 Developmental Robotics Architectures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 724.4 Hybrid Cognitive Architectures. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 734.4.1 Neural versus Symbolic; Global versus Local . . . . . . . . . . . . . . . . . . . . . . . . . 754.5 Globalist versus Localist Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 784.5.1 CLARION . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 794.5.2 The Society of Mind and the Emotion Machine . . . . . . . . . . . . . . . . . . . . . . 804.5.3 DUAL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 804.5.4 4D/RCS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 814.5.5 PolyScheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 824.5.6 Joshua Blue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 834.5.7 LIDA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 844.5.8 The Global Workspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 844.5.9 The LIDA Cognitive Cycle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 854.5.10 Psi and MicroPsi . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 884.5.11 The Emergence of Emotion in the Psi Model . . . . . . . . . . . . . . . . . . . . . . . . . 914.5.12 Knowledge Representation, Action Selection and Planning in Psi . . . . . . . 93Contentsxv4.5.13 Psi versus CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 945 A Generic Architecture of Human-Like Cognition . . . . . . . . . . . . . . . . . . . . . . . . 955.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 955.2 Key Ingredients of the Integrative Human-Like Cognitive Architecture Diagram 965.3 An Architecture Diagram for Human-Like General Intelligence . . . . . . . . . . . . . . . 975.4 Interpretation and Application of the Integrative Diagram . . . . . . . . . . . . . . . . . . . 1046 A Brief Overview of CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1076.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1076.2 High-Level Architecture of CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1076.3 Current and Prior Applications of OpenCog . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1086.3.1 Transitioning from Virtual Agents to a Physical Robot . . . . . . . . . . . . . . . . 1106.4 Memory Types and Associated Cognitive Processes in CogPrime . . . . . . . . . . . . . 1106.4.1 Cognitive Synergy in PLN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1116.5 Goal-Oriented Dynamics in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1136.6 Analysis and Synthesis Processes in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1146.7 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116Section II Toward a General Theory of General Intelligence7 A Formal Model of Intelligent Agents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1297.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1297.2 A Simple Formal Agents Model (SRAM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1307.2.1 Goals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1317.2.2 Memory Stores . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1327.2.3 The Cognitive Schematic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1337.3 Toward a Formal Characterization of Real-World General Intelligence . . . . . . . . . 1357.3.1 Biased Universal Intelligence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1367.3.2 Connecting Legg and Hutter’s Model of Intelligent Agents to the RealWorld . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1377.3.3 Pragmatic General Intelligence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1387.3.4 Incorporating Computational Cost . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1397.3.5 Assessing the Intelligence of Real-World Agents . . . . . . . . . . . . . . . . . . . . . . 1397.4 Intellectual Breadth: Quantifying the Generality of an Agent’s Intelligence . . . . . 1417.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1428 Cognitive Synergy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1438.1 Cognitive Synergy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1438.2 Cognitive Synergy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1448.3 Cognitive Synergy in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1468.3.1 Cognitive Processes in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1468.4 Some Critical Synergies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1498.5 The Cognitive Schematic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1518.6 Cognitive Synergy for Procedural and Declarative Learning . . . . . . . . . . . . . . . . . . 1538.6.1 Cognitive Synergy in MOSES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1538.6.2 Cognitive Synergy in PLN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1558.7 Is Cognitive Synergy Tricky? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157xviContents8.7.1 The Puzzle: Why Is It So Hard to Measure Partial Progress TowardHuman-Level AGI? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1578.7.2 A Possible Answer: Cognitive Synergy is Tricky! . . . . . . . . . . . . . . . . . . . . . . 1588.7.3 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1599 General Intelligence in the Everyday Human World . . . . . . . . . . . . . . . . . . . . . . 1619.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1619.2 Some Broad Properties of the Everyday World That Help Structure Intelligence 1629.3 Embodied Communication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1639.3.1 Generalizing the Embodied Communication Prior . . . . . . . . . . . . . . . . . . . . 1669.4 Naive Physics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1669.4.1 Objects, Natural Units and Natural Kinds . . . . . . . . . . . . . . . . . . . . . . . . . . 1679.4.2 Events, Processes and Causality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1689.4.3 Stuffs, States of Matter, Qualities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1689.4.4 Surfaces, Limits, Boundaries, Media . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1689.4.5 What Kind of Physics Is Needed to Foster Human-like Intelligence? . . . . . 1699.5 Folk Psychology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1709.5.1 Motivation, Requiredness, Value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1719.6 Body and Mind . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1719.6.1 The Human Sensorium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1719.6.2 The Human Body’s Multiple Intelligences . . . . . . . . . . . . . . . . . . . . . . . . . . . 1729.7 The Extended Mind and Body . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1769.8 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17610 A Mind-World Correspondence Principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17710.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17710.2 What Might a General Theory of General Intelligence Look Like? . . . . . . . . . . . . 17810.3 Steps Toward A (Formal) General Theory of General Intelligence . . . . . . . . . . . . . 17910.4 The Mind-World Correspondence Principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18010.5 How Might the Mind-World Correspondence Principle Be Useful? . . . . . . . . . . . . 18110.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182Section III Cognitive and Ethical Development11 Stages of Cognitive Development . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18711.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18711.2 Piagetan Stages in the Context of a General Systems Theory of Development . . 18811.3 Piaget’s Theory of Cognitive Development . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18811.3.1 Perry’s Stages. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19211.3.2 Keeping Continuity in Mind . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19211.4 Piaget’s Stages in the Context of Uncertain Inference . . . . . . . . . . . . . . . . . . . . . . . 19311.4.1 The Infantile Stage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19511.4.2 The Concrete Stage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19611.4.3 The Formal Stage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20011.4.4 The Reflexive Stage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202Contentsxvii12 The Engineering and Development of Ethics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20512.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20512.2 Review of Current Thinking on the Risks of AGI . . . . . . . . . . . . . . . . . . . . . . . . . . . 20612.3 The Value of an Explicit Goal System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20912.4 Ethical Synergy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21012.4.1 Stages of Development of Declarative Ethics . . . . . . . . . . . . . . . . . . . . . . . . . 21112.4.2 Stages of Development of Empathic Ethics . . . . . . . . . . . . . . . . . . . . . . . . . . 21412.4.3 An Integrative Approach to Ethical Development . . . . . . . . . . . . . . . . . . . . . 21512.4.4 Integrative Ethics and Integrative AGI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21612.5 Clarifying the Ethics of Justice: Extending the Golden Rule in to aMultifactorial Ethical Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21912.5.1 The Golden Rule and the Stages of Ethical Development . . . . . . . . . . . . . 22212.5.2 The Need for Context-Sensitivity and Adaptiveness in DeployingEthical Principles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22312.6 The Ethical Treatment of AGIs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22612.6.1 Possible Consequences of Depriving AGIs of Freedom . . . . . . . . . . . . . . . . . 22812.6.2 AGI Ethics as Boundaries Between Humans and AGIs Become Blurred . 22912.7 Possible Benefits of Closely Linking AGIs to the Global Brain . . . . . . . . . . . . . . . . 23012.7.1 The Importance of Fostering Deep, Consensus-Building InteractionsBetween People with Divergent Views . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23112.8 Possible Benefits of Creating Societies of AGIs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23312.9 AGI Ethics As Related to Various Future Scenarios . . . . . . . . . . . . . . . . . . . . . . . . 23412.9.1 Capped Intelligence Scenarios . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23412.9.2 Superintelligent AI: Soft-Takeoff Scenarios . . . . . . . . . . . . . . . . . . . . . . . . . . . 23512.9.3 Superintelligent AI: Hard-Takeoff Scenarios . . . . . . . . . . . . . . . . . . . . . . . . . 23512.9.4 Global Brain Mindplex Scenarios . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23712.10Conclusion: Eight Ways to Bias AGI Toward Friendliness . . . . . . . . . . . . . . . . . . . . 23912.10.1Encourage Measured Co-Advancement of AGI Software and AGI EthicsTheory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24112.10.2Develop Advanced AGI Sooner Not Later . . . . . . . . . . . . . . . . . . . . . . . . . . . 241Section IV Networks for Explicit and Implicit Knowledge Representation13 Local, Global and Glocal Knowledge Representation . . . . . . . . . . . . . . . . . . . . . . 24513.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24513.2 Localized Knowledge Representation using Weighted, Labeled Hypergraphs . . . . 24613.2.1 Weighted, Labeled Hypergraphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24613.3 Atoms: Their Types and Weights . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24713.3.1 Some Basic Atom Types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24713.3.2 Variable Atoms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24913.3.3 Logical Links . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25113.3.4 Temporal Links . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25213.3.5 Associative Links . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25313.3.6 Procedure Nodes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25413.3.7 Links for Special External Data Types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25413.3.8 Truth Values and Attention Values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25513.4 Knowledge Representation via Attractor Neural Networks . . . . . . . . . . . . . . . . . . . 256xviiiContents13.4.1 The Hopfield neural net model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25613.4.2 Knowledge Representation via Cell Assemblies . . . . . . . . . . . . . . . . . . . . . . 25713.5 Neural Foundations of Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25813.5.1 Hebbian Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25813.5.2 Virtual Synapses and Hebbian Learning Between Assemblies . . . . . . . . . . 25813.5.3 Neural Darwinism . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25913.6 Glocal Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26013.6.1 A Semi-Formal Model of Glocal Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26213.6.2 Glocal Memory in the Brain . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26313.6.3 Glocal Hopfield Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26813.6.4 Neural-Symbolic Glocality in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26914 Representing Implicit Knowledge via Hypergraphs . . . . . . . . . . . . . . . . . . . . . . . 27114.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27114.2 Key Vertex and Edge Types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27114.3 Derived Hypergraphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27214.3.1 SMEPH Vertices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27214.3.2 SMEPH Edges . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27314.4 Implications of Patternist Philosophy for Derived Hypergraphs of IntelligentSystems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27414.4.1 SMEPH Principles in CogPrime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27615 Emergent Networks of Intelligence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27915.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27915.2 Small World Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28015.3 Dual Network Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28115.3.1 Hierarchical Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28115.3.2 Associative, Heterarchical Networks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28215.3.3 Dual Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284Section V A Path to Human-Level AGI16 AGI Preschool . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28916.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28916.1.1 Contrast to Standard AI Evaluation Methodologies . . . . . . . . . . . . . . . . . . . 29016.2 Elements of Preschool Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29116.3 Elements of Preschool Curriculum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29216.3.1 Preschool in the Light of Intelligence Theory . . . . . . . . . . . . . . . . . . . . . . . . 29316.4 Task-Based Assessment in AGI Preschool . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29516.5 Beyond Preschool . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29816.6 Issues with Virtual Preschool Engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29816.6.1 Integrating Virtual Worlds with Robot Simulators . . . . . . . . . . . . . . . . . . . . 30116.6.2 BlocksNBeads World . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30117 A Preschool-Based Roadmap to Advanced AGI . . . . . . . . . . . . . . . . . . . . . . . . . . . 30717.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30717.2 Measuring Incremental Progress Toward Human-Level AGI . . . . . . . . . . . . . . . . . . 30817.3 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315Contentsxix18 Advanced Self-Modification: A Possible Path to Superhuman AGI . . . . . . . . 31718.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31718.2 Cognitive Schema Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31818.3 Self-Modification via Supercompilation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31918.3.1 Three Aspects of Supercompilation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32118.3.2 Supercompilation for Goal-Directed Program Modification . . . . . . . . . . . . . 32218.4 Self-Modification via Theorem-Proving . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323A Glossary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325A.1 List of Specialized Acronyms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325A.2 Glossary of Specialized Terms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 326References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 343
Chapter 1Introduction1.1 AI Returns to Its RootsOur goal in this book is straightforward, albeit ambitious: to present a conceptual and technicaldesign for a thinking machine, a software program capable of the same qualitative sort of generalintelligence as human beings. It’s not certain exactly how far the design outlined here will beable to take us, but it seems plausible that once fully implemented, tuned and tested, it will beable to achieve general intelligence at the human level and in some respects beyond.Our ultimate aim is Artificial General Intelligence construed in the broadest sense, includingartificial creativity and artificial genius. We feel it is important to emphasize the extremelybroad potential of Artificial General Intelligence systems. The human brain is not built to bemodified, except via the slow process of evolution. Engineered AGI systems, built according todesigns like the one outlined here, will be much more susceptible to rapid improvement fromtheir initial state. It seems reasonable to us to expect that, relatively shortly after achieving thefirst roughly human-level AGI system, AGI systems with various sorts of beyond-human-levelcapabilities will be achieved.Though these long-term goals are core to our motivations, we will spend much of our time hereexplaining how we think we can make AGI systems do relatively simple things, like the thingshuman children do in preschool. The penultimate chapter of (Part 2 of) the book describes athought-experiment involving a robot playing with blocks, responding to the request "Build mesomething I haven’t seen before." We believe that preschool creativity contains the seeds of,and the core structures and dynamics underlying, adult human level genius ... and new, as yetunforeseen forms of artificial innovation.Much of the book focuses on a specific AGI architecture, which we call CogPrime, and whichis currently in the midst of implementation using the OpenCog software framework. CogPrimeis large and complex and embodies a host of specific decisions regarding the various aspects ofintelligence. We don’t view CogPrime as the unique path to advanced AGI, nor as the ultimateend-all of AGI research. We feel confident there are multiple possible paths to advanced AGI,and that in following any of these paths, multiple theoretical and practical lessons will belearned, leading to modifications of the ideas possessed while along the early stages of the path.But our goal here is to articulate one path that we believe makes sense to follow, one overalldesign that we believe can work.12 1 Introduction1.2 AGI versus Narrow AIAn outsider to the AI field might think this sort of book commonplace in the research literature,but insiders know that’s far from the truth. The field of Artificial Intelligence (AI) was foundedin the mid 1950s with the aim of constructing “thinking machines” - that is, computer systemswith human-like general intelligence, including humanoid robots that not only look but actand think with intelligence equal to and ultimately greater than human beings. But in theintervening years, the field has drifted far from its ambitious roots, and this book representspart of a movement aimed at restoring the initial goals of the AI field, but in a manner poweredby new tools and new ideas far beyond those available half a century ago.After the first generation of AI researchers found the task of creating human-level AGI verydifficult given the technology of their time, the AI field shifted focus toward what Ray Kurzweilhas called "narrow AI" – the understanding of particular specialized aspects of intelligence; andthe creation of AI systems displaying intelligence regarding specific tasks in relatively narrowdomains. In recent years, however, the situation has been changing. More and more researchershave recognized the necessity – and feasibility – of returning to the original goals of the field.In the decades since the 1950s, cognitive science and neuroscience have taught us a lot aboutwhat a cognitive architecture needs to look like to support roughly human-like general intelligence.Computer hardware has advanced to the point where we can build distributed systemscontaining large amounts of RAM and large numbers of processors, carrying out complex tasksin real time. The AI field has spawned a host of ingenious algorithms and data structures, whichhave been successfully deployed for a huge variety of purposes.Due to all this progress, increasingly, there has been a call for a transition from the currentfocus on highly specialized “narrow AI” problem solving systems, back to confronting the moredifficult issues of “human level intelligence” and more broadly “artificial general intelligence(AGI).” Recent years have seen a growing number of special sessions, workshops and conferencesdevoted specifically to AGI, including the annual BICA (Biologically Inspired CognitiveArchitectures) AAAI Symposium, and the international AGI conference series (one in 2006,and annual since 2008). And, even more exciting, as reviewed in Chapter 4, there are a numberof contemporary projects focused directly and explicitly on AGI (sometimes under the name"AGI", sometimes using related terms such as "Human Level Intelligence").In spite of all this progress, however, we feel that no one has yet clearly articulated a detailed,systematic design for an AGI, with potential to yield general intelligence at the human leveland ultimately beyond. In this spirit, our main goal in this lengthy two-part book is to outlinea novel design for a thinking machine – an AGI design which we believe has the capability toproduce software systems with intelligence at the human adult level and ultimately beyond.Many of the technical details of this design have been previously presented online in a wikibook[Goe10b]; and the basic ideas of the design have been presented briefly in a series of conferencepapers [GPSL03, GPPG06, Goe09c]. But the overall design has not been presented in a coherentand systematic way before this book. In order to frame this design properly, we also presenta considerable number of broader theoretical and conceptual ideas here, some more and someless technical in nature.1.4 The Secret Sauce 31.3 CogPrimeThe AGI design presented here has not previously been granted a name independently of itsparticular software implementations, but for the purposes of this book it needs one, so we’vechristened it CogPrime . This fits with the name “OpenCogPrime” that has already beenused to describe the software implementation of CogPrime within the open-source OpenCogAGI software framework. The OpenCogPrime software, right now, implements only a smallfraction of the CogPrime design as described here. However, OpenCog was designed specificallyto enable efficient, scalable implementation of the full CogPrime design (as well as to serve as amore general framework for AGI R&D); and work currently proceeds in this direction, thoughthere is a lot of work still to be done and many challenges remain. 1The CogPrime design is more comprehensive and thorough than anything that has beenpresented in the literature previously, including the work of others reviewed in Chapter 4. Itcovers all the key aspects of human intelligence, and explains how they interoperate and howthey can be implemented in digital computer software. Part 1 of this work outlines CogPrime ata high level, and makes a number of more general points about artificial general intelligence andthe path thereto; then Part 2 digs deeply into the technical particulars of CogPrime. Even Part2, however, doesn’t explain all the details of CogPrime that have been worked out so far, andit definitely doesn’t explain all the implementation details that have gone into designing andbuilding OpenCogPrime. Creating a thinking machine is a large task, and even the intermediatelevel of detail takes up a lot of pages.1.4 The Secret SauceThere is no consensus on why all the related technological and scientific progress mentionedabove has not yet yielded AI software systems with human-like general intelligence (or evengreater levels of brilliance!). However, we hypothesize that the core reason boils down to thefollowing three points:• Intelligence depends on the emergence of certain high-level structures and dynamics acrossa system’s whole knowledge base;• We have not discovered any one algorithm or approach capable of yielding the emergenceof these structures;• Achieving the emergence of these structures within a system formed by integrating a numberof different AI algorithms and structures requires careful attention to the manner in which1 This brings up a terminological note: At several places in this Volume and the next we will refer to the currentCogPrime or OpenCog implementation; in all cases this refers to OpenCog as of late 2013. We realize the riskof mentioning the state of our software system at time of writing: for future readers this may give the wrongimpression, because if our project goes well, more and more of CogPrime will get implemented and tested astime goes on (e.g. within the OpenCog framework, under active development at time of writing). However, notmentioning the current implementation at all seems an even worse course to us, since we feel readers will beinterested to know which of our ideas – at time of writing – have been honed via practice and which have not.Online resources such as http://opencog.org may be consulted by readers curious about the current stateof the main OpenCog implementation; though in future forks of the code may be created, or other systems maybe built using some or all of the ideas in this book, etc.4 1 Introductionthese algorithms and structures are integrated; and so far the integration has not been donein the correct way.The human brain appears to be an integration of an assemblage of diverse structures anddynamics, built using common components and arranged according to a sensible cognitive architecture.However, its algorithms and structures have been honed by evolution to work closelytogether – they are very tightly inter-adapted, in the same way that the different organs ofthe body are adapted to work together. Due to their close interoperation they give rise to theoverall systemic behaviors that characterize human-like general intelligence. We believe thatthe main missing ingredient in AI so far is cognitive synergy: the fitting-together of differentintelligent components into an appropriate cognitive architecture, in such a way that thecomponents richly and dynamically support and assist each other, interrelating very closely ina similar manner to the components of the brain or body and thus giving rise to appropriateemergent structures and dynamics. This leads us to one of the central hypotheses underlyingthe CogPrime approach to AGI: that the cognitive synergy ensuing from integratingmultiple symbolic and subsymbolic learning and memory components in an appropriatecognitive architecture and environment, can yield robust intelligence at thehuman level and ultimately beyond.The reason this sort of intimate integration has not yet been explored much is that it’s difficulton multiple levels, requiring the design of an architecture and its component algorithms witha view toward the structures and dynamics that will arise in the system once it is coupledwith an appropriate environment. Typically, the AI algorithms and structures correspondingto different cognitive functions have been developed based on divergent theoretical principles,by disparate communities of researchers, and have been tuned for effective performance ondifferent tasks in different environments. Making such diverse components work together in atruly synergetic and cooperative way is a tall order, yet we believe that this – rather than someparticular algorithm, structure or architectural principle – is the “secret sauce” needed to createhuman-level AGI based on technologies available today.1.5 Extraordinary Proof?There is a saying that “extraordinary claims require extraordinary proof” and by that standard,if one believes that having a design for an advanced AGI is an extraordinary claim, thisbook must be rated a failure. We don’t offer extraordinary proof that CogPrime, once fullyimplemented and educated, will be capable of human-level general intelligence and more.It would be nice if we could offer mathematical proof that CogPrime has the potential wethink it does, but at the current time mathematics is simply not up to the job. We’ll pursuethis direction briefly in Chapter 7 and other chapters, where we’ll clarify exactly what kindof mathematical claim “CogPrime has the potential for human-level intelligence” turns out tobe. Once this has been clarified, it will be clear that current mathematical knowledge does notyet let us evaluate, or even fully formalize, this kind of claim. Perhaps one day rigorous anddetailed analyses of practical AGI designs will be feasible – and we look forward to that day –but it’s not here yet.Also, it would of course be profoundly exciting if we could offer dramatic practical demonstrationsof CogPrime’s capabilities. We do have a partial software implementation, in theOpenCogPrime system, but currently the things OpenCogPrime does are too simple to really1.5 Extraordinary Proof? 5serve as proofs of CogPrime’s power for advanced AGI. We have used some CogPrime ideas inthe OpenCog framework to do things like natural language understanding and data mining, andto control virtual dogs in online virtual worlds; and this has been very useful work in multiplesenses. It has taught us more about the CogPrime design; it has produced some useful softwaresystems; and it constitutes fractional work building toward a full OpenCog based implementationof CogPrime. However, to date, the things OpenCogPrime has done are all things thatcould have been done in different ways without the CogPrime architecture (though perhaps notas elegantly nor with as much room for interesting expansion).The bottom line is that building an AGI is a big job. Software companies like Microsoft spenddozens to hundreds of man-years building software products like word processors and operatingsystems, so it should be no surprise that creating a digital intelligence is also a relatively largescalesoftware engineering project. As time advances and software tools improve, the number ofman-hours required to develop advanced AGI gradually decreases – but right now, as we writethese words, it’s still a rather big job. In the OpenCogPrime project we are making a seriousattempt to create a CogPrime based AGI using an open-source development methodology,with the open-source Linux operating system as one of our inspirations. But the open-sourcemethodology doesn’t work magic either, and it remains a large project, currently at an earlystage. I emphasize this point so that readers lacking software engineering expertise don’t takethe currently fairly limited capabilities of OpenCogPrime as somehow a damning indictment ofthe potential of the CogPrime design. The design is one thing, the implementation another –and the OpenCogPrime implementation currently encompasses perhaps one third to one halfof the key ideas in this book.So we don’t have extraordinary proof to offer. What we aim to offer instead are clearlyconstructedconceptual and technical arguments as to why we think the CogPrime design hasdramatic AGI potential.It is also possible to push back a bit on the common intuition that having a design for humanlevelAGI is such an “extraordinary claim.” It may be extraordinary relative to contemporaryscience and culture, but we have a strong feeling that the AGI problem is not difficult in thesame ways that most people (including most AI researchers) think it is. We suspect that inhindsight, after human-level AGI has been achieved, people will look back in shock that it tookhumanity so long to come up with a workable AGI design. As you’ll understand once you’vefinished Part 1 of the book, we don’t think general intelligence is nearly as “extraordinary”and mysterious as it’s commonly made out to be. Yes, building a thinking machine is hard –but humanity has done a lot of other hard things before. It may seem difficult to believe thathuman-level general intelligence could be achieved by something as simple as a collection ofalgorithms linked together in an appropriate way and used to control an agent. But we suggestthat, once the first powerful AGI systems are produced, it will become apparent that engineeringhuman-level minds is not so profoundly different from engineering other complex systems.All in all, we’ll consider the book successful if a significant percentage of open-minded,appropriately-educated readers come away from it scratching their chins and pondering: “Hmm.You know, that just might work.” and a small percentage come away thinking "Now that’s aninitiative I’d really like to help with!".6 1 Introduction1.6 Potential Approaches to AGIIn principle, there is a large number of approaches one might take to building an AGI, startingfrom the knowledge, software and machinery now available. This is not the place to reviewthem in detail, but a brief list seems apropos, including commentary on why these are not theapproaches we have chosen for our own research. Our intent here is not to insult or dismissthese other potential approaches, but merely to indicate why, as researchers with limited timeand resources, we have made a different choice regarding where to focus our own energies.1.6.1 Build AGI from Narrow AIMost of the AI programs around today are “narrow AI” programs – they carry out one particularkind of task intelligently. One could try to make an advanced AGI by combining a bunch ofenhanced narrow AI programs inside some kind of overall framework.However, we’re rather skeptical of this approach because none of these narrow AI programshave the ability to generalize across domains – and we don’t see how combining them or extendingthem is going to cause this to magically emerge.1.6.2 Enhancing ChatbotsOne could seek to make an advanced AGI by taking a chatbot, and trying to improve its codeto make it actually understand what it’s talking about. We have some direct experience withthis route, as in 2010 our AI consulting firm was contracted to improve Ray Kurzweil’s onlinechatbot "Ramona". Our new Ramona understands a lot more than the previous Ramona versionor a typical chatbot, due to using Wikipedia and other online resources, but still it’s far froman AGI.A more ambitious attempt in this direction was Jason Hutchens’ a-i.com project, whichsought to create a human child level AGI via development and teaching of a statistical learningbased chatbot (rather than the typical rule-based kind). The difficulty with this approach,however, is that the architecture of a chatbot is fundamentally different from the architectureof a generally intelligent mind. Much of what’s important about the human mind is not directlyobservable in conversations, so if you start from conversation and try to work toward an AGIarchitecture from there, you’re likely to miss many critical aspects.1.6.3 Emulating the BrainOne can approach AGI by trying to figure out how the brain works, using brain imaging andother tools from neuroscience, and then emulating the brain in hardware or software.One rather substantial problem with this approach is that we don’t really understand howthe brain works yet, because our software for measuring the brain is still relatively crude. Thereis no brain scanning method that combines high spatial and temporal accuracy, and none is1.6 Potential Approaches to AGI 7likely to come about for a decade or two. So to do brain-emulation AGI seriously, one needs towait a while until brain scanning technology improves.Current AI methods like neural nets that are loosely based on the brain, are really not brainlikeenough to make a serious claim at emulating the brain’s approach to general intelligence.We don’t yet have any real understanding of how the brain represents abstract knowledge, forexample, or how it does reasoning (though the authors, like many others, have made somespeculations in this regard [GMIH08]).Another problem with this approach is that once you’re done, what you get is somethingwith a very humanlike mind, and we already have enough of those! However, this is perhapsnot such a serious objection, because a digital-computer-based version of a human mind couldbe studied much more thoroughly than a biology-based human mind. We could observe itsdynamics in real-time in perfect precision, and could then learn things that would allow us tobuild other sorts of digital minds.1.6.4 Evolve an AGIAnother approach is to try to run an evolutionary process inside the computer, and wait foradvanced AGI to evolve.One problem with this is that we don’t know how evolution works all that well. There’s afield of artificial life, but so far its results have been fairly disappointing. It’s not yet clear howmuch one can vary on the chemical structures that underly real biology, and still get powerfulevolution like we see in real biology. If we need good artificial chemistry to get good artificialbiology, then do we need good artificial physics to get good artificial chemistry?Another problem with this approach, of course, is that it might take a really long time.Evolution took billions of years on Earth, using a massive amount of computational power. Tomake the evolutionary approach to AGI effective, one would need some radical innovations tothe evolutionary process (such as, perhaps, using probabilistic methods like BOA [Pel05] orMOSES [Loo06] in place of traditional evolution).1.6.5 Derive an AGI design mathematicallyOne can try to use the mathematical theory of intelligence to figure out how to make advancedAGI.This interests us greatly, but there’s a huge gap between the rigorous math of intelligenceas it exists today and anything of practical value. As we’ll discuss in Chapter 7, most of therigorous math of intelligence right now is about how to make AI on computers with dramaticallyunrealistic amounts of memory or processing power. When one tries to create a theoreticalunderstanding of real-world general intelligence, one arrives at quite different sorts of considerations,as we will roughly outline in Chapter 10. Ideally we would like to be able to study theCogPrime design using a rigorous mathematical theory of real-world general intelligence, but atthe moment that’s not realistic. The best we can do is to conceptually analyze CogPrime andits various components in terms of relevant mathematical and theoretical ideas; and performanalysis of CogPrime’s individual structures and components at varying levels of rigor.8 1 Introduction1.6.6 Use heuristic computer science methodsThe computer science field contains a number of abstract formalisms, algorithms and structuresthat have relevance beyond specific narrow AI applications, yet aren’t necessarily understoodas thoroughly as would be required to integrate them into the rigorous mathematical theory ofintelligence. Based on these formalisms, algorithms and structures, a number of "single formalism/algorithmfocused" AGI approaches have been outlined, some of which will be reviewed inChapter 4. For example Pei Wang’s NARS (”Non-Axiomatic Reasoning System”) approach isbased on a specific logic which he argues to be the "logic of general intelligence" – so, while hissystem contains many other aspects than this logic, he considers this logic to be the crux of thesystem and the source of its potential power as an AGI system.The basic intuition on the part of these "single formalism/algorithm focused" researchersseems to be that there is one key formalism or algorithm underlying intelligence, and if youachieve this key aspect in your AGI program, you’re going to get something that fundamentallythinks like a person, even if it has some differences due to its different implementation andembodiment. On the other hand, it’s also possible that this idea is philosophically incorrect:that there is no one key formalism, algorithm, structure or idea underlying general intelligence.The CogPrime approach is based on the intuition that to achieve human-level, roughly humanlikegeneral intelligence based on feasible computational resources, one needs an appropriateheterogeneous combination of algorithms and structures, each coping with different types ofknowledge and different aspects of the problem of achieving goals in complex environments.1.6.7 Integrative Cognitive ArchitectureFinally, to create advanced AGI one can try to build some sort of integrative cognitive architecture:a software system with multiple components that each carry out some cognitive function,and that connect together in a specific way to try to yield overall intelligence.Cognitive science gives us some guidance about the overall architecture, and computer scienceand neuroscience give us a lot of ideas about what to put in the different components. But stillthis approach is very complex and there is a lot of need for creative invention.This is the approach we consider most “serious” at present (at least until neuroscience advancesfurther). And, as will be discussed in depth in these pages, this is the approach we’vechosen: CogPrime is an integrative AGI architecture.1.6.8 Can Digital Computers Really Be Intelligent?All the AGI approaches we’ve just mentioned assume that it’s possible to make AGI on digitalcomputers. While we suspect this is correct, we must note that it isn’t proven.It might be that – as Penrose [Pen96], Hameroff [Ham87] and others have argued – we needquantum computers or quantum gravity computers to make AGI. However, there is no evidenceof this at this stage. Of course the brain like all matter is described by quantum mechanics,but this doesn’t imply that the brain is a “macroscopic quantum system” in a strong sense(like, say, a Bose-Einstein condensate). And even if the brain does use quantum phenomena in1.7 Five Key Words 9a dramatic way to carry out some of its cognitive processes (a hypothesis for which there is nocurrent evidence), this doesn’t imply that these quantum phenomena are necessary in order tocarry out the given cognitive processes. For example there is evidence that birds use quantumnonlocal phenomena to carry out navigation based on the Earth’s magnetic fields [GRM + 11];yet scientists have built instruments that carry out the same functions without using any specialquantum effects. The importance of quantum phenomena in biology (except via their obviousrole in giving rise to biological phenomena describable via classical physics) remains a subjectof debate [AGBD + 08].Quantum “magic” aside, it is also conceivable that building AGI is fundamentally impossiblefor some other reason we don’t understand. Without getting religious about it, it is rationallyquite possible that some aspects of the universe are beyond the scope of scientific methods.Science is fundamentally about recognizing patterns in finite sets of bits (e.g. finite sets offinite-precision observations), whereas mathematics recognizes many sets much larger than this.Selmer Bringsjord [BZ03], and other advocates of “hypercomputing” approaches to intelligence,argue that the human mind depends on massively large infinite sets and therefore can never besimulated on digital computers nor understood via finite sets of finite-precision measurementssuch as science deals with.But again, while this sort of possibility is interesting to speculate about, there’s no real reasonto believe it at this time. Brain science and AI are both very young sciences and the “workinghypothesis” that digital computers can manifest advanced AGI has hardly been explored atall yet, relative to what will be possible in the next decades as computers get more and morepowerful and our understanding of neuroscience and cognitive science gets more and morecomplete. The CogPrime AGI design presented here is based on this working hypothesis.Many of the ideas in the book are actually independent of the “mind can be implementeddigitally” working hypothesis, and could apply to AGI systems built on analog, quantum orother non-digital frameworks – but we will not pursue these possibilities here. For the moment,outlining an AGI design for digital computers is hard enough! Regardless of speculations aboutquantum computing in the brain, it seems clear that AGI on quantum computers is part of ourfuture and will be a powerful thing; but the description of a CogPrime analogue for quantumcomputers will be left for a later work.1.7 Five Key WordsAs noted, the CogPrime approach lies squarely in the integrative cognitive architecture camp.But it is not a haphazard or opportunistic combination of algorithms and data structures. Atbottom it is motivated by the patternist philosophy of mind laid out in Ben Goertzel’s bookThe Hidden Pattern [Goe06a], which was in large part a summary and reformulation of ideaspresented in a series of books published earlier by the same author [Goe94], [Goe93a], [Goe93b],[Goe97], [Goe01]. A few of the core ideas of this philosophy are laid out in Chapter 3, thoughthat chapter is by no means a thorough summary.One way to summarize some of the most important yet commonsensical parts of the patternistphilosophy of mind, in an AGI context, is to list five words: perception, memory, prediction,action, goals.In a phrase: “A mind uses perception and memory to make predictions aboutwhich actions will help it achieve its goals.”10 1 IntroductionThis ties in with the ideas of many other thinkers, including Jeff Hawkins’ “memory/prediction”theory [HB06], and it also speaks directly to the formal characterization of intelligencepresented in Chapter 7: general intelligence as “the ability to achieve complex goals in complexenvironments.”Naturally the goals involved in the above phrase may be explicit or implicit to the intelligentagent, and they may shift over time as the agent develops.Perception is taken to mean pattern recognition: the recognition of (novel or familiar) patternsin the environment or in the system itself. Memory is the storage of already-recognizedpatterns, enabling recollection or regeneration of these patterns as needed. Action is the formationof patterns in the body and world. Prediction is the utilization of temporal patterns toguess what perceptions will be seen in the future, and what actions will achieve what effects inthe future – in essence, prediction consists of temporal pattern recognition, plus the (implicitor explicit) assumption that the universe possesses a "habitual tendency" according to whichpreviously observed patterns continue to apply.1.7.1 Memory and Cognition in CogPrimeEach of these five concepts has a lot of depth to it, and we won’t say too much about them inthis brief introductory overview; but we will take a little time to say something about memoryin particular.As we’ll see in Chapter 7, one of the things that the mathematical theory of general intelligencemakes clear is that, if you assume your AI system has a huge amount of computationalresources, then creating general intelligence is not a big trick. Given enough computing power,a very brief and simple program can achieve any computable goal in any computable environment,quite effectively. Marcus Hutter’s AIXI tl design [Hut05] gives one way of doing this,backed up by rigorous mathematics. Put informally, what this means is: the problem of AGI isreally a problem of coping with inadequate compute resources, just as the problem of naturalintelligence is really a problem of coping with inadequate energetic resources.One of the key ideas underlying CogPrime is a principle called cognitive synergy, whichexplains how real-world minds achieve general intelligence using limited resources, by appropriatelyorganizing and utilizing their memories.This principle says that there are many different kinds of memory in the mind: sensory,episodic, procedural, declarative, attentional, intentional. Each of them has certain learningprocesses associated with it; for example, reasoning is associated with declarative memory.Synergy arises here in the way the learning processes associated with each kind of memory havegot to help each other out when they get stuck, rather than working at cross-purposes.Cognitive synergy is a fundamental principle of general intelligence – it doesn’t tend to playa central role when you’re building narrow-AI systems.In the CogPrime approach all the different kinds of memory are linked together in a singlemeta-representation, a sort of combined semantic/neural network called the AtomSpace. Itrepresents everything from perceptions and actions to abstract relationships and concepts andeven a system’s model of itself and others. When specialized representations are used for othertypes of knowledge (e.g. program trees for procedural knowledge, spatiotemporal hierarchiesfor perceptual knowledge) then the knowledge stored outside the AtomSpace is represented via1.8 Virtually and Robotically Embodied AI 11tokens (Atoms) in the AtomSpace, allowing it to be located by various cognitive processes, andassociated with other memory items of any type.So for instance an OpenCog AI system has an AtomSpace, plus some specialized knowledgestores linked into the AtomSpace; and it also has specific algorithms acting on the AtomSpaceand appropriate specialized stores corresponding to each type of memory. Each of these algorithmsis complex and has its own story; for instance (an incomplete list, for more detail seethe following section of this Introduction):• Declarative knowledge is handled using Probabilistic Logic Networks, described in Chapter34 and others;• Procedural knowledge is handled using MOSES, a probabilistic evolutionary learning algorithmdescribed in Chapter 21 and others;• Attentional knowledge is handled by ECAN (economic attention allocation), described inChapter 23 and others;• OpenCog contains a language comprehension system called RelEx that takes English sentencesand turns them into nodes and links in the AtomSpace. It’s currently being extendedto handle Chinese. RelEx handles mostly declarative knowledge but also involvessome procedural knowledge for linguistic phenomena like reference resolution and semanticdisambiguation.But the crux of the CogPrime cognitive architecture is not any particular cognitive process,but rather the way they all work together using cognitive synergy.1.8 Virtually and Robotically Embodied AIAnother issue that will arise frequently in these pages is embodiment. There’s a lot of debate inthe AI community over whether embodiment is necessary for advanced AGI or not. Personally,we doubt it’s necessary but we think it’s extremely convenient, and are thus considerablyinterested in both virtual world and robotic embodiment. The CogPrime architecture itself isneutral on the issue of embodiment, and it could be used to build a mathematical theoremprover or an intelligent chat bot just as easily as an embodied AGI system. However, most ofour attention has gone into figuring out how to use CogPrime to control embodied agents invirtual worlds, or else (to a lesser extent) physical robots. For instance, during 2011-2012 weare involved in a Hong Kong government funded project using OpenCog to control video gameagents in a simple game world modeled on the game Minecraft [GPC + 11].Current virtual world technology has significant limitations that make them far less thanideal from an AGI perspective, and in Chapter 16 we will discuss how they can be remedied.However, for the medium-term future virtual worlds are not going to match the natural worldin terms of richness and complexity – and so there’s also something to be said for physicalrobots that interact with all the messiness of the real world.With this in mind, in the Artificial Brain Lab at Xiamen University in 2009-2010, we conductedsome experiments using OpenCog to control the Nao humanoid robot [GD09]. The goalof that work was to take the same code that controls the virtual dog and use it to control thephysical robot. But it’s harder because in this context we need to do real vision processingand real motor control. A similar project is being undertaken in Hong Kong at time of writing,involving a collaboration between OpenCog AI developers and David Hanson’s robotics12 1 Introductiongroup. One of the key ideas involved in this project is explicit integration of subsymbolic andmore symbolic subsystems. For instance, one can use a purely subsymbolic, hierarchical patternrecognition network for vision processing, and then link its internal structures into the nodesand links in the AtomSpace that represent concepts. So the subsymbolic and symbolic systemscan work harmoniously and productively together, a notion we will review in more detail inChapter 26.1.9 Language LearningOne of the subtler aspects of our current approach to teaching CogPrime is language learning.Three relatively crisp and simple approaches to language learning would be:• Build a language processing system using hand-coded grammatical rules, based on linguistictheory;• Train a language processing system using supervised, unsupervised or semisupervised learning,based on computational linguistics;• Have an AI system learn language via experience, based on imitation and reinforcement andexperimentation, without any built-in distinction between linguistic behaviors and otherbehaviors.While the third approach is conceptually appealing, our current approach in CogPrime (describedin a series of chapters in Part 2) is none of the above, but rather a combination of theabove. OpenCog contains a natural language processing system built using a combination ofthe rule-based and statistical approaches, which has reasonably adequate functionality; and ourplan is to use it as an initial condition for ongoing adaptive improvement based on embodiedcommunicative experience.1.10 AGI EthicsWhen discussing AGI work with the general public, ethical concerns often arise. Science fictionfilms like the Terminator series have raised public awareness of the possible dangers ofadvanced AGI systems without correspondingly advanced ethics. Non-profit organizations likethe Singularity Institute for AI ((http://singinst.org) have arisen specifically to raise attentionabout, and foster research on, these potential dangers.Our main focus here is on how to create AGI, not how to teach an AGI human ethicalprinciples. However, we will address the latter issue explicitly in Chapter 12, and we do think it’simportant to emphasize that AGI ethics has been at the center of the design process throughoutthe conception and development of CogPrime and OpenCog.Broadly speaking there are (at least) two major threats related to advanced AGI. One isthat people might use AGIs for bad ends; and the other is that, even if an AGI is made withthe best intentions, it might reprogram itself in a way that causes it to do something terrible.If it’s smarter than us, we might be watching it carefully while it does this, and have no ideawhat’s going on.1.12 Key Claims of the Book 13The best way to deal with this second “bad AGI” problem is to build ethics into your AGIarchitecture – and we have done this with CogPrime, via creating a goal structure that explicitlysupports ethics-directed behavior, and via creating an overall architecture that supports “ethicalsynergy” along with cognitive synergy. In short, the notion of ethical synergy is that there aredifferent kinds of ethical thinking associated with the different kinds of memory and you wantto be sure your AGI has all of them, and that it uses them together effectively.In order to create AGI that is not only intelligent but beneficial to other sentient beings,ethics has got to be part of the design and the roadmap. As we teach our AGI systems, we needto lead them through a series of instructional and evaluative tasks that move from a primitivelevel to the mature human level – in intelligence, but also in ethical judgment.1.11 Structure of the BookThe book is divided into two parts. The technical particulars of CogPrime are discussed in Part2; what we deal with in Part 1 are important preliminary and related matters such as:• The nature of real-world general intelligence, both conceptually and from the perspectiveof formal modeling (Section I).• The nature of cognitive and ethical development for humans and AGIs (Section III).• The high-level properties of CogPrime, including the overall architecture and the varioussorts of memory involved (Section IV).• What kind of path may viably lead us from here to AGI, with focus laid on preschool-typeenvironments that easily foster humanlike cognitive development. Various advanced aspectsof AGI systems, such as the network and algebraic structures that may emerge from them,the ways in which they may self-modify, and the degree to which their initial design mayconstrain or guide their future state even after long periods of radical self-improvement(Section V).One point made repeatedly throughout Part 1, which is worth emphasizing here, is the currentlack of a really rigorous and thorough general technical theory of general intelligence. Such atheory, if complete, would be incredibly helpful for understanding complex AGI architectureslike CogPrime. Lacking such a theory, we must work on CogPrime and other such systems usinga combination of theory, experiment and intuition. This is not a bad thing, but it will be veryhelpful if the theory and practice of AGI are able to grow collaboratively together.1.12 Key Claims of the BookWe will wrap up this Introduction with a systematic list of some of the key claims to be arguedfor in these pages. Not all the terms and ideas in these claims have been mentioned in thepreceding portions of this Introduction, but we hope they will be reasonably clear to the readeranyway, at least in a general sense. This list of claims will be revisited in Chapter 49 near theend of Part 2, where we will look back at the ideas and arguments that have been put forth infavor of them, in the intervening chapters.14 1 IntroductionIn essence this is a list of claims such that, if the reader accepts these claims, they shouldprobably accept that the CogPrime approach to AGI is a viable one. On the other hand if thereader rejects one or more of these claims, they may find one or more aspects of CogPrimeunacceptable for some reason.Without further ado, now, the claims:1. General intelligence (at the human level and ultimately beyond) can be achieved via creatinga computational system that seeks to achieve its goals, via using perception and memoryto predict which actions will achieve its goals in the contexts in which it finds itself.2. To achieve general intelligence in the context of human-intelligence-friendly environmentsand goals using feasible computational resources, it’s important that an AGI system canhandle different kinds of memory (declarative, procedural, episodic, sensory, intentional,attentional) in customized but interoperable ways.3. Cognitive synergy: It’s important that the cognitive processes associated with different kindsof memory can appeal to each other for assistance in overcoming bottlenecks in a mannerthat enables each cognitive process to act in a manner that is sensitive to the particularitiesof each others’ internal representations, and that doesn’t impose unreasonable delays onthe overall cognitive dynamics.4. As a general principle, neither purely localized nor purely global memory is sufficient forgeneral intelligence under feasible computational resources; “glocal” memory will be required.5. To achieve human-like general intelligence, it’s important for an intelligent agent to havesensory data and motoric affordances that roughly emulate those available to humans.We don’t know exactly how close this emulation needs to be, which means that our AGIsystems and platforms need to support fairly flexible experimentation with virtual-worldand/or robotic infrastructures.6. To work toward adult human-level, roughly human-like general intelligence, one fairly easilycomprehensible path is to use environments and goals reminiscent of human childhood, andseek to advance one’s AGI system along a path roughly comparable to that followed byhuman children.7. It is most effective to teach an AGI system aimed at roughly human-like general intelligencevia a mix of spontaneous learning and explicit instruction, and to instruct it via acombination of imitation, reinforcement and correction, and a combination of linguistic andnonlinguistic instruction.8. One effective approach to teaching an AGI system human language is to supply it withsome in-built linguistic facility, in the form of rule-based and statistical-linguistics-basedNLP systems, and then allow it to improve and revise this facility based on experience.9. An AGI system with adequate mechanisms for handling the key types of knowledge mentionedabove, and the capability to explicitly recognize large-scale patterns in itself, should,upon sustained interaction with an appropriate environment in pursuit of appropriategoals, emerge a variety of complex structures in its internal knowledge network,including, but not limited to:• a hierarchical network, representing both a spatiotemporal hierarchy and an approximate“default inheritance” hierarchy, cross-linked• a heterarchical network of associativity, roughly aligned with the hierarchical network• a self network which is an approximate micro image of the whole network1.12 Key Claims of the Book 15• inter-reflecting networks modeling self and others, reflecting a “mirrorhouse” designpattern10. Given the strengths and weaknesses of current and near-future digital computers,a. A (loosely) neural-symbolic network is a good representation for directly storing manykinds of memory, and interfacing between those that it doesn’t store directly;b. Uncertain logic is a good way to handle declarative knowledge. To deal with the problemsfacing a human-level AGI, an uncertain logic must integrate imprecise probabilityand fuzziness with a broad scope of logical constructs. PLN is one good realization.c. Programs are a good way to represent procedures (both cognitive and physical-action,but perhaps not including low-level motor-control procedures).d. Evolutionary program learning is a good way to handle difficult program learning problems.Probabilistic learning on normalized programs is one effective approach to evolutionaryprogram learning. MOSES is one good realization of this approach.e. Multistart hill-climbing, with a strong Occam prior, is a good way to handle relativelystraightforward program learning problems.f. Activation spreading and Hebbian learning comprise a reasonable way to handle attentionalknowledge (though other approaches, with greater overhead cost, may providebetter accuracy and may be appropriate in some situations).• Artificial economics is an effective approach to activation spreading and Hebbianlearning in the context of neural-symbolic networks;• ECAN is one good realization of artificial economics;• A good trade-off between comprehensiveness and efficiency is to focus on two kindsof attention: processor attention (represented in CogPrime by ShortTermImportance)and memory attention (represented in CogPrime by LongTermImportance).g. Simulation is a good way to handle episodic knowledge (remembered and imagined).Running an internal world simulation engine is an effective way to handle simulation.h. Hybridization of one’s integrative neural-symbolic system with a spatiotemporally hierarchicaldeep learning system is an effective way to handle representation and learningof low-level sensorimotor knowledge. DeSTIN is one example of a deep learning systemof this nature that can be effective in this context.i. One effective way to handle goals is to represent them declaratively, and allocate attentionamong them economically. CogPrime’s PLN/ECAN based framework for handlingintentional knowledge is one good realization.11. It is important for an intelligent system to have some way of recognizing large-scale patternsin itself, and then embodying these patterns as new, localized knowledge items inits memory. Given the use of a neural-symbolic network for knowledge representation, agraph-mining based “map formation” heuristic is one good way to do this.12. Occam’s Razor: Intelligence is closely tied to the creation of procedures that achieve goalsin environments in the simplest possible way. Each of an AGI system’s cognitive algorithmsshould embody a simplicity bias in some explicit or implicit form.13. An AGI system, if supplied with a commonsensically ethical goal system and an intentionalcomponent based on rigorous uncertain inference, should be able to reliably achieve a muchhigher level of commonsensically ethical behavior than any human being.14. Once sufficiently advanced, an AGI system with a logic-based declarative knowledge approachand a program-learning-based procedural knowledge approach should be able to16 1 Introductionradically self-improve via a variety of methods, including supercompilation and automatedtheorem-proving.Section IArtificial and Natural General Intelligence
Chapter 2What Is Human-Like General Intelligence?2.1 IntroductionCogPrime, the AGI architecture on which the bulk of this book focuses, is aimed at the creationof artificial general intelligence that is vaguely human-like in nature, and possesses capabilitiesat the human level and ultimately beyond.Obviously this description begs some foundational questions, such as, for starters: What is"general intelligence"? What is "human-like general intelligence"? What is "intelligence" at all?Perhaps in the future there will exist a rigorous theory of general intelligence which appliesusefully to real-world biological and digital intelligences. In later chapters we will give someideas in this direction. But such a theory is currently nascent at best. So, given the presentstate of science, these two questions about intelligence must be handled via a combination offormal and informal methods. This brief, informal chapter attempts to explain our view on thenature of intelligence in sufficient detail to place the discussion of CogPrime in appropriatecontext, without trying to resolve all the subtleties.Psychologists sometimes define human general intelligence using IQ tests and related instruments– so one might wonder: why not just go with that? But these sorts of intelligence testingapproaches have difficulty even extending to humans from diverse cultures [HHPO12] [Fis01].So it’s clear that to ground AGI approaches that are not based on precise modeling of humancognition, one requires a more fundamental understanding of the nature of general intelligence.On the other hand, if one conceives intelligence too broadly and mathematically, there’s a riskof leaving the real human world too far behind. In this chapter (followed up in Chapters 9 and7 with more rigor), we present a highly abstract understanding of intelligence-in-general, andthen portray human-like general intelligence as a (particularly relevant) special case.2.1.1 What Is General Intelligence?Many attempts to characterize general intelligence have been made; Legg and Hutter [LH07a]review over 70! Our preferred abstract characterization of intelligence is: the capability of asystem to choose actions maximizing its goal-achievement, based on its perceptionsand memories, and making reasonably efficient use of its computational resources1920 2 What Is Human-Like General Intelligence?[Goe10c]. A general intelligence is then understood as one that can do this for a variety ofcomplex goals in a variety of complex environments.However, apart from positing definitions, it is difficult to say anything nontrivial about generalintelligence in general. Marcus Hutter [Hut05] has demonstrated, using a characterizationof general intelligence similar to the one above, that a very simple algorithm called AIXI tl candemonstrate arbitrarily high levels of general intelligence, if given sufficiently immense computationalresources. This is interesting because it shows that (if we assume the universe caneffectively be modeled as a computational system) general intelligence is basically a problem ofcomputational efficiency. The particular structures and dynamics that characterize real-worldgeneral intelligences like humans arise because of the need to achieve reasonable levels of intelligenceusing modest space and time resources.The “patternist” theory of mind presented in [Goe06a] and briefly summarized in Chapter3 below presents a number of emergent structures and dynamics that are hypothesized tocharacterize pragmatic general intelligence, including such things as system-wide hierarchicaland heterarchical knowledge networks, and a dynamic and self-maintaining self-model. Much ofthe thinking underlying CogPrime has centered on how to make multiple learning componentscombine to give rise to these emergent structures and dynamics.2.1.2 What Is Human-like General Intelligence?General principles like “complex goals in complex environments” and patternism are not sufficientto specify the nature of human-like general intelligence. Due to the harsh reality ofcomputational resource restrictions, real-world general intelligences are necessarily biased toparticular classes of environments. Human intelligence is biased toward the physical, social andlinguistic environments in which humanity evolved, and if AI systems are to possess humanlikegeneral intelligence they must to some extent share these biases.But what are these biases, specifically? This is a large and complex question, which we seekto answer in a theoretically grounded way in Chapter 9. However, before turning to abstracttheory, one may also approach the question in a pragmatic way, by looking at the categories ofthings that humans do to manifest their particular variety of general intelligence. This is thetask of the following section.2.2 Commonly Recognized Aspects of Human-like IntelligenceIt would be nice if we could give some sort of “standard model of human intelligence” in thischapter, to set the context for our approach to artificial general intelligence – but the truth isthat there isn’t any. What the cognitive science field has produced so far is better described as:a broad set of principles and platitudes, plus a long, loosely-organized list of ideas and results.Chapter 5 below constitutes an attempt to present an integrative architecture diagram forhuman-like general intelligence, synthesizing the ideas of a number of different AGI and cognitivetheorists. However, though the diagram given there attempts to be inclusive, it nonethelesscontains many features that are accepted by only a plurality of the research community.2.2 Commonly Recognized Aspects of Human-like Intelligence 21The following list of key aspects of human-like intelligence has a better claim at truly beinggeneric and representing the consensus understanding of contemporary science. It was producedby a very simple method: starting with the Wikipedia page for cognitive psychology, and thenadding a few items onto it based on scrutinizing the tables of contents of some top-rankedcognitive psychology textbooks. There is some redundancy among list items, and perhaps alsosome minor omissions (depending on how broadly one construes some of the items), but thepoint is to give a broad indication of human mental functions as standardly identified in thepsychology field:• Perception– General perception– Psychophysics– Pattern recognition (the ability to correctly interpret ambiguous sensory information)– Object and event recognition– Time sensation (awareness and estimation of the passage of time)• Motor Control– Motor planning– Motor execution– Sensorimotor integration• Categorization– Category induction and acquisition– Categorical judgement and classification– Category representation and structure– Similarity• Memory– Aging and memory– Autobiographical memory– Constructive memory– Emotion and memory– False memories– Memory biases– Long-term memory– Episodic memory– Semantic memory– Procedural memory– Short-term memory– Sensory memory– Working memory• Knowledge representation– Mental imagery– Propositional encoding– Imagery versus propositions as representational mechanisms22 2 What Is Human-Like General Intelligence?– Dual-coding theories– Mental models• Language– Grammar and linguistics– Phonetics and phonology– Language acquisition• Thinking– Choice– Concept formation– Judgment and decision making– Logic, formal and natural reasoning– Problem solving– Planning– Numerical cognition– Creativity• Consciousness– Attention and Filtering (the ability to focus mental effort on specific stimuli whilstexcluding other stimuli from consideration)– Access consciousness– Phenomenal consciousness• Social Intelligence– Distributed Cognition– EmpathyIf there’s nothing surprising to you in the above list, I’m not surprised! If you’ve read abit in the modern cognitive science literature, the list may even seem trivial. But it’s worthreflecting that 50 years ago, no such list could have been produced with the same level of broadacceptance. And less than 100 years ago, the Western world’s scientific understanding of themind was dominated by Freudian thinking; and not too long after that, by behaviorist thinking,which argued that theorizing about what went on inside the mind made no sense, and scienceshould focus entirely on analyzing external behavior. The progress of cognitive science hasn’tmade as many headlines as contemporaneous progress in neuroscience or computing hardwareand software, but it’s certainly been dramatic. One of the reasons that AGI is more achievablenow than in the 1950s and 60s when the AI field began, is that now we understand the structuresand processes characterizing human thinking a lot better.In spite of all the theoretical and empirical progress in the cognitive science field, however,there is still no consensus among experts on how the various aspects of intelligence in the above“human intelligence feature list” are achieved and interrelated. In these pages, however, forthe purpose of motivating CogPrime, we assume a broad integrative understanding roughly asfollows:• Perception: There is significant evidence that human visual perception occurs using aspatiotemporal hierarchy of pattern recognition modules, in which higher-level modules2.2 Commonly Recognized Aspects of Human-like Intelligence 23deal with broader spacetime regions, roughly as in the DeSTIN AGI architecture discussedin Chapter 4. Further, there is evidence that each module carries out temporal predictivepattern recognition as well as static pattern recognition. Audition likely utilizes a similarhierarchy. Olfaction may use something more like a Hopfield attractor neural network, asdescribed in Chapter 13. The networks corresponding to different sense modalities havemultiple cross-linkages, more at the upper levels than the lower, and also link richly intothe parts of the mind dealing with other functions.• Motor Control: This appears to be handled by a spatiotemporal hierarchy as well, in whicheach level of the hierarchy corresponds to higher-level (in space and time) movements. Thehierarchy is very tightly linked in with the perceptual hierarchies, allowing sensorimotorlearning and coordination.• Memory: There appear to be multiple distinct but tightly cross-linked memory systems,corresponding to different sorts of knowledge such as declarative (facts and beliefs), procedural,episodic, sensorimotor, attentional and intentional (goals).• Knowledge Representation: There appear to be multiple base-level representationalsystems; at least one corresponding to each memory system, but perhaps more than that.Additionally there must be the capability to dynamically create new context-specific representationalsystems founded on the base representational system.• Language: While there is surely some innate biasing in the human mind toward learningcertain types of linguistic structure, it’s also notable that language shares a great deal ofstructure with other aspects of intelligence like social roles [CB00] and the physical world[Cas07]. Language appears to be learned based on biases toward learning certain types ofrelational role systems; and language processing seems a complex mix of generic reasoningand pattern recognition processes with specialized acoustic and syntactic processingroutines.• Consciousness is pragmatically well-understood using Baars’ “global workspace” theory,in which a small subset of the mind’s content is summoned at each time into a “workingmemory” aka “workspace” aka “attentional focus” where it is heavily processed and used toguide action selection.• Thinking is a diverse combination of processes encompassing things like categorization,(crisp and uncertain) reasoning, concept creation, pattern recognition, and others; theseprocesses must work well with all the different types of memory and must effectively integrateknowledge in the global workspace with knowledge in long-term memory.• Social Intelligence seems closely tied with language and also with self-modeling; we modelourselves in large part using the same specialized biases we use to help us model others.None of the points in the above bullet list is particularly controversial, but neither are anyof them universally agreed-upon by experts. However, in order to make any progress on AGIdesign one must make some commitments to particular cognition-theoretic understandings, atthis level and ultimately at more precise levels as well. Further, general philosophical analyseslike the patternist philosophy to be reviewed in the following chapter only provide limitedguidance here. Patternism provides a filter for theories about specific cognitive functions – itrules out assemblages of cognitive-function-specific theories that don’t fit together to yield amind that could act effectively as a pattern-recognizing, goal-achieving system with the rightinternal emergent structures. But it’s not a precise enough filter to serve as a sole guide forcognitive theory even at the high level.The above list of points leads naturally into the integrative architecture diagram presentedin Chapter 5. But that generic architecture diagram is fairly involved, and before presenting24 2 What Is Human-Like General Intelligence?it, we will go through some more background regarding human-like intelligence (in the restof this chapter), philosophy of mind (in Chapter 3) and contemporary AGI architectures (inChapter4).2.3 Further Characterizations of Humanlike IntelligenceWe now present a few complementary approaches to characterizing the key aspects of humanlikeintelligence, drawn from different perspectives in the psychology and AI literature. Thesedifferent approaches all overlap substantially, which is good, yet each gives a slightly differentslant.2.3.1 Competencies Characterizing Human-like IntelligenceFirst we give a list of key competencies characterizing human level intelligence resulting fromthe the AGI Roadmap Workshop held at the University of Knoxville in October 2008 1 , whichwas organized by Ben Goertzel and Itamar Arel. In this list, each broad competency area islisted together with a number of specific competencies sub-areas within its scope:1. Perception: vision, hearing, touch, proprioception, crossmodal2. Actuation: physical skills, navigation, tool use3. Memory: episodic, declarative, behavioral4. Learning: imitation, reinforcement, interactive verbal instruction, written media, experimentation5. Reasoning: deductive, abductive, inductive, causal, physical, associational, categorization6. Planning: strategic, tactical, physical, social7. Attention: visual, social, behavioral8. Motivation: subgoal creation, affect-based motivation, control of emotions9. Emotion: expressing emotion, understanding emotion10. Self: self-awareness, self-control, other-awareness11. Social: empathy, appropriate social behavior, social communication, social inference, groupplay, theory of mind12. Communication: gestural, pictorial, verbal, language acquisition, cross-modal13. Quantitative: counting, grounded arithmetic, comparison, measurement14. Building/Creation: concept formation, verbal invention, physical construction, socialgroup formationClearly this list is getting at the same things as the textbook headings given in Section 2.2,but with a different emphasis due to its origin among AGI researchers rather than cognitive1 See http://www.ece.utk.edu/~itamar/AGI_Roadmap.html; participants included: Sam Adams, IBMResearch; Ben Goertzel, Novamente LLC; Itamar Arel, University of Tennessee; Joscha Bach, Institute of CognitiveScience, University of Osnabruck, Germany; Robert Coop, University of Tennessee; Rod Furlan, SingularityInstitute; Matthias Scheutz, Indiana University; J. Storrs Hall, Foresight Institute; Alexei Samsonovich, GeorgeMason University; Matt Schlesinger, Southern Illinois University; John Sowa, Vivomind Intelligence, Inc.; StuartC. Shapiro, University at Buffalo2.3 Further Characterizations of Humanlike Intelligence 25psychologists. As part of the AGI Roadmap project, specific tasks were created correspondingto each of the sub-areas in the above list; we will describe some of these tasks in Chapter 17.2.3.2 Gardner’s Theory of Multiple IntelligencesThe diverse list of human-level “competencies” given above is reminiscent of Gardner’s [Gar99]multiple intelligences (MI) framework – a psychological approach to intelligence assessmentbased on the idea that different people have mental strengths in different high-level domains,so that intelligence tests should contain aspects that focus on each of these domains separately.MI does not contradict the “complex goals in complex environments” view of intelligence, butrather may be interpreted as making specific commitments regarding which complex tasks andwhich complex environments are most important for roughly human-like intelligence.MI does not seek an extreme generality, in the sense that it explicitly focuses on domainsin which humans have strong innate capability as well as general-intelligence capability; therecould easily be non-human intelligences that would exceed humans according to both the commonsensehuman notion of “general intelligence” and the generic “complex goals in complexenvironments” or Hutter/Legg-style definitions, yet would not equal humans on the MI criteria.This strong anthropocentrism of MI is not a problem from an AGI perspective so long asone uses MI in an appropriate way, i.e. only for assessing the extent to which an AGI systemdisplays specifically human-like general intelligence. This restrictiveness is the price one paysfor having an easily articulable and relatively easily implementable evaluation framework.Table ?? summarizes the types of intelligence included in Gardner’s MI theory.Intelligence TypeLinguisticLogical-MathematicalMusicalBodily-KinestheticSpatial-VisualInterpersonalAspectsWords and language, written and spoken; retention, interpretationand explanation of ideas and information via language;understands relationship between communicationand meaningLogical thinking, detecting patterns, scientific reasoningand deduction; analyse problems, perform mathematicalcalculations, understands relationship between cause andeffect towards a tangible outcomeMusical ability, awareness, appreciation and use of sound;recognition of tonal and rhythmic patterns, understandsrelationship between sound and feelingBody movement control, manual dexterity, physical agilityand balance; eye and body coordinationVisual and spatial perception; interpretation and creationof images; pictorial imagination and expression; understandsrelationship between images and meanings, and betweenspace and effectPerception of other people’s feelings; relates to others; interpretationof behaviour and communications; understandsrelationships between people and their situationsTable 2.1: Types of Intelligence in Gardner’s Multiple Intelligence Theory26 2 What Is Human-Like General Intelligence?2.3.3 Newell’s Criteria for a Human Cognitive ArchitectureFinally, another related perspective is given by Alan Newell’s “functional criteria for a humancognitive architecture” [New90], which require that a humanlike AGI system should:1. Behave as an (almost) arbitrary function of the environment2. Operate in real time3. Exhibit rational, i.e., effective adaptive behavior4. Use vast amounts of knowledge about the environment5. Behave robustly in the face of error, the unexpected, and the unknown6. Integrate diverse knowledge7. Use (natural) language8. Exhibit self-awareness and a sense of self9. Learn from its environment10. Acquire capabilities through development11. Arise through evolution12. Be realizable within the brainIn our view, Newell’s criterion 1 is poorly-formulated, for while universal Turing computingpower is easy to come by, any finite AI system must inevitably be heavily adapted to someparticular class of environments for straightforward mathematical reasons [Hut05, GPI + 10].On the other hand, his criteria 11 and 12 are not relevant to the CogPrime approach as we arenot doing biological modeling but rather AGI engineering. However, Newell’s criteria 2-10 areessential in our view, and all will be covered in the following chapters.2.3.4 intelligence and CreativityCreativity is a key aspect of intelligence. While sometimes associated especially with geniuslevelintelligence in science or the arts, actually creativity is pervasive throughout intelligence,at all levels. When a child makes a flying toy car by pasting paper bird wings on his toy car, andwhen a bird figures out how to use a curved stick to get a piece of food out of a difficult corner– this is creativity, just as much as the invention of a new physics theory or the design of a newfashion line. The very nature of intelligence – achieving complex goals in complex environments– requires creativity for its achievement, because the nature of complex environments and goalsis that they are always unveiling new aspects, so that dealing with them involves inventingthings beyond what worked for previously known aspects.CogPrime contains a number of cognitive dynamics that are especially effective at creatingnew ideas, such as: concept creation (which synthesizes new concepts via combining aspectsof previous ones), probabilistic evolutionary learning (which simulates evolution by naturalselection, creating new procedures via mutation, combination and probabilistic modeling basedon previous ones), and analogical inference (an aspect of the Probabilistic Logic Networkssubsystems). But ultimately creativity is about how a system combines all the processes at itsdisposal to synthesize novel solutions to the problems posed by its goals in its environment.There are times, of course, when the same goal can be achieved in multiple ways – somemore creative than others. In CogPrime this relates to the existence of multiple top-level goals,one of which may be novelty. A system with novelty as one of its goals, alongside other more2.4 Preschool as a View into Human-like General Intelligence 27specific goals, will have a tendency to solve other problems in creative ways, thus fulfilling itsnovelty goal along with its other goals. This can be seen at the level of childlike behaviors, andalso at a much more advanced level. Salvador Dali wanted to depict his thoughts and feelings,but he also wanted to do so in a striking and unusual way; this combination of aspirationsspurred him to produce his amazing art. A child who is asked to draw a house, but has agoal of novelty, may draw a tower with a swimming pool on the roof rather than a typicalColonial structure. A physical motivated by novelty will seek a non-obvious solution to theequation at hand, rather than just applying tried and true methods, and perhaps discoversome new phenomenon. Novelty can be measured formally in terms of information-theoreticsurprisingness based upon a given basis of knowledge and experience [Sch06]; something thatis novel and creative to a child may be familiar to the adult world, and a solution that seemsnovel and creative to a brilliant scientist today, may seem like cliche’ elementary school levelwork 100 years from now.Measuring creativity is even more difficult and subjective than measuring intelligence. Qualitatively,however, we humans can recognize it; and we suspect that the qualitative emergenceof dramatic, multidisciplinary computational creativity will be one of the things that makes thehuman population feel emotionally that advanced AGI has finally arrived.2.4 Preschool as a View into Human-like General IntelligenceOne issue that arises when pursuing the grand goal of human-level general intelligence is howto measure partial progress. The classic Turing Test of imitating human conversation remainstoo difficult to usefully motivate immediate-term AI research (see [HF95] [Fre90] for argumentsthat it has been counterproductive for the AI field). The same holds true for comparable alternativeslike the Robot College Test of creating a robot that can attend a semester of universityand obtain passing grades. However, some researchers have suggested intermediary goals, thatconstitute partial progress toward the grand goal and yet are qualitatively different from thehighly specialized problems to which most current AI systems are applied.In this vein, Sam Adams and his team at IBM have outlined a so-called “Toddler TuringTest,” in which one seeks to use AI to control a robot qualitatively displaying similar cognitivebehaviors to a young human child (say, a 3 year old) [AABL02]. In fact this sort of idea has along and venerable history in the AI field – Alan Turing’s original 1950 paper on AI [Tur50],where he proposed the Turing Test, contains the suggestion that"Instead of trying to produce a programme to simulate the adult mind,why not rather try to produce one which simulates the child’s?"We find this childlike cognition based approach promising for many reasons, including its integrativenature: what a young child does involves a combination of perception, actuation, linguisticand pictorial communication, social interaction, conceptual problem solving and creativeimagination. Specifically, inspired by these ideas, in Chapter 16 we will suggest the approachof teaching and testing early-stage AGI systems in environments that emulate the preschoolsused for teaching human children.Human intelligence evolved in response to the demands of richly interactive environments,and a preschool is specifically designed to be a richly interactive environment with the capabilityto stimulate diverse mental growth. So, we are currently exploring the use of CogPrime to control28 2 What Is Human-Like General Intelligence?virtual agents in preschool-like virtual world environments, as well as commercial humanoidrobot platforms such as the Nao (see Figure 2.1) or Robokind (2.2) in physical preschool-likerobot labs.Another advantage of focusing on childlike cognition is that child psychologists have createda variety of instruments for measuring child intelligence. In Chapter 17, we will discuss anapproach to evaluating the general intelligence of human childlike AGI systems via combiningtests typically used to measure the intelligence of young human children, with additional testscrafted based on cognitive science and the standard preschool curriculum.To put it differently: While our long-term goal is the creation of genius machines with generalintelligence at the human level and beyond, we believe that every young child has a certaingenius; and by beginning with this childlike genius, we can built a platform capable of developinginto a genius machine with far more dramatic capabilities.2.4.1 Design for an AGI PreschoolMore precisely, we don’t suggest to place a CogPrime system in an environment that is anexact imitation of a human preschool – this would be inappropriate since current robotic orvirtual bodies are very differently abled than the body of a young human child. But we aim toplace CogPrime in an environment emulating the basic diversity and educational character ofa typical human preschool. We stress this now, at this early point in the book, because we willuse running examples throughout the book drawn from the preschool context.The key notion in modern preschool design is the “learning center,” an area designed andoutfitted with appropriate materials for teaching a specific skill. Learning centers are designed toencourage learning by doing, which greatly facilitates learning processes based on reinforcement,imitation and correction; and also to provide multiple techniques for teaching the same skills,to accommodate different learning styles and prevent overfitting and overspecialization in thelearning of new skills.Centers are also designed to cross-develop related skills. A “manipulatives center,” for example,provides physical objects such as drawing implements, toys and puzzles, to facilitatedevelopment of motor manipulation, visual discrimination, and (through sequencing and classificationgames) basic logical reasoning. A “dramatics center” cross-trains interpersonal andempathetic skills along with bodily-kinesthetic, linguistic, and musical skills. Other centers,such as art, reading, writing, science and math centers are also designed to train not just onearea, but to center around a primary intelligence type while also cross-developing related areas.For specific examples of the learning centers associated with particular contemporary preschools,see [Nei98]. In many progressive, student-centered preschools, students are left largely to theirown devices to move from one center to another throughout the preschool room. Generally,each center will be staffed by an instructor at some points in the day but not others, providinga variety of learning experiences.To imitate the general character of a human preschool, we will create several centers in ourrobot lab. The precise architecture will be adapted via experience but initial centers will likelybe:• a blocks center: a table with blocks on it• a language center: a circle of chairs, intended for people to sit around and talk with therobot2.5 Integrative and Synergetic Approaches to Artificial General Intelligence 29• a manipulatives center, with a variety of different objects of different shapes and sizes,intended to teach visual and motor skills• a ball play center: where balls are kept in chests and there is space for the robot to kickthe balls around• a dramatics center where the robot can observe and enact various movementsOne Running ExampleAs we proceed through the various component structures and dynamics of CogPrime in thefollowing chapters, it will be useful to have a few running examples to use to explain how thevarious parts of the system are supposed to work. One example we will use fairly frequently isdrawn from the preschool context: the somewhat open-ended task of Build me somethingout of blocks, that you haven’t built for me before, and then tell me what it is. Thisis a relatively simple task that combines multiple aspects of cognition in a richly interconnectedway, and is the sort of thing that young children will naturally do in a preschool setting.2.5 Integrative and Synergetic Approaches to Artificial GeneralIntelligenceIn Chapter 1 we characterized CogPrime as an integrative approach. And we suggest that thenaturalness of integrative approaches to AGI follows directly from comparing above lists ofcapabilities and criteria to the array of available AI technologies. No single known algorithmor data structure appears easily capable of carrying out all these functions, so if one wantsto proceed now with creating a general intelligence that is even vaguely humanlike, one mustintegrate various AI technologies within some sort of unifying architecture.For this reason and others, an increasing amount of work in the AI community these daysis integrative in one sense or another. Estimation of Distribution Algorithms integrate probabilisticreasoning with evolutionary learning [Pel05]. Markov Logic Networks [RD06] integrateformal logic and probabilistic inference, as does the Probabilistic Logic Networks framework[GIGH08] utilized in CogPrime and explained further in the book, and other works in the“Progic” area such as [WW06]. Leslie Pack Kaelbling has synthesized low-level robotics methods(particle filtering) with logical inference [ZPK07]. Dozens of further examples could be given.The construction of practical robotic systems like the Stanley system that won the DARPAGrand Challenge [Tea06] involve the integration of numerous components based on differentprinciples. These algorithmic and pragmatic innovations provide ample raw materials for theconstruction of integrative cognitive architectures and are part of the reason why childlike AGIis more approachable now than it was 50 or even 10 years ago.Further, many of the cognitive architectures described in the current AI literature are “integrative”in the sense of combining multiple, qualitatively different, interoperating algorithms.Chapter 4 gives a high-level overview of existing cognitive architectures, dividing them intosymbolic, emergentist (e.g. neural network) and hybrid architectures. The hybrid architecturesgenerally integrate symbolic and neural components, often with multiple subcomponents withineach of these broad categories. However, we believe that even these excellent architectures arenot integrative enough, in the sense that they lack sufficiently rich and nuanced interactions30 2 What Is Human-Like General Intelligence?between the learning components associated with different kinds of memory, and hence are unlikelyto give rise to the emergent structures and dynamics characterizing general intelligence.One of the central ideas underlying CogPrime is that with an integrative cognitive architecturethat combines multiple aspects of intelligence, achieved by diverse structures and algorithms,within a common framework designed specifically to support robust synergetic interactionsbetween these aspects.The simplest way to create an integrative AI architecture is to loosely couple multiple componentscarrying out various functions, in such a way that the different components pass inputsand outputs amongst each other but do not interfere with or modulate each others’ internalfunctioning in real-time. However, the human brain appears to be integrative in a much tightersense, involving rich real-time dynamical coupling between various components with distinctbut related functions. In [Goe09a] we have hypothesized that the brain displays a property ofcognitive synergy, according to which multiple learning processes can not only dispatchsubproblems to each other, but also share contextual understanding in real-time, sothat each one can get help from the others in a contextually savvy way. By imbuing AI architectureswith cognitive synergy, we hypothesize, one can get past the bottlenecks that haveplagued AI in the past. Part of the reasoning here, as elaborated in Chapter 9 and [Goe09b], isthat real physical and social environments display a rich dynamic interconnection between theirvarious aspects, so that richly dynamically interconnected integrative AI architectures will beable to achieve goals within them more effectively.And this brings us to the patternist perspective on intelligent systems, alluded to above andfleshed out further in Chapter 3 with its focus on the emergence of hierarchically and heterarchicallystructured networks of patterns, and pattern-systems modeling self and others. Ultimatelythe purpose of cognitive synergy in an AGI system is to enable the various AI algorithms andstructures composing the system to work together effectively enough to give rise to the rightsystem-wide emergent structures characterizing real-world general intelligence. The underlyingtheory is that intelligence is not reliant on any particular structure or algorithm, but is relianton the emergence of appropriately structured networks of patterns, which can then be used toguide ongoing dynamics of pattern recognition and creation. And the underlying hypothesis isthat the emergence of these structures cannot be achieved by a loosely interconnected assemblageof components, no matter how sensible the architecture; it requires a tightly connected,synergetic system.It is possible to make these theoretical ideas about cognition mathematically rigorous; forinstance, Appendix ?? briefly presents a formal definition of cognitive synergy that has beenanalyzed as part of an effort to prove theorems about the importance of cognitive synergy forgiving rise to emergent system properties associated with general intelligence. However, whilewe have found such formal analyses valuable for clarifying our designs and understanding theirqualitative properties, we have concluded that, for the present, the best way to explore ourhypotheses about cognitive synergy and human-like general intelligence is empirically – viabuilding and testing systems like CogPrime.2.5.1 Achieving Humanlike Intelligence via Cognitive SynergySumming up: at the broadest level, there are four primary challenges in constructing an integrative,cognitive synergy based approach to AGI:2.5 Integrative and Synergetic Approaches to Artificial General Intelligence 311. choosing an overall cognitive architecture that possesses adequate richness and flexibilityfor the task of achieving childlike cognition.2. Choosing appropriate AI algorithms and data structures to fulfill each of the functionsidentified in the cognitive architecture (e.g. visual perception, audition, episodic memory,language generation, analogy,...)3. Ensuring that these algorithms and structures, within the chosen cognitive architecture,are able to cooperate in such a way as to provide appropriate coordinated, synergeticintelligent behavior (a critical aspect since childlike cognition is an integrated functionalresponse to the world, rather than a loosely coupled collection of capabilities.)4. Embedding one’s system in an environment that provides sufficiently rich stimuli andinteractions to enable the system to use this cooperation to ongoingly, creatively developan intelligent internal world-model and self-model.We argue that CogPrimeprovides a viable way to address these challenges.32 2 What Is Human-Like General Intelligence?Fig. 2.1: The Nao humanoid robot2.5 Integrative and Synergetic Approaches to Artificial General Intelligence 33Fig. 2.2: The Nao humanoid robot
Chapter 3A Patternist Philosophy of Mind3.1 IntroductionIn the last chapter we discussed human intelligence from a fairly down-to-earth perspective,looking at the particular intelligent functions that human beings carry out in their everydaylives. And we strongly feel this practical perspective is important: Without this concreteness, it’stoo easy for AGI research to get distracted by appealing (or frightening) abstractions of varioussorts. However, it’s also important to look at the nature of mind and intelligence from a moregeneral and conceptual perspective, to avoid falling into an approach that follows the particularsof human capability but ignores the deeper structures and dynamics of mind that ultimatelyallow human minds to be so capable. In this chapter we very briefly review some ideas from thepatternist philosophy of mind, a general conceptual framework on intelligence which hasbeen inspirational for many key aspects of the CogPrime design, and which has been ongoinglydeveloped by one of the authors (Ben Goertzel) during the last two decades (in a series ofpublications beginning in 1991, most recently The Hidden Pattern [Goe06a]). Some of the ideasdescribed are quite broad and conceptual, and are related to CogPrime only via serving asgeneral inspirations; others are more concrete and technical, and are actually utilized withinthe design itself.CogPrime is an integrative design formed via the combination of a number of differentphilosophical, scientific and engineering ideas. The success or failure of the design doesn’t dependon any particular philosophical understanding of intelligence. In that sense, the more abstractnotions presented in this chapter should be considered “optional” rather than critical in aCogPrime context. However, due to the core role patternism has played in the development ofCogPrime, understanding a few things about general patternist philosophy will be helpful forunderstanding CogPrime, even for those readers who are not philosophically inclined. Thosereaders who are philosophically inclined, on the other hand, are urged to read The HiddenPattern and then interpret the particulars of CogPrime in this light.3.2 Some Patternist PrinciplesThe patternist philosophy of mind is a general approach to thinking about intelligent systems.It is based on the very simple premise that mind is made of pattern – and that a mind is a3536 3 A Patternist Philosophy of Mindsystem for recognizing patterns in itself and the world, critically including patterns regardingwhich procedures are likely to lead to the achievement of which goals in which contexts.Pattern as the basis of mind is not in itself is a very novel idea; this concept is present, forinstance, in the 19th-century philosophy of Charles Peirce [Pei34], in the writings of contemporaryphilosophers Daniel Dennett [Den91] and Douglas Hofstadter [Hof79, Hof96], in BenjaminWhorf’s [Who64] linguistic philosophy and Gregory Bateson’s [Bat79] systems theory of mindand nature. Bateson spoke of the Metapattern: “that it is pattern which connects.” In Goertzel’swritings on philosophy of mind, an effort has been made to pursue this theme more thoroughlythan has been done before, and to articulate in detail how various aspects of human mind andmind in general can be well-understood by explicitly adopting a patternist perspective. 1In the patternist perspective, "pattern" is generally defined as "representation as somethingsimpler." Thus, for example, if one measures simplicity in terms of bit-count, then a programcompressing an image would be a pattern in that image. But if one uses a simplicity measureincorporating run-time as well as bit-count, then the compressed version may or may not be apattern in the image, depending on how one’s simplicity measure weights the two factors. Thisdefinition encompasses simple repeated patterns, but also much more complex ones. Whilepattern theory has typically been elaborated in the context of computational theory, it is notintrinsically tied to computation; rather, it can be developed in any context where there is anotion of "representation" or "production" and a way of measuring simplicity. One just needsto be able to assess the extent to which f represents or produces X, and then to compare thesimplicity of f and X; and then one can assess whether f is a pattern in X. A formalization ofthis notion of pattern is given in [Goe06a] and briefly summarized at the end of this chapter.Next, in patternism the mind of an intelligent system is conceived as the (fuzzy) set ofpatterns in that system, and the set of patterns emergent between that system and othersystems with which it interacts. The latter clause means that the patternist perspective isinclusive of notions of distributed intelligence [Hut96]. Basically, the mind of a system is thefuzzy set of different simplifying representations of that system that may be adopted.Intelligence is conceived, similarly to in Marcus Hutter’s [Hut05] recent work (and as elaboratedinformally in Chapter 2 above, and formally in Chapter 7 below), as the ability to achievecomplex goals in complex environments; where complexity itself may be defined as the possessionof a rich variety of patterns. A mind is thus a collection of patterns that is associatedwith a persistent dynamical process that achieves highly-patterned goals in highly-patternedenvironments.An additional hypothesis made within the patternist philosophy of mind is that reflection iscritical to intelligence. This lets us conceive an intelligent system as a dynamical system thatrecognizes patterns in its environment and itself, as part of its quest to achieve complex goals.While this approach is quite general, it is not vacuous; it gives a particular structure to thetasks of analyzing and synthesizing intelligent systems. About any would-be intelligent system,we are led to ask questions such as:• How are patterns represented in the system? That is, how does the underlying infrastructureof the system give rise to the displaying of a particular pattern in the system’s behavior?• What kinds of patterns are most compactly represented within the system?• What kinds of patterns are most simply learned?1 In some prior writings the term “psynet model of mind” has been used to refer to the application of patternistphilosophy to cognitive theory, but this term has been "deprecated" in recent publications as it seemed tointroduce more confusion than clarification.3.2 Some Patternist Principles 37• What learning processes are utilized for recognizing patterns?• What mechanisms are used to give the system the ability to introspect (so that it canrecognize patterns in itself)?Now, these same sorts of questions could be asked if one substituted the word “pattern” withother words like “knowledge” or “information”. However, we have found that asking these questionsin the context of pattern leads to more productive answers, avoiding unproductive bywaysand also tying in very nicely with the details of various existing formalisms and algorithms forknowledge representation and learning.Among the many kinds of patterns in intelligent systems, semiotic patterns are particularlyinteresting ones. Peirce decomposed these into three categories:• iconic patterns, which are patterns of contextually important internal similarity betweentwo entities (e.g. an iconic pattern binds a picture of a person to that person)• indexical patterns, which are patterns of spatiotemporal co-occurrence (e.g. an indexicalpattern binds a wedding dress and a wedding)• symbolic patterns, which are patterns indicating that two entities are often involved inthe same relationships (e.g. a symbolic pattern between the number “5” (the symbol) andvarious sets of 5 objects (the entities that the symbol is taken to represent))Of course, some patterns may span more than one of these semiotic categories; and thereare also some patterns that don’t fall neatly into any of these categories. But the semioticpatterns are particularly important ones; and symbolic patterns have played an especially largerole in the history of AI, because of the radically different approaches different researchers havetaken to handling them in their AI systems. Mathematical logic and related formalisms providesophisticated mechanisms for combining and relating symbolic patterns (“symbols”), and someAI approaches have focused heavily on these, sometimes more so than on the identification ofsymbolic patterns in experience or the use of them to achieve practical goals. We will look fairlycarefully at these differences in Chapter 4.Pursuing the patternist philosophy in detail leads to a variety of particular hypotheses andconclusions about the nature of mind. Following from the view of intelligence in terms ofachieving complex goals in complex environments, comes a view in which the dynamics ofa cognitive system are understood to be governed by two main forces:• self-organization, via which system dynamics cause existing system patterns to give rise tonew ones• goal-oriented behavior, which will be defined more rigorously in Chapter 7, but basicallyamounts to a system interacting with its environment in a way that appears like an attemptto maximize some reasonably simple functionSelf-organized and goal-oriented behavior must be understood as cooperative aspects. If anagent is asked to build a surprising structure out of blocks and does so, this is goal-oriented.But the agent’s ability to carry out this goal-oriented task will be greater if it has previouslyplayed around with blocks a lot in an unstructured, spontaneous way. And the “nudge towardcreativity” given to it by asking it to build a surprising blocks structure may cause it to exploresome novel patterns, which then feed into its future unstructured blocks play.Based on these concepts, as argued in detail in [Goe06a], several primary dynamical principlesmay be posited, including:38 3 A Patternist Philosophy of Mind• Evolution , conceived as a general process via which patterns within a large populationthereof are differentially selected and used as the basis for formation of new patterns, basedon some “fitness function” that is generally tied to the goals of the agent– Example: If trying to build a blocks structure that will surprise Bob, an agent maysimulate several procedures for building blocks structures in its “mind’s eye”, assessingfor each one the expected degree to which it might surprise Bob. The search throughprocedure space could be conducted as a form of evolution, via an algorithm such asMOSES.• Autopoiesis: the process by which a system of interrelated patterns maintains its integrity,via a dynamic in which whenever one of the patterns in the system begins to decrease inintensity, some of the other patterns increase their intensity in a manner that causes thetroubled pattern to increase in intensity again– Example: An agent’s set of strategies for building the base of a tower, and its set ofstrategies for building the middle part of a tower, are likely to relate autopoietically. Ifthe system partially forgets how to build the base of a tower, then it may regeneratethis missing knowledge via using its knowledge about how to build the middle part(i.e., it knows it needs to build the base in a way that will support good middle parts).Similarly if it partially forgets how to build the middle part, then it may regenerate thismissing knowledge via using its knowledge about how to build the base (i.e. it knows agood middle part should fit in well with the sorts of base it knows are good).– This same sort of interdependence occurs between pattern-sets containing more thantwo elements– Sometimes (as in the above example) autopoietic interdependence in the mind is tiedto interdependencies in the physical world, sometimes not.• Association. Patterns, when given attention, spread some of this attention to other patternsthat they have previously been associated with in some way. Furthermore, there isPeirce’s law of mind [Pei34], which could be paraphrased in modern terms as stating thatthe mind is an associative memory network, whose dynamics dictate that every idea inthe memory is an active agent, continually acting on those ideas with which the memoryassociates it.– Example: Building a blocks structure that resembles a tower, spreads attention to memoriesof prior towers the agents has seen, and also to memories of people the agent knowshave seen towers, and structures it has built at the same time as towers, structures thatresemble towers in various respects, etc.• Differential attention allocation / credit assignment. Patterns that have been valuablefor goal-achievement are given more attention, and are encouraged to participate ingiving rise to new patterns.– Example: Perhaps in a prior instance of the task “build me a surprising structure out ofblocks,” searching through memory for non-blocks structures that the agent has playedwith has proved a useful cognitive strategy. In that case, when the task is posed to theagent again, it should tend to allocate disproportionate resources to this strategy.• Pattern creation. Patterns that have been valuable for goal-achievement are mutated andcombined with each other to yield new patterns.3.2 Some Patternist Principles 39– Example: Building towers has been useful in a certain context, but so has buildingstructures with a large number of triangles. Why not build a tower out of triangles?Or maybe a vaguely tower-like structure that uses more triangles than a tower easilycould?– Example: Building an elongated block structure resembling a table was successful in thepast, as was building a structure resembling a very flat version of a chair. Generalizing,maybe building distorted versions of furniture is good. Or maybe it is building distortedversion of any previously perceived objects that is good. Or maybe both, to differentdegrees....Next, for a variety of reasons outlined in [Goe06a] it becomes appealing to hypothesize that thenetwork of patterns in an intelligent system must give rise to the following large-scale emergentstructures• Hierarchical network. Patterns are habitually in relations of control over other patterns thatrepresent more specialized aspects of themselves.– Example: The pattern associated with “tall building” has some control over the patternassociated with “tower”, as the former represents a more general concept ... and “tower”has some control over “Eiffel tower”, etc.• Heterarchical network. The system retains a memory of which patterns have previouslybeen associated with each other in any way.– Example: “Tower” and “snake” are distant in the natural pattern hierarchy, but may beassociatively/heterarchically linked due to having a common elongated structure. Thisheterarchical linkage may be used for many things, e.g. it might inspire the creativeconstruction of a tower with a snake’s head.• Dual network. Hierarchical and heterarchical structures are combined, with the dynamicsof the two structures working together harmoniously. Among many possible ways to hierarchicallyorganize a set of patterns, the one used should be one that causes hierarchicallynearby patterns to have many meaningful heterarchical connections; and of course, thereshould be a tendency to search for heterarchical connections among hierarchically nearbypatterns.– Example: While the set of patterns hierarchically nearby “tower” and the set of patternsheterarchically nearby “tower” will be quite different, they should still have more overlapthan random pattern-sets of similar sizes. So, if looking for something else heterarchicallynear “tower”, using the hierarchical information about “tower” should be of some use,and vice versa.– In PLN, hierarchical relationships correspond to Atoms A and B so that InheritanceABand InheritanceBA have highly dissimilar strength; and heterarchical relationships correspondto IntensionalSimilarity relationships. The dual network structure then ariseswhen intensional and extensional inheritance approximately correlate with each other,so that inference about either kind of inheritance assists with figuring out about theother kind.• Self structure. A portion of the network of patterns forms into an approximate image of theoverall network of patterns.40 3 A Patternist Philosophy of Mind– Example: Each time the agent builds a certain structure, it observes itself buildingthe structure, and its role as “builder of a tall tower” (or whatever the structure is)becomes part of its self-model. Then when it is asked to build something new, it mayconsult its self-model to see if it believes itself capable of building that sort of thing (forinstance, if it is asked to build something very large, its self-model may tell it that itlacks persistence for such projects, so it may reply “I can try, but I may wind up notfinishing it”).As we proceed through the CogPrime design in the following pages, we will see how eachof these abstract concepts arises concretely from CogPrime’s structures and algorithms. If thetheory of [Goe06a] is correct, then the success of CogPrime as a design will depend largely onwhether these high-level structures and dynamics can be made to emerge from the synergeticinteraction of CogPrime’s representation and algorithms, when they are utilized to control anappropriate agent in an appropriate environment.3.3 Cognitive SynergyNow we dig a little deeper and present a different sort of “general principle of feasible generalintelligence”, already hinted in earlier chapters: the cognitive synergy principle 2 , which is botha conceptual hypothesis about the structure of generally intelligent systems in certain classes ofenvironments, and a design principle used to guide the design of CogPrime. Chapter 8 presentsa mathematical formalization of the notion of cognitive synergy; here we present the conceptualidea informally, which makes it more easily digestible but also more vague-sounding.We will focus here on cognitive synergy specifically in the case of “multi-memory systems,”which we define as intelligent systems whose combination of environment, embodiment andmotivational system make it important for them to possess memories that divide into partiallybut not wholly distinct components corresponding to the categories of:• Declarative memory– Examples of declarative knowledge: Towers on average are taller than buildings. I generallyam better at building structures I imagine, than at imitating structures I’m shownin pictures.• Procedural memory (memory about how to do certain things)– Examples of procedural knowledge: Practical know-how regarding how to pick up anelongated rectangular block, or a square one. Know-how regarding when to approacha problem by asking “What would one of my teachers do in this situation” versus bythinking through the problem from first principles.• Sensory and episodic memory– Example of sensory knowledge: memory of Bob’s face; memory of what a specific tallblocks tower looked like2 While these points are implicit in the theory of mind given in [Goe06a], they are not articulated in thisspecific form there. So the material presented in this section is a new development within patternist philosophy,developed since [Goe06a] in a series of conference papers such as [Goe09a].3.3 Cognitive Synergy 41– Example of episodic knowledge: memory of the situation in which the agent first metBob; memory of a situation in which a specific tall blocks tower was built• Attentional memory (knowledge about what to pay attention to in what contexts)– Example of attentional knowledge: When involved with a new person, it’s useful to payattention to whatever that person looks at• Intentional memory (knowledge about the system’s own goals and subgoals)– Example of intentional knowledge: If my goal is to please some person whom I don’tknow that well, then a subgoal may be figuring out what makes that person smile.In Chapter 9 below we present a detailed argument as to how the requirement for a multimemoryunderpinning for general intelligence emerges from certain underlying assumptionsregarding the measurement of the simplicity of goals and environments. Specifically we arguethat each of these memory types corresponds to certain modes of communication, so that intelligentagents which have to efficiently handle a sufficient variety of types of communication withother agents, are going to have to handle all these types of memory. These types of communicationoverlap and are often used together, which implies that the different memories and theirassociated cognitive processes need to work together. The points made in this section do notrely on that argument regarding the relation of multiple memory types to the environmentalsituation of multiple communication types. What they do rely on is the assumption that, inthe intelligence agent in question, the different components of memory are significantly but notwholly distinct. That is, there are significant “family resemblances” between the memories of asingle type, yet there are also thoroughgoing connections between memories of different types.Repeating the above points in a slightly more organized manner and then extending them, theessential idea of cognitive synergy, in the context of multi-memory systems, may be expressedin terms of the following points1. Intelligence, relative to a certain set of environments, may be understood as the capabilityto achieve complex goals in these environments.2. With respect to certain classes of goals and environments, an intelligent system requires a“multi-memory” architecture, meaning the possession of a number of specialized yet interconnectedknowledge types, including: declarative, procedural, attentional, sensory, episodicand intentional (goal-related). These knowledge types may be viewed as different sorts ofpatterns that a system recognizes in itself and its environment.3. Such a system must possess knowledge creation (i.e. pattern recognition / formation) mechanismscorresponding to each of these memory types. These mechanisms are also called“cognitive processes.”4. Each of these cognitive processes, to be effective, must have the capability to recognize whenit lacks the information to perform effectively on its own; and in this case, to dynamicallyand interactively draw information from knowledge creation mechanisms dealing with othertypes of knowledge5. This cross-mechanism interaction must have the result of enabling the knowledge creationmechanisms to perform much more effectively in combination than they would if operatednon-interactively. This is “cognitive synergy.”Interactions as mentioned in Points 4 and 5 in the above list are the real conceptual meatof the cognitive synergy idea. One way to express the key idea here, in an AI context, is that42 3 A Patternist Philosophy of Mindmost AI algorithms suffer from combinatorial explosions: the number of possible elements tobe combined in a synthesis or analysis is just too great, and the algorithms are unable tofilter through all the possibilities, given the lack of intrinsic constraint that comes along witha “general intelligence” context (as opposed to a narrow-AI problem like chess-playing, wherethe context is constrained and hence restricts the scope of possible combinations that needsto be considered). In an AGI architecture based on cognitive synergy, the different learningmechanisms must be designed specifically to interact in such a way as to palliate each others’combinatorial explosions - so that, for instance, each learning mechanism dealing with a certainsort of knowledge, must synergize with learning mechanisms dealing with the other sorts ofknowledge, in a way that decreases the severity of combinatorial explosion.One prerequisite for cognitive synergy to work is that each learning mechanism must recognizewhen it is “stuck,” meaning it’s in a situation where it has inadequate information tomake a confident judgment about what steps to take next. Then, when it does recognize thatit’s stuck, it may request help from other, complementary cognitive mechanisms.3.4 The General Structure of Cognitive Dynamics: Analysis andSynthesisWe have discussed the need for synergetic interrelation between cognitive processes correspondingto different types of memory ... and the general high-level cognitive dynamics that a mindmust possess (evolution, autopoiesis). The next step is to dig further into the nature of the cognitiveprocesses associated with different memory types and how they give rise to the neededhigh-level cognitive dynamics. In this section we present a general theory of cognitive processesbased on a decomposition of cognitive processes into the two categories of analysis and synthesis,and a general formulation of each of these categories 3 .Specifically we focus here on what we call focused cognitive processes; that is, cognitiveprocesses that selectively focus attention on a subset of the patterns making up a mind. Ingeneral these are not the only kind, there may also be global cognitive processes that act onevery pattern in a mind. An example of a global cognitive process in CogPrime is the basicattention allocation process, which spreads “importance” among all knowledge in the system’smemory. Global cognitive processes are also important, but focused cognitive processes aresubtler to understand which is why we spend more time on them here.3.4.1 Component-Systems and Self-Generating SystemsWe begin with autopoesis – and, more specifically, with the concept of a “component-system”,as described in George Kampis’s book Self-Modifying Systems in Biology and Cognitive Science[Kam91], and as modified into the concept of a “self-generating system” or SGS in Goertzel’sbook Chaotic Logic [Goe94]. Roughly speaking, a Kampis-style component-system consists ofa set of components that combine with each other to form other compound components. The3 While these points are highly compatible with theory of mind given in [Goe06a], they are not articulated there.The material presented in this section is a new development within patternist philosophy, presented previouslyonly in the article [GPPG06].3.4 The General Structure of Cognitive Dynamics: Analysis and Synthesis 43metaphor Kampis uses is that of Lego blocks, combining to form bigger Lego structures. Compoundstructures may in turn be combined together to form yet bigger compound structures.A self-generating system is basically the same concept as a component-system, but understoodto be computable, whereas Kampis claims that component-systems are uncomputable.Next, in SGS theory there is also a notion of reduction (not present in the Lego metaphor):sometimes when components are combined in a certain way, a “reaction” happens, which maylead to the elimination of some of the components. One relevant metaphor here is chemistry.Another is abstract algebra: for instance, if we combine a component f with its “inverse” componentf −1 , both components are eliminated. Thus, we may think about two stages in theinteraction of sets of components: combination, and reduction. Reduction may be thought ofas algebraic simplification, governed by a set of rules that apply to a newly created compoundcomponent, based on the components that are assembled within it.Formally, suppose C 1 , C 2 , ... is the set of components present in a discrete-time componentsystemat time t. Then, the components present at time t+1 are a subset of the set of componentsof the formReduce(Join(C i (1), ..., C i (r)))where Join is a joining operation, and Reduce is a reduction operator. The joining operationis assumed to map tuples of components into components, and the reduction operator is assumedto map the space of components into itself. Of course, the specific nature of a component systemis totally dependent on the particular definitions of the reduction and joining operators; infollowing chapters we will specify these for the CogPrime system, but for the purpose of thebroader theoretical discussion in this section they may be left general.What is called the “cognitive equation” in Chaotic Logic [Goe94] is the case of a SGS wherethe patterns in the system at time t have a tendency to correspond to components of the systemat future times t + s. So, part of the action of the system is to transform implicit knowledge(patterns among system components) into explicit knowledge (specific system components). Wewill see one version of this phenomenon in Chapter 14 where we model implicit knowledge usingmathematical structures called “derived hypergraphs”; and we will also later review several waysin which CogPrime’s dynamics explicitly encourage cognitive-equation type dynamics, e.g.:• inference, which takes conclusions implicit in the combination of logical relationships, andmakes them implicit by deriving new logical relationships from them• map formation, which takes concepts that have often been active together, and creates newconcepts grouping them• association learning, which creates links representing patterns of association between entities• probabilistic procedure learning, which creates new models embodying patterns regardingwhich procedures tend to perform well according to particular fitness functions3.4.2 Analysis and SynthesisNow we move on to the main point of this section: the argument that all or nearly all focusedcognitive processes are expressible using two general process-schemata we call synthesis and44 3 A Patternist Philosophy of Mindanalysis 4 . The notion of “focused cognitive process” will be exemplified more thoroughly below,but in essence what is meant is a cognitive process that begins with a small number of items(drawn from memory) as its focus, and has as its goal discovering something about theseitems, or discovering something about something else in the context of these items or in a waystrongly biased by these items. This is different from a global cognitive process whose goal ismore broadly-based and explicitly involves all or a large percentage of the knowledge in anintelligent system’s memory store.Among the focused cognitive processes are those governed by the so-called cognitive schematicimplicationContext ∧ P rocedure → Goalwhere the Context involves sensory, episodic and/or declarative knowledge; and attentionalknowledge is used to regulate how much resource is given to each such schematic implication inmemory. Synergy among the learning processes dealing with the context, the procedure and thegoal is critical to the adequate execution of the cognitive schematic using feasible computationalresources. This sort of explicitly goal-driven cognition plays a significant though not necessarilydominant role in CogPrime, and is also related to production rules systems and other traditionalAI systems, as will be articulated in Chapter 4.The synthesis and analysis processes as we conceive them, in the general framework of SGStheory, are as follows. First, synthesis, as shown in Figure 3.1, is defined assynthesis: Iteratively build compounds from the initial component pool using the combinators,greedily seeking compounds that seem likely to achieve the goal.Or in more detail:1. Begin with some initial components (the initial “current pool”), an additional set of componentsidentified as “combinators” (combination operators), and a goal function2. Combine the components in the current pool, utilizing the combinators, to form productcomponents in various ways, carrying out reductions as appropriate, and calculating relevantquantities associated with components as needed3. Select the product components that seem most promising according to the goal function,and add these to the current pool (or else simply define these as the current pool)4. Return to Step 2And analysis, as shown in Figure 3.2, is defined asanalysis: Iteratively search (the system’s long-term memory) for component-sets that combineusing the combinators to form the initial component pool (or subsets thereof), greedilyseeking component-sets that seem likely to achieve the goalor in more detail:1. Begin with some components (the initial “current pool”) and a goal function2. Seek components so that, if one combines them to form product components using thecombinators and then performs appropriate reductions, one obtains (as many as possibleof) the components in the current pool4 In [GPPG06], what is here called “analysis” was called “backward synthesis”, a name which has some advantagessince it indicated that what’s happening is a form of creation; but here we have opted for the more traditionalanalysis/synthesis terminology3.4 The General Structure of Cognitive Dynamics: Analysis and Synthesis 45Fig. 3.1: The General Process of Synthesis3. Use the newly found constructions of the components in the current pool, to update thequantitative properties of the components in the current pool, and also (via the currentpool) the quantitative properties of the components in the initial pool4. Out of the components found in Step 2, select the ones that seem most promising accordingto the goal function, and add these to the current pool (or else simply define these as thecurrent pool)5. Return to Step 2More formally, synthesis may be specified as follows. Let X denote the set of combinators,and let Y 0 denote the initial pool of components (the initial focus of the cognitive process).Given Y i , let Z i denote the setReduce(Join(C i (1), ..., C i (r)))where the C i are drawn from Y i or from X. We may then sayY i+1 = F ilter(Z i )where F ilter is a function that selects a subset of its arguments.Analysis, on the other hand, begins with a set W of components, and a set X of combinators,and tries to find a series Y i so that according to the process of synthesis, Y n =W .In practice, of course, the implementation of a synthesis process need not involve the explicitconstruction of the full set Z i . Rather, the filtering operation takes place implicitly during theconstruction of Y i+1 . The result, however, is that one gets some subset of the compounds produciblevia joining and reduction from the set of components present in Y i plus the combinatorsX.46 3 A Patternist Philosophy of MindFig. 3.2: The General Process of AnalysisConceptually one may view synthesis as a very generic sort of “growth process,” and analysisas a very generic sort of “figuring out how to grow something.” The intuitive idea underlyingthe present proposal is that these forward-going and backward-going “growth processes” areamong the essential foundations of cognitive control, and that a conceptually sound design forcognitive control should explicitly make use of this fact. To abstract away from the details,what these processes are about is:• taking the general dynamic of compound-formation and reduction as outlined in Kampisand Chaotic Logic• introducing goal-directed pruning (“filtering”) into this dynamic so as to account for thelimitations of computational resources that are a necessary part of pragmatic intelligence3.4.3 The Dynamic of Iterative Analysis and SynthesisWhile synthesis and analysis are both very useful on their own, they achieve their greatest powerwhen harnessed together. It is my hypothesis that the dynamic pattern of alternating synthesisand analysis has a fundamental role in cognition. Put simply, synthesis creates new mentalforms by combining existing ones. Then, analysis seeks simple explanations for the forms in themind, including the newly created ones; and, this explanation itself then comprises additionalnew forms in the mind, to be used as fodder for the next round of synthesis. Or, to put it yetmore simply:3.4 The General Structure of Cognitive Dynamics: Analysis and Synthesis 47⇒ Combine ⇒ Explain ⇒ Combine ⇒ Explain ⇒ Combine ⇒It is not hard to express this alternating dynamic more formally, as well.• Let X denote any set of components.• Let F(X) denote a set of components which is the result of synthesis on X.• Let B(X) denote a set of components which is the result of analysis of X. We assume alsoa heuristic biasing the synthesis process toward simple constructs.• Let S(t) denote a set of components at time t, representing part of a system’s knowledgebase.• Let I(t) denote components resulting from the external environment at time t.Then, we may consider a dynamical iteration of the formS(t + 1) = B(F (S(t) + I(t)))This expresses the notion of alternating synthesis and analysis formally, as a dynamicaliteration on the space of sets of components. We may then speak about attractors of thisiteration: fixed points, limit cycles and strange attractors. One of the key hypotheses I wishto put forward here is that some key emergent cognitive structures are strange attractors ofthis equation. The iterative dynamic of combination and explanation leads to the emergenceof certain complex structures that are, in essence, maintained when one recombines their partsand then seeks to explain the recombinations. These structures are built in the first placethrough iterative recombination and explanation, and then survive in the mind because theyare conserved by this process. They then ongoingly guide the construction and destruction ofvarious other temporary mental structures that are not so conserved.3.4.4 Self and Focused Attention as Approximate Attractors of theDynamic of Iterated Forward-AnalysisAs noted above, patternist philosophy argues that two key aspects of intelligence are emergentstructures that may be called the “self” and the “attentional focus.” These, it is suggested, areaspects of intelligence that may not effectively be wired into the infrastructure of an intelligentsystem, though of course the infrastructure may be configured in such a way as to encouragetheir emergence. Rather, these aspects, by their nature, are only likely to be effective if theyemerge from the cooperative activity of various cognitive processes acting within a broad baseof knowledge.Above we have described the pattern of ongoing habitual oscillation between synthesis andanalysis as a kind of “dynamical iteration.” Here we will argue that both self and attentionalfocus may be viewed as strange attractors of this iteration. The mode of argument is relativelyinformal. The essential processes under consideration are ones that are poorly understood froman empirical perspective, due to the extreme difficulty involved in studying them experimentally.For understanding self and attentional focus, we are stuck in large part with introspection, whichis famously unreliable in some contexts, yet still dramatically better than having no informationat all. So, the philosophical perspective on self and attentional focus given here is a synthesis ofempirical and introspective notions, drawn largely from the published thinking and research of48 3 A Patternist Philosophy of Mindothers but with a few original twists. From a CogPrime perspective, its use has been to guidethe design process, to provide a grounding for what otherwise would have been fairly arbitrarychoices.3.4.4.1 SelfAnother high-level intelligent system pattern mentioned above is the “self”, which we here will tiein with analysis and synthesis processes. The term “self” as used here refers to the “phenomenalself” [Met04] or “self-model”. That is, the self is the model that a system builds internally,reflecting the patterns observed in the (external and internal) world that directly pertain tothe system itself. As is well known in everyday human life, self-models need not be completelyaccurate to be useful; and in the presence of certain psychological factors, a more accurateself-model may not necessarily be advantageous. But a self-model that is too badly inaccuratewill lead to a badly-functioning system that is unable to effectively act toward the achievementof its own goals.The value of a self-model for any intelligent system carrying out embodied agentive cognitionis obvious. And beyond this, another primary use of the self is as a foundation for metaphorsand analogies in various domains. Patterns recognized pertaining to the self are analogicallyextended to other entities. In some cases this leads to conceptual pathologies, such as the anthropomorphizationof trees, rocks and other such objects that one sees in some precivilizedcultures. But in other cases this kind of analogy leads to robust sorts of reasoning - for instance,in reading Lakoff and Nunez’s [LN00] intriguing explorations of the cognitive foundations ofmathematics, it is pretty easy to see that most of the metaphors on which they hypothesizemathematics to be based, are grounded in the mind’s conceptualization of itself as a spatiotemporallyembedded entity, which in turn is predicated on the mind’s having a conceptualizationof itself (a self) in the first place.A self-model can in many cases form a self-fulfilling prophecy (to make an obvious doubleentendre’!).Actions are generated based on one’s model of what sorts of actions one can and/orshould take; and the results of these actions are then incorporated into one’s self-model. If aself-model proves a generally bad guide to action selection, this may never be discovered, unlesssaid self-model includes the knowledge that semi-random experimentation is often useful.In what sense, then, may it be said that self is an attractor of iterated analysis? Analysisinfers the self from observations of system behavior. The system asks: What kind of systemmight I be, in order to give rise to these behaviors that I observe myself carrying out? Basedon asking itself this question, it constructs a model of itself, i.e. it constructs a self. Then, thisself guides the system’s behavior: it builds new logical relationships its self-model and variousother entities, in order to guide its future actions oriented toward achieving its goals. Based onthe behaviors newly induced via this constructive, forward-synthesis activity, the system maythen engage in analysis again and ask: What must I be now, in order to have carried out thesenew actions? And so on.Our hypothesis is that after repeated iterations of this sort, in infancy, finally during earlychildhood a kind of self-reinforcing attractor occurs, and we have a self-model that is resilientand doesn’t change dramatically when new instances of action- or explanation-generation occur.This is not strictly a mathematical attractor, though, because over a long period of time the selfmay well shift significantly. But, for a mature self, many hundreds of thousands or millions offorward-analysis cycles may occur before the self-model is dramatically modified. For relatively3.4 The General Structure of Cognitive Dynamics: Analysis and Synthesis 49long periods of time, small changes within the context of the existing self may suffice to allowthe system to control itself intelligently.Humans can also develop what are known as subselves [Row90]. A subself is a partiallyautonomous self-network focused on particular tasks, environments or interactions. It containsa unique model of the whole organism, and generally has its own set of episodic memories,consisting of memories of those intervals during which it was the primary dynamic mode controllingthe organism. One common example is the creative subself – the subpersonality thattakes over when a creative person launches into the process of creating something. In thesetimes, a whole different personality sometimes emerges, with a different sort of relationshipto the world. Among other factors, creativity requires a certain open-ness that is not alwaysproductive in an everyday life context, so it’s natural for the self-system of a highly creativeperson to bifurcate into one self-system for everyday life, and another for the protected contextof creative activity. This sort of phenomenon might emerge naturally in CogPrime systems aswell if they were exposed to appropriate environments and social situations.Finally, it is interesting to speculate regarding how self may differ in future AI systems asopposed to in humans. The relative stability we see in human selves may not exist in AI systemsthat can self-improve and change more fundamentally and rapidly than humans can. There maybe a situation in which, as soon as a system has understood itself decently, it radically modifiesitself and hence violates its existing self-model. Thus: intelligence without a long-term stable self.In this case the “attractor-ish” nature of the self holds only over much shorter time scales thanfor human minds or human-like minds. But the alternating process of synthesis and analysisfor self-construction is still critical, even though no reasonably stable self-constituting attractorever emerges. The psychology of such intelligent systems will almost surely be beyond humanbeings’ capacity for comprehension and empathy.3.4.4.2 Attentional FocusFinally, we turn to the notion of an “attentional focus” similar to Baars’ [Baa97] notion of aGlobal Workspace, which will be reviewed in more detail in Chapter 4: a collection of mentalentities that are, at a given moment, receiving far more than the usual share of an intelligentsystem’s computational resources. Due to the amount of attention paid to items in the attentionalfocus, at any given moment these items are in large part driving the cognitive processesgoing on elsewhere in the mind as well - because the cognitive processes acting on the items inthe attentional focus are often involved in other mental items, not in attentional focus, as well(and sometimes this results in pulling these other items into attentional focus). An intelligentsystem must constantly shift its attentional focus from one set of entities to another based onchanges in its environment and based on its own shifting discoveries.In the human mind, there is a self-reinforcing dynamic pertaining to the collection of entitiesin the attentional focus at any given point in time, resulting from the observation that: If Ais in the attentional focus, and A and B have often been associated in the past, then oddsare increased that B will soon be in the attentional focus. This basic observation has beenrefined tremendously via a large body of cognitive psychology work; and neurologically it followsnot only from Hebb’s [Heb49] classic work on neural reinforcement learning, but also fromnumerous more modern refinements [SB98]. But it implies that two items A and B, if both inthe attentional focus, can reinforce each others’ presence in the attentional focus, hence forminga kind of conspiracy to keep each other in the limelight. But of course, this kind of dynamic50 3 A Patternist Philosophy of Mindmust be counteracted by a pragmatic tendency to remove items from the attentional focus ifgiving them attention is not providing sufficient utility in terms of the achievement of systemgoals.The synthesis and analysis perspective provides a more systematic perspective on this selfreinforcingdynamic. Synthesis occurs in the attentional focus when two or more items in thefocus are combined to form new items, new relationships, new ideas. This happens continually,as one of the main purposes of the attentional focus is combinational. On the other hand,Analysis then occurs when a combination that has been speculatively formed is then linkedin with the remainder of the mind (the “unconscious”, the vast body of knowledge that is notin the attentional focus at the given moment in time). Analysis basically checks to see whatsupport the new combination has within the existing knowledge store of the system. Thus,forward-analysis basically comes down to “generate and test”, where the testing takes the formof attempting to integrate the generated structures with the ideas in the unconscious longtermmemory. One of the most obvious examples of this kind of dynamic is creative thinking(Boden, 2003; Goertzel, 1997), where the attentional focus continually combinationally createsnew ideas, which are then tested via checking which ones can be validated in terms of (built upfrom) existing knowledge.The analysis stage may result in items being pushed out of the attentional focus, to bereplaced by others. Likewise may the synthesis stage: the combinations may overshadow andthen replace the things combined. However, in human minds and functional AI minds, theattentional focus will not be a complete chaos with constant turnover: Sometimes the same set ofideas – or a shifting set of ideas within the same overall family of ideas – will remain in focus for awhile. When this occurs it is because this set or family of ideas forms an approximate attractorfor the dynamics of the attentional focus, in particular for the forward-analysis dynamic ofspeculative combination and integrative explanation. Often, for instance, a small “core set” ofideas will remain in the attentional focus for a while, but will not exhaust the attentional focus:the rest of the attentional focus will then, at any point in time, be occupied with other ideasrelated to the ones in the core set. Often this may mean that, for a while, the whole of theattentional focus will move around quasi-randomly through a “strange attractor” consisting ofthe set of ideas related to those in the core set.3.4.5 ConclusionThe ideas presented above (the notions of synthesis and analysis, and the hypothesis of self andattentional focus as attractors of the iterative forward-analysis dynamic) are quite generic andare hypothetically proposed to be applicable to any cognitive system, natural or artificial. Laterchapters will discuss the manifestation of the above ideas in the context of CogPrime. We havefound that the analysis/synthesis approach is a valuable tool for conceptualizing CogPrime’scognitive dynamics, and we conjecture that a similar utility may be found more generally.Next, so as not to end the section on too blasé of a note, we will also make a strongerhypothesis: that, in order for a physical or software system to achieve intelligence that is roughlyhuman-level in both capability and generality, using computational resources on the same orderof magnitude as the human brain, this system must• manifest the dynamic of iterated synthesis and analysis, as modes of an underlying “selfgeneratingsystem” dynamic3.5 Perspectives on Machine Consciousness 51• do so in such a way as to lead to self and attentional focus as emergent structures that serveas approximate attractors of this dynamic, over time periods that are long relative to thebasic “cognitive cycle time” of the system’s forward-analysis dynamicsTo prove the truth of a hypothesis of this nature would seem to require mathematics fairlyfar beyond anything that currently exists. Nonetheless, however, we feel it is important toformulate and discuss such hypotheses, so as to point the way for future investigations boththeoretical and pragmatic.3.5 Perspectives on Machine ConsciousnessFinally, we can’t let a chapter on philosophy – even a brief one – end without some discussionof the thorniest topic in the philosophy of mind: consciousness. Rather than seeking to resolveor comprehensively review this most delicate issue, we will restrict ourselves to giving it inAppendix ?? an overview of many of the common views on the subject; and here in the main textdiscussing the relationship between consciousness theory and patternist philosophy of cognition,the practical work of designing and building AGI.One fairly concrete idea about consciousness, that relates closely to certain aspects of theCogPrime design, is that the subjective experience of being conscious of some entity X, is correlatedwith the presence of a very intense pattern in one’s overall mind-state, corresponding to X.This simple idea is also the essence of neuroscientist Susan Greenfield’s theory of consciousness[Gre01] (but in her theory, "overall mind-state" is replaced with "brain-state"), and has muchdeeper historical roots in philosophy of mind which we shall not venture to unravel here.This observation relates to the idea of "moving bubbles of awareness" in intelligent systems.If an intelligent system consists of multiple processing or data elements, and during each (sufficientlylong) interval of time some of these elements get much more attention than others,then one may view the system as having a certain "attentional focus" during each interval. Theattentional focus is itself a significant pattern in the system (the pattern being "these elementshabitually get more processor and memory", roughly speaking). As the attentional focus shiftsover time one has a "moving bubble of pattern" which then corresponds experientially to a"moving bubble of awareness."This notion of a "moving bubble of awareness" ties in very closely to global workspacetheory [Baa97] (briefly mentioned above), a cognitive theory that has broad support fromneuroscience and cognitive science and has also served as the motivation for Stan Franklin’sLIDA AI system [BF09], to be discussed in Chapter ??. The global workspace theory views themind as consisting of a large population of small, specialized processes – a society of agents.These agents organize themselves into coalitions, and coalitions that are relevant to contextuallynovel phenomena, or contextually important goals, are pulled into the global workspace (whichis identified with consciousness). This workspace broadcasts the message of the coalition to allthe unconscious agents, and recruits other agents into consciousness. Various sorts of contexts– e.g. goal contexts, perceptual contexts, conceptual contexts and cultural contexts – play arole in determining which coalitions are relevant, and form the unconscious "background" ofthe conscious global workspace. New perceptions are often, but not necessarily, pushed into theworkspace. Some of the agents in the global workspace are concerned with action selection, i.e.with controlling and passing parameters to a population of possible actions. The contents ofthe workspace at any given time have a certain cohesiveness and interdependency, the so-called52 3 A Patternist Philosophy of Mind"unity of consciousness." In essence the contents of the global workspace form a moving bubbleof attention or awareness.In CogPrime, this moving bubble is achieved largely via economic attention network (ECAN)equations [GPI + 10] that propagate virtual currency between nodes and links representing elementsof memories, so that the attentional focus consists of the wealthiest nodes and links.Figures 3.3 and 3.4 illustrate the existence and flow of attentional focus in OpenCog. On theother hand, in Hameroff’s recent model of the brain [Ham10], the brain’s moving bubble ofattention is achieved through dendro-dendritic connections and the emergent dendritic web.Fig. 3.3: Graphical depiction of the momentary bubble of attention in the memory of anOpenCog AI system. Circles and lines represent nodes and links in OpenCogPrimes memory,and stars denote those nodes with a high level of attention (represented in OpenCog bythe ShortTermImportance node variable) at the particular point in time.In this perspective, self, free will and reflective consciousness are specific phenomena occurringwithin the moving bubble of awareness. They are specific ways of experiencing awareness,corresponding to certain abstract types of physical structures and dynamics, which we shallendeavor to identify in detail in Appendix ??.3.6 Postscript: Formalizing Pattern 53Fig. 3.4: Graphical depiction of the momentary bubble of attention in the memory of anOpenCog AI system, a few moments after the bubble shown in Figure 3.3, indicating the movingof the bubble of attention. Depictive conventions are the same as in Figure 1. This showsan idealized situation where the declarative knowledge remains invariant from one moment tothe next but only the focus of attention shifts. In reality both will evolve together.3.6 Postscript: Formalizing PatternFinally, before winding up our very brief tour through patternist philosophy of mind, we willbriefly visit patternism’s more formal side. Many of the key aspects of patternism have beenrigorously formalized. Here we give only a few very basic elements of the relevant mathematics,which will be used later on in the exposition of CogPrime. (Specifically, the formal definition ofpattern emerges in the CogPrime design in the definition of a fitness function for “pattern mining”algorithms and Occam-based concept creation algorithms, and the definition of intensionalinheritance within PLN.)We give some definitions, drawn from Appendix 1 of [Goe06a]:Definition 1 Given a metric space (M, d), and two functions c : M → [0, ∞] (the “simplicitymeasure”) and F : M → M (the “production relationship”), we say that P ∈ M is a patternin X ∈ M to the degree54 3 A Patternist Philosophy of Mindι P X =((1 −d(F (P), X)c(X)) c(X) − c(P)c(X)This degree is called the pattern intensity of P in X. It quantifies the extent to which Pis a pattern in X. Supposing that F (P) = X, then the first factor in the definition equals 1,and we are left with only the second term, which measures the degree of compression obtainedvia representing X as the result of P rather than simply representing X directly. The greaterthe compression ratio obtained via using P to represent X, the greater the intensity of P as apattern in X. The first time, in the case F (P) ≠ X, adjusts the pattern intensity downwards toaccount for the amount of error with which F (P) approximates ≠ X. If one holds the secondfactor fixed and thinks about varying the first factor, then: The greater the error, the lossierthe compression, and the lower the pattern intensity.For instance, if one wishes one may take c to denote algorithmic information measured onsome reference Turing machine, and F (X) to denote what appears on the second tape of atwo-tape Turing machine t time-steps after placing X on its first tape. Other more naturalisticcomputational models are also possible here and are discussed extensively in Appendix 1 of[Goe06a].) +Definition 2 The structure of X ∈ M is the fuzzy set St Xfunctionχ StX (P) = ι P Xdefined via the membershipThis lets us formalize our definition of “mind” alluded to above: the mind of X as the setof patterns associated with X. We can formalize this, for instance, by considering P to belongto the mind of X if it is a pattern in some Y that includes X. There are then two numbersto look at: ι P X and P (Y |X) (the percentage of Y that is also contained in X). To define thedegree to which P belongs to the mind of X we can then combine these two numbers usingsome function f that is monotone increasing in both arguments. This highlights the somewhatarbitrary semantics of “of” in the phrase “the mind of X.” Which of the patterns binding X toits environment are part of X’s mind, and which are part of the world? This isn’t necessarilya good question, and the answer seems to depend on what perspective you choose, representedformally in the present framework by what combination function f you choose (for instance iff(a, b) = a r b 2−r then it depends on the choice of 0 < r < 1).Next, we can formalize the notion of a “pattern space” by positing a metric on patterns, thusmaking pattern space a metric space, which will come in handy in some places in later chapters:Definition 3 Assuming M is a countable space, the structural distance is a metric d Stdefined on M viad St (X, Y ) = T (χ StX , χ StY )where T is the Tanimoto distance.The Tanimoto distance between two real vectors A and B is defined asT (A, B) =A · B‖A‖ 2 + ‖B‖ 2 − A · Band since M is countable this can be applied to fuzzy sets such as St X via considering thelatter as vectors. (As an aside, this can be generalized to uncountable M as well, but we willnot require this here.)3.6 Postscript: Formalizing Pattern 55Using this definition of pattern, combined with the formal theory of intelligence given inChapter 7, one may formalize the various hypotheses made in the previous section, regardingthe emergence of different kinds of networks and structures as patterns in intelligent systems.However, it appears quite difficult to prove the formal versions of these hypotheses given currentmathematical tools, which renders such formalizations of limited use.Finally, consider the case where the metric space M has a partial ordering < on it; we maythen defineDefinition 3.1. R ∈ M is a subpattern in X ∈ M to the degree∫κ R P∈MX =true(R < P )dιP X∫P∈M dιP XThis degree is called the subpattern intensity of P in X.Roughly speaking, the subpattern intensity measures the percentage of patterns in X thatcontain R (where "containment" is judged by the partial ordering <). But the percentage ismeasured using a weighted average, where each pattern is weighted by its intensity as a patternin X. A subpattern may or may not be a pattern on its own. A nonpattern that happens tooccur within many patterns may be an intense subpattern.Whether the subpatterns in X are to be considered part of the "mind" of X is a somewhatsuperfluous question of semantics. Here we choose to extend the definition of mind given in[Goe06a] to include subpatterns as well as patterns, because this makes it simpler to describethe relationship between hypersets and minds, as we will do in Appendix ??.
Chapter 4Brief Survey of Cognitive Architectures4.1 IntroductionWhile we believe CogPrime is the most thorough attempt at an architecture for advanced AGI,to date, we certainly recognize there have been many valuable attempts in the past with similaraims; and we also have great respect for other AGI efforts occurring in parallel with Cog-Prime development, based on alternative, sometimes overlapping, theoretical presuppositionsand practical choices. In most of this book we will ignore these other current and historicalefforts except where they are directly useful for CogPrime – there are many literature reviewsalready published, and this is a research treatise not a textbook. In this chapter, however, wewill break from this pattern and give a rough high-level overview of the various AGI architecturesat play in the field today. The overview definitely has a bias toward other work withsome direct relevance to CogPrime, but not an overwhelming bias; we also discuss a number ofapproaches that are unrelated to, and even in some cases conceptually orthogonal to, our own.CogPrime builds on prior AI efforts in a variety of ways. Most of the specific algorithmsand structures in CogPrime have their roots in prior AI work; and in addition, the CogPrimecognitive architecture has been heavily inspired by some other holistic cognitive architectures,especially (but not exclusively) MicroPsi [Bac09], LIDA [BF09] and DeSTIN [ARK09a, ARC09].In this chapter we will briefly review some existing cognitive architectures, with especial butnot exclusive emphasis on the latter three.We will articulate some rough mappings between elements of these other architectures andelements of CogPrime – some in this chapter, and some in Chapter 5. However, these mappingswill mostly be left informal and very incompletely specified. The articulation of detailed interarchitecturemappings is an important project, but would be a substantial additional projectgoing well beyond the scope of this book. We will not give a thorough review of the similaritiesand differences between CogPrime and each of these architectures, but only mention some ofthe highlights.The reader desiring a more thorough review of cognitive architectures is referred to WlodekDuch’s review paper from the AGI-08 conference [DOP08]; and also to Alexei Samsonovich’sreview paper [Sam10], which compares a number of cognitive architectures in terms of a featurechecklist, and was created collaboratively with the creators of the architectures.Duch, in his survey of cognitive architectures [DOP08], divides existing approaches into threeparadigms – symbolic, emergentist and hybrid – as broadly indicated in Figure 4.1. Drawing onhis survey and updating slightly, we give here some key examples of each, and then explain why5758 4 Brief Survey of Cognitive ArchitecturesCogPrime represents a significantly more effective approach to embodied human-like generalintelligence. In our treatment of emergentist architectures, we pay particular attention to developmentalrobotics architectures, which share considerably with CogPrime in terms of underlyingphilosophy, but differ via not integrating a symbolic “language and inference” component suchas CogPrime includes.In brief, we believe that the hybrid approach is the most pragmatic one given the current stateof AI technology, but that the emergentist approach gets something fundamentally right, byfocusing on the emergence of complex dynamics and structures from the interactions of simplecomponents. So CogPrime is a hybrid architecture which (according to the cognitive synergyprinciple) binds its components together very tightly dynamically, allowing the emergence ofcomplex dynamics and structures in the integrated system. Most other hybrid architectures areless tightly coupled and hence seem ill-suited to give rise to the needed emergent complexity. Theother hybrid architectures that do possess the needed tight coupling, such as MicroPsi [Bac09],strike us as underdeveloped and founded on insufficiently powerful learning algorithms.Fig. 4.1: Duch’s simplified taxonomy of cognitive architectures. CogPrime falls into the “hybrid”category, but differs from other hybrid architectures in its focus on synergetic interactionsbetween components and their potential to give rise to appropriate system-wide emergent structuresenabling general intelligence.4.2 Symbolic Cognitive ArchitecturesA venerable tradition in AI focuses on the physical symbol system hypothesis [New90], whichstates that minds exist mainly to manipulate symbols that represent aspects of the world orthemselves. A physical symbol system has the ability to input, output, store and alter symbolicentities, and to execute appropriate actions in order to reach its goals. Generally, symboliccognitive architectures focus on “working memory” that draws on long-term memory as needed,and utilize a centralized control over perception, cognition and action. Although in principlesuch architectures could be arbitrarily capable (since symbolic systems have universal repre-4.2 Symbolic Cognitive Architectures 59sentational and computational power, in theory), in practice symbolic architectures tend to beweak in learning, creativity, procedure learning, and episodic and associative memory. Decadesof work in this tradition have not resolved these issues, which has led many researchers toexplore other options. A few of the more important symbolic cognitive architectures are:• SOAR [LRN87], a classic example of expert rule-based cognitive architecture designed tomodel general intelligence. It has recently been extended to handle sensorimotor functions,though in a somewhat cognitively unnatural way; and is not yet strong in areas such asepisodic memory, creativity, handling uncertain knowledge, and reinforcement learning.• ACT-R [AL03] is fundamentally a symbolic system, but Duch classifies it as a hybrid systembecause it incorporates connectionist-style activation spreading in a significant role; andthere is an experimental thoroughly connectionist implementation to complement the primarymainly-symbolic implementation. Its combination of SOAR-style “production rules”with large-scale connectionist dynamics allows it to simulate a variety of human psychologicalphenomena, but abstract reasoning, creativity and transfer learning are still missing.• EPIC [RCK01], a cognitive architecture aimed at capturing human perceptual, cognitiveand motor activities through several interconnected processors working in parallel. Thesystem is controlled by production rules for cognitive processors and a set of perceptual(visual, auditory, tactile) and motor processors operating on symbolically coded featuresrather than raw sensory data. It has been connected to SOAR for problem solving, planningand learning,• ICARUS [Lan05], an integrated cognitive architecture for physical agents, with knowledgespecified in the form of reactive skills, each denoting goal-relevant reactions to a class ofproblems. The architecture includes a number of modules: a perceptual system, a planningsystem, an execution system, and several memory systems. Concurrent processing is absent,attention allocation is fairly crude, and uncertain knowledge is not thoroughly handled.• SNePS (Semantic Network Processing System) [SE07] is a logic, frame and network-basedknowledge representation, reasoning, and acting system that has undergone over threedecades of development. While it has been used for some interesting prototype experimentsin language processing and virtual agent control, it has not yet been used for anylarge-scale or real-world application.• Cyc [LG90] is an AGI architecture based on predicate logic as a knowledge representation,and using logical reasoning techniques to answer questions and derive new knowledge fromold. It has been connected to a natural language engine, and designs have been createdfor the connection of Cyc with Albus’s 4D-RCS [AM01]. Cyc’s most unique aspect is thelarge database of commonsense knowledge that Cycorp has accumulated (millions of piecesof knowledge, entered by specially trained humans in predicate logic format); part of thephilosophy underlying Cyc is that once a sufficient quantity of knowledge is accumulated inthe knowledge base, the problem of creating human-level general intelligence will becomemuch less difficult due to the ability to leverage this knowledge.While these architectures contain many valuable ideas and have yielded some interesting results,we feel they are incapable on their own of giving rise to the emergent structures and dynamicsrequired to yield humanlike general intelligence using feasible computational resources. However,we are more sanguine about the possibility of ideas and components from symbolic architecturesplaying a role in human-level AGI via incorporation in hybrid architectures.We now review a few symbolic architectures in slightly more detail.60 4 Brief Survey of Cognitive Architectures4.2.1 SOARThe cognitive architectures best known among AI academics are probably Soar and ACT-R,both of which are explicitly being developed with the dual goals of creating human-level AGIand modeling all aspects of human psychology. Neither the Soar nor ACT-R communities feelthemselves particularly near these long-term goals, yet they do take them seriously.Soar is based on IF-THEN rules, otherwise known as “production rules.” On the surface thismakes it similar to old-style expert systems, but Soar is much more than an expert system; it’sat minimum a sophisticated problem-solving engine. Soar explicitly conceives problem solvingas a search through solution space for a “goal state” representing a (precise or approximate)problem solution. It uses a methodology of incremental search, where each step is supposed tomove the system a little closer to its problem-solving goal, and each step involves a potentiallycomplex “decision cycle.”In the simplest case, the decision cycle has two phases:• Gathering appropriate information from the system’s long-term memory (LTM) into itsworking memory (WM)• A decision procedure that uses the gathered information to decide an actionIf the knowledge available in LTM isn’t enough to solve the problem, then the decisionprocedure invokes search heuristics like hill-climbing, which try to create new knowledge (newproduction rules) that will help move the system closer to a solution. If a solution is found bychaining together multiple production rules, then a chunking mechanism is used to combinethese rules together into a single rule for future use. One could view the chunking mechanismas a way of converting explicit knowledge into implicit knowledge, similar to “map formation”in CogPrime (see Chapter 42 of Part 2), but in the current Soar design and implementation itis a fairly crude mechanism.In recent years Soar has acquired a number of additional methods and modalities, includingsome visual reasoning methods and some mechanisms for handling episodic and proceduralknowledge. These expand the scope of the system but the basic production rule and chunkingmechanisms as briefly described above remain the core “cognitive algorithm” of the system.From a CogPrime perspective, what Soar offers is certainly valuable, e.g.• heuristics for transferring knowledge from LTM into WM• chaining and chunking of implications• methods for interfacing between other forms of knowledge and implicationsHowever, a very short and very partial list of the major differences between Soar and Cog-Prime would include• CogPrime contains a variety of other core cognitive mechanisms beyond the managementand chunking of implications• the variety of “chunking” type methods in CogPrime goes far beyond the sort of localizedchunking done in Soar• CogPrime is committed to representing uncertainty at the base level whereas Soar’s productionrules are crisp• The mechanisms for LTM-WM interaction are rather different in CogPrime, being basedon complex nonlinear dynamics as represented in Economic Attention Allocation (ECAN)• Currently Soar does not contain creativity-focused heuristics like blending or evolutionarylearning in its core cognitive dynamic.4.2 Symbolic Cognitive Architectures 614.2.2 ACT-RIn the grand scope of cognitive architectures, ACT-R is quite similar to Soar, but there aremany micro-level differences. ACT-R is defined in terms of declarative and procedural knowledge,where procedural knowledge takes the form of Soar-like production rules, and declarativeknowledge takes the form of chunks. It contains a variety of mechanisms for learning new rulesand chunks from old; and also contains sophisticated probabilistic equations for updating theactivation levels associated with items of knowledge (these equations being roughly analogousin function to, though quite different from, the ECAN equations in CogPrime).Figure 4.2 displays the current architecture of ACT-R. The flow of cognition in the system isin response to the current goal, currently active information from declarative memory, informationattended to in perceptual modules (vision and audition are implemented), and the currentstate of motor modules (hand and speech are implemented). The early work with ACT-R wasbased on comparing system performance to human behavior, using only behavioral measures,such as the timing of keystrokes or patterns of eye movements. Using such measures, it was notpossible to test detailed assumptions about which modules were active in the performance ofa task. More recently the ACT-R community has been engaged in a process of using imagingdata to provide converging data on module activity. Figure 4.3 illustrates the associations theyhave made between the modules in Figure 4.2 and brain regions. Coordination among all ofthese components occurs through actions of the procedural module, which is mapped to thebasal ganglia.Fig. 4.2: High-level architecture of ACT-RIn practice ACT-R, even more so than Soar, seems to be used more as a programmingframework for cognitive modeling than as an AI system. One can fairly easily use ACT-Rto program models of specific human mental behaviors, which may then be matched against62 4 Brief Survey of Cognitive ArchitecturesFig. 4.3: Conjectured Mapping Between ACT-R and the Brainpsychological data. Opinions differ as to whether this sort of modeling is valuable for achievingAGI goals. CogPrime is not designed to support this kind of modeling, as it intentionally doesmany things very differently from humans.ACT-R in its original form did not say much about perceptual and motor operations, butrecent versions have incorporated EPIC, an independent cognitive architecture focused on modelingthese aspects of human behavior.4.2.3 Cyc and TexaiOur review of cognitive architectures would be incomplete without mentioning Cyc [LG90],one of the best known and best funded AGI-oriented projects in history. While the main focusof the Cyc project has been on the hand-coding of large amounts of declarative knowledge,there is also a cognitive architecture of sorts there. The center of Cyc is an engine for logicaldeduction, acting on knowledge represented in predicate logic. A natural language engine hasbeen associated with the logic engine, which enables one to ask English questions and getEnglish replies.Stephen Reed, while an engineer at Cycorp, designed a perceptual-motor front end for Cycbased on James Albus’ Reference Model Architecture; the ensuing system, called Cognitive-Cyc, would have been the first full-fledged cognitive architecture based on Cyc, but was notimplemented. Reed left Cycorp and is now building a system called Texai, which has manysimilarities to Cyc (and relies upon the OpenCyc knowledge base, a subset of Cyc’s overallknowledge base), but incorporates a CognitiveCyc style cognitive architecture.4.2 Symbolic Cognitive Architectures 634.2.4 NARSPei Wang’s NARS logic [Wan06] played a large role in the development of PLN, CogPrime’suncertain logic component, a relationship that is discussed in depth in [GMIH08] and won’tbe re-emphasized here. However, NARS is more than just an uncertain logic, it is also anoverall cognitive architecture (which is centered on NARS logic, but also includes other aspects).CogPrime bears little relation to NARS except in the specific similarities between PLN logicand NARS logic, but, the other aspects of NARS are worth briefly recounting here.NARS is formulated as a system for processing tasks, where a task consists of a question or apiece of new knowledge. The architecture is focused on declarative knowledge, but some piecesof knowledge may be associated with executable procedures, which allows NARS to carry outcontrol activities (in roughly the same way that a Prolog program can).At any given time a NARS system contains• working memory: a small set of tasks which are active, kept for a short time, and closelyrelated to new questions and new knowledge• long-term memory: a huge set of knowledge which is passive, kept for a long time, and notnecessarily related to current questions and knowledgeThe working and long term memory spaces of NARS may each be thought of as a set ofchunks, where each chunk consists of a set of tasks and a set of knowledge. NARS’s basiccognitive process is:1. choose a chunk2. choose a task from that chunk3. choose a piece of knowledge from that chunk4. use the task and knowledge to do inference5. send the new tasks to corresponding chunksDepending on the nature of the task and knowledge, the inference involved may be one ofthe following:• if the task is a question, and the knowledge happens to be an answer to the question, acopy of the knowledge is generated as a new task• backward inference• revision (merging two pieces of knowledge with the same form but different truth value)• forward inference• execution of a procedure associated with a piece of knowledgeUnlike many other systems, NARS doesn’t decide what type of inference is used to processa task when the task is accepted, but works in a data-driven way – that is, it is the task andknowledge that dynamically determine what type of inference will be carried outThe “choice” processes mentioned above are done via assigning relative priorities to• chunks (where they are called activity)• tasks (where they are called urgency)• knowledge (where they are called importance)64 4 Brief Survey of Cognitive Architecturesand then distributing the system’s resources accordingly, based on a probabilistic algorithm.(It’s interesting to note that while NARS uses probability theory as part of its control mechanism,the logic it uses to represent its own knowledge about the world is nonprobabilistic. Thisis considered conceptually consistent, in the context of NARS theory, because system controlis viewed as a domain where the system’s knowledge is more complete, thus more amenable toprobabilistic reasoning.)4.2.5 GLAIR and SNePSAnother logic-focused cognitive architecture, very different from NARS in detail, is StuartShapiro’s GLAIR cognitive architecture, which is centered on the SNePS paraconsistent logic[SE07].Like NARS, the core “cognitive loop” of GLAIR is based on reasoning: either thinking aboutsome percept (e.g. linguistic input, or sense data from the virtual or physical world), or answeringsome question. This inference based cognition process is turned into an intelligent agentcontrol process via coupling it with an acting component, which operates according to a set ofpolicies, each one of which tells the system when to take certain internal or external actions(including internal reasoning actions) in response to its observed internal and external situation.GLAIR contains multiple layers:• the Knowledge Layer (KL), which contains the beliefs of the agent, and is where reasoning,planning, and act selection are performed• the Sensori-Actuator Layer (SAL), contains the controllers of the sensors and effectors ofthe hardware or software robot.• the Perceptuo-Motor Layer (PML), which grounds the KL symbols in perceptual structuresand subconscious actions, contains various registers for providing the agent’s sense of situatednessin the environment, and handles translation and communication between the KLand the SAL.The logical Knowledge Layer incorporates multiple memory types using a common representation(including declarative, procedural, episodic, attentional and intentional knowledge, andmeta-knowledge). To support this broad range of knowledge types, a broad range of logical inferencemechanisms are used, so that the KL may be variously viewed as predicate logic based,frame based, semantic network based, or from other perspectives.What makes GLAIR more robust than most logic based AI approaches is the novel paraconsistentlogical formalism used in the knowledge base, which means (among other things)that uncertain, speculative or erroneous knowledge may exist in the system’s memory withoutleading the system to create a broadly erroneous view of the world or carry out egregiouslyunintelligent actions. CogPrime is not thoroughly logic-focused like GLAIR is, but in its logicalaspect it seeks a similar robustness through its use of PLN logic, which embodies propertiesrelated to paraconsistency.Compared to CogPrime, we see that GLAIR has a similarly integrative approach, but thatthe integration of different sorts of cognition is done more strictly within the framework oflogical knowledge representation.4.3 Emergentist Cognitive Architectures 654.3 Emergentist Cognitive ArchitecturesAnother species of cognitive architecture expects abstract symbolic processing to emerge fromlower-level “subsymbolic” dynamics, which sometimes (but not always) are designed to simulateneural networks or other aspects of human brain function. These architectures are typicallystrong at recognizing patterns in high-dimensional data, reinforcement learning and associativememory; but no one has yet shown how to achieve high-level functions such as abstract reasoningor complex language processing using a purely subsymbolic approach. A few of the moreimportant subsymbolic, emergentist cognitive architectures are:• DeSTIN [ARK09a, ARC09], which is part of CogPrime, may also be considered as anautonomous AGI architecture, in which case it is emergentist and contains mechanismsto encourage language, high-level reasoning and other abstract aspects of intelligent toemerge from hierarchical pattern recognition and related self-organizing network dynamics.In CogPrime DeSTIN is used as part of a hybrid architecture, which greatly reduces thereliance on DeSTIN’s emergent properties.• Hierarchical Temporal Memory (HTM) [HB06] is a hierarchical temporal patternrecognition architecture, presented as both an AI approach and a model of the cortex. Sofar it has been used exclusively for vision processing and we will discuss its shortcomingslater in the context of our treatment of DeSTIN.• SAL [JL08], based on the earlier and related IBCA (Integrated Biologically-based CognitiveArchitecture) is a large-scale emergent architecture that seeks to model distributedinformation processing in the brain, especially the posterior and frontal cortex and thehippocampus. So far the architectures in this lineage have been used to simulate varioushuman psychological and psycholinguistic behaviors, but haven’t been shown to give rise tohigher-level behaviors like reasoning or subgoaling.• NOMAD (Neurally Organized Mobile Adaptive Device) automata and its successors[KE06] are based on Edelman’s “Neural Darwinism” model of the brain, and feature largenumbers of simulated neurons evolving by natural selection into configurations that carryout sensorimotor and categorization tasks. The emergence of higher-level cognition fromthis approach seems rather unlikely.• Ben Kuipers and his colleagues [MK07, MK08, MK09]have pursued an extremely innovativeresearch program which combines qualitative reasoning and reinforcement learning to enablean intelligent agent to learn how to act, perceive and model the world. Kuipers’ notion of“bootstrap learning” involves allowing the robot to learn almost everything about its world,including for instance the structure of 3D space and other things that humans and otheranimals obtain via their genetic endowments. Compared to Kuipers’ approach, CogPrimefalls in line with most other approaches which provide more “hard-wired” structure, followingthe analogy to biological organisms that are born with more innate biases.There is also a set of emergentist architectures focused specifically on developmental robotics,which we will review below in a separate subsection, as all of these share certain commoncharacteristics.Our general perspective on the emergentist approach is that it is philosophically correctbut currently pragmatically inadequate. Eventually, some emergentist approach could surelysucceed at giving rise to humanlike general intelligence – the human brain, after all, is plainlyan emergentist system. However, we currently lack understanding of how the brain gives riseto abstract reasoning and complex language, and none of the existing emergentist systems66 4 Brief Survey of Cognitive Architecturesseem remotely capable of giving rise to such phenomena. It seems to us that the creation ofa successful emergentist AGI will have to wait for either a detailed understanding of how thebrain gives rise to abstract thought, or a much more thorough mathematical understanding ofthe dynamics of complex self-organizing systems.The concept of cognitive synergy is more relevant to emergentist than to symbolic architectures.In a complex emergentist architecture with multiple specialized components, much ofthe emergence is expected to arise via synergy between different richly interacting components.Symbolic systems, at least in the forms currently seen in the literature, seem less likely to giverise to cognitive synergy as their dynamics tend to be simpler. And hybrid systems, as we shallsee, are somewhat diverse in this regard: some rely heavily on cognitive synergies and othersconsist of more loosely coupled components.We now review the DeSTIN emergentist architecture in more detail, and then turn to thedevelopmental robotics architectures.4.3.1 DeSTIN: A Deep Reinforcement Learning Approach to AGIThe DeSTIN architecture, created by Itamar Arel and his colleagues, addresses the problemof general intelligence using hierarchical spatiotemporal networks designed to enable scalableperception, state inference and reinforcement-learning-guided action in real-world environments.DeSTIN has been developed with the plan of gradually extending it into a complete system forhumanoid robot control, founded on the same qualitative information-processing principles asthe human brain (though without striving for detailed biological realism). However, the practicalwork with DeSTIN to date has focused on visual and auditory processing; and in the context ofthe present proposal, the intention is to utilize DeSTIN for perception and actuation orientedprocessing, hybridizing it with CogPrime which will handle abstract cognition and language.Here we will discuss DeSTIN primarily in the perception context, only briefly mentioning theapplication to actuation which is conceptually similar.In DeSTIN (see Figure 4.4), perception is carried out by a deep spatiotemporal inferencenetwork, which is connected to a similarly architected critic network that provides feedback onthe inference network’s performance, and an action network that controls actuators based on theactivity in the inference network (Figure 4.5 depicts a standard action hierarchy, of which thehierarchy in DeSTIN is an example). The nodes in these networks perform probabilistic patternrecognition according to algorithms to be described below; and the nodes in each of the networksmay receive states of nodes in the other networks as inputs, providing rich interconnectivityand synergetic dynamics.4.3.1.1 Deep versus Shallow Learning for Perceptual Data ProcessingThe most critical feature of DeSTIN is its uniquely robust approach to modeling the worldbased on perceptual data. Mimicking the efficiency and robustness by which the human brainanalyzes and represents information has been a core challenge in AI research for decades. Forinstance, humans are exposed to massive amounts of visual and auditory data every secondof every day, and are somehow able to capture critical aspects of it in a way that allows forappropriate future recollection and action selection. For decades, it has been known that the4.3 Emergentist Cognitive Architectures 67Fig. 4.4: High-level architecture of DeSTINbrain is a massively parallel fabric, in which computation processes and memory storage arehighly distributed. But massive parallelism is not in itself a solution – one also needs the rightarchitecture; which DeSTIN provides, building on prior work in the area of deep learning.Humanlike intelligence is heavily adapted to the physical environments in which humansevolved; and one key aspect of sensory data coming from our physical environments is itshierarchical structure. However, most machine learning and pattern recognition systems are“shallow” in structure, not explicitly incorporating the hierarchical structure of the world intheir architecture. In the context of perceptual data processing, the practical result of this isthe need to couple each shallow learner with a pre-processing stage, wherein high-dimensionalsensory signals are reduced to a lower-dimension feature space that can be understood by theshallow learner. The hierarchical structure of the world is thus crudely captured in the hierarchyof “preprocessor plus shallow learner.” In this sort of approach, much of the intelligence of thesystem shifts to the feature extraction process, which is often imperfect and always applicationdomainspecific.Deep machine learning has emerged as a more promising framework for dealing with complex,high-dimensional real-world data. Deep learning systems possess a hierarchical structure thatintrinsically biases them to recognize the hierarchical patterns present in real-world data. Thus,they hierarchically form a feature space that is driven by regularities in the observations, ratherthan by hand-crafted techniques. They also offer robustness to many of the distortions andtransformations that characterize real-world signals, such as noise, displacement, scaling, etc.Deep belief networks [HOT06] and Convolutional Neural Networks [LBDE90] have beendemonstrated to successfully address pattern inference in high dimensional data (e.g. images).They owe their success to their underlying paradigm of partitioning large data structures intosmaller, more manageable units, and discovering the dependencies that may or may not exist68 4 Brief Survey of Cognitive ArchitecturesFig. 4.5: A standard, general-purpose hierarchical control architecture. DeSTIN’s control hierarchyexemplifies this architecture, with the difference lying mainly in the DeSTIN controlhierarchy’s tight integration with the state inference (perception) and critic (reinforcement)hierarchies.between such units. However, this paradigm has its limitations; for instance, these approachesdo not represent temporal information with the same ease as spatial structure. Moreover, somekey constraints are imposed on the learning schemes driving these architectures, namely theneed for layer-by-layer training, and oftentimes pre-training. DeSTIN overcomes the limitationsof prior deep learning approaches to perception processing, and also extends beyond perceptionto action and reinforcement learning.4.3.1.2 DeSTIN for Perception ProcessingThe hierarchical architecture of DeSTIN’s spatiotemporal inference network comprises an arrangementinto multiple layers of “nodes” comprising multiple instantiations of an identicalcortical circuit. Each node corresponds to a particular spatiotemporal region, and uses a statisticallearning algorithm to characterize the sequences of patterns that are presented to it bynodes in the layer beneath it. More specifically,• At the very lowest layer of the hierarchy nodes receive as input raw data (e.g. pixels of animage) and continuously construct a belief state that attempts to characterize the sequencesof patterns viewed.4.3 Emergentist Cognitive Architectures 69• The second layer, and all those above it, receive as input the belief states of nodes at theircorresponding lower layers, and attempt to construct belief states that capture regularitiesin their inputs.• Each node also receives as input the belief state of the node above it in the hierarchy (whichconstitutes “contextual” information)Fig. 4.6: Small-scale instantiation of the DeSTIN perceptual hierarchy. Each box represents anode, which corresponds to a spatiotemporal region (nodes higher in the hierarchy correspondingto larger regions). O denotes the current observation in the region, C is the state of the higherlayernode, and S and S ′ denote state variables pertaining to two subsequent time steps. Ineach node, a statistical learning algorithm is used to predict subsequent states based on priorstates, current observations, and the state of the higher-layer node.More specifically, each of the DeSTIN nodes, referring to a specific spacetime region, containsa set of state variables conceived as clusters, each corresponding to a set of previously-observedsequences of events. These clusters are characterized by centroids (and are hence assumedroughly spherical in shape), and each of them comprises a certain "spatiotemporal form" recognizedby the system in that region. Each node then contains the task of predicting the likelihoodof a certain centroid being most apropos in the near future, based on the past history of observationsin the node. This prediction may be done by simple probability tabulation, or via70 4 Brief Survey of Cognitive Architecturesapplication of supervised learning algorithms such as recurrent neural networks. These clusteringand prediction processes occur separately in each node, but the nodes are linked togethervia bidirectional dynamics: each node feeds input to its parents, and receives "advice" from itsparents that is used to condition its probability calculations in a contextual way.These processes are executed formally by the following basic belief update rule, which governsthe learning process and is identical for every node in the architecture. The belief state is aprobability mass function over the sequences of stimuli that the nodes learns to represent.Consequently, each node is allocated a predefined number of state variables each denoting adynamic pattern, or sequence, that is autonomously learned. The DeSTIN update rule mapsthe current observation (o), belief state (b), and the belief state of a higher-layer node or context(c), to a new (updated) belief state (b ′ ), such thatalternatively expressed asb ′ (s ′ ) = Pr (s ′ |o, b, c) = Pr (s′ ∩ o ∩ b ∩ c), (4.1)Pr (o ∩ b ∩ c)b ′ (s ′ ) = Pr(o|s′ , b, c) Pr (s ′ |b, c) Pr (b, c). (4.2)Pr (o|b, c) Pr (b, c)Under the assumption that observations depend only on the true state, or Pr(o|s ′ , b, c) =Pr(o|s ′ ), we can further simplify the expression such thatb ′ (s ′ ) = Pr(o|s′ ) Pr (s ′ |b, c), (4.3)Pr (o|b, c)where Pr (s ′ |b, c) = ∑ Pr (s ′ |s, c) b (s), yielding the belief update rules∈Sb ′ (s ′ ) =Pr (o|s ′ ) ∑ Pr (s ′ |s, c) b (s)s∈S∑Pr (o|s ′′ ) ∑ Pr (s ′′ |s, c) b (s) , (4.4)s ′′ ∈Swhere S denotes the sequence set (i.e. belief dimension) such that the denominator term is anormalization factor.One interpretation of eq. (4.4) would be that the static pattern similarity metric, Pr (o|s ′ ) ,is modulated by a construct that reflects the system dynamics, Pr (s ′ |s, c). As such, the beliefstate inherently captures both spatial and temporal information. In our implementation, thebelief state of the parent node, c, is chosen using the selection rules∈Sc = arg max b p (s), (4.5)swhere b p is the belief distribution of the parent node.A close look at eq. (4.4) reveals that there are two core constructs to be learned, Pr(o|s ′ )and Pr(s ′ |s, c). In the current DeSTIN design, the former is learned via online clustering whilethe latter is learned based on experience by inductively learning a rule that predicts the nextstate s ′ given the prior state s and c.The overall result is a robust framework that autonomously (i.e. with no human engineeredpre-processing of any type) learns to represent complex data patterns, and thus serves the4.3 Emergentist Cognitive Architectures 71critical role of building and maintaining a model of the state of the world. In a vision processingcontext, for example, it allows for powerful unsupervised classification. If shown a variety ofreal-world scenes, it will automatically form internal structures corresponding to the variousnatural categories of objects shown in the scenes, such as trees, chairs, people, etc.; and alsothe various natural categories of events it sees, such as reaching, pointing, falling. And, as willbe discussed below, it can use feedback from DeSTIN’s action and critic networks to furthershape its internal world-representation based on reinforcement signals.Benefits of DeSTIN for Perception ProcessingDeSTIN’s perceptual network offers multiple key attributes that render it more powerful thanother deep machine learning approaches to sensory data processing:1. The belief space that is formed across the layers of the perceptual network inherentlycaptures both spatial and temporal regularities in the data. Given that many applicationsrequire that temporal information be discovered for robust inference, this is a key advantageover existing schemes.2. Spatiotemporal regularities in the observations are captured in a coherent manner (ratherthan being represented via two separate mechanisms)3. All processing is both top-down and bottom-up, and both hierarchical and heterarchical,based on nonlinear feedback connections directing activity and modulating learning in multipledirections through DeSTIN’s cortical circuits4. Support for multi-modal fusing is intrinsic within the framework, yielding a powerful stateinference system for real-world, partially-observable settings.5. Each node is identical, which makes it easy to map the design to massively parallel platforms,such as graphics processing units.Points 2-4 in the above list describe how DeSTIN’s perceptual network displays its own“cognitive synergy” in a way that fits naturally into the overall synergetic dynamics of the overallCogPrime architecture. Using this cognitive synergy, DeSTIN’s perceptual network addressesa key aspect of general intelligence: the ability to robustly infer the state of the world, withwhich the system interacts, in an accurate and timely manner.4.3.1.3 DeSTIN for Action and ControlDeSTIN’s perceptual network performs unsupervised world-modeling, which is a critical aspectof intelligence but of course is not the whole story. DeSTIN’s action network, coupled with theperceptual network, orchestrates actuator commands into complex movements, but also carriesout other functions that are more cognitive in nature.For instance, people learn to distinguish between cups and bowls in part via hearing otherpeople describe some objects as cups and others as bowls. To emulate this kind of learning,DeSTIN’s critic network provides positive or negative reinforcement signals based on whetherthe action network has correctly identified a given object as a cup or a bowl, and this signalthen impacts the nodes in the action network. The critic network takes a simple external “degreeof success or failure” signal and turns it into multiple reinforcement signals to be fed into themultiple layers of the action network. The result is that the action network self-organizes so72 4 Brief Survey of Cognitive Architecturesas to include an implicit “cup versus bowl” classifier, whose inputs are the outputs of some ofthe nodes in the higher levels of the perceptual network. This classifier belongs in the actionnetwork because it is part of the procedure by which the DeSTIN system carries out the actionof identifying an object as a cup or a bowl.This example illustrates how the learning of complex concepts and procedures is dividedfluidly between the perceptual network, which builds a model of the world in an unsupervisedway, and the action network, which learns how to respond to the world in a manner that willreceive positive reinforcement from the critic network.4.3.2 Developmental Robotics ArchitecturesA particular subset of emergentist cognitive architectures are sufficiently important that weconsider them separately here: these are developmental robotics architectures, focused on controllingrobots without significant “hard-wiring” of knowledge or capabilities, allowing robotsto learn (and learn how to learn, etc.) via their engagement with the world. A significant focusis often placed here on “intrinsic motivation,” wherein the robot explores the world guided byinternal goals like novelty or curiosity, forming a model of the world as it goes along, basedon the modeling requirements implied by its goals. Many of the foundations of this researcharea were laid by Juergen Schmidhuber’s work in the 1990s [Sch91b, Sch91a, Sch95, Sch02], butnow with more powerful computers and robots the area is leading to more impressive practicaldemonstrations.We mention here a handful of the important initiatives in this area:• Juyang Weng’s Dav [HZT + 02] and SAIL [WHZ + 00] projects involve mobile robots thatexplore their environments autonomously, and learn to carry out simple tasks by building uptheir own world-representations through both unsupervised and teacher-driven processingof high-dimensional sensorimotor data. The underlying philosophy is based on human childdevelopment [WH06], the knowledge representations involved are neural network based,and a number of novel learning algorithms are involved, especially in the area of visionprocessing.• FLOWERS [BO09], an initiative at the French research institute INRIA, led by Pierre-Yves Oudeyer, is also based on a principle of trying to reconstruct the processes of developmentof the human child’s mind, spontaneously driven by intrinsic motivations. Kaplan[Kap08] has taken this project in a direction closely related to our own via the creationof a “robot playroom.” Experiential language learning has also been a focus of the project[OK06], driven by innovations in speech understanding.• IM-CLEVER 1 , a new European project coordinated by Gianluca Baldassarre and conductedby a large team of researchers at different institutions, is focused on creating softwareenabling an iCub [MSV + 08] humanoid robot to explore the environment and learn to carryout human childlike behaviors based on its own intrinsic motivations. As this project is theclosest to our own we will discuss it in more depth below.Like CogPrime, IM-CLEVER is a humanoid robot intelligence architecture guided by intrinsicmotivations, and using hierarchical architectures for reinforcement learning and sensory ab-1 http://im-clever.noze.it/project/project-description4.4 Hybrid Cognitive Architectures 73straction. IM-CLEVER’s motivational structure is based in part on Schmidhuber’s informationtheoreticmodel of curiosity [Sch06]; and CogPrime’s Psi-based motivational structure utilizesprobabilistic measures of novelty, which are mathematically related to Schmidhuber’s measures.On the other hand, IM-CLEVER’s use of reinforcement learning follows Schmidhuber’searlier work RL for cognitive robotics [BS04, BZGS06], Barto’s work on intrinsically motivatedreinforcement learning [SB06, SM05], and Lee’s [LMC07b, LMC07a] work on developmentalreinforcement learning; whereas CogPrime’s assemblage of learning algorithms is more diverse,including probabilistic logic, concept blending and other symbolic methods (in the OCP component)as well as more conventional reinforcement learning methods (in the DeSTIN component).In many respects IM-CLEVER bears a moderately strong resemblance to DeSTIN, whoseintegration with CogPrime is discussed in Chapter 26 of Part 2 (although IM-CLEVER hasmuch more focus on biological realism than DeSTIN). Apart from numerous technical differences,the really big distinction between IM-CLEVER and CogPrime is that in the latter weare proposing to hybridize a hierarchical-abstraction/reinforcement-learning system (such asDeSTIN) with a more abstract symbolic cognition engine that explicitly handles probabilisticlogic and language. IM-CLEVER lacks the aspect of hybridization with a symbolic system, takingmore of a pure emergentist strategy. Like DeSTIN considered as a standalone architectureIM-CLEVER does entail a high degree of cognitive synergy, between components dealing withperception, world-modeling, action and motivation. However, the “emergentist versus hybrid”is a large qualitative difference between the two approaches.In all, while we largely agree with the philosophy underlying developmental robotics, ourintuition is that the learning and representational mechanisms underlying the current systemsin this area are probably not powerful enough to lead to human child level intelligence. Weexpect that these systems will develop interesting behaviors but fall short of robust preschoollevel competency, especially in areas like language and reasoning where symbolic systems havetypically proved more effective. This intuition is what impels us to pursue a hybrid approach,such as CogPrime. But we do feel that eventually, once the mechanisms underlying brains arebetter understood and robotic bodies are richer in sensation and more adept in actuation, somesort of emergentist, developmental-robotics approach can be successful at creating humanlike,human-level AGI.4.4 Hybrid Cognitive ArchitecturesIn response to the complementary strengths and weaknesses of the symbolic and emergentistapproaches, in recent years a number of researchers have turned to integrative, hybrid architectures,which combine subsystems operating according to the two different paradigms. Thecombination may be done in many different ways, e.g. connection of a large symbolic subsystemwith a large subsymbolic system, or the creation of a population of small agents each of whichis both symbolic and subsymbolic in nature.Nils Nilsson expressed the motivation for hybrid AGI systems very clearly in his article atthe AI-50 conference (which celebrated the 50’th anniversary of the AI field) [Nil09]. Whileaffirming the value of the Physical Symbol System Hypothesis that underlies symbolic AI, heargues that “the PSSH explicitly assumes that, whenever necessary, symbols will be groundedin objects in the environment through the perceptual and effector capabilities of a physicalsymbol system.” Thus, he continues,74 4 Brief Survey of Cognitive Architectures“I grant the need for non-symbolic processes in some intelligent systems, but I think they supplementrather than replace symbol systems. I know of no examples of reasoning, understandinglanguage, or generating complex plans that are best understood as being performed by systemsusing exclusively non-symbolic processes....AI systems that achieve human-level intelligence will involve a combination of symbolic andnon-symbolic processing.”A few of the more important hybrid cognitive architectures are:• CLARION [SZ04] is a hybrid architecture that combines a symbolic component for reasoningon “explicit knowledge” with a connectionist component for managing “implicit knowledge.”Learning of implicit knowledge may be done via neural net, reinforcement learning,or other methods. The integration of symbolic and subsymbolic methods is powerful, but agreat deal is still missing such as episodic knowledge and learning and creativity. Learningin the symbolic and subsymbolic portions is carried out separately rather than dynamicallycoupled, minimizing “cognitive synergy” effects.• DUAL [NK04] is the most impressive system to come out of Marvin Minsky’s “Society ofMind” paradigm. It features a population of agents, each of which combines symbolic andconnectionist representation, self-organizing to collectively carry out tasks such as perception,analogy and associative memory. The approach seems innovative and promising, butit is unclear how the approach will scale to high-dimensional data or complex reasoningproblems due to the lack of a more structured high-level cognitive architecture.• LIDA [BF09] is a comprehensive cognitive architecture heavily based on Bernard Baars’“Global Workspace Theory”. It articulates a “cognitive cycle” integrating various forms ofmemory and intelligent processing in a single processing loop. The architecture ties in wellwith both neuroscience and cognitive psychology, but it deals most thoroughly with “lowerlevel” aspects of intelligence, handling more advanced aspects like language and reasoningonly somewhat sketchily. There is a clear mapping between LIDA structures and processesand corresponding structures and processing in OCP; so that it’s only a mild stretch to viewCogPrime as an instantiation of the general LIDA approach that extends further both inthe lower level (to enable robot action and sensation via DeSTIN) and the higher level (toenable advanced language and reasoning via OCP mechanisms that have no direct LIDAanalogues).• MicroPsi [Bac09] is an integrative architecture based on Dietrich Dorner’s Psi model of motivation,emotion and intelligence. It has been tested on some practical control applications,and also on simulating artificial agents in a simple virtual world. MicroPsi’s comprehensivenessand basis in neuroscience and psychology are impressive, but in the current versionof MicroPsi, learning and reasoning are carried out by algorithms that seem unlikely toscale. OCP incorporates the Psi model for motivation and emotion, so that MicroPsi andCogPrime may be considered very closely related systems. But similar to LIDA, MicroPsicurrently focuses on the “lower level” aspects of intelligence, not yet directly handling advancedprocesses like language and abstract reasoning.• PolyScheme [Cas07] integrates multiple methods of representation, reasoning and inferenceschemes for general problem solving. Each Polyscheme “specialist” models a differentaspect of the world using specific representation and inference techniques, interacting withother specialists and learning from them. Polyscheme has been used to model infant reasoningincluding object identity, events, causality, and spatial relations. The integration of4.4 Hybrid Cognitive Architectures 75reasoning methods is powerful, but the overall cognitive architecture is simplistic comparedto other systems and seems focused more on problem-solving than on the broader problemof intelligent agent control.• Shruti [SA93] is a fascinating biologically-inspired model of human reflexive inference,which represents in connectionist architecture relations, types, entities and causal rulesusing focal-clusters. However, much like Hofstadter’s earlier Copycat architecture [Hof95],Shruti seems more interesting as a prototype exploration of ideas than as a practical AGIsystem; at least, after a significant time of development it has not proved significantlyeffective in any applications• James Albus’s 4D/RCS robotics architecture shares a great deal with some of the emergentistarchitectures discussed above, e.g. it has the same hierarchical pattern recognitionstructure as DeSTIN and HTM, and the same three cross-connected hierarchies as DeSTIN,and shares with the developmental robotics architectures a focus on real-time adaptation tothe structure of the world. However, 4D/RCS is not foundationally learning-based but relieson hard-wired architecture and algorithms, intended to mimic the qualitative structure ofrelevant parts of the brain (and intended to be augmented by learning, which differentiatesit from emergentist approaches.As our own CogPrime approach is a hybrid architecture, it will come as no surprise thatwe believe several of the existing hybrid architectures are fundamentally going in the rightdirection. However, nearly all the existing hybrid architectures have severe shortcomings whichwe feel will prevent them from achieving robust humanlike AGI.Many of the hybrid architectures are in essence “multiple, disparate algorithms carrying outseparate functions, encapsulated in black boxes and communicating results with each other.”For instance, PolyScheme, ACT-R and CLARION all display this “modularity” property to asignificant extent. These architectures lack the rich, real-time interaction between the internaldynamics of various memory and learning processes that we believe is critical to achievinghumanlike general intelligence using realistic computational resources. On the other hand, thosearchitectures that feature richer integration – such as DUAL, Shruti, LIDA and MicroPsi – havethe flaw of relying (at least in their current versions) on overly simplistic learning algorithms,which drastically limits their scalability.It does seem plausible to us that some of these hybrid architectures could be dramaticallyextended or modified so as to produce humanlike general intelligence. For instance, one couldreplace LIDA’s learning algorithms with others that interrelate with each other in a nuancedsynergetic way; or one could replace MicroPsi’s simple learning and reasoning methods withmuch more powerful and scalable ones acting on the same data structures. However, makingthese changes would dramatically alter the cognitive architectures in question on multiple levels.4.4.1 Neural versus Symbolic; Global versus LocalThe “symbolic versus emergentist” dichotomy that we have used to structure our review of cognitivearchitectures is not absolute nor fully precisely defined; it is more of a heuristic distinction.In this section, before plunging into the details of particular hybrid cognitive architectures, wereview two other related dichotomies that are useful for understanding hybrid systems: neuralversus symbolic systems, and globalist versus localist knowledge representation.76 4 Brief Survey of Cognitive Architectures4.4.1.1 Neural-Symbolic IntegrationThe distinction between neural and symbolic systems has gotten fuzzier and fuzzier in recentyears, with developments such as• Logic-based systems being used to control embodied agents (hence using logical terms todeal with data that is apparently perception or actuation-oriented in nature, rather thanbeing symbolic in the semiotic sense), see [SS03a] and [GMIH08].• Hybrid systems combining neural net and logical parts, or using logical or neural net componentsinterchangeably in the same role [LAon].• Neural net systems being used for strongly symbolic tasks such as automated grammarlearning ([Elm91], [Elm91], plus more recent work.)Figure 4.7 presents a schematic diagram of a generic neural-symbolic system, generalizingfrom [BH05], a paper that gives an elegant categorization of neural-symbolic AI systems. Figure4.8 depicts several broad categories of neural-symbolic architecture.Fig. 4.7: Generic neural-symbolic architectureBader and Hitzler categorize neural-symbolic systems according to three orthogonal axes:interrelation, language and usage. “Language” refers to the type of language used in the symboliccomponent, which may be logical, automata-based, formal grammar-based, etc. “Usage” refersto the purpose to which the neural-symbolic interrelation is put. We tend to use “learning” asan encompassing term for all forms of ongoing knowledge-creation, whereas Bader and Hitzlerdistinguish learning from reasoning.Of Bader and Hitzler’s three axes the one that interests us most here is “interrelation”, whichrefers to the way the neural and symbolic components of the architecture intersect with eachother. They distinguish “hybrid” architectures which contain separate but equal, interactingneural and symbolic components; versus “integrative” architectures in which the symbolic componentessentially rides piggyback on the neural component, extracting information from it andhelping it carry out its learning, but playing a clearly derived and secondary role. We preferSun’s (2001) term “monolithic” to Bader and Hitzler’s “integrative” to describe this type ofsystem, as the latter term seems best preserved in its broader meaning.4.4 Hybrid Cognitive Architectures 77Fig. 4.8: Broad categories of neural-symbolic architectureWithin the scope of hybrid neural-symbolic systems, there is another axis which Bader andHitzler do not focus on, because the main interest of their review is in monolithic systems. Wecall this axis "interactivity"’, and what we are referring to is the frequency of high-informationcontent,high-influence interaction between the neural and symbolic components in the hybridsystem. In a low-interaction hybrid system, the neural and symbolic components don’t exchangelarge amounts of mutually influential information all that frequently, and basically act likeindependent system components that do their learning/reasoning/thinking periodically sendingeach other their conclusions. In some cases, interaction may be asymmetric: one component mayfrequently send a lot of influential information to the other, but not vice versa. However, ourhypothesis is that the most capable neural-symbolic systems are going to be the symmetricallyhighly interactive ones.In a symmetric high-interaction hybrid neural-symbolic system, the neural and symboliccomponents exchange influential information sufficiently frequently that each one plays a majorrole in the other one’s learning/reasoning/thinking processes. Thus, the learning processes ofeach component must be considered as part of the overall dynamic of the hybrid system. Thetwo components aren’t just feeding their outputs to each other as inputs, they’re mutuallyguiding each others’ internal processing.One can make a speculative argument for the relevance of this kind of architecture to neuroscience.It seems plausible that this kind of neural-symbolic system roughly emulates the kindof interaction that exists between the brain’s neural subsystems implementing localist symbolicprocessing, and the brain’s neural subsystems implementing globalist, classically “connectionist”processing. It seems most likely that, in the brain, symbolic functionality emerges froman underlying layer of neural dynamics. However, it is also reasonable to conjecture that thissymbolic functionality is confined to a functionally distinct subsystem of the brain, which then78 4 Brief Survey of Cognitive Architecturesinteracts with other subsystems in the brain much in the manner that the symbolic and neuralcomponents of a symmetric high-interaction neural-symbolic system interact.Neuroscience speculations aside, however, our key conjecture regarding neural-symbolic integrationis that this sort of neural-symbolic system presents a promising direction for artificialgeneral intelligence research. In Chapter 26 of Volume 2 we will give a more concrete idea ofwhat a symmetric high-interaction hybrid neural-symbolic architecture might look like, exploringthe potential for this sort of hybridization between the OpenCogPrime AGI architecture(which is heavily symbolic in nature) and hierarchical attractor neural net based architecturessuch as DeSTIN.4.5 Globalist versus Localist RepresentationsAnother interesting distinction, related to but different from “symbolic versus emergentist”and “neural versus symbolic”, may be drawn between cognitive systems (or subsystems) wherememory is essentially global, and those where memory is essentially local. In this sectionwe will pursue this distinction in various guises, along with the less familiar notion of glocalmemory.This globalist/localist distinction is most easily conceptualized by reference to memoriescorresponding to categories of entities or events in an external environment. In an AI systemthat has an internal notion of “activation” – i.e. in which some of its internal elements are moreactive than others, at any given point in time – one can define the internal image of an externalevent or entity as the fuzzy set of internal elements that tend to be active when that event orentity is presented to the system’s sensors. If one has a particular set S of external entities orevents of interest, then, the degree of memory localization of such an AI system relative to Smay be conceived as the percentage of the system’s internal elements that have a high degreeof membership in the internal image of an average element of S.Of course, this characterization of localization has its limitations, such as the possibility ofambiguity regarding what are the “system elements” of a given AI system; and the exclusivefocus on internal images of external phenomena rather than representation of internal abstractconcepts. However, our goal here is not to formulate an ultimate, rigorous and thorough ontologyof memory systems, but only to pose a “rough and ready” categorization so as to properly frameour discussion of some specific AGI issues relevant to CogPrime. Clearly the ideas pursued herewill benefit from further theoretical exploration and elaboration.In this sense, a Hopfield neural net [Ami89] would be considered “globalist” since it has a lowdegree of memory localization (most internal images heavily involve a large number of systemelements); whereas Cyc would be considered “localist” as it has a very high degree of memorylocalization (most internal images are heavily focused on a small set of system elements).However, although Hopfield nets and Cyc form handy examples, the “globalist vs. localist”distinction as described above is not identical to the “neural vs. symbolic” distinction. For it isin principle quite possible to create localist systems using formal neurons, and also to createglobalist systems using formal logic. And “globalist-localist” is not quite identical to “symbolic vsemergentist” either, because the latter is about coordinated system dynamics and behavior notjust about knowledge representation. CogPrime combines both symbolic and (loosely) neuralrepresentations, and also combines globalist and localist representations in a way that we willcall “glocal” and analyze more deeply in Chapter 13; but there are many other ways these various4.5 Globalist versus Localist Representations 79properties could be manifested by AI systems. Rigorously studying the corpus of existing (orhypothetical!) cognitive architectures using these ideas would be a large task, which we do notundertake here.In the next sections we review several hybrid architectures in more detail, focusing mostdeeply on LIDA and MicroPsi which have been directly inspirational for CogPrime.4.5.1 CLARIONRon Sun’s CLARION architecture (see Figure 4.9) is interesting in its combination of symbolicand neural aspects – a combination that is used in a sophisticated way to embody the distinctionand interaction between implicit and explicit mental processes. From a CLARION perspective,architectures like Soar and ACT-R are severely limited in that they deal only with explicitknowledge and associated learning processes.CLARION consists of a number of distinct subsystems, each of which contains a dual representationalstructure, including a “rules and chunks” symbolic knowledge store somewhatsimilar to ACT-R, and a neural net knowledge store embodying implicit knowledge. The mainsubsystems are:• An action-centered subsystem to control actions;• A non-action-centered subsystem to maintain general knowledge;• A motivational subsystem to provide underlying motivations for perception, action, andcognition;• A meta-cognitive subsystem to monitor, direct, and modify the operations of all the othersubsystems.Fig. 4.9: The CLARION cognitive architecture.80 4 Brief Survey of Cognitive Architectures4.5.2 The Society of Mind and the Emotion MachineIn his influential but controversial book The Society of Mind [Min88], Marvin Minsky describeda model of human intelligence as something that is built up from the interactions of numeroussimple agents. He spells out in great detail how various particular cognitive functions may beachieved via agents and their interactions. He leaves no room for any central algorithms orstructures of thought, famously arguing: “What magical trick makes us intelligent? The trickis that there is no trick. The power of intelligence stems from our vast diversity, not from anysingle, perfect principle.”This perspective was extended in the more recent work The Emotion Machine [Min07], whereMinsky argued that emotions are “ways to think” evolved to handle different “problem types”that exist in the world. The brain is posited to have rule-based mechanisms (selectors) thatturns on emotions to deal with various problems.Overall, both of these works serve better as works of speculative cognitive science than asworks of AI or cognitive architecture per se. As neurologist Richard Restak said in his reviewof Emotion Machine, “Minsky does a marvelous job parsing other complicated mental activitiesinto simpler elements. ... But he is less effective in relating these emotional functions to what’sgoing on in the brain.” As Restak added, he is also not so effective at relating these emotionalfunctions to straightforwardly implementable algorithms or data structures.Push Singh, in his PhD thesis and followup work [SBC05], did the best job so far of creatinga concrete AI design based on Minsky’s ideas. While Singh’s system was certainly interesting,it was also noteworthy for its lack of any learning mechanisms, and its exclusive focus onexplicit rather than implicit knowledge. Due to Singh’s tragic death, his work was never broughtanywhere near completion. It seems fair to say that there has not yet been a serious cognitivearchitecture posed based closely on Minsky’s ideas.4.5.3 DUALThe closest thing to a Minsky-ish cognitive architecture is probably DUAL, which takes theSociety of Mind concept and adds to it a number of other interesting ideas. DUAL integratessymbolic and connectionist approaches at a deeper level than CLARION, and has been usedto model various cognitive functions such as perception, analogy and judgment. Computationsin DUAL emerge from the self-organized interaction of many micro-agents, each of which isa hybrid symbolic/connectionist device. Each DUAL agent plays the role of a neural networknode, with an activation level and activation spreading dynamics; but also plays the role ofa symbol, manipulated using formal rules. The agents exchange messages and activation vialinks that can be learned and modified, and they form coalitions which collectively representconcepts, episodes, and facts.The structure of the model is sketchily depicted in Figure 4.10, which covers the applicationof DUAL to a toy environment called TextWorld. The visual input corresponding to a stimulusis presented on a two-dimensional visual array representing the front end of the system.Perceptual primitives like blobs and terminations are immediately generated by cheap parallelcomputations. Attention is controlled at each time by an object which allocates it selectivelyto some area of the stimulus. A detailed symbolic representation is constructed for this areawhich tends to fade away as attention is withdrawn from it and allocated to another one. Cate-4.5 Globalist versus Localist Representations 81gorization of visual memory contents takes place by retrieving object and scene categories fromDUAL’s semantic memory and mapping them onto current visual memory representations.Fig. 4.10: The three main components of the DUAL model: the retinotopic visual array (RVA),the visual working memory (VWM) and DUAL’s semantic memory. Attention is allocated toan area of the visual array by the object in VWM controlling attention, while scene and objectcategories corresponding to the contents of VWM are retrieved from the semantic memory.In principle the DUAL framework seems quite powerful; using the language of CogPrime,however, it seems to us that the learning mechanisms of DUAL have not been formulated insuch a way as to give rise to powerful, scalable cognitive synergy. It would likely be possibleto create very powerful AGI systems within DUAL, and perhaps some very CogPrime -likesystems as well. But the systems that have been created or designed for use within DUAL sofar seem not to be that powerful in their potential or scope.4.5.4 4D/RCSIn a rather different direction, James Albus, while at the National Bureau of Standards, developeda very thorough and impressive architecture for intelligent robotics called 4D/RCS,which was implemented in a number of machines including unmanned automated vehicles. Thisarchitecture lacks critical aspects of intelligence such as learning and creativity, but combinesperception, action, planning and world-modeling in a highly effective and tightly-integratedfashion.The architecture has three hierarchies of memory/processing units: one for perception, onefor action and one for modeling and guidance. Each unit has a certain spatiotemporal scope,82 4 Brief Survey of Cognitive Architecturesand (except for the lowest level) supervenes over children whose spatiotemporal scope is a subsetof its own. The action hierarchy takes care of decomposing tasks into subtasks; whereas thesensation hierarchy takes care of grouping signals into entities and events. The modeling/guidancehierarchy mediates interactions between perception and action based on its understandingof the world and the system’s goals.In his book [AM01] Albus describes methods for extending 4D/RCS into a complete cognitivearchitecture, but these extensions have not been elaborated in full detail nor implemented.Fig. 4.11: Albus’s 4D-RCS architecture for a single vehicle4.5.5 PolySchemeNick Cassimatis’s PolyScheme architecture [Cas07] shares with GLAIR the use of multiplelogical reasoning methods on a common knowledge store. While its underlying ideas are quitegeneral, currently PolyScheme is being developed in the context of the “object tracking” domain(construed very broadly). As a logic framework PolyScheme is fairly conventional (unlike GLAIRor NARS with their novel underlying formalisms), but PolyScheme has some unique conceptualaspects, for instance its connection with Cassimatis’s theory of mind, which holds that the samecore set of logical concepts and relationships underlies both language and physical reasoning[Cas04]. This ties in with the use of a common knowledge store for multiple cognitive processes;for instance it suggests that• the same core relationships can be used for physical reasoning and parsing, but that eachof these domains may involve some additional relationships.• language processing may be done via physical-reasoning-based cognitive processes, plus theadditional activity of some language-specific processes4.5 Globalist versus Localist Representations 83Fig. 4.12: Albus’s perceptual, motor and modeling hierarchies4.5.6 Joshua BlueSam Adams and his colleagues at IBM have created a cognitive architecture called Joshua Blue[AABL02], which has some significant similarities to CogPrime. Similar to our current researchdirection with CogPrime, Joshua Blue was created with loose emulation of child cognitivedevelopment in mind; and, also similar to CogPrime, it features a number of cognitive processesacting on a common neural-symbolic knowledge store. The specific cognitive processes involvedin Joshua Blue and CogPrime are not particularly similar, however. At time of writing (2012)84 4 Brief Survey of Cognitive ArchitecturesJoshua Blue is not under active development and has not been for some time; however, theproject may be reanimated in future.Joshua Blue’s core knowledge representation is a semantic network of nodes connected bylinks along which activation spreads. Although many of the nodes have specific semantic referents,as in a classical semantic net, the spread of activation through the network is designed tolead to the emergence of “assemblies” (which could also be thought of as dynamical attractors)in a manner more similar to an attractor neural network.A major difference from typical semantic or neural network models is the central role thataffect plays in the system’s dynamics. The weights of the links in the knowledge base are adjusteddynamically based on the emotional context – a very direct way of ensuring that cognitiveprocesses and mental representations are continuously influenced by affect. Qualitatively, thismimics the way that particular emotions in the human brain correlate with the disseminationthroughout the brain of particular neurotransmitters, which then affect synaptic activity.A result of this architecture is that in Joshua Blue, emotion directs attention in a very directway: affective weighting is important in determining which associated objects will become part ofthe focus of attention, or will be retained from memory. A notable similarity between CogPrimeand Joshua Blue is that in both systems, nodes are assigned two quantitative attention values,one governing allocation of current system resources (mainly processor time; this is CogPrime’sShortTermImportance) and one governing the long-term allocation of memory (CogPrime’sLongTermImportance).The concrete work done with Joshua Blue involved using it to control a simple agent in a simulatedworld, with the goal that via human interaction, the agent would develop a complex andhumanlike emotional and motivational structure from its simple in-built emotions and drives,and would then develop complex cognitive capabilities as part of this development process.4.5.7 LIDAThe LIDA architecture developed by Stan Franklin and his colleagues [BF09] is based on theconcept of the “cognitive cycle” - a notion that is important to nearly every BICA (BiologicallyInspired Cognitive Architectures) and also to the brain, but that plays a particularly centralrole in LIDA. As Franklin says, "as a matter of principle, every autonomous agent, be it human,animal, or artificial, must frequently sample (sense) its environment, process (make sense of)this input, and select an appropriate response (action). The agent’s “life” can be viewed asconsisting of a continual sequence of iterations of these cognitive cycles. Such cycles constitutethe indivisible elements of attention, the least sensing and acting to which we can attend. Acognitive cycle can be thought of as a moment of cognition, a cognitive "moment"."4.5.8 The Global WorkspaceLIDA is heavily based on the “global workspace” concept developed by Bernard Baars. As thisconcept is also directly relevant to CogPrime it is worth briefly describing here.In essence Baars’ Global Workspace Theory (GWT) is a particular hypothesis about howworking memory works and the role it plays in the mind. Baars conceives working memory as the4.5 Globalist versus Localist Representations 85“inner domain in which we can rehearse telephone numbers to ourselves or, more interestingly,in which we carry on the narrative of our lives. It is usually thought to include inner speechand visual imagery.” Baars uses the term “consciousness” to refer to the contents of workingmemory – a theoretical commitment that is not part of the CogPrime design. In this sectionwe will use the term “consciousness” in Baars’ way, but not throughout the rest of the book.Baars conceives working memory and consciousness in terms of a “theater metaphor” – accordingto which, in the “theater of consciousness” a “spotlight of selective attention” shinesa bright spot on stage. The bright spot reveals the global workspace – the contents of consciousness,which may be metaphorically considered as a group of actors moving in and out ofconsciousness, making speeches or interacting with each other. The unconscious is representedby the audience watching the play ... and there is also a role for the director (the mind’s executiveprocesses) behind the scenes, along with a variety of helpers like stage hands, scriptwriters, scene designers, etc.GWT describes a fleeting memory with a duration of a few seconds. This is much shorterthan the 10-30 seconds of classical working memory – according to GWT there is a very brief“cognitive cycle” in which the global workspace is refreshed, and the time period an item remainsin working memory generally spans a large number of these elementary “refresh” actions. GWTcontents are proposed to correspond to what we are conscious of, and are said to be broadcastto a multitude of unconscious cognitive brain processes. Unconscious processes, operating inparallel, can form coalitions which can act as input processes to the global workspace. Eachunconscious process is viewed as relating to certain goals, and seeking to get involved withcoalitions that will get enough importance to become part of the global workspace – becauseonce they’re in the global workspace they’ll be allowed to broadcast out across the mind as awhole, which include broadcasting to the internal and external actuators that allow the mindto do things. Getting into the global workspace is a process’s best shot at achieving its goals.Obviously, the theater metaphor used to describe the GWT is evocative but limited; forinstance, the unconscious in the mind does a lot more than the audience in a theater. Theunconscious comes up with complex creative ideas sometimes, which feed into consciousness –almost as if the audience is also the scriptwriter. Baars’ theory, with its understanding of unconsciousdynamics in terms of coalition-building, fails to describe the subtle dynamics occurringwithin the various forms of long-term memory, which result in subtle nonlinear interactionsbetween long term memory and working memory. But nevertheless, GWT successfully modelsa number of characteristics of consciousness, including its role in handling novel situations, itslimited capacity, its sequential nature, and its ability to trigger a vast range of unconsciousbrain processes. It is the framework on which LIDA’s theory of the cognitive cycle is built.4.5.9 The LIDA Cognitive CycleThe simplest cognitive cycle is that of an animal, which senses the world, compares sensation tomemory, and chooses an action, all in one fluid subjective moment. But the same cognitive cyclestructure/process applies to higher-level cognitive processes as well. The LIDA architecture isbased on the LIDA model of the cognitive cycle, which posits a particular structure underlyingthe cognitive cycle that possess the generality to encompass both simple and complex cognitivemoments.86 4 Brief Survey of Cognitive ArchitecturesThe LIDA cognitive cycle itself is a theoretical construct that can be implemented in manyways, and indeed other BICAs like CogPrime and Psi also manifest the LIDA cognitive cyclein their dynamics, though utilizing different particular structures to do so.Figure 4.13 shows the cycle pictorially, starting in the upper left corner and proceedingclockwise. At the start of a cycle, the LIDA agent perceives its current situation and allocatesattention differentially to various parts of it. It then broadcasts information about the mostimportant parts (which constitute the agent’s consciousness), and this information gets featuresextracted from it, when then get passed along to episodic and semantic memory, that interactin the “global workspace” to create a model for the agent’s current situation. This model then,in interaction with procedural memory, enables the agent to choose an appropriate action andexecute it - the critical “action-selection” phase!Fig. 4.13: The LIDA Cognitive CycleThe LIDA Cognitive Cycle in More Depth2We now run through the cognitive cycle in more detail. It begins with sensory stimuli fromthe agent’s external internal environment. Low-level feature detectors in sensory memory beginthe process of making sense of the incoming stimuli. These low-level features are passed toperceptual memory where higher-level features, objects, categories, relations, actions, situations,2 This section paraphrases heavily from [Fra06]4.5 Globalist versus Localist Representations 87etc. are recognized. These recognized entities, called percepts, are passed to the workspace,where a model of the agent’s current situation is assembled.Workspace structures serve as cues to the two forms of episodic memory, yielding both shortand long term remembered local associations. In addition to the current percept, the workspacecontains recent percepts that haven’t yet decayed away, and the agent’s model of the thencurrentsituation previously assembled from them. The model of the agent’s current situation isupdated from the previous model using the remaining percepts and associations. This updatingprocess will typically require looking back to perceptual memory and even to sensory memory,to enable the understanding of relations and situations. This assembled new model constitutesthe agent’s understanding of its current situation within its world. Via constructing the model,the agent has made sense of the incoming stimuli.Now attention allocation comes into play, because a real agent lacks the computational resourcesto work with all parts of its world-model with maximal mental focus. Portions of themodel compete for attention. These competing portions take the form of (potentially overlapping)coalitions of structures comprising parts the model. Once one such coalition wins thecompetition, the agent has decided what to focus its attention on.And now comes the purpose of all this processing: to help the agent to decide what to donext. The winning coalition passes to the global workspace, the namesake of Global WorkspaceTheory, from which it is broadcast globally. Though the contents of this conscious broadcastare available globally, the primary recipient is procedural memory, which stores templates ofpossible actions including their context and possible results.Procedural memory also stores an activation value for each such template – a value thatattempts to measure the likelihood of an action taken within its context producing the expectedresult. It’s worth noting that LIDA makes a rather specific assumption here. LIDA’s“activation” values are like the probabilistic truth values of the implications in CogPrime’sContext ∧ Procedure → Goal triples. However, in CogPrime this probability is not the same asthe ShortTermImportance “attention value” associated with the Implication link representingthat implication. Here LIDA merges together two concepts that in CogPrime are separate.Templates whose contexts intersect sufficiently with the contents of the conscious broadcastinstantiate copies of themselves with their variables specified to the current situation. Theseinstantiations are passed to the action selection mechanism, which chooses a single action fromthese instantiations and those remaining from previous cycles. The chosen action then goes tosensorimotor memory, where it picks up the appropriate algorithm by which it is then executed.The action so taken affects the environment, and the cycle is complete.The LIDA model hypothesizes that all human cognitive processing is via a continuing iterationof such cognitive cycles. It acknowledges that other cognitive processes may also occur,refining and building on the knowledge used in the cognitive cycle (for instance, the cognitivecycle itself doesn’t mention abstract reasoning or creativity). But the idea is that these otherprocesses occur in the context of the cognitive cycle, which is the main loop driving the internaland external activities of the organism.4.5.9.1 Avoiding Combinatorial Explosion via Adaptive Attention AllocationLIDA avoids combinatorial explosions in its inference processes via two methods, both of whichare also important in CogPrime :• combining reasoning via association with reasoning via deduction88 4 Brief Survey of Cognitive Architectures• foundational use of uncertainty in reasoningOne can create an analogy between LIDA’s workspace structures and codelets and a logicbasedarchitecture’s assertions and functions. However, LIDA’s codelets only operate on thestructures that are active in the workspace during any given cycle. This includes recent perceptions,their closest matches in other types of memory, and structures recently created by othercodelets. The results with the highest estimate of success, i.e. activation, will then be selected.Uncertainty plays a role in LIDA’s reasoning in several ways, most notably through the baseactivation of its behavior codelets, which depend on the model’s estimated probability of thecodelet’s success if triggered. LIDA observes the results of its behaviors and updates the baseactivation of the responsible codelets dynamically.We note that for this kind of uncertain inference/activation interplay to scale well, somelevel of cognitive synergy must be present; and based on our understanding of LIDA it is notclear to us whether the particular inference and association algorithms used in LIDA possessthe requisite synergy.4.5.9.2 LIDA versus CogPrimeThe LIDA cognitive cycle, broadly construed, exists in CogPrime as in other cognitive architectures.To see how, it suffices to map the key LIDA structures into corresponding CogPrimestructures, as is done in Table 4.1. Of course this table does not cover all CogPrime processes,as LIDA does not constitute a thorough explanation of CogPrime structure and dynamics. Andin most cases the corresponding CogPrime and LIDA processes don’t work in exactly the sameway; for instance, as noted above, LIDA’s action selection relies solely on LIDA’s “activation”values, whereas CogPrime’s action selection process is more complex, relying on aspects ofCogPrime that lack LIDA analogues.4.5.10 Psi and MicroPsiWe have saved for last the architecture that has the most in common with CogPrime : JoschaBach’s MicroPsi architecture, closely based on Dietrich Dorner’s Psi theory. CogPrime hasborrowed substantially from Psi in its handling of emotion and motivation; but Psi also hasother aspects that differ considerably from CogPrime. Here we will focus more heavily on thepoints of overlap, but will mention the key points of difference as well.The overall Psi cognitive architecture, which is centered on the Psi model of the motivationalsystem, is roughly depicted in Figure 4.14.Psi’s motivational system begins with Demands, which are the basic factors that motivatethe agent. For an animal these would include things like food, water, sex, novelty, socialization,protection of one’s children, and so forth. For an intelligent robot they might include thingslike electrical power, novelty, certainty, socialization, well-being of others and mental growth.Psi also specifies two fairly abstract demands and posits them as psychologically fundamental(see Figure 4.15):• competence, the effectiveness of the agent at fulfilling its Urges• certainty, the confidence of the agent’s knowledge4.5 Globalist versus Localist Representations 89LIDACogPrimeDeclarative memory Atomspaceattentional codelets Schema that adjust importance of Atoms explicitlycoalitionsmapsglobal workspaceattentional focusbehavior codeletsschemaprocedural memory (scheme net) procedures in ProcedureRepository; and network ofSchemaNodes in the Atomspaceaction selection (behavior net) propagation of STICurrency from goals to actions, andaction selection processtransient episodic memory perceptual atoms entering AT with high STI, whichrapidly decreases in most caseslocal workspacesbubbles of interlinked Atoms with moderate importance,focused on by a subset of MindAgents (definedin Chapter 19 of Part 2) for a period of timeperceptual associative memory HebbianLinks in the ATsensory memoryspaceserver/timeserver, plus auxiliary stores for othersensessensorimotor memory Atoms storing record of actions taken, linked in withAtoms indexed in sensory memoryTable 4.1: CogPrime Analogues of Key LIDA FeaturesEach demand is assumed to come with a certain “target level” or “target range” (and thesemay fluctuate over time, or may change as a system matures and develops). An Urge is said todevelop when a demand deviates from its target range: the urge then seeks to return the demandto its target range. For instance, in an animal-like agent the demand related to food is moreclearly described as “fullness,” and there is a target range indicating that the agent is neither toohungry nor too full of food. If the agent’s fullness deviates from this range, an Urge to returnthe demand to its target range arises. Similarly, if an agent’s novelty deviates from its targetrange, this means the agent’s life has gotten either too boring or too disconcertingly weird, andthe agent gets an Urge for either more interesting activities (in the case of below-range novelty)or more familiar ones (in the case of above-range novelty).There is also a primitive notion of Pleasure (and its opposite, displeasure), which is consideredas different from the complex emotion of “happiness.” Pleasure is understood as associatedwith Urges: pleasure occurs when an Urge is (at least partially) satisfied, whereas displeasureoccurs when an urge gets increasingly severe. The degree to which an Urge is satisfied is notnecessarily defined instantaneously; it may be defined, for instance, as a time-decaying weightedaverage of the proximity of the demand to its target range over the recent past.So, for instance if an agent is bored and gets a lot of novel stimulation, then it experiencessome pleasure. If it’s bored and then the monotony of its stimulation gets even more extreme,then it experiences some displeasure.Note that, according to this relatively simplistic approach, any decrease in the amount ofdissatisfaction causes some pleasure; whereas if everything always continues within its acceptablerange, there isn’t any pleasure. This may seem a little counterintuitive, but it’s importantto understand that these simple definitions of “pleasure” and “displeasure” are not intended tofully capture the natural language concepts associated with those words. The natural languageterms are used here simply as heuristics to convey the general character of the processes in-90 4 Brief Survey of Cognitive ArchitecturesFig. 4.14: High-Level Architecture of the Psi Modelvolved. These are very low level processes whose analogues in human experience are largelybelow the conscious level.A Goal is considered as a statement that the system may strive to make true at some futuretime. A Motive is an (urge, goal) pair, consisting of a goal whose satisfaction is predicted toimply the satisfaction of some urge. In fact one may consider Urges as top-level goals, and theagent’s other goals as their subgoals.In Psi an agent has one “ruling motive” at any point in time, but this seems an oversimplificationmore applicable to simple animals than to human-like or other advanced AI systems.In general one may think of different motives having different weights indicating the amount ofresources that will be spent on pursuing them.Emotions in Psi are considered as complex systemic response-patterns rather than explicitlyconstructed entities. An emotion is the set of mental entities activated in response to a certainset of urges. Dorner conceived theories about how various common emotions emerge from thedynamics of urges and motives as described in the Psi model. “Intentions” are also considered ascomposite entities: an intention at a given point in time consists of the active motives, togetherwith their related goals, behavior programs and so forth.4.5 Globalist versus Localist Representations 91The basic logic of action in Psi is carried out by “triples” that are very similar to CogPrime’sContext ∧ Procedure → Goal triples. However, an important role is played by four modulatorsthat control how the processes of perception, cognition and action selection are regulated at agiven time:• activation, which determines the degree to which the agent is focused on rapid, intensiveactivity versus reflective, cognitive activity• resolution level, which determines how accurately the system tries to perceive the world• certainty, which determines how hard the system tries to achieve definite, certain knowledge• selection threshold, which determines how willing the system is to change its choice of whichgoals to focus onThese modulators characterize the system’s emotional and cognitive state at a very abstractlevel; they are not emotions per se, but they have a large effect on the agent’s emotions. Theirintended interaction is depicted in Figure 4.15.Fig. 4.15: Primary Interrelationships Between Psi Modulators4.5.11 The Emergence of Emotion in the Psi ModelWe now briefly review the specifics of how Psi models the emergence of emotion. The basic idea isto define a small set of proto-emotional dimensions in terms of basic Urges and modulators.Then, emotions are identified with regions in the space spanned by these dimensions.The simplest approach uses a six-dimensional continuous space:1. pleasure92 4 Brief Survey of Cognitive Architectures2. arousal3. resolution level4. selection threshold (i.e. degree of dominance of the leading motive)5. level of background checks (the rate of the securing behavior)6. level of goal-directed behaviorFigure 4.16 shows how the latter 5 of these dimensions are derived from underlying urges andmodulators. Note that these dimensions are not orthogonal; for instance resolution is mainly inverselyrelated to arousal. Additional dimensions are also discussed, for instance it is postulatedthat to deal with social emotions one may wish to introduce two more demands correspondingto inner and outer obedience to social norms, and then define dimensions in terms of these.Fig. 4.16: Five Proto-Emotional Dimensions Implicit in the Psi ModelSpecific emotions are then characterized in terms of these dimensions. According to [Bac09],for instance, “Anger ... is characterized by high arousal, low resolution, strong motive dominance,few background checks and strong goal-orientedness; sadness by low arousal, high resolution,strong dominance, few background-checks and low goal-orientedness.”I’m a bit skeptical of the contention that these dimensions fully characterize the relevantemotions. Anger for instance seems to have some particular characteristics not implied by theabove list of dimensional values. The list of dimensional values associated with anger doesn’ttell us that an angry person is more likely to punch someone than to bounce up and down,for example. However, it does seem that the dimensional values associated with an emotion are4.5 Globalist versus Localist Representations 93informative about the emotion, so that positioning an emotion on the given dimensions tellsone a lot.4.5.12 Knowledge Representation, Action Selection and Planning inPsiIn addition to the basic motivation/emotion architecture of Psi, which has been adopted (withsome minor changes) for use in CogPrime, Psi has a number of other aspects that are somewhatdifferent from their CogPrime analogues.First of all, on the micro level, Psi represents knowledge using structures called “quads.” Eachquad is a cluster of 5 neurons containing a core neuron, and four other neurons representingbefore/after and part-of/has-part relationships in regard to that core neuron. Quads are naturallyassembled into spatiotemporal hierarchies, though they are not required to form part ofsuch a structure.Psi stores knowledge using quads arranged in three networks, which are conceptually similarto the networks in Albus’s 4D/RCS and Arel’s DeSTIN architectures:• A sensory network, which stores declarative knowledge: schemas representing images, objects,events and situations as hierarchical structures.• A motor network, which contains procedural knowledge by way of hierarchical behaviorprograms• A motivational network handling demandsPerception in Psi, which is centered in the sensory network, follows principles similar toDeSTIN (which are shared also by other systems), for instance the principle of perception asprediction. Psi’s “HyPercept” mechanism performs hypothesis-based perception: it attempts topredict what is there to be perceived and then attempts to verify these predictions using sensationand memory. Furthermore HyPercept is intimately coupled with actions in the externalworld, according to the concept of “Neisser’s perceptual cycle,” the cycle between explorationand representation of reality. Perceptually acquired information is translated into schemas capableof guiding behaviors, and these are enacted (sometimes affecting the world in significantways) and in the process used to guide further perception. Imaginary perceptions are handledvia a “mental stage” analogous to CogPrime’s internal simulation world.Action selection in Psi works based on what are called “triplets,” each of which consists of• a sensor schema (pre-conditions, “condition schema”; like CogPrime’s “context”)• a subsequent motor schema (action, effector; like CogPrime’s “procedure”)• a final sensor schema (post-conditions, expectations; like an CogPrime predicate or goal)What distinguishes these triplets from classic production rules as used in (say) Soar andACT-R is that the triplets may be partial (some of the three elements may be missing) andmay be uncertain. However, there seems no fundamental difference between these triplets andCogPrime’s concept/procedure/goal triplets, at a high level; the difference lies in the underlyingknowledge representation used for the schemata, and the probabilistic logic used to representthe implication.The work of figuring out what schema to execute to achieve the chosen goal in the currentcontext is done in Psi using a combination of processes called the “Rasmussen ladder” (named94 4 Brief Survey of Cognitive Architecturesafter Danish psychologist Jens Rasmussen). The Rasmussen ladder describes the organizationof action as a movement between the stages of skill-based behavior, rule-based behavior andknowledge-based behavior, as follows:• If a given task amounts to a trained routine, an automatism or skill is activated; it canusually be executed without conscious attention and deliberative control.• If there is no automatism available, a course of action might be derived from rules; before aknown set of strategies can be applied, the situation has to be analyzed and the strategieshave to be adapted.• In those cases where the known strategies are not applicable, a way of combining theavailable manipulations (operators) into reaching a given goal has to be explored at first.This stage usually requires a recomposition of behaviors, that is, a planning process.The planning algorithm used in the Psi and MicroPsi implementations is a fairly simplehill-climbing planner. While it’s hypothesized that a more complex planner may be needed foradvanced intelligence, part of the Psi theory is the hypothesis that most real-life planning anorganism needs to do is fairly simple, once the organism has the right perceptual representationsand goals.4.5.13 Psi versus CogPrimeOn a high level, the similarities between Psi and CogPrime are quite strong:• interlinked declarative, procedural and intentional knowledge structures, represented usingneural-symbolic methods (though, the knowledge structures have somewhat different highlevelstructures and low-level representational mechanisms in the two systems)• perception via prediction and perception/action integration• action selection via triplets that resemble uncertain, potentially partial production rules• similar motivation/emotion framework, since CogPrime incorporates a variant of Psi forthisOn the nitty-gritty level there are many differences between the systems, but on the bigpicturelevel the main difference lies in the way the cognitive synergy principle is pursued inthe two different approaches. Psi and MicroPsi rely on very simple learning algorithms that areclosely tied to the “quad” neurosymbolic knowledge representation, and hence interoperate ina fairly natural way without need for subtle methods of “synergy engineering.” CogPrime usesmuch more diverse and sophisticated learning algorithms which thus require more sophisticatedmethods of interoperation in order to achieve cognitive synergy.Chapter 5A Generic Architecture of Human-Like Cognition5.1 IntroductionWhen writing the first draft of this book, some years ago, we had the idea to explain CogPrimeby aligning its various structures and processes with the ones in the "standard architecturediagram" of the human mind. After a bit of investigation, though, we gradually came to therealization that no such thing existed. There was no standard flowchart or other sort of diagramexplaining the modern consensus on how human thought works. Many such diagramsexisted, but each one seemed to represent some particular focus or theory, rather than an overallintegrative understanding.Since there are multiple opinions regarding nearly every aspect of human intelligence, itwould be difficult to get two cognitive scientists to fully agree on every aspect of an overallhuman cognitive architecture diagram. Prior attempts to outline detailed mind architectureshave tended to follow highly specific theories of intelligence, and hence have attracted onlymoderate interest from researchers not adhering to those theories. An example is Minsky’s workpresented in The Emotion Machine [Min07], which arguably does constitute an architecturediagram for the human mind, but which is only loosely grounded in current empirical knowledgeand stands more as a representation of Minsky’s own intuitive understanding.But nevertheless, it seemed to us that a reasonable attempt at an integrative, relativelytheory-neutral "human cognitive architecture diagram" would be better than nothing. So naturally,we took it on ourselves to create such a diagram. This chapter is the result – it draws onthe thinking of a number of cognitive science and AGI researchers, integrating their perspectivesin a coherent, overall architecture diagram for human, and human-like, general intelligence. Thespecific architecture diagram of CogPrime, given in Chapter 6 below, may then be understoodas a particular instantiation of this generic architecture diagram of human-like cognition.There is no getting around the fact that, to a certain extent, the diagram presented herereflects our particular understanding of how the mind works. However, it was intentionallyconstructed with the goal of not being just an abstracted version of the CogPrime architecturediagram! It does not reflect our own idiosyncratic understanding of human intelligence, as muchas a combination of understandings previously presented by multiple researchers (includingourselves), arranged according to our own taste in a manner we find conceptually coherent.With this in mind, we call it the "Integrative Human-Like Cognitive Architecture Diagram," orfor short "the integrative diagram." We have made an effort to ensure that as many pieces ofthe integrative diagram as possible are well grounded in psychological and even neuroscientific9596 5 A Generic Architecture of Human-Like Cognitiondata, rather than mainly embodying speculative notions; however, given the current state ofknowledge, this could not be done to a complete extent, and there is still some speculationinvolved here and there.While based on understandings of human intelligence, the integrative diagram is intended toserve as an architectural outline for human-like general intelligence more broadly. For example,CogPrime is explicitly not intended as a precise emulation of human intelligence, and does manythings quite differently than the human mind, yet can still fairly straightforwardly be mappedinto the integrative diagram.The integrative diagram focuses on structure, but this should not be taken to represent avaluation of structure over dynamics in our approach to intelligence. Following chapters treatvarious dynamical phenomena in depth.5.2 Key Ingredients of the Integrative Human-Like CognitiveArchitecture DiagramThe main ingredients we’ve used in assembling the integrative diagram are as follows:• Our own views on the various types of memory critical for human-like cognition, and theneed for tight, "synergetic" interactions between the cognitive processes focused on these• Aaron Sloman’s high-level architecture diagram of human intelligence [Slo01], drawn fromhis CogAff architecture, which strikes me as a particularly clear embodiment of "moderncommon sense" regarding the overall architecture of the human mind. We have added onlya couple items to Sloman’s high-level diagram, which we felt deserved an explicit high-levelrole that he did not give them: emotion, language and reinforcement.• The LIDA architecture diagram presented by Stan Franklin and Bernard Baars [BF09].We think LIDA is an excellent model of working memory and what Sloman calls "reactiveprocesses", with well-researched grounding in the psychology and neuroscience literature.We have adapted the LIDA diagram only very slightly for use here, changing some ofthe terminology on the arrows, and indicating where parts of the LIDA diagram indicateprocesses elaborated in more detail elsewhere in the integrative diagram.• The architecture diagram of the Psi model of motivated cognition, presented by JoschaBach in [Bac09] based on prior work by Dietrich Dorner [Dör02]. This diagram is presentedwithout significant modification; however it should be noted that Bach and Dorner presentthis diagram in the context of larger and richer cognitive models, the other aspects of whichare not all incorporated in the integrative diagram.• James Albus’s three-hierarchy model of intelligence [AM01], involving coupled perception,action and reinforcement hierarchies. Albus’s model, utilized in the creation of intelligentunmanned automated vehicles, is a crisp embodiment of many ideas emergent from the fieldof intelligent control systems.• Deep learning networks as a model of perception (and action and reinforcement learning),as embodied for example in the work of Itamar Arel [ARC09] and Jeff Hawkins [HB06]. Theintegrative diagram adopts this as the basic model of the perception and action subsystemsof human intelligence. Language understanding and generation are also modeled accordingto this paradigm.5.3 An Architecture Diagram for Human-Like General Intelligence 97One possible negative reaction to the integrative diagram might be to say that it’s a kindof Frankenstein monster diagram, piecing together aspects of different theories in a way thatviolates the theoretical notions underlying all of them! For example, the integrative diagramtakes LIDA as a model of working memory and reactive processing, but from the papers onLIDA it’s unclear whether the creators of LIDA construe it more broadly than that. The deeplearning community tends to believe that the architecture of current deep learning networks,in itself, is close to sufficient for human-level general intelligence – whereas the integrativediagram appropriates the ideas from this community mainly for handling perception, actionand language, etc.On the other hand, in a more positive perspective, one could view the integrative diagramas consistent with LIDA, but merely providing much more detail on some of the boxes in theLIDA diagram (e.g. dealing with perception and long-term memory). And one could view theintegrative diagram as consistent with the deep learning paradigm – via viewing it, not asa description of components to be explicitly implemented in an AGI system, but rather as adescription of the key structures and processes that must emerge in deep learning network, basedon its engagement with the world, in order for it to achieve human-like general intelligence.Our own view, underlying the creation of the integrative diagram, is that different communitiesof cognitive science researchers have focused on different aspects of intelligence, and havethus each created models that are more fully fleshed out in some aspects than others. But thesevarious models all link together fairly cleanly, which is not surprising as they are all groundedin the same data regarding human intelligence. Many judgment calls must be made in fusingmultiple models in the way that the integrative diagram does, but we feel these can be madewithout violating the spirit of the component models. In assembling the integrative diagram, wehave made these judgment calls as best we can, but we’re well aware that different judgmentswould also be feasible and defensible. Revisions are likely as time goes on, not only due tonew data about human intelligence but also to evolution of understanding regarding the bestapproach to model integration.Another possible argument against the ideas presented here is that there’s nothing new – allthe ingredients presented have been given before elsewhere. To this our retort is to quote Pascal:"Let no one say that I have said nothing new ... the arrangement of the subject is new." Thevarious architecture diagrams incorporated into the integrative diagram are either extremelyhigh level (Sloman’s diagram) or focus primarily on one aspect of intelligence, treating theothers very concisely by summarizing large networks of distinction structures and processes insmall boxes. The integrative diagram seeks to cover all aspects of human-like intelligence at aroughly equal granularity – a different arrangement.This kind of high-level diagramming exercise is not precise enough, nor dynamics-focusedenough, to serve as a guide for creating human-level or more advanced AGI. But it can be auseful tool for explaining and interpreting a concrete AGI design, such as CogPrime.5.3 An Architecture Diagram for Human-Like General IntelligenceThe integrative diagram is presented here in a series of seven Figures.Figure 5.1 gives a high-level breakdown into components, based on Sloman’s high-levelcognitive-architectural sketch [Slo01]. This diagram represents, roughly speaking, "modern commonsense" about how a human-like mind is architected. The separation between structures98 5 A Generic Architecture of Human-Like CognitionFig. 5.1: High-Level Architecture of a Human-Like Mindand processes, embodied in having separate boxes for Working Memory vs. Reactive Processes,and for Long Term Memory vs. Deliberative Processes, could be viewed as somewhat artificial,since in the human brain and most AGI architectures, memory and processing are closely integrated.However, the tradition in cognitive psychology is to separate out Working Memory andLong Term Memory from the cognitive processes acting thereupon, so we have adhered to thatconvention. The other changes from Sloman’s diagram are the explicit inclusion of language,representing the hypothesis that language processing is handled in a somewhat special way inthe human brain; and the inclusion of a reinforcement component parallel to the perception andaction hierarchies, as inspired by intelligent control systems theory (e.g. Albus as mentionedabove) and deep learning theory. Of course Sloman’s high level diagram in its original form isintended as inclusive of language and reinforcement, but we felt it made sense to give themmore emphasis.Figure 5.2, modeling working memory and reactive processing, is essentially the LIDA diagramas given in prior papers by Stan Franklin, Bernard Baars and colleagues [BF09]. Theboxes in the upper left corner of the LIDA diagram pertain to sensory and motor processing,which LIDA does not handle in detail, and which are modeled more carefully by deep learningtheory. The bottom left corner box refers to action selection, which in the integrative diagramis modeled in more detail by Psi. The top right corner box refers to Long-Term Memory, whichthe integrative diagram models in more detail as a synergetic multi-memory system (Figure5.4).The original LIDA diagram refers to various "codelets", a key concept in LIDA theory. Wehave replaced "attention codelets" here with "attention flow", a more generic term. We suggestone can think of an attention codelet as: a piece of information stating that, for a certain groupof items, it’s currently pertinent to pay attention to this group as a collective.5.3 An Architecture Diagram for Human-Like General Intelligence 99Fig. 5.2: Architecture of Working Memory and Reactive Processing, closely modeled on theLIDA architectureFigure 5.3, modeling motivation and action selection, is a lightly modified version of thePsi diagram from Joscha Bach’s book Principles of Synthetic Intelligence [Bac09]. The maindifference from Psi is that in the integrative diagram the Psi motivated action framework isembedded in a larger, more complex cognitive model. Psi comes with its own theory of workingand long-term memory, which is related to but different from the one given in the integrativediagram – it views the multiple memory types distinguished in the integrative diagram asemergent from a common memory substrate. Psi comes with its own theory of perception andaction, which seems broadly consistent with the deep learning approach incorporated in theintegrative diagram. Psi’s handling of working memory lacks the detailed, explicit workflow ofLIDA, though it seems broadly conceptually consistent with LIDA.In Figure 5.3, the box labeled "Other portions of working memory" is labeled "Protocol andsituation memory" in the original Psi diagram. The Perception, Action Execution and ActionSelection boxes have fairly similar semantics to the similarly labeled boxes in the LIDA-likeFigure 5.2, so that these diagrams may be viewed as overlapping. The LIDA model doesn’texplain action selection and planning in as much detail as Psi, so the Psi-like Figure 5.3 couldbe viewed as an elaboration of the action-selection portion of the LIDA-like Figure 5.2. InPsi, reinforcement is considered as part of the learning process involved in action selection andplanning; in Figure 5.3 an explicit "reinforcement box" has been added to the original Psidiagram, to emphasize this.Figure 5.4, modeling long-term memory and deliberative processing, is derived from our ownprior work studying the "cognitive synergy" between different cognitive processes associatedwith different types of memory. The division into types of memory is fairly standard. Declarative,procedural, episodic and sensorimotor memory are routinely distinguished; we like to distinguishattentional memory and intentional (goal) memory as well, and view these as the interfacebetween long-term memory and the mind’s global control systems. One focus of our AGI designwork has been on designing learning algorithms, corresponding to these various types of memory,100 5 A Generic Architecture of Human-Like CognitionFig. 5.3: Architecture of Motivated ActionFig. 5.4: Architecture of Long-Term Memory and Deliberative and Metacognitive Thinkingthat interact with each other in a synergetic way [Goe09c], helping each other to overcometheir intrinsic combinatorial explosions. There is significant evidence that these various typesof long-term memory are differently implemented in the brain, but the degree of structure anddynamical commonality underlying these different implementations remains unclear.5.3 An Architecture Diagram for Human-Like General Intelligence 101Each of these long-term memory types has its analogue in working memory as well. In somecognitive models, the working memory and long-term memory versions of a memory type andcorresponding cognitive processes, are basically the same thing. CogPrime is mostly like this– it implements working memory as a subset of long-term memory consisting of items withparticularly high importance values. The distinctive nature of working memory is enforced viausing slightly different dynamical equations to update the importance values of items withimportance above a certain threshold. On the other hand, many cognitive models treat workingand long term memory as more distinct than this, and there is evidence for significant functionaland anatomical distinctness in the brain in some cases. So for the purpose of the integrativediagram, it seemed best to leave working and long-term memory subcomponents as parallel butdistinguished.Figure 5.4 also encompasses metacognition, under the hypothesis that in human beings andhuman-like minds, metacognitive thinking is carried out using basically the same processes asplain ordinary deliberative thinking, perhaps with various tweaks optimizing them for thinkingabout thinking. If it turns out that humans have, say, a special kind of reasoning facultyexclusively for metacognition, then the diagram would need to be modified. Modeling of selfand others is understood to occur via a combination of metacognition and deliberative thinking,as well as via implicit adaptation based on reactive processing.Fig. 5.5: Architecture for Multimodal PerceptionFigure 5.5 models perception, according to the basic ideas of deep learning theory. Vision andaudition are modeled as deep learning hierarchies, with bottom-up and top-down dynamics. Thelower layers in each hierarchy refer to more localized patterns recognized in, and abstracted from,sensory data. Output from these hierarchies to the rest of the mind is not just through the toplayers, but via some sort of sampling from various layers, with a bias toward the top layers. Thedifferent hierarchies cross-connect, and are hence to an extent dynamically coupled together. Itis also recognized that there are some sensory modalities that aren’t strongly hierarchical, e.g102 5 A Generic Architecture of Human-Like Cognitiontouch and smell (the latter being better modeled as something like an asymmetric Hopfield net,prone to frequent chaotic dynamics [LLW + 05]) – these may also cross-connect with each otherand with the more hierarchical perceptual subnetworks. Of course the suggested architecturecould include any number of sensory modalities; the diagram is restricted to four just forsimplicity.The self-organized patterns in the upper layers of perceptual hierarchies may become quitecomplex and may develop advanced cognitive capabilities like episodic memory, reasoning, languagelearning, etc. A pure deep learning approach to intelligence argues that all the aspectsof intelligence emerge from this kind of dynamics (among perceptual, action and reinforcementhierarchies). Our own view is that the heterogeneity of human brain architecture argues againstthis perspective, and that deep learning systems are probably better as models of perceptionand action than of general cognition. However, the integrative diagram is not committed toour perspective on this – a deep-learning theorist could accept the integrative diagram, butargue that all the other portions besides the perceptual, action and reinforcement hierarchiesshould be viewed as descriptions of phenomena that emerge in these hierarchies due to theirinteraction.Fig. 5.6: Architecture for Action and ReinforcementFigure 5.6 shows an action subsystem and a reinforcement subsystem, parallel to the perceptionsubsystem. Two action hierarchies, one for an arm and one for a leg, are shown for5.3 An Architecture Diagram for Human-Like General Intelligence 103concreteness, but of course the architecture is intended to be extended more broadly. In thehierarchy corresponding to an arm, for example, the lowest level would contain control patternscorresponding to individual joints, the next level up to groupings of joints (like fingers), thenext level up to larger parts of the arm (hand, elbow). The different hierarchies correspondingto different body parts cross-link, enabling coordination among body parts; and they also connectat multiple levels to perception hierarchies, enabling sensorimotor coordination. Finallythere is a module for motor planning, which links tightly with all the motor hierarchies, andalso overlaps with the more cognitive, inferential planning activities of the mind, in a mannerthat is modeled different ways by different theorists. Albus [AM01] has elaborated this kind ofhierarchy quite elaborately.The reward hierarchy in Figure 5.6 provides reinforcement to actions at various levels onthe hierarchy, and includes dynamics for propagating information about reinforcement up anddown the hierarchy.Fig. 5.7: Architecture for Language ProcessingFigure 5.7 deals with language, treating it as a special case of coupled perception and action.The traditional architecture of a computational language comprehension system is a pipeline[JM09] [Goe10d], which is equivalent to a hierarchy with the lowest-level linguistic features (e.g.sounds, words) at the bottom, and the highest level features (semantic abstractions) at the top,and syntactic features in the middle. Feedback connections enable semantic and cognitive modulationof lower-level linguistic processing. Similarly, language generation is commonly modeledhierarchically, with the top levels being the ideas needing verbalization, and the bottom levelcorresponding to the actual sentence produced. In generation the primary flow is top-down,with bottom-up flow providing modulation of abstract concepts by linguistic surface forms.So, that’s it – an integrative architecture diagram for human-like general intelligence, splitamong seven different pictures, formed by judiciously merging together architecture diagramsproduced via a number of cognitive theorists with different, overlapping foci and researchparadigms.Is anything critical left out of the diagram? A quick perusal of the table of contents ofcognitive psychology textbooks suggests to me that if anything major is left out, it’s alsounknown to current cognitive psychology. However, one could certainly make an argument forexplicit inclusion of certain other aspects of intelligence, that in the integrative diagram are104 5 A Generic Architecture of Human-Like Cognitionleft as implicit emergent phenomena. For instance, creativity is obviously very important tointelligence, but, there is no "creativity" box in any of these diagrams – because in our view,and the view of the cognitive theorists whose work we’ve directly drawn on here, creativityis best viewed as a process emergent from other processes that are explicitly included in thediagrams.5.4 Interpretation and Application of the Integrative DiagramA tongue-partly-in-cheek definition of a biological pathway is "a subnetwork of a biologicalnetwork, that fits on a single journal page." Cognitive architecture diagrams have a similarproperty – they are crude abstractions of complex structures and dynamics, sculpted in accordancewith the size of the printed page, and the tolerance of the human eye for absorbingdiagrams, and the tolerance of the human author for making diagrams.However, sometimes constraints – even arbitrary ones – are useful for guiding creative efforts,due to the fact that they force choices. Creating an architecture for human-like generalintelligence that fits in a few (okay, seven) fairly compact diagrams, requires one to make manychoices about what features and relationships are most essential. In constructing the integrativediagram, we have sought to make these choices, not purely according to our own tastes in cognitivetheory or AGI system design, but according to a sort of blend of the taste and judgmentof a number of scientists whose views we respect, and who seem to have fairly compatible,complementary perspectives.What is the use of a cognitive architecture diagram like this? It can help to give newcomersto the field a basic idea about what is known and suspected about the nature of human-likegeneral intelligence. Also, it could potentially be used as a tool for cross-correlating differentAGI architectures. If everyone who authored an AGI architecture would explain how their architectureaccounts for each of the structures and processes identified in the integrative diagram,this would give a means of relating the various AGI designs to each other.The integrative diagram could also be used to help connect AGI and cognitive psychologyto neuroscience in a more systematic way. In the case of LIDA, a fairly careful correspondencehas been drawn up between the LIDA diagram nodes and links and various neural structuresand processes [FB08]. Similar knowledge exists for the rest of the integrative diagram, thoughnot organized in such a systematic fashion. A systematic curation of links between the nodesand links in the integrative diagram and current neuroscience knowledge, would constitute aninteresting first approximation of the holistic cognitive behavior of the human brain.Finally (and harking forward to later chapters), the big omission in the integrative diagramis dynamics. Structure alone will only get you so far, and you could build an AGI system withreasonable-looking things in each of the integrative diagram’s boxes, interrelating according tothe given arrows, and yet still fail to make a viable AGI system. Given the limitations thereal world places on computing resources, it’s not enough to have adequate representationsand algorithms in all the boxes, communicating together properly and capable doing the rightthings given sufficient resources. Rather, one needs to have all the boxes filled in properlywith structures and processes that, when they act together using feasible computing resources,will yield appropriately intelligent behaviors via their cooperative activity. And this has to dowith the complex interactive dynamics of all the processes in all the different boxes – which is5.4 Interpretation and Application of the Integrative Diagram 105something the integrative diagram doesn’t touch at all. This brings us again to the network ofideas we’ve discussed under the name of "cognitive synergy," to be discussed later on.It might be possible to make something similar to the integrative diagram on the level ofdynamics rather than structures, complementing the structural integrative diagram given here;but this would seem significantly more challenging, because we lack a standard set of tools fordepicting system dynamics. Most cognitive theorists and AGI architects describe their structuralideas using boxes-and-lines diagrams of some sort, but there is no standard method for depictingcomplex system dynamics. So to make a dynamical analogue to the integrative diagram, viaa similar integrative methodology, one would first need to create appropriate diagrammaticformalizations of the dynamics of the various cognitive theories being integrated – a fascinatingbut onerous task.When we first set out to make an integrated cognitive architecture diagram, via combiningthe complementary insights of various cognitive science and AGI theorists, we weren’t sure howwell it would work. But now we feel the experiment was generally a success – the resultantintegrated architecture seems sensible and coherent, and reasonably complete. It doesn’t comeclose to telling you everything you need to know to understand or implement a human-likemind – but it tells you the various processes and structures you need to deal with, and which oftheir interrelations are most critical. And, perhaps just as importantly, it gives a concrete wayof understanding the insights of a specific but fairly diverse set of cognitive science and AGItheorists as complementary rather than contradictory. In a CogPrime context, it provides away of tying in the specific structures and dynamics involved in CogPrime, with a more genericportrayal of the structures and dynamics of human-like intelligence.
Chapter 6A Brief Overview of CogPrime6.1 IntroductionJust as there are many different approaches to human flight – airplanes, helicopters, balloons,spacecraft, and doubtless many methods no person has thought of yet – similarly, there are likelymany different approaches to advanced artificial general intelligence. All the different approachesto flight exploit the same core principles of aerodynamics in different ways; and similarly, thevarious different approaches to AGI will exploit the same core principles of general intelligencein different ways.In the chapters leading up to this one, we have taken a fairly broad view of the projectof engineering AGI. We have presented a conception and formal model of intelligence, anddescribed environments, teaching methodologies and cognitive and developmental pathwaysthat we believe are collectively appropriate for the creation of AGI at the human level andultimately beyond, and with a roughly human-like bias to its intelligence. These ideas standalone and may be compatible with a variety of approaches to engineering AGI systems. However,they also set the stage for the presentation of CogPrime, the particular AGI design on whichwe are currently working.The thorough presentation of the CogPrime design is the job of Part 2 of this book – where,not only are the algorithms and structures involved in CogPrime reviewed in more detailed,but their relationship to the theoretical ideas underlying CogPrime is pursued more deeply.The job of this chapter is a smaller one: to give a high-level overview of some key aspects theCogPrime architecture at a mostly nontechnical level, so as to enable you to approach Part2 with a little more idea of what to expect. The remainder of Part 1, following this chapter,will present various theoretical notions enabling the particulars, intent and consequences of theCogPrime design to be more thoroughly understood.6.2 High-Level Architecture of CogPrimeFigures 6.1, 6.2 , 6.4 and 6.5 depict the high-level architecture of CogPrime, which involvesthe use of multiple cognitive processes associated with multiple types of memory to enablean intelligent agent to execute the procedures that it believes have the best probability ofworking toward its goals in its current context. In a robot preschool context, for example, the107108 6 A Brief Overview of CogPrimetop-level goals will be simple things such as pleasing the teacher, learning new informationand skills, and protecting the robot’s body. Figure 6.3 shows part of the architecture via whichcognitive processes interact with each other, via commonly acting on the AtomSpace knowledgerepository.Comparing these diagrams to the integrative human cognitive architecture diagrams givenin Chapter 5, one sees the main difference is that the CogPrime diagrams commit to specificstructures (e.g. knowledge representations) and processes, whereas the generic integrative architecturediagram refers merely to types of structures and processes. For instance, the integrativediagram refers generally to declarative knowledge and learning, whereas the CogPrime diagramrefers to PLN, as a specific system for reasoning and learning about declarative knowledge. Table6.1 articulates the key connections between the components of the CogPrime diagram andthose of the integrative diagram, thus indicating the general cognitive functions instantiated byeach of the CogPrime components.6.3 Current and Prior Applications of OpenCogBefore digging deeper into the theory, and elaborating some of the dynamics underlying theabove diagrams, we pause to briefly discuss some of the practicalities of work done with theOpenCog system currently implementing parts of the CogPrime architecture.OpenCog, the open-source software framework underlying the “OpenCogPrime” (currentlypartial) implementation of the CogPrime architecture, has been used for commercial applicationsin the area of natural language processing and data mining; for instance, see [GPPG06]where OpenCogPrime’s PLN reasoning and RelEx language processing are combined to doautomated biological hypothesis generation based on information gathered from PubMed abstracts.Most relevantly to the present work, it has also been used to control virtual agents invirtual worlds [GEA08].Prototype work done during 2007-2008 involved using an OpenCog variant called the Open-PetBrain to control virtual dogs in a virtual world (see Figure 6.6 for a screenshot of anOpenPetBrain-controlled virtual dog). While these OpenCog virtual dogs did not display intelligenceclosely comparable to that of real dogs (or human children), they did demonstrate avariety of interesting and relevant functionalities including:• learning new behaviors based on imitation and reinforcement• responding to natural language commands and questions, with appropriate actions andnatural language replies• spontaneous exploration of their world, remembering their experiences and using them tobias future learning and linguistic interactionOne current OpenCog initiative involves extending the virtual dog work via using OpenCogto control virtual agents in a game world inspired by the game Minecraft. These agents areinitially specifically concerned with achieving goals in a game world via constructing structureswith blocks and carrying out simple English communications. Representative example taskswould be:• Learning to build steps or ladders to get desired objects that are high up• Learning to build a shelter to protect itself from aggressors6.3 Current and Prior Applications of OpenCog 109Fig. 6.1: High-Level Architecture of CogPrime. This is a conceptual depiction, not adetailed flowchart (which would be too complex for a single image). Figures 6.2 , 6.4 and 6.5highlight specific aspects of this diagram.• Learning to build structures resembling structures that it’s shown (even if the availablematerials are a bit different)• Learning how to build bridges to cross chasmsOf course, the AI significance of learning tasks like this all depends on what kind of feedbackthe system is given, and how complex its environment is. It would be relatively simple to makean AI system do things like this in a trivial and highly specialized way, but that is not the intentof the project the goal is to have the system learn to carry out tasks like this using generallearning mechanisms and a general cognitive architecture, based on embodied experience and110 6 A Brief Overview of CogPrimeonly scant feedback from human teachers. If successful, this will provide an outstanding platformfor ongoing AGI development, as well as a visually appealing and immediately meaningful demofor OpenCog.Specific, particularly simple tasks that are the focus of this project team’s current work attime of writing include:• Watch another character build steps to reach a high-up object• Figure out via imitation of this that, in a different context, building steps to reach a highup object may be a good idea• Also figure out that, if it wants a certain high-up object but there are no materials forbuilding steps available, finding some other way to get elevated will be a good idea thatmay help it get the object6.3.1 Transitioning from Virtual Agents to a Physical RobotPreliminary experiments have also been conducted using OpenCog to control a Nao robot as wellas a virtual dog [GdG08]. This involves hybridizing OpenCog with a separate (but interlinked)subsystem handling low-level perception and action. In the experiments done so far, this hasbeen accomplished in an extremely simplistic way. How to do this right is a topic treated indetail in Chapter 26 of Part 2.We suspect that reasonable level of capability will be achievable by simply interposing DeS-TIN (or some other system in its place) as a perception/action “black box” between OpenCogand a robot. Some preliminary experiments in this direction have already been carried out, connectingthe OpenPetBrain to a Nao robot using simpler, less capable software than DeSTIN inthe intermediary role (off-the-shelf speech-to-text, text-to-speech and visual object recognitionsoftware).However, we also suspect that to achieve robustly intelligent robotics we must go beyond thisapproach, and connect robot perception and actuation software with OpenCogPrime in a “whitebox” manner that allows intimate dynamic feedback between perceptual, motoric, cognitiveand linguistic functions. We will achieve this via the creation and real-time utilization of linksbetween the nodes in CogPrime’s and DeSTIN’s internal networks (a topic to be explored inmore depth later in this chapter).6.4 Memory Types and Associated Cognitive Processes in CogPrimeNow we return to the basic description of the CogPrime approach, turning to aspects of therelationship between structure and dynamics. Architecture diagrams are all very well, but,ultimately it is dynamics that makes an architecture come alive. Intelligence is all about learning,which is by definition about change, about dynamical response to the environment and internalself-organizing dynamics.CogPrime relies on multiple memory types and, as discussed above, is founded on the premisethat the right course in architecting a pragmatic, roughly human-like AGI system is to handledifferent types of memory differently in terms of both structure and dynamics.6.4 Memory Types and Associated Cognitive Processes in CogPrime 111CogPrime’s memory types are the declarative, procedural, sensory, and episodic memorytypes that are widely discussed in cognitive neuroscience [TC05], plus attentional memory forallocating system resources generically, and intentional memory for allocating system resourcesin a goal-directed way. Table 6.2 overviews these memory types, giving key references and indicatingthe corresponding cognitive processes, and also indicating which of the generic patternistcognitive dynamics each cognitive process corresponds to (pattern creation, association, etc.).Figure 6.7 illustrates the relationships between several of the key memory types in the contextof a simple situation involving an OpenCogPrime-controlled agent in a virtual world.In terms of patternist cognitive theory, the multiple types of memory in CogPrime should beconsidered as specialized ways of storing particular types of patterns, optimized for spacetimeefficiency. The cognitive processes associated with a certain type of memory deal with creatingand recognizing patterns of the type for which the memory is specialized. While in principle allthe different sorts of pattern could be handled in a unified memory and processing architecture,the sort of specialization used in CogPrime is necessary in order to achieve acceptable efficientgeneral intelligence using currently available computational resources. And as we have arguedin detail in Chapter 7, efficiency is not a side-issue but rather the essence of real-world AGI(since as Hutter has shown, if one casts efficiency aside, arbitrary levels of general intelligencecan be achieved via a trivially simple program).The essence of the CogPrime design lies in the way the structures and processes associatedwith each type of memory are designed to work together in a closely coupled way, yielding cooperativeintelligence going beyond what could be achieved by an architecture merely containingthe same structures and processes in separate “black boxes.”The inter-cognitive-process interactions in OpenCog are designed so that• conversion between different types of memory is possible, though sometimes computationallycostly (e.g. an item of declarative knowledge may with some effort be interpretedprocedurally or episodically, etc.)• when a learning process concerned centrally with one type of memory encounters a situationwhere it learns very slowly, it can often resolve the issue by converting some of the relevantknowledge into a different type of memory: i.e. cognitive synergy6.4.1 Cognitive Synergy in PLNTo put a little meat on the bones of the "cognitive synergy" idea, discussed repeatedly in priorchapters and more extensively in latter chapters, we now elaborate a little on the role it playsin the interaction between procedural and declarative learning.While MOSES handles much of CogPrime’s procedural learning, and CogPrime’s internalsimulation engine handles most episodic knowledge, CogPrime’s primary tool for handlingdeclarative knowledge is an uncertain inference framework called Probabilistic Logic Networks(PLN). The complexities of PLN are the topic of a lengthy technical monograph [GMIH08], andare summarized in Chapter 34; here we will eschew most details and focus mainly on pointingout how PLN seeks to achieve efficient inference control via integration with other cognitiveprocesses.As a logic, PLN is broadly integrative: it combines certain term logic rules with more standardpredicate logic rules, and utilizes both fuzzy truth values and a variant of imprecise probabilitiescalled indefinite probabilities. PLN mathematics tells how these uncertain truth values propagate112 6 A Brief Overview of CogPrimethrough its logic rules, so that uncertain premises give rise to conclusions with reasonablyaccurately estimated uncertainty values. This careful management of uncertainty is critical forthe application of logical inference in the robotics context, where most knowledge is abstractedfrom experience and is hence highly uncertain.PLN can be used in either forward or backward chaining mode; and in the language introducedabove, it can be used for either analysis or synthesis. As an example, we will considerbackward chaining analysis, exemplified by the problem of a robot preschool-student trying todetermine whether a new playmate “Bob” is likely to be a regular visitor to is preschool or not(evaluating the truth value of the implication Bob → regular_visitor). The basic backwardchaining process for PLN analysis looks like:1. Given an implication L ≡ A → B whose truth value must be estimated (for instanceL ≡ Concept ∧ Procedure → Goal as discussed above), create a list (A 1 , ..., A n ) of (inferencerule, stored knowledge) pairs that might be used to produce L2. Using analogical reasoning to prior inferences, assign each A i a probability of success• If some of the A i are estimated to have reasonable probability of success at generatingreasonably confident estimates of L’s truth value, then invoke Step 1 with A i in placeof L (at this point the inference process becomes recursive)• If none of the A i looks sufficiently likely to succeed, then inference has “gotten stuck”and another cognitive process should be invoked, e.g.– Concept creation may be used to infer new concepts related to A and B, and thenStep 1 may be revisited, in the hope of finding a new, more promising A i involvingone of the new concepts– MOSES may be invoked with one of several special goals, e.g. the goal of findinga procedure P so that P (X) predicts whether X → B. If MOSES finds such aprocedure P then this can be converted to declarative knowledge understandableby PLN and Step 1 may be revisited....– Simulations may be run in CogPrime’s internal simulation engine, so as to observethe truth value of A → B in the simulations; and then Step 1 may be revisited....The combinatorial explosion of inference control is combatted by the capability to defer toother cognitive processes when the inference control procedure is unable to make a sufficientlyconfident choice of which inference steps to take next. Note that just as MOSES may relyon PLN to model its evolving populations of procedures, PLN may rely on MOSES to createcomplex knowledge about the terms in its logical implications. This is just one example of themultiple ways in which the different cognitive processes in CogPrime interact synergetically; amore thorough treatment of these interactions is given in [Goe09a].In the “new playmate” example, the interesting case is where the robot initially seems notto know enough about Bob to make a solid inferential judgment (so that none of the A i seemparticularly promising). For instance, it might carry out a number of possible inferences and notcome to any reasonably confident conclusion, so that the reason none of the A i seem promisingis that all the decent-looking ones have been tried already. So it might then recourse to MOSES,simulation or concept creation.For instance, the PLN controller could make a list of everyone who has been a regularvisitor, and everyone who has not been, and pose MOSES the task of figuring out a procedurefor distinguishing these two categories. This procedure could then be used directly to make theneeded assessment, or else be translated into logical rules to be used within PLN inference. For6.5 Goal-Oriented Dynamics in CogPrime 113example, perhaps MOSES would discover that older males wearing ties tend not to becomeregular visitors. If the new playmate is an older male wearing a tie, this is directly applicable.But if the current playmate is wearing a tuxedo, then PLN may be helpful via reasoning thateven though a tuxedo is not a tie, it’s a similar form of fancy dress – so PLN may extend theMOSES-learned rule to the present case and infer that the new playmate is not likely to be aregular visitor.6.5 Goal-Oriented Dynamics in CogPrimeCogPrime’s dynamics has both goal-oriented and “spontaneous” aspects; here for simplicity’ssake we will focus on the goal-oriented ones. The basic goal-oriented dynamic of the CogPrimesystem, within which the various types of memory are utilized, is driven by implications knownas “cognitive schematics”, which take the formContext ∧ P rocedure → Goal < p >(summarized C ∧ P → G). Semi-formally, this implication may be interpreted to mean: “If thecontext C appears to hold currently, then if I enact the procedure P , I can expect to achieve thegoal G with certainty p.” Cognitive synergy means that the learning processes corresponding tothe different types of memory actively cooperate in figuring out what procedures will achievethe system’s goals in the relevant contexts within its environment.CogPrime’s cognitive schematic is significantly similar to production rules in classical architectureslike SOAR and ACT-R (as reviewed in Chapter 4; however, there are significantdifferences which are important to CogPrime’s functionality. Unlike with classical productionrules systems, uncertainty is core to CogPrime’s knowledge representation, and each CogPrimecognitive schematic is labeled with an uncertain truth value, which is critical to its utilization byCogPrime’s cognitive processes. Also, in CogPrime, cognitive schematics may be incomplete,missing one or two of the terms, which may then be filled in by various cognitive processes(generally in an uncertain way). A stronger similarity is to MicroPsi’s triplets; the differencesin this case are more low-level and technical and have already been mentioned in Chapter 4.Finally, the biggest difference between CogPrime’s cognitive schematics and production rulesor other similar constructs, is that in CogPrime this level of knowledge representation is notthe only important one. CLARION [SZ04], as reviewed above, is an example of a cognitivearchitecture that uses production rules for explicit knowledge representation and then uses atotally separate subsymbolic knowledge store for implicit knowledge. In CogPrimeboth explicit and implicit knowledge are stored in the same graph of nodes and links, with• explicit knowledge stored in probabilistic logic based nodes and links such as cognitiveschematics (see Figure 6.8 for a depiction of some explicit linguistic knowledge.)• implicit knowledge stored in patterns of activity among these same nodes and links, definedvia the activity of the “importance” values (see Figure 6.9 for an illustrative example thereof)associated with nodes and links and propagated by the ECAN attention allocation processThe meaning of a cognitive schematic in CogPrime is hence not entirely encapsulated in itsexplicit logical form, but resides largely in the activity patterns that ECAN causes its activationor exploration to give rise to. And this fact is important because the synergetic interactionsof system components are in large part modulated by ECAN activity. Without the real-time114 6 A Brief Overview of CogPrimecombination of explicit and implicit knowledge in the system’s knowledge graph, the synergeticinteraction of different cognitive processes would not work so smoothly, and the emergence ofeffective high-level hierarchical, heterarchical and self structures would be less likely.6.6 Analysis and Synthesis Processes in CogPrimeWe now return to CogPrime’s fundamental cognitive dynamics, using examples from the “virtualdog” application to motivate the discussion.The cognitive schematic Context ∧ Procedure → Goal leads to a conceptualization of theinternal action of an intelligent system as involving two key categories of learning:• Analysis: Estimating the probability p of a posited C ∧ P → G relationship• Synthesis: Filling in one or two of the variables in the cognitive schematic, given assumptionsregarding the remaining variables, and directed by the goal of maximizing theprobability of the cognitive schematicMore specifically, where synthesis is concerned,• The MOSES probabilistic evolutionary program learning algorithm is applied to find P ,given fixed C and G. Internal simulation is also used, for the purpose of creating a simulationembodying C and seeing which P lead to the simulated achievement of G.– Example: A virtual dog learns a procedure P to please its owner (the goal G) in thecontext C where there is a ball or stick present and the owner is saying “fetch”.• PLN inference, acting on declarative knowledge, is used for choosing C, given fixed P andG (also incorporating sensory and episodic knowledge as appropriate). Simulation may alsobe used for this purpose.– Example: A virtual dog wants to achieve the goal G of getting food, and it knows thatthe procedure P of begging has been successful at this before, so it seeks a context Cwhere begging can be expected to get it food. Probably this will be a context involving afriendly person.• PLN-based goal refinement is used to create new subgoals G to sit on the right hand sideof instances of the cognitive schematic.– Example: Given that a virtual dog has a goal of finding food, it may learn a subgoal offollowing other dogs, due to observing that other dogs are often heading toward theirfood.• Concept formation heuristics are used for choosing G and for fueling goal refinement, butespecially for choosing C (via providing new candidates for C). They are also used forchoosing P , via a process called “predicate schematization” that turns logical predicates(declarative knowledge) into procedures.– Example: At first a virtual dog may have a hard time predicting which other dogs aregoing to be mean to it. But it may eventually observe common features among a numberof mean dogs, and thus form its own concept of “pit bull,” without anyone ever teachingit this concept explicitly.6.6 Analysis and Synthesis Processes in CogPrime 115Where analysis is concerned:• PLN inference, acting on declarative knowledge, is used for estimating the probability ofthe implication in the cognitive schematic, given fixed C, P and G. Episodic knowledgeis also used in this regard, via enabling estimation of the probability via simple similaritymatching against past experience. Simulation is also used: multiple simulations may be run,and statistics may be captured therefrom.– Example: To estimate the degree to which asking Bob for food (the procedure P is “askingfor food”, the context C is “being with Bob”) will achieve the goal G of getting food, thevirtual dog may study its memory to see what happened on previous occasions where itor other dogs asked Bob for food or other things, and then integrate the evidence fromthese occasions.• Procedural knowledge, mapped into declarative knowledge and then acted on by PLN inference,can be useful for estimating the probability of the implication C ∧ P → G, in caseswhere the probability of C ∧ P 1 → G is known for some P 1 related to P .– Example: knowledge of the internal similarity between the procedure of asking for foodand the procedure of asking for toys, allows the virtual dog to reason that if asking Bobfor toys has been successful, maybe asking Bob for food will be successful too.• Inference, acting on declarative or sensory knowledge, can be useful for estimating theprobability of the implication C ∧ P → G, in cases where the probability of C 1 ∧ P → G isknown for some C 1 related to C.– Example: if Bob and Jim have a lot of features in common, and Bob often respondspositively when asked for food, then maybe Jim will too.• Inference can be used similarly for estimating the probability of the implication C ∧P → G,in cases where the probability of C ∧ P → G 1 is known for some G 1 related to G. Conceptcreation can be useful indirectly in calculating these probability estimates, via providingnew concepts that can be used to make useful inference trails more compact and henceeasier to construct.– Example: The dog may reason that because Jack likes to play, and Jack and Jill are bothchildren, maybe Jill likes to play too. It can carry out this reasoning only if its conceptcreation process has invented the concept of “child” via analysis of observed data.In these examples we have focused on cases where two terms in the cognitive schematic arefixed and the third must be filled in; but just as often, the situation is that only one of theterms is fixed. For instance, if we fix G, sometimes the best approach will be to collectivelylearn C and P . This requires either a procedure learning method that works interactively with adeclarative-knowledge-focused concept learning or reasoning method; or a declarative learningmethod that works interactively with a procedure learning method. That is, it requires the sortof cognitive synergy built into the CogPrime design.116 6 A Brief Overview of CogPrime6.7 ConclusionTo thoroughly describe a comprehensive, integrative AGI architecture in a brief chapter wouldbe an impossible task; all we have attempted here is a brief overview, to be elaborated on inthe 800-odd pages of Part 2 of this book. We do not expect this brief summary to be enough toconvince the skeptical reader that the approach described here has a reasonable odds of successat achieving its stated goals, or even of fulfilling the conceptual notions outlined in the precedingchapters. However, we hope to have given the reader at least a rough idea of what sort of AGIdesign we are advocating, and why and in what sense we believe it can lead to advanced artificialgeneral intelligence. For more details on the structure, dynamics and underlying concepts ofCogPrime, the reader is encouraged to proceed to Part 2– after completing Part 1, of course.Please be patient – building a thinking machine is a big topic, and we have a lot to say aboutit!6.7 Conclusion 117Fig. 6.2: Key Explicitly Implemented Processes of CogPrime . The large box at thecenter is the Atomspace, the system’s central store of various forms of (long-term and working)memory, which contains a weighted labeled hypergraph whose nodes and links are "Atoms" ofvarious sorts. The hexagonal boxes at the bottom denote various hierarchies devoted to recognitionand generation of patterns: perception, action and linguistic. Intervening between theserecognition/generation hierarchies and the Atomspace, we have a pattern mining/imprintingcomponent (that recognizes patterns in the hierarchies and passes them to the Atomspace; andimprints patterns from the Atomspace on the hierarchies); and also OpenPsi, a special dynamicalframework for choosing actions based on motivations. Above the Atomspace we have ahost of cognitive processes, which act on the Atomspace, some continually and some only ascontext dictates, carrying out various sorts of learning and reasoning (pertinent to various sortsof memory) that help the system fulfill its goals and motivations.118 6 A Brief Overview of CogPrimeFig. 6.3: MindAgents and AtomSpace in OpenCog. This is a conceptual depiction ofone way cognitive processes may interact in OpenCog – they may be wrapped in MindAgentobjects, which interact via cooperatively acting on the AtomSpace.6.7 Conclusion 119Fig. 6.4: Links Between Cognitive Processes and the Atomspace. The cognitive processesdepicted all act on the Atomspace, in the sense that they operate by observing certainAtoms in the Atomspace and then modifying (or in rare cases deleting) them, and potentiallyadding new Atoms as well. Atoms represent all forms of knowledge, but some forms of knowledgeare additionally represented by external data stores connected to the Atomspace, such asthe Procedure Repository; these are also shown as linked to the Atomspace.120 6 A Brief Overview of CogPrimeFig. 6.5: Invocation of Atom Operations By Cognitive Processes. This diagram depictssome of the Atom modification, creation and deletion operations carried out by the abstractcognitive processes in the CogPrime architecture.6.7 Conclusion 121CogPrimeComponentInt. Diag.Sub-DiagramInt. Diag. ComponentProcedure Repository Long-Term Memory ProceduralProcedure Repository Working Memory Active ProceduralAssociative EpisodicMemoryLong-Term MemoryEpisodicAssociative EpisodicMemoryWorking Memory Transient EpisodicBackup Store Long-Term Memoryno correlate: a function notnecessarily possessed by the humanmindSpacetime Server Long-Term Memory Declarative and SensorimotorDimensionalEmbedding Spaceno clear correlate: atool for helpingmultiple types of LTMDimensionalEmbedding Agentno clear correlateBlendingLong-Term andWorking MemoryConcept FormationClusteringLong-Term andWorking MemoryConcept FormationPLN ProbabilisticInferenceMOSES / HillclimbingWorld SimulationEpisodic Encoding /RecallEpisodic Encoding /RecallForgetting / Freezing/ DefrostingMap FormationLong-Term andWorking MemoryLong-Term andWorking MemoryLong-Term andWorking MemoryLong-Term g MemoryWorking MemoryLong-Term andWorking MemoryLong-Term MemoryReasoning and PlanLearning/OptimizationProcedure LearningSimulationStory-tellingConsolidationno correlate: a function notnecessarily possessed by the humanmindConcept Formation and PatternMiningAttention AllocationLong-Term andWorking MemoryHebbian/Attentional LearningAttention AllocationHigh-Level MindArchitectureReinforcementAttention Allocation Working MemoryPerceptual Associative Memory andLocal AssociationAtomSpaceHigh-Level MindArchitectureno clear correlate: a general tool forrepresenting memory includinglong-term and working, plus some ofperception and actionAtomSpace Working MemoryGlobal Workspace (the high-STIportion of AtomSpace) & otherWorkspacesDeclarative AtomsLong-Term andWorking MemoryDeclarative and SensorimotorProcedure AtomsLong-Term andWorking MemoryProceduralHebbian AtomsLong-Term andWorking MemoryAttentionalGoal AtomsLong-Term andWorking MemoryIntentionalFeeling AtomsLong-Term andWorking Memoryspanning Declarative, Intentional andSensorimotorOpenPsiHigh-Level MindArchitectureMotivation / Action SelectionOpenPsi Working Memory Action SelectionPattern MinerHigh-Level MindArchitecturearrows between perception andworking and long-term memoryPattern Miner Working Memoryarrows between sensory memory andperceptual associative and transientepisodic memoryarrows between action selection and122 6 A Brief Overview of CogPrimeFig. 6.6: Screenshot of OpenCog-controlled virtual dogFig. 6.7: Relationship Between Multiple Memory Types. The bottom left corner showsa program tree, constituting procedural knowledge. The upper left shows declarative nodes andlinks in the Atomspace. The upper right corner shows a relevant system goal. The lower rightcorner contains an image symbolizing relevant episodic and sensory knowledge. All the varioustypes of knowledge link to each other and can be approximatively converted to each other.6.7 Conclusion 123Memory TypeDeclarativeProceduralEpisodicAttentionalIntentionalSensorySpecific Cognitive ProcessesProbabilistic Logic Networks (PLN)[GMIH08]; conceptual blending[FT02]MOSES (a novel probabilisticevolutionary program learningalgorithm) [Loo06]internal simulation engine [GEA08]Economic Attention Networks(ECAN) [GPI + 10]probabilistic goal hierarchy refined byPLN and ECAN, structuredaccording to MicroPsi [Bac09]In CogBot, this will be supplied bythe DeSTIN componentGeneral CognitiveFunctionspattern creationpattern creationassociation, patterncreationassociation, creditassignmentcredit assignment,pattern creationassociation, attentionallocation, patterncreation, creditassignmentTable 6.2: Memory Types and Cognitive Processes in CogPrime. The third column indicatesthe general cognitive function that each specific cognitive process carries out, according to thepatternist theory of cognition.124 6 A Brief Overview of CogPrimeFig. 6.8: Example of Explicit Knowledge in the Atomspace. One simple example ofexplicitly represented knowledge in the Atomspace is linguistic knowledge, such as words andthe concepts directly linked to them. Not all of a CogPrime system’s concepts correlate towords, but some do.6.7 Conclusion 125Fig. 6.9: Example of Implicit Knowledge in the Atomspace. A simple example of implicitknowledge in the Atomspace. The "chicken" and "food" concepts are represented by "maps"of ConceptNodes interconnected by HebbianLinks, where the latter tend to form between ConceptNodesthat are often simultaneously important. The bundle of links between nodes in thechicken map and nodes in the food map, represents an "implicit, emergent link" between thetwo concept maps. This diagram also illustrates "glocal" knowledge representation, in that thechicken and food concepts are each represented by individual nodes, but also by distributedmaps. The "chicken" ConceptNode, when important, will tend to make the rest of the mapimportant – and vice versa. Part of the overall chicken concept possessed by the system is expressedby the explicit links coming out of the chicken ConceptNode, and part is representedonly by the distributed chicken map as a whole.
Section IIToward a General Theory of General Intelligence
Chapter 7A Formal Model of Intelligent Agents7.1 IntroductionThe artificial intelligence field is full of sophisticated mathematical models and equations, butmost of these are highly specialized in nature – e.g. formalizations of particular logic systems,analyzes of the dynamics of specific sorts of neural nets, etc. On the other hand, a number ofhighly general models of intelligent systems also exist, including Hutter’s recent formalizationof universal intelligence [Hut05] and a large body of work in the disciplines of systems scienceand cybernetics – but these have tended not to yield many specific lessons useful for engineeringAGI systems, serving more as conceptual models in mathematical form.It would be fantastic to have a mathematical theory bridging these extremes – a real "generaltheory of general intelligence," allowing the derivation and analysis of specific structures andprocesses playing a role in practical AGI systems, from broad mathematical models of generalintelligence in various situations and under various constraints. However, the path to such atheory is not entirely clear at present; and, as valuable as such a theory would be, we don’tbelieve such a thing to be necessary for creating advanced AGI. One possibility is that thedevelopment of such a theory will occur contemporaneously and synergetically with the adventof practical AGI technology.Lacking a mature, pragmatically useful "general theory of general intelligence," however, wehave still found it valuable to articulate certain theoretical ideas about the nature of generalintelligence, with a level of rigor a bit greater than the wholly informal discussions of the previouschapters. The chapters in this section of the book articulate some ideas we have developed inpursuit of a general theory of general intelligence; ideas that, even in their current relativelyundeveloped form, have been very helpful in guiding our concrete work on the CogPrime design.This chapter presents a more formal version of the notion of intelligence as “achieving complexgoals in complex environments,” based on a formal model of intelligent agents. These formalizationsof agents and intelligence will be used in later chapters as a foundation for formalizingother concepts like inference and cognitive synergy. Chapters 8 and 9 pursue the notion of cognitivesynergy a little more thoroughly than was done in previous chapters. Chapter 10 sketchesa general theory of general intelligence using tools from category theory – not bringing it to thelevel where one can use it to derive specific AGI algorithms and structures; but still, presentingideas that will be helpful in interpreting and explaining specific aspects of the CogPrime designin Part 2. Finally, Appendix ?? explores an additional theoretical direction, in which the mindof an intelligent system is viewed in terms of certain curved spaces – a novel way of thinking129130 7 A Formal Model of Intelligent Agentsabout the dynamics of general intelligence, which has been useful in guiding development of theECAN component of CogPrime, and we expect will have more general value in future.Despite the intermittent use of mathematical formalism, the ideas presented in this sectionare fairly speculative, and we do not propose them as constituting a well-demonstrated theoryof general intelligence. Rather, we propose them as an interesting way of thinking about generalintelligence, which appears to be consistent with available data, and which has proved inspirationalto us in conceiving concrete structures and dynamics for AGI, as manifested for examplein the CogPrime design. Understanding the way of thinking described in these chapters is valuablefor understanding why the CogPrime design is the way it is, and for relating CogPrime toother practical and intellectual systems, and extending and improving CogPrime.7.2 A Simple Formal Agents Model (SRAM)We now present a formalization of the concept of “intelligent agents” – beginning with a formalizationof “agents” in general.Drawing on [Hut05, LH07a], we consider a class of active agents which observe and exploretheir environment and also take actions in it, which may affect the environment. Formally,the agent sends information to the environment by sending symbols from some finite alphabetcalled the action space Σ; and the environment sends signals to the agent with symbols froman alphabet called the perception space, denoted P. Agents can also experience rewards, whichlie in the reward space, denoted R, which for each agent is a subset of the rational unit interval.The agent and environment are understood to take turns sending signals back and forth,yielding a history of actions, observations and rewards, which may be denotedor elsea 1 o 1 r 1 a 2 o 2 r 2 ...a 1 x 1 a 2 x 2 ...if x is introduced as a single symbol to denote both an observation and a reward. Thecomplete interaction history up to and including cycle t is denoted ax 1:t ; and the history beforecycle t is denoted ax <t = ax 1:t−1 .The agent is represented as a function π which takes the current history as input, and producesan action as output. Agents need not be deterministic, an agent may for instance induce aprobability distribution over the space of possible actions, conditioned on the current history. Inthis case we may characterize the agent by a probability distribution π(a t |ax <t ). Similarly, theenvironment may be characterized by a probability distribution µ(x k |ax <k a k ). Taken together,the distributions π and µ define a probability measure over the space of interaction sequences.Next, we extend this model in a few ways, intended to make it better reflect the realities ofintelligent computational agents. The first modification is to allow agents to maintain memories(of finite size), via adding memory actions drawn from a set M into the history of actions,observations and rewards. The second modification is to introduce the notion of goals.7.2 A Simple Formal Agents Model (SRAM) 1317.2.1 GoalsWe define goals as mathematical functions (to be specified below) associated with symbolsdrawn from the alphabet G; and we consider the environment as sending goal-symbols to theagent along with regular observation-symbols. (Note however that the presentation of a goalsymbolto an agent does not necessarily entail the explicit communication to the agent of thecontents of the goal function. This must be provided by other, correlated observations.) We alsointroduce a conditional distribution γ(g, µ) that gives the weight of a goal g in the context ofa particular environment µ.In this extended framework, an interaction sequence looks likeor elsea 1 o 1 g 1 r 1 a 2 o 2 g 2 r 2 ...a 1 y 1 a 2 y 2 ...where g i are symbols corresponding to goals, and y is introduced as a single symbol to denotethe combination of an observation, a reward and a goal.Each goal function maps each finite interaction sequence I g,s,t = ay s:t with g s to g t correspondingto g, into a value r g (I g,s,t ) ∈ [0, 1] indicating the value or “raw reward” of achievingthe goal during that interaction sequence. The total reward r t obtained by the agent is the sumof the raw rewards obtained at time t from all goals whose symbols occur in the agent’s historybefore t.This formalism of goal-seeking agents allows us to formalize the notion of intelligence as“achieving complex goals in complex environments” – a direction that is pursued in Section 7.3below.Note that this is an external perspective of system goals, which is natural from the perspectiveof formally defining system intelligence in terms of system behavior, but is not necessarily verynatural in terms of system design. From the point of view of AGI design, one is generally moreconcerned with the (implicit or explicit) representation of goals inside an AGI system, as inCogPrime’s Goal Atoms to be reviewed in Chapter 22 below.Further, it is important to also consider the case where an AGI system has no explicit goals,and the system’s environment has no immediately identifiable goals either. But in this case, wedon’t see any clear way to define a system’s intelligence, except via approximating the system interms of other theoretical systems which do have explicit goals. This approximation approachis developed in Section 7.3.5 below.The awkwardness of linking the general formalism of intelligence theory presented here, withthe practical business of creating and designing AGI systems, may indicate a shortcoming onthe part of contemporary intelligence theory or AGI designs. On the other hand, this sort ofsituation often occurs in other domains as well – e.g. the leap from quantum theory to theanalysis of real-world systems like organic molecules involves a lot of awkwardness and largeleaps a well.132 7 A Formal Model of Intelligent Agents7.2.2 Memory StoresAs well as goals, we introduce into the model a long-term memory and a workspace. Regardinglong-term memory we assume the agent’s memory consists of multiple memory stores correspondingto various types of memory, e.g.: procedural (K P roc ), declarative (K Dec ), episodic(K Ep ), attentional (K Att ) and Intentional (K Int ). In Appendix ?? a category-theoretic modelof these memory stores is introduced; but for the moment, we need only assume the existenceof• an injective mapping Θ Ep : K Ep → H where H is the space of fuzzy sets of subhistories(subhistories being “episodes” in this formalism)• an injective mapping Θ P roc : K P roc × M × W → A, where M is the set of memory states,W is the set of (observation, goal, reward) triples, and A is the set of actions (this mapseach procedure object into a function that enacts actions in the environment or memory,based on the memory state and current world-state)• an injective mapping Θ Dec : K Dec → L, where L is the set of expressions in some formal language(which may for example be a logical language), which possesses words correspondingto the observations, goals, reward values and actions in our agent formalism• an injective mapping Θ Int : K Int → G, where G is the space of goals mentioned above• an injective mapping Θ Att : K Int ∪ K Ep ∪ K P roc ∪ K Ec → V, where V is the space of“attention values” (structures that gauge the importance of paying attention to an item ofknowledge over various time-scales or in various contexts)We also assume that the vocabulary of actions contains memory-actions corresponding to theoperations of inserting the current observation, goal, reward or action into the episodic and/ordeclarative memory store. And, we assume that the activity of the agent, at each time-step,includes the enaction of one or more of the procedures in the procedural memory store. If severalprocedures are enacted at once, then the end result is still formally modeled as a single actiona = a [1] ∗ ... ∗ a [k] where ∗ is an operator on action-space that composes multiple actions into asingle one.Finally, we assume that, at each time-step, the agent may carry out an external action a ion the environment, a memory action m i on the (long-term) memory, and an action b i on itsinternal workspace. Among the actions that can be carried out on the workspace, are theability to insert or delete observations, goals, actions or reward-values from the workspace.The workspace can be thought of as a sort of short-term memory or else in terms of Baars’“global workspace” concept mentioned above. The workspace provides a medium for interactionbetween the different memory types.The workspace provides a mechanism by which declarative, episodic and procedural memorymay interact with each other. For this mechanism to work, we must assume that there areactions corresponding to query operations that allow procedures to look into declarative andepisodic memory. The nature of these query operations will vary among different agents, butwe can assume that in general an agent has• one or more procedures Q Dec (x) serving as declarative queries, meaning that when Q Dec isenacted on some x that is an ordered set of items in the workspace, the result is that oneor more items from declarative memory is entered into the workspace• one or more procedures Q Ep (x) serving as episodic queries, meaning that when Q Ep isenacted on some x that is an ordered set of items in the workspace, the result is that oneor more items from episodic memory is entered into the workspace7.2 A Simple Formal Agents Model (SRAM) 133One additional aspect of CogPrime’s knowledge representation that is important to PLN isthe attachment of nonnegative weights n i corresponding to elementary observations o i . Theseweights denote the amount of evidence contained in the observation. For instance, in the contextof a robotic agent, one could use these values to encode the assumption that an elementary visualobservation has more evidential value than an elementary olfactory observation.We now have a model of an agent with long-term memory comprising procedural, declarativeand episodic aspects, an internal cognitive workspace, and the capability to use procedures todrive actions based on items in memory and the workspace, and to move items between longtermmemory and the workspace.7.2.2.1 Modeling CogPrimeOf course, this formal model may be realized differently in various real-world AGI systems. InCogPrime we have• a weighted, labeled hypergraph structure called the AtomSpace used to store declarativeknowledge (this is the representation used by PLN)• a collection of programs in a LISP-like language called Combo, stored in a ProcedureRepositorydata structure, used to store procedural knowledge• a collection of partial “movies” of the system’s experience, played back using an internalsimulation engine, used to store episodic knowledge• AttentionValue objects, minimally containing ShortTermImportance (STI) and LongTermImportance(LTI) values used to store attentional knowledge• Goal Atoms for intentional knowledge, stored in the same format as declarative knowledgebut whose dynamics involve a special form of artificial currency that is used to govern actionselectionThe AtomSpace is the central repository and procedures and episodes are linked to Atomsin the AtomSpace which serve as their symbolic representatives. The “workspace” in CogPrimeexists only virtually: each item in the AtomSpace has a “short term importance” (STI) level, andthe workspace consists of those items in the AtomSpace with highest STI, and those proceduresand episodes whose symbolic representatives in the AtomSpace have highest STI.On the other hand, as we saw above, the LIDA architecture uses separate representations forprocedural, declarative and episodic memory, but also has an explicit workspace component,where the most currently contextually relevant items from all different types of memory aregathered and used together in the course of actions. However, compared to CogPrime, it lackscomparably fine-grained methods for integrating the different types of memory.Systematically mapping various existing cognitive architectures, or human brain structure,into this formal agents model would be a substantial though quite plausible exercise; but wewill not undertake this here.7.2.3 The Cognitive SchematicNext we introduce an additional specialization into SRAM: the cognitive schematic, writteninformally as134 7 A Formal Model of Intelligent AgentsContext & P rocedure → Goaland considered more formally as holds(C) & ex(P ) → h i where h may be an externally specifiedgoal g i or an internally specified goal h derived as a (possibly uncertain) subgoal of one of moreg i ; C is a piece of declarative or episodic knowledge and P is a procedure that the agent caninternally execute to generate a series of actions. ex(P ) is the proposition that P is successfullyexecuted. If C is episodic then holds(C) may be interpreted as the current context (i.e. somefinite slice of the agent’s history) being similar to C; if C is declarative then holds(C) may beinterpreted as the truth value of C evaluated at the current context. Note that C may refer tosome part of the world quite distant from the agent’s current sensory observations; but it maystill be formally evaluated based on the agent’s history.In the standard CogPrime notation as introduced formally in Chapter 20 (where indentationhas function-argument syntax similar to that in Python, and relationship types are prependedto their relata without parentheses), for the case C is declarative this would be written asPredictiveExtensionalImplicationANDCExecution PGand in the case C is episodic one replaces C in this formula with a predicate expressing C’ssimilarity to the current context. The semantics of the PredictiveExtensionalInheritance relationwill be discussed below. The Execution relation simply denotes the proposition that procedureP has been executed.For the class of SRAM agents who (like CogPrime) use the cognitive schematic to governmany or all of their actions, a significant fragment of agent intelligence boils down to estimatingthe truth values of PredictiveExtensionalImplication relationships. Action selection procedurescan be used, which choose procedures to enact based on which ones are judged most likelyto achieve the current external goals g i in the current context. Rather than enter into theparticularities of action selection or other cognitive architecture issues, we will restrict ourselvesto PLN inference, which in the context of the present agent model is a method for handlingPredictiveImplication in the cognitive schematic.Consider an agent in a virtual world, such as a virtual dog, one of whose external goals is toplease its owner. Suppose its owner has asked it to find a cat, and it can translate this into asubgoal “find cat.” If the agent operates according to the cognitive schematic, it will search forP so thatPredictiveExtensionalImplicationANDCExecution PEvaluationfoundcatholds.7.3 Toward a Formal Characterization of Real-World General Intelligence 1357.3 Toward a Formal Characterization of Real-World GeneralIntelligenceHaving defined what we mean by an agent acting in an environment, we now turn to thequestion of what it means for such an agent to be “intelligent.”As we have reviewed extensively in Chapter 2 above, “intelligence” is a commonsense, “folkpsychology” concept, with all the imprecision and contextuality that this generally entails.One cannot expect any compact, elegant formalism to capture all of its meanings. Even inthe psychology and AI research communities, divergent definitions abound; Legg and Hutter[LH07a] lists and organizes 70+ definitions from the literature.Practical study of natural intelligence in humans and other organisms, and practical design,creation and instruction of artificial intelligences, can proceed perfectly well without anagreed-upon formalization of the “intelligence” concept. Some researchers may conceive theirown formalisms to guide their own work, others may feel no need for any such thing.But nevertheless, it is of interest to seek formalizations of the concept of intelligence, whichcapture useful fragments of the commonsense notion of intelligence, and provide guidance forpractical research in cognitive science and AI. A number of such formalizations have been givenin recent decades, with varying degrees of mathematical rigor. Perhaps the most carefullywroughtformalization of intelligence so far is the theory of “universal intelligence” presented byShane Legg and Marcus Hutter in [LH07b], which draws on ideas from algorithmic informationtheory.Universal intelligence captures a certain aspect of the “intelligence” concept very well, andhas the advantage of connecting closely with ideas in learning theory, decision theory andcomputation theory. However, the kind of general intelligence it captures best, is a kind whichis in a sense more general in scope than human-style general intelligence. Universal intelligencedoes capture the sense in which humans are more intelligent than worms, which are moreintelligent than rocks; and the sense in which theoretical AGI systems like Hutter’s AIXI orAIXI tl [Hut05] would be much more intelligent than humans. But it misses essential aspectsof the intelligence concept as it is used in the context of intelligent natural systems like humansor real-world AI systems.Our main goal in this section is to present variants of universal intelligence that bettercapture the notion of intelligence as it is typically understood in the context of real-worldnatural and artificial systems. The first variant we describe is pragmatic general intelligence,which is inspired by the intuitive notion of intelligence as “the ability to achieve complex goalsin complex environments,” given in [Goe93a]. After assuming a prior distribution over thespace of possible environments, and one over the space of possible goals, one then defines thepragmatic general intelligence as the expected level of goal-achievement of a system relativeto these distributions. Rather than measuring truly broad mathematical general intelligence,pragmatic general intelligence measures intelligence in a way that’s specifically biased towardcertain environments and goals.Another variant definition is then presented, the efficient pragmatic general intelligence,which takes into account the amount of computational resources utilized by the system inachieving its intelligence. Some argue that making efficient use of available resources is a definingcharacteristic of intelligence, see e.g. [Wan06].A critical question left open is the characterization of the prior distributions correspondingto everyday human reality; we give a semi-formal sketch of some ideas on this in Chapter 9below, where we present the notion of a “communication prior,” which assigns a probability136 7 A Formal Model of Intelligent Agentsweight to a situation S based on the ease with which one agent in a society can communicateS to another agent in that society, using multimodal communication (including verbalization,demonstration, dramatic and pictorial depiction, etc.).Finally, we present a formal measure of the “generality” of an intelligence, which precisiatesthe informal distinction between “general AI” and “narrow AI.”7.3.1 Biased Universal IntelligenceTo define universal intelligence, Legg and Hutter consider the class of environments that arereward-summable, meaning that the total amount of reward they return to any agent is boundedby 1. Where r i denotes the reward experienced by the agent from the environment at time i,the expected total reward for the agent π from the environment µ is defined as∞∑Vµ π ≡ E( r i ) ≤ 1To extend their definition in the direction of greater realism, we first introduce a second-orderprobability distribution ν, which is a probability distribution over the space of environmentsµ. The distribution ν assigns each environment a probability. One such distribution ν is theSolomonoff-Levin universal distribution in which one sets ν = 2 −K(µ) ; but this is not the onlydistribution ν of interest. In fact a great deal of real-world general intelligence consists of theadaptation of intelligent systems to particular distributions ν over environment-space, differingfrom the universal distribution.We then defineDefinition 4 The biased universal intelligence of an agent π is its expected performancewith respect to the distribution ν over the space of all computable reward-summable environments,E, that is,1Υ (π) ≡ ∑ µ∈Eν(µ)V π µLegg and Hutter’s universal intelligence is obtained by setting ν equal to the universaldistribution.This framework is more flexible than it might seem. E.g. suppose one wants to incorporateagents that die. Then one may create a special action, say a 666 , corresponding to the state ofdeath, to create agents that• in certain circumstances output action a 666• have the property that if their previous action was a 666 , then all of their subsequent actionsmust be a 666and to define a reward structure so that actions a 666 always bring zero reward. It then followsthat death is generally a bad thing if one wants to maximize intelligence. Agents that die willnot get rewarded after they’re dead; and agents that live only 70 years, say, will be restrictedfrom getting rewards involving long-term patterns and will hence have specific limits on theirintelligence.7.3 Toward a Formal Characterization of Real-World General Intelligence 1377.3.2 Connecting Legg and Hutter’s Model of Intelligent Agents tothe Real WorldA notable aspect of the Legg and Hutter formalism is the separation of the reward mechanismfrom the cognitive mechanisms of the agent. While commonplace in the reinforcement learningliterature, this seems psychologically unrealistic in the context of biological intelligences andmany types of machine intelligences. Not all human intelligent activity is specifically rewardseekingin nature; and even when it is, humans often pursue complexly constructed rewards,that are defined in terms of their own cognitions rather than separately given. Suppose a certainhuman’s goals are true love, or world peace, and the proving of interesting theorems – then thesegoals are defined by the human herself, and only she knows if she’s achieved them. An externallyprovidedreward signal doesn’t capture the nature of this kind of goal-seeking behavior, whichcharacterizes much human goal-seeking activity (and will presumably characterize much of thegoal-seeking activity of advanced engineered intelligences also) ... let alone human behavior thatis spontaneous and unrelated to explicit goals, yet may still appear commonsensically intelligent.One could seek to bypass this complaint about the reward mechanisms via a sort of “neo-Freudian” argument, via• associating the reward signal, not with the “external environment” as typically conceived,but rather with a portion of the intelligent agent’s brain that is separate from the cognitivecomponent• viewing complex goals like true love, world peace and proving interesting theorems as indirectways of achieving the agent’s “basic goals”, created within the agent’s memory viasubgoaling mechanismsbut it seems to us that a general formalization of intelligence should not rely on such strongassumptions about agents’ cognitive architectures. So below, after introducing the pragmaticand efficient pragmatic general intelligence measures, we will propose an alternate interpretationwherein the mechanism of external rewards is viewed as a theoretical test framework forassessing agent intelligence, rather than a hypothesis about intelligent agent architecture.In this alternate interpretation, formal measures like the universal, pragmatic and efficientpragmatic general intelligence are viewed as not directly applicable to real-world intelligences,because they involve the behaviors of agents over a wide variety of goals and environments,whereas in real life the opportunities to observe agents are more limited. However, they areviewed as being indirectly applicable to real-world agents, in the sense that an external intelligencecan observe an agent’s real-world behavior and then infer its likely intelligence accordingto these measures.In a sense, this interpretation makes our formalized measures of intelligence the opposite ofreal-world IQ tests. An IQ test is a quantified, formalized test which is designed to approximatelypredict the informal, qualitative achievement of humans in real life. On the other hand,the formal definitions of intelligence we present here are quantified, formalized tests that aredesigned to capture abstract notions of intelligence, but which can be approximately evaluatedon a real-world intelligent system by observing what it does in real life.138 7 A Formal Model of Intelligent Agents7.3.3 Pragmatic General IntelligenceThe above concept of biased universal intelligence is perfectly adequate for many purposes, butit is also interesting to explicitly introduce the notion of a goal into the calculation. This allowsus to formally capture the notion presented in [Goe93a] of intelligence as “the ability to achievecomplex goals in complex environments.”If the agent is acting in environment µ, and is provided with g s corresponding to g at thestart and the end of the time-interval T = {i ∈ (s, ..., t)}, then the expected goal-achievementof the agent, relative to g, during the interval is the expectationV π µ,g,T ≡ E(t∑r g (I g,s,i ))i=swhere the expectation is taken over all interaction sequences I g,s,i drawn according to µ. Wethen proposeDefinition 5 The pragmatic general intelligence of an agent π, relative to the distributionν over environments and the distribution γ over goals, is its expected performance with respectto goals drawn from γ in environments drawn from ν, over the time-scales natural to the goals;that is,∑Π(π) ≡ ν(µ)γ(g, µ)Vµ,g,Tπµ∈E,g∈G,T(in those cases where this sum is convergent).This definition formally captures the notion that “intelligence is achieving complex goals incomplex environments,” where “complexity” is gauged by the assumed measures ν and γ.If ν is taken to be the universal distribution, and γ is defined to weight goals according tothe universal distribution, then pragmatic general intelligence reduces to universal intelligence.Furthermore, it is clear that a universal algorithmic agent like AIXI [Hut05] would alsohave a high pragmatic general intelligence, under fairly broad conditions. As the interactionhistory grows longer, the pragmatic general intelligence of AIXI would approach the theoreticalmaximum; as AIXI would implicitly infer the relevant distributions via experience. However,if significant reward discounting is involved, so that near-term rewards are weighted muchhigher than long-term rewards, then AIXI might compare very unfavorably in pragmatic generalintelligence, to other agents designed with prior knowledge of ν, γ and τ in mind.The most interesting case to consider is where ν and γ are taken to embody some particularbias in a real-world space of environments and goals, and this bias is appropriately reflectedin the internal structure of an intelligent agent. Note that an agent needs not lack universalintelligence in order to possess pragmatic general intelligence with respect to some non-universaldistribution over goals and environments. However, in general, given limited resources, theremay be a tradeoff between universal intelligence and pragmatic intelligence. Which leads to thenext point: how to encompass resource limitations into the definition.One might argue that the definition of Pragmatic General Intelligence is already encompassedby Legg and Hutter’s definition because one may bias the distribution of environments withinthe latter by considering different Turing machines underlying the Kolmogorov complexity.However this is not a general equivalence because the Solomonoff-Levin measure intrinsically7.3 Toward a Formal Characterization of Real-World General Intelligence 139decays exponentially, whereas an assumptive distribution over environments might decay atsome other rate. This issue seems to merit further mathematical investigation.7.3.4 Incorporating Computational CostLet η π,µ,g,T be a probability distribution describing the amount of computational resources consumedby an agent π while achieving goal g over time-scale T . This is a probability distributionbecause we want to account for the possibility of nondeterministic agents. So, η π,µ,g,T (Q) tellsthe probability that Q units of resources are consumed. For simplicity we amalgamate spaceand time resources, energetic resources, etc. into a single number Q, which is assumed to livein some subset of the positive reals. Space resources of course have to do with the size of thesystem’s memory. Then we may defineDefinition 6 The efficient pragmatic general intelligence of an agent π with resourceconsumption η π,µ,g,T , relative to the distribution ν over environments and the distribution γover goals, is its expected performance with respect to goals drawn from γ in environments drawnfrom ν, over the time-scales natural to the goals, normalized by the amount of computationaleffort expended to achieve each goal; that is,Π Eff (π) ≡∑µ∈E,g∈G,Q,T(in those cases where this sum is convergent).ν(µ)γ(g, µ)η π,µ,g,T (Q)Vµ,g,Tπ QThis is a measure that rates an agent’s intelligence higher if it uses fewer computationalresources to do its business. Roughly, it measures reward achieved per spacetime computationunit.Note that, by abandoning the universal prior, we have also abandoned the proof of convergencethat comes with it. In general the sums in the above definitions need not converge; andexploration of the conditions under which they do converge is a complex matter.7.3.5 Assessing the Intelligence of Real-World AgentsThe pragmatic and efficient pragmatic general intelligence measures are more “realistic” thanthe Legg and Hutter universal intelligence measure, in that they take into account the innatebiasing and computational resource restrictions that characterize real-world intelligence. But asdiscussed earlier, they still live in “fantasy-land” to an extent – they gauge the intelligence of anagent via a weighted average over a wide variety of goals and environments; and they presumea simplistic relationship between agents and rewards that does not reflect the complexitiesof real-world cognitive architectures. It is not obvious from the foregoing how to apply thesemeasures to real-world intelligent systems, which lack the ability to exist in such a wide varietyof environments within their often brief lifespans, and mostly go about their lives doing thingsother than pursuing quantified external rewards. In this brief section we describe an approachto bridging this gap. The treatment is left semi-formal in places.140 7 A Formal Model of Intelligent AgentsWe suggest to view the definitions of pragmatic and efficient pragmatic general intelligencein terms of a “possible worlds” semantics – i.e. to view them as asking, counterfactually, howan agent would perform, hypothetically, on a series of tests (the tests being goals, defined inrelation to environments and reward signals).Real-world intelligent agents don’t normally operate in terms of explicit goals and rewards;these are abstractions that we use to think about intelligent agents. However, this is no objectionto characterizing various sorts of intelligence in terms of counterfactuals like: how would systemS operate if it were trying to achieve this or that goal, in this or that environment, in order toseek reward? We can characterize various sorts of intelligence in terms of how it can be inferredan agent would perform on certain tests, even though the agent’s real life does not consist oftaking these tests.This conceptual approach may seem a bit artificial but we don’t currently see a betteralternative, if one wishes to quantitatively gauge intelligence (which is, in a sense, an “artificial”thing to do in the first place). Given a real-world agent X and a mandate to assess its intelligence,the obvious alternative to looking at possible worlds in the manner of the above definitions,is just looking directly at the properties of the things X has achieved in the real world duringits lifespan. But this isn’t an easy solution, because it doesn’t disambiguate which aspects ofX’s achievements were due to its own actions versus due to the rest of the world that X wasinteracting with when it made its achievements. To distinguish the amount of achievement thatX “caused” via its own actions requires a model of causality, which is a complex can of worms initself; and, critically, the standard models of causality also involve counterfactuals (asking “whatwould have been achieved in this situation if the agent X hadn’t been there”, etc.) [MW07].Regardless of the particulars, it seems impossible to avoid counterfactual realities in assessingintelligence.The approach we suggest – given a real-world agent X with a history of actions in a particularworld, and a mandate to assess its intelligence – is to introduce an additional player, an inferenceagent δ, into the picture. The agent π modeled above is then viewed as π X : the model of X thatδ constructs, in order to explore X’s inferred behaviors in various counterfactual environments.In the test situations embodied in the definitions of pragmatic and efficient pragmatic generalintelligence, the environment gives π X rewards, based on specifically configured goals. In X’sreal life, the relation between goals, rewards and actions will generally be significantly subtlerand perhaps quite different.We model the real world similarly to the “fantasy world” of the previous section, but withthe omission of goals and rewards. We define a naturalistic context as one in which all goals andrewards are constant, i.e. g i = g 0 and r i = r 0 for all i. This is just a mathematical conventionfor stating that there are no precisely-defined external goals and rewards for the agent. In anaturalistic context, we then have a situation where agents create actions based on the pasthistory of actions and perceptions, and if there is any relevant notion of reward or goal, itis within the cognitive mechanism of some agent. A naturalistic agent X is then an agent πwhich is restricted to one particular naturalistic context, involving one particular environmentµ (formally, we may achieve this within the framework of agents described above via dictatingthat X issues constant “null actions” a 0 in all environments except µ).Next, we posit a metric space (Σ µ , d) of naturalistic agents defined on a naturalistic contextinvolving environment µ, and a subspace ∆ ∈ Σ µ of inference agents, which are naturalisticagents that output predictions of other agents’ behaviors (a notion we will not fully formalizehere). If agents are represented as program trees, then d may be taken as edit distance on treespace [Bil05]. Then, for each agent δ ∈ ∆, we may assess7.4 Intellectual Breadth: Quantifying the Generality of an Agent’s Intelligence 141• the prior probability θ(δ) according to some assumed distribution θ• the effectiveness p(δ, X) of δ at predicting the actions of an agent X ∈ Σ µWe may then defineDefinition 7 The inference ability of the agent δ, relative to µ and X, is∑Y ∈Σq µ,X (δ) = θ(δ)µsim(X, Y )p(δ, Y )∑Y ∈Σ µsim(X, Y )where sim is a specified decreasing function of d(X, Y ), such as sim(X, Y ) =11+d(X,Y ) .To construct π X , we may then use the model of X created by the agent δ ∈ ∆ with thehighest inference ability relative to µ and X (using some specified ordering, in case of a tie).Having constructed π X , we can then say thatDefinition 8 The inferred pragmatic general intelligence (relative to ν and γ) of a naturalisticagent X defined relative to an environment µ, is defined as the pragmatic general intelligenceof the model π X of X produced by the agent δ ∈ ∆ with maximal inference ability relative to µ(and in the case of a tie, the first of these in the ordering defined over ∆). The inferred efficientpragmatic general intelligence of X relative to µ is defined similarly.This provides a precise characterization of the pragmatic and efficient pragmatic intelligenceof real-world systems, based on their observed behaviors. It’s a bit messy; but the real worldtends to be like that.7.4 Intellectual Breadth: Quantifying the Generality of an Agent’sIntelligenceWe turn now to a related question: How can one quantify the degree of generality that anintelligent agent possesses? Above we have discussed the qualitative distinction between AGIand “Narrow AI”, and intelligence as we have formalized it above is specifically intended asa measure of general intelligence. But quantifying intelligence is different than quantifyinggenerality versus narrowness.To make the discussion simpler, we introduce the term “context” as a shorthand for “environment/intervaltriple (µ, g, T ).” Given a context (µ, g, T ), and a set Σ of agents, one mayconstruct a fuzzy set Ag µ,g,T gathering those agents that are intelligent relative to the context;and given a set of contexts, one may also define a fuzzy set Con π gathering those contexts withrespect to which a given agent π is intelligent. The relevant formulas are:χ Agµ,g,T (π) = χ Conπ (µ, g, T ) = 1 N∑Qη µ,g,T (Q)V π µ,g,TQwhere N = N(µ, g, T ) is a normalization factor defined appropriately, e.g. via N(µ, g, T ) =maxπ V π µ,g,T .One could make similar definitions leaving out the computational cost factor Q, but wesuspect that incorporating Q is a more promising direction. We then propose142 7 A Formal Model of Intelligent AgentsDefinition 9 The intellectual breadth of an agent π, relative to the distribution ν overenvironments and the distribution γ over goals, iswhere H is the entropy andH(χ P Con π(µ, g, T ))χ P Con π(µ, g, T ) =∑(µ α,g β .T ω)ν(µ)γ(g, µ)χ Conπ (µ, g, T )ν(µ α )γ(g β , µ α )χ Conπ (µ α , g β , T ω )is the probability distribution formed by normalizing the fuzzy set χ Conπ (µ, g, T ).A similar definition of the intellectual breadth of a context (µ, g, T ), relative to the distributionσ over agents, may be posited. A weakness of these definitions is that they don’t try toaccount for dependencies between agents or contexts; perhaps more refined formulations maybe developed that account explicitly for these dependencies.Note that the intellectual breadth of an agent as defined here is largely independent ofthe (efficient or not) pragmatic general intelligence of that agent. One could have a rather(efficiently or not) pragmatically generally intelligent system with little breadth: this would bea system very good at solving a fair number of hard problems, yet wholly incompetent on alarger number of hard problems. On the other hand, one could also have a terribly (efficiently ornot) pragmatically generally stupid system with great intellectual breadth: i.e a system roughlyequally dumb in all contexts!Thus, one can characterize an intelligent agent as “narrow” with respect to distribution ν overenvironments and the distribution γ over goals, based on evaluating it as having low intellectualbreadth. A “narrow AI” relative to ν and γ would then be an AI agent with a relatively highefficient pragmatic general intelligence but a relatively low intellectual breadth.7.5 ConclusionOur main goal in this chapter has been to push the formal understanding of intelligence in a morepragmatic direction. Much more work remains to be done, e.g. in specifying the environment,goal and efficiency distributions relevant to real-world systems, but we believe that the ideaspresented here constitute nontrivial progress.If the line of research suggested in this chapter succeeds, then eventually, one will be able todo AGI research as follows: Specify an AGI architecture formally, and then use the mathematicsof general intelligence to derive interesting results about the environments, goals and hardwareplatforms relative to which the AGI architecture will display significant pragmatic or efficientpragmatic general intelligence, and intellectual breadth. The remaining chapters in this sectionpresent further ideas regarding how to work toward this goal. For the time being, such a modeof AGI research remains mainly for the future, but we have still found the formalism given inthese chapters useful for formulating and clarifying various aspects of the CogPrime design aswill be presented in later chapters.Chapter 8Cognitive Synergy8.1 Cognitive SynergyAs we have seen, the formal theory of general intelligence, in its current form, doesn’t reallytell us much that’s of use for creating real-world AGI systems. It tells us that creating extraordinarilypowerful general intelligence is almost trivial if one has unrealistically huge amountsof computational resources; and that creating moderately powerful general intelligence usingfeasible computational resources is all about creating AI algorithms and data structures that(explicitly or implicitly) match the restrictions implied by a certain class of situations, to whichthe general intelligence is biased.We’ve also described, in various previous chapters, some non-rigorous, conceptual principlesthat seem to explain key aspects of feasible general intelligence: the complementary reliance onevolution and autopoiesis, the superposition of hierarchical and heterarchical structures, and soforth. These principles can be considered as broad strategies for achieving general intelligencein certain broad classes of situations. Although, a lot of research needs to be done to figure outnice ways to describe, for instance, in what class of situations evolution is an effective learningstrategy, in what class of situations dual hierarchical/heterarchical structure is an effective wayto organize memory, etc.In this chapter we’ll dig deeper into one of the “general principle of feasible general intelligences”briefly alluded to earlier: the cognitive synergy principle, which is both a conceptualhypothesis about the structure of generally intelligent systems in certain classes of environments,and a design principle used to guide the architecting of CogPrime.We will focus here on cognitive synergy specifically in the case of “multi-memory systems,”which we define as intelligent systems (like CogPrime) whose combination of environment,embodiment and motivational systems make it important for them to possess memories thatdivide into partially but not wholly distinct components corresponding to the categories of:• Declarative memory• Procedural memory (memory about how to do certain things)• Sensory and episodic memory• Attentional memory (knowledge about what to pay attention to in what contexts• Intentional memory (knowledge about the system’s own goals and subgoals)In Chapter 9 below we present a detailed argument as to how the requirement for a multimemoryunderpinning for general intelligence emerges from certain underlying assumptions143144 8 Cognitive Synergyregarding the measurement of the simplicity of goals and environments; but the points madehere do not rely on that argument. What they do rely on is the assumption that, in theintelligence in question, the different components of memory are significantly but not whollydistinct. That is, there are significant “family resemblances” between the memories of a singletype, yet there are also thoroughgoing connections between memories of different types.The cognitive synergy principle, if correct, applies to any AI system demonstrating intelligencein the context of embodied, social communication. However, one may also take the theoryas an explicit guide for constructing AGI systems; and of course, the bulk of this book describesone AGI architecture, CogPrime, designed in such a way.It is possible to cast these notions in mathematical form, and we make some efforts in thisdirection in Appendix ??, using the languages of category theory and information geometry.However, this formalization has not yet led to any rigorous proof of the generality of cognitivesynergy nor any other exciting theorems; with luck this will come as the mathematics is furtherdeveloped. In this chapter the presentation is kept on the heuristic level, which is all that iscritically needed for motivating the CogPrime design.8.2 Cognitive SynergyThe essential idea of cognitive synergy, in the context of multi-memory systems, may be expressedin terms of the following points:1. Intelligence, relative to a certain set of environments, may be understood as the capabilityto achieve complex goals in these environments.2. With respect to certain classes of goals and environments (see Chapter 9 for a hypothesisin this regard), an intelligent system requires a “multi-memory” architecture, meaningthe possession of a number of specialized yet interconnected knowledge types, including:declarative, procedural, attentional, sensory, episodic and intentional (goal-related). Theseknowledge types may be viewed as different sorts of patterns that a system recognizes initself and its environment. Knowledge of these various different types must be interlinked,and in some cases may represent differing views of the same content (see Figure ??)3. Such a system must possess knowledge creation (i.e. pattern recognition / formation) mechanismscorresponding to each of these memory types. These mechanisms are also called“cognitive processes.”4. Each of these cognitive processes, to be effective, must have the capability to recognize whenit lacks the information to perform effectively on its own; and in this case, to dynamicallyand interactively draw information from knowledge creation mechanisms dealing with othertypes of knowledge5. This cross-mechanism interaction must have the result of enabling the knowledge creationmechanisms to perform much more effectively in combination than they would if operatednon-interactively. This is “cognitive synergy.”While these points are implicit in the theory of mind given in [Goe06a], they are not articulatedin this specific form there.Interactions as mentioned in Points 4 and 5 in the above list are the real conceptual meatof the cognitive synergy idea. One way to express the key idea here is that most AI algorithmssuffer from combinatorial explosions: the number of possible elements to be combined in a8.2 Cognitive Synergy 145Fig. 8.1: Illustrative example of the interactions between multiple types of knowledge, in representinga simple piece of knowledge. Generally speaking, one type of knowledge can be convertedto another, at the cost of some loss of information. The synergy between cognitive processesassociated with corresponding pieces of knowledge, possessing different type, is a critical aspectof general intelligence.synthesis or analysis is just too great, and the algorithms are unable to filter through all thepossibilities, given the lack of intrinsic constraint that comes along with a “general intelligence”context (as opposed to a narrow-AI problem like chess-playing, where the context is constrainedand hence restricts the scope of possible combinations that needs to be considered). In an AGIarchitecture based on cognitive synergy, the different learning mechanisms must be designedspecifically to interact in such a way as to palliate each others’ combinatorial explosions - sothat, for instance, each learning mechanism dealing with a certain sort of knowledge, mustsynergize with learning mechanisms dealing with the other sorts of knowledge, in a way thatdecreases the severity of combinatorial explosion.One prerequisite for cognitive synergy to work is that each learning mechanism must recognizewhen it is “stuck,” meaning it’s in a situation where it has inadequate information tomake a confident judgment about what steps to take next. Then, when it does recognize thatit’s stuck, it may request help from other, complementary cognitive mechanisms.A theoretical notion closely related to cognitive synergy is the cognitive schematic, formalizedin Chapter 7 above, which states that the activity of the different cognitive processes involvedin an intelligent system may be modeled in terms of the schematic implicationContext ∧ P rocedure → Goal146 8 Cognitive Synergywhere the Context involves sensory, episodic and/or declarative knowledge; and attentionalknowledge is used to regulate how much resource is given to each such schematic implication inmemory. Synergy among the learning processes dealing with the context, the procedure and thegoal is critical to the adequate execution of the cognitive schematic using feasible computationalresources.Finally, drilling a little deeper into Point 3 above, one arrives at a number of possible knowledgecreation mechanisms (cognitive processes) corresponding to each of the key types of knowledge.Figure ?? below gives a high-level overview of the main types of cognitive process consideredin the current version of Cognitive Synergy Theory, categorized according to the typeof knowledge with which each process deals.8.3 Cognitive Synergy in CogPrimeDifferent cognitive systems will use different processes to fulfill the various roles identified inFigure ?? above. Here we briefly preview the basic cognitive processes that the CogPrime AGIdesign uses for these roles, and the synergies that exist between these.8.3.1 Cognitive Processes in CogPrime: a Cognitive Synergy Based Architecture..." from ICCI 2009Table 8.1: defaultTable will go hereTable 8.2: The OpenCogPrime data structures used to represent the key knowledge types involvedTable 8.3: defaultTable will go hereTable 8.4: Key cognitive processes, and the algorithms that play their roles in CogPrimeTables 8.1 and 8.3 present the key structures and processes involved in CogPrime, identifyingeach one with a certain memory/process type as considered in cognitive synergy theory. Thatis: each of these cognitive structures or processes deals with one or more types of memory –declarative, procedural, sensory, episodic or attentional. Table 8.5 describes the key CogPrime8.3 Cognitive Synergy in CogPrime 147Fig. 8.2: High-level overview of the key cognitive dynamics considered here in the context ofcognitive synergy. The cognitive synergy principle describes the behavior of a system as itpursues a set of goals (which in most cases may be assumed to be supplied to the system“a priori”, but then refined by inference and other processes). The assumed intelligent agentmodel is roughly as follows: At each time the system chooses a set of procedures to execute,based on its judgments regarding which procedures will best help it achieve its goals in thecurrent context. These procedures may involve external actions (e.g. involving conversation,or controlling an agent in a simulated world) and/or internal cognitive actions. In order tomake these judgments it must effectively manage declarative, procedural, episodic, sensoryand attentional memory, each of which is associated with specific algorithms and structuresas depicted in the diagram. There are also global processes spanning all the forms of memory,including the allocation of attention to different memory items and cognitive processes, and theidentification and reification of system-wide activity patterns (the latter referred to as “mapformation”)Table 8.5: defaultTable will go hereTable 8.6: Key OpenCogPrime cognitive processes categorized according to knowledge type andprocess type148 8 Cognitive Synergyprocesses in terms of the “analysis vs. synthesis” distinction. Finally, Tables ?? and ?? exemplifythese structures and processes in the context of embodied virtual agent control.In the CogPrime context, a procedure in this cognitive schematic is a program tree storedin the system’s procedural knowledge base; and a context is a (fuzzy, probabilistic) logicalpredicate stored in the AtomSpace, that holds, to a certain extent, during each interval of time.A goal is a fuzzy logical predicate that has a certain value at each interval of time, as well.Attentional knowledge is handled in CogPrime by the ECAN artificial economics mechanism,that continually updates ShortTermImportance and LongTerm Importance values associatedwith each item in the CogPrime system’s memory, which control the amount of attention othercognitive mechanisms pay to the item, and how much motive the system has to keep theitem in memory. HebbianLinks are then created between knowledge items that often possessShortTermImportance at the same time; this is CogPrime’s version of traditional Hebbianlearning.ECAN has deep interactions with other cognitive mechanisms as well, which are essentialto its efficient operation; for instance, PLN inference may be used to help ECAN extrapolateconclusions about what is worth paying attention to, and MOSES may be used to recognizesubtle attentional patterns. ECAN also handles “assignment of credit”, the figuring-out of thecauses of an instance of successful goal-achievement, drawing on PLN and MOSES as neededwhen the causal inference involved here becomes difficult.The synergies between CogPrime’s cognitive processes are well summarized below, which isa 16x16 matrix summarizing a host of interprocess interactions generic to CST.One key aspect of how CogPrime implements cognitive synergy is PLN’s sophisticated managementof the confidence of judgments. This ties in with the way OpenCogPrime’s PLN inferenceframework represents truth values in terms of multiple components (as opposed to thesingle probability values used in many probabilistic inference systems and formalisms): eachitem in OpenCogPrime’s declarative memory has a confidence value associated with it, whichtells how much weight the system places on its knowledge about that memory item. This assistswith cognitive synergy as follows: A learning mechanism may consider itself “stuck”, generallyspeaking, when it has no high-confidence estimates about the next step it should take.Without reasonably accurate confidence assessment to guide it, inter-component interactioncould easily lead to increased rather than decreased combinatorial explosion. And of coursethere is an added recursion here, in that confidence assessment is carried out partly via PLNinference, which in itself relies upon these same synergies for its effective operation.To illustrate this point further, consider one of the synergetic aspects described in ?? below:the role cognitive synergy plays in deductive inference. Deductive inference is a hard problemin general - but what is hard about it is not carrying out inference steps, but rather “inferencecontrol” (i.e., choosing which inference steps to carry out). Specifically, what must happen fordeduction to succeed in CogPrime is:1. the system must recognize when its deductive inference process is “stuck”, i.e. when thePLN inference control mechanism carrying out deduction has no clear idea regarding whichinference step(s) to take next, even after considering all the domain knowledge at is disposal2. in this case, the system must defer to another learning mechanism to gather more informationabout the different choices available - and the other learning mechanism chosen must,a reasonable percentage of the time, actually provide useful information that helps PLN toget “unstuck” and continue the deductive process8.4 Some Critical Synergies 149For instance, deduction might defer to the “attentional knowledge” subsystem, and makea judgment as to which of the many possible next deductive steps are most associated withthe goal of inference and the inference steps taken so far, according to the HebbianLinks constructedby the attention allocation subsystem, based on observed associations. Or, if this fails,deduction might ask MOSES (running in supervised categorization mode) to learn predicatescharacterizing some of the terms involving the possible next inference steps. Once MOSES providesthese new predicates, deduction can then attempt to incorporate these into its inferenceprocess, hopefully (though not necessarily) arriving at a higher-confidence next step.8.4 Some Critical SynergiesReferring back to Figure ??, and summarizing many of the ideas in the previous section, Table?? enumerates a number of specific ways in which the cognitive processes mentioned in theFigure may synergize with one another, potentially achieving dramatically greater efficiencythan would be possible on their own.Of course, realizing these synergies on the practical algorithmic level requires significantinventiveness and may be approached in many different ways. The specifics of how CogPrimemanifests these synergies are discussed in many following chapters.Fig. 8.3: This table, and the following ones, show some of the synergies between the primarycognitive processes explicitly used in CogPrime.150 8 Cognitive Synergy8.5 The Cognitive Schematic 1518.5 The Cognitive SchematicNow we return to the “cognitive schematic” notion, according to which various cognitive processesinvolved in intelligence may be understood to work together via the implicationContext ∧ P rocedure → Goal < p >(summarized C ∧ P → G). Semi-formally, this implication may be interpreted to mean: “If thecontext C appears to hold currently, then if I enact the procedure P , I can expect to achievethe goal G with certainty p.”The cognitive schematic leads to a conceptualization of the internal action of an intelligentsystem as involving two key categories of learning:• Analysis: Estimating the probability p of a posited C ∧ P → G relationship• Synthesis: Filling in one or two of the variables in the cognitive schematic, given assumptionsregarding the remaining variables, and directed by the goal of maximizing theprobability of the cognitive schematicMore specifically, where synthesis is concerned, some key examples are:• The MOSES probabilistic evolutionary program learning algorithm is applied to find P ,given fixed C and G. Internal simulation is also used, for the purpose of creating a simulationembodying C and seeing which P lead to the simulated achievement of G.– Example: A virtual dog learns a procedure P to please its owner (the goal G) in thecontext C where there is a ball or stick present and the owner is saying “fetch”.• PLN inference, acting on declarative knowledge, is used for choosing C, given fixed P andG (also incorporating sensory and episodic knowledge as appropriate). Simulation may alsobe used for this purpose.152 8 Cognitive Synergy– Example: A virtual dog wants to achieve the goal G of getting food, and it knows thatthe procedure P of begging has been successful at this before, so it seeks a context Cwhere begging can be expected to get it food. Probably this will be a context involving afriendly person.• PLN-based goal refinement is used to create new subgoals G to sit on the right hand sideof instances of the cognitive schematic.– Example: Given that a virtual dog has a goal of finding food, it may learn a subgoal offollowing other dogs, due to observing that other dogs are often heading toward theirfood.• Concept formation heuristics are used for choosing G and for fueling goal refinement, butespecially for choosing C (via providing new candidates for C). They are also used forchoosing P , via a process called “predicate schematization” that turns logical predicates(declarative knowledge) into procedures.– Example: At first a virtual dog may have a hard time predicting which other dogs aregoing to be mean to it. But it may eventually observe common features among a numberof mean dogs, and thus form its own concept of “pit bull,” without anyone ever teachingit this concept explicitly.Where analysis is concerned:• PLN inference, acting on declarative knowledge, is used for estimating the probability ofthe implication in the cognitive schematic, given fixed C, P and G. Episodic knowledgeis also used this regard, via enabling estimation of the probability via simple similaritymatching against past experience. Simulation is also used: multiple simulations may berun, and statistics may be captured therefrom.– Example: To estimate the degree to which asking Bob for food (the procedure P is “askingfor food”, the context C is “being with Bob”) will achieve the goal G of getting food, thevirtual dog may study its memory to see what happened on previous occasions where itor other dogs asked Bob for food or other things, and then integrate the evidence fromthese occasions.• Procedural knowledge, mapped into declarative knowledge and then acted on by PLN inference,can be useful for estimating the probability of the implication C ∧ P → G, in caseswhere the probability of C ∧ P 1 → G is known for some P 1 related to P .– Example: knowledge of the internal similarity between the procedure of asking for foodand the procedure of asking for toys, allows the virtual dog to reason that if asking Bobfor toys has been successful, maybe asking Bob for food will be successful too.• Inference, acting on declarative or sensory knowledge, can be useful for estimating theprobability of the implication C ∧ P → G, in cases where the probability of C 1 ∧ P → G isknown for some C 1 related to C.– Example: if Bob and Jim have a lot of features in common, and Bob often respondspositively when asked for food, then maybe Jim will too.• Inference can be used similarly for estimating the probability of the implication C ∧P → G,in cases where the probability of C ∧ P → G 1 is known for some G 1 related to G. Concept8.6 Cognitive Synergy for Procedural and Declarative Learning 153creation can be useful indirectly in calculating these probability estimates, via providingnew concepts that can be used to make useful inference trails more compact and henceeasier to construct.– Example: The dog may reason that because Jack likes to play, and Jack and Jill are bothchildren, maybe Jill likes to play too. It can carry out this reasoning only if its conceptcreation process has invented the concept of “child” via analysis of observed data.In these examples we have focused on cases where two terms in the cognitive schematic arefixed and the third must be filled in; but just as often, the situation is that only one of theterms is fixed. For instance, if we fix G, sometimes the best approach will be to collectivelylearn C and P . This requires either a procedure learning method that works interactively with adeclarative-knowledge-focused concept learning or reasoning method; or a declarative learningmethod that works interactively with a procedure learning method. That is, it requires the sortof cognitive synergy built into the CogPrime design.8.6 Cognitive Synergy for Procedural and Declarative LearningWe now present a little more algorithmic detail regarding the operation and synergetic interactionof CogPrime’s two most sophisticated components: the MOSES procedure learningalgorithm (see Chapter 33), and the PLN uncertain inference framework (see Chapter 34). Thetreatment is necessarily quite compact, since we have not yet reviewed the details of eitherMOSES or PLN; but as well as illustrating the notion of cognitive synergy more concretely,perhaps the high-level discussion here will make clearer how MOSES and PLN fit into the bigpicture of CogPrime.8.6.1 Cognitive Synergy in MOSESMOSES, CogPrime’s primary algorithm for learning procedural knowledge, has been tested ona variety of application problems including standard GP test problems, virtual agent control,biological data analysis and text classification [Loo06]. It represents procedures internally asprogram trees. Each node in a MOSES program tree is supplied with a “knob,” comprising aset of values that may potentially be chosen to replace the data item or operator at that node.So for instance a node containing the number 7 may be supplied with a knob that can takeon any integer value. A node containing a while loop may be supplied with a knob that cantake on various possible control flow operators including conditionals or the identity. A nodecontaining a procedure representing a particular robot movement, may be supplied with a knobthat can take on values corresponding to multiple possible movements. Following a metaphorsuggested by Douglas Hofstadter [Hof96], MOSES learning covers both “knob twiddling” (settingthe values of knobs) and “knob creation.”MOSES is invoked within CogPrime in a number of ways, but most commonly for finding aprocedure P satisfying a probabilistic implication C&P → G as described above, where C is anobserved context and G is a system goal. In this case the probability value of the implicationprovides the “scoring function” that MOSES uses to assess the quality of candidate procedures.154 8 Cognitive SynergyFig. 8.4: High-Level Control Flow of MOSES AlgorithmFor example, suppose an CogPrime -controlled robot is trying to learn to play the gameof “tag." (I.e. a multi-agent game in which one agent is specially labeled "it", and runs afterthe other player agents, trying to touch them. Once another agent is touched, it becomes thenew "it" and the previous "it" becomes just another player agent.) Then its context C is thatothers are trying to play a game they call “tag” with it; and we may assume its goals are toplease them and itself, and that it has figured out that in order to achieve this goal it shouldlearn some procedure to follow when interacting with others who have said they are playing“tag.” In this case a potential tag-playing procedure might contain nodes for physical actionslike step_f orward(speed s), as well as control flow nodes containing operators like if else(for instance, there would probably be a conditional telling the robot to do something differentdepending on whether someone seems to be chasing it). Each of these program tree nodes wouldhave an appropriate knob assigned to it. And the scoring function would evaluate a procedureP in terms of how successfully the robot played tag when controlling its behaviors according toP (noting that it may also be using other control procedures concurrently with P ). It’s worthnoting here that evaluating the scoring function in this case involves some inference already,because in order to tell if it is playing tag successfully, in a real-world context, it must watchand understand the behavior of the other players.MOSES follows the high-level control flow depicted in Figure 8.4, which corresponds to thefollowing process for evolving a metapopulation of “demes“ of programs (each deme being a setof relatively similar programs, forming a sort of island in program space):1. Construct an initial set of knobs based on some prior (e.g., based on an empty program;or more interestingly, using prior knowledge supplied by PLN inference based on thesystem’s memory) and use it to generate an initial random sampling of programs. Add thisdeme to the metapopulation.2. Select a deme from the metapopulation and update its sample, as follows:8.6 Cognitive Synergy for Procedural and Declarative Learning 155a. Select some promising programs from the deme’s existing sample to use for modeling,according to the scoring function.b. Considering the promising programs as collections of knob settings, generate new collectionsof knob settings by applying some (competent) optimization algorithm. For bestperformance on difficult problems, it is important to use an optimization algorithm thatmakes use of the system’s memory in its choices, consulting PLN inference to helpestimate which collections of knob settings will work best.c. Convert the new collections of knob settings into their corresponding programs, reducethe programs to normal form, evaluate their scores, and integrate them into thedeme’s sample, replacing less promising programs. In the case that scoring is expensive,score evaluation may be preceded by score estimation, which may use PLN inference,enaction of procedures in an internal simulation environment, and/or similaritymatching against episodic memory.3. For each new program that meet the criterion for creating a new deme, if any:a. Construct a new set of knobs (a process called “representation-building”) to define aregion centered around the program (the deme’s exemplar), and use it to generate anew random sampling of programs, producing a new deme.b. Integrate the new deme into the metapopulation, possibly displacing less promisingdemes.4. Repeat from step 2.MOSES is a complex algorithm and each part plays its role; if any one part is removed theperformance suffers significantly [Loo06]. However, the main point we want to highlight here isthe role played by synergetic interactions between MOSES and other cognitive components suchas PLN, simulation and episodic memory, as indicated in boldface in the above pseudocode.MOSES is a powerful procedure learning algorithm, but used on its own it runs into scalabilityproblems like any other such algorithm; the reason we feel it has potential to play a major rolein a human-level AI system is its capacity for productive interoperation with other cognitivecomponents.Continuing the “tag” example, the power of MOSES’s integration with other cognitive processeswould come into play if, before learning to play tag, the robot has already played simplergames involving chasing. If the robot already has experience chasing and being chased by otheragents, then its episodic and declarative memory will contain knowledge about how to pursueand avoid other agents in the context of running around an environment full of objects, and thisknowledge will be deployable within the appropriate parts of MOSES’s Steps 1 and 2. Crossprocessand cross-memory-type integration make it tractable for MOSES to act as a “transferlearning” algorithm, not just a task-specific machine-learning algorithm.8.6.2 Cognitive Synergy in PLNWhile MOSES handles much of CogPrime’s procedural learning, and OpenCogPrimes internalsimulation engine handles most episodic knowledge, CogPrime’s primary tool for handlingdeclarative knowledge is an uncertain inference framework called Probabilistic Logic Networks(PLN). The complexities of PLN are the topic of a lengthy technical monograph [GMIH08], and156 8 Cognitive Synergyhere we will eschew most details and focus mainly on pointing out how PLN seeks to achieveefficient inference control via integration with other cognitive processes.As a logic, PLN is broadly integrative: it combines certain term logic rules with more standardpredicate logic rules, and utilizes both fuzzy truth values and a variant of imprecise probabilitiescalled indefinite probabilities. PLN mathematics tells how these uncertain truth values propagatethrough its logic rules, so that uncertain premises give rise to conclusions with reasonablyaccurately estimated uncertainty values. This careful management of uncertainty is critical forthe application of logical inference in the robotics context, where most knowledge is abstractedfrom experience and is hence highly uncertain.PLN can be used in either forward or backward chaining mode; and in the language introducedabove, it can be used for either analysis or synthesis. As an example, we will considerbackward chaining analysis, exemplified by the problem of a robot preschool-student trying todetermine whether a new playmate “Bob” is likely to be a regular visitor to its preschool or not(evaluating the truth value of the implication Bob → regular_visitor). The basic backwardchaining process for PLN analysis looks like:1. Given an implication L ≡ A → B whose truth value must be estimated (for instanceL ≡ C&P → G as discussed above), create a list (A 1 , ..., A n ) of (inference rule, storedknowledge) pairs that might be used to produce L2. Using analogical reasoning to prior inferences, assign each A i a probability of success• If some of the A i are estimated to have reasonable probability of success at generatingreasonably confident estimates of L’s truth value, then invoke Step 1 with A i in placeof L (at this point the inference process becomes recursive)• If none of the A i looks sufficiently likely to succeed, then inference has “gotten stuck”and another cognitive process should be invoked, e.g.– Concept creation may be used to infer new concepts related to A and B, and thenStep 1 may be revisited, in the hope of finding a new, more promising A i involvingone of the new concepts– MOSES may be invoked with one of several special goals, e.g. the goal of findinga procedure P so that P (X) predicts whether X → B. If MOSES finds such aprocedure P then this can be converted to declarative knowledge understandableby PLN and Step 1 may be revisited....– Simulations may be run in CogPrime’s internal simulation engine, so as to observethe truth value of A → B in the simulations; and then Step 1 may be revisited....The combinatorial explosion of inference control is combatted by the capability to defer toother cognitive processes when the inference control procedure is unable to make a sufficientlyconfident choice of which inference steps to take next. Note that just as MOSES may relyon PLN to model its evolving populations of procedures, PLN may rely on MOSES to createcomplex knowledge about the terms in its logical implications. This is just one example of themultiple ways in which the different cognitive processes in CogPrime interact synergetically; amore thorough treatment of these interactions is given in Chapter 49.In the “new playmate” example, the interesting case is where the robot initially seems notto know enough about Bob to make a solid inferential judgment (so that none of the A i seemparticularly promising). For instance, it might carry out a number of possible inferences and notcome to any reasonably confident conclusion, so that the reason none of the A i seem promisingis that all the decent-looking ones have been tried already. So it might then recourse to MOSES,simulation or concept creation.8.7 Is Cognitive Synergy Tricky? 157For instance, the PLN controller could make a list of everyone who has been a regularvisitor, and everyone who has not been, and pose MOSES the task of figuring out a procedurefor distinguishing these two categories. This procedure could then used directly to make theneeded assessment, or else be translated into logical rules to be used within PLN inference. Forexample, perhaps MOSES would discover that older males wearing ties tend not to becomeregular visitors. If the new playmate is an older male wearing a tie, this is directly applicable.But if the current playmate is wearing a tuxedo, then PLN may be helpful via reasoning thateven though a tuxedo is not a tie, it’s a similar form of fancy dress – so PLN may extend theMOSES-learned rule to the present case and infer that the new playmate is not likely to be aregular visitor.8.7 Is Cognitive Synergy Tricky?1In this section we use the notion of cognitive synergy to explore a question that arisesfrequently in the AGI community: the well-known difficulty of measuring intermediate progresstoward human-level AGI. We explore some potential reasons underlying this, via extending thenotion of cognitive synergy to a more refined notion of "tricky cognitive synergy." These ideasare particularly relevant to the problem of creating a roadmap toward AGI, as we’ll explore inChapter 17 below.8.7.1 The Puzzle: Why Is It So Hard to Measure Partial ProgressToward Human-Level AGI?It’s not entirely straightforward to create tests to measure the final achievement of human-levelAGI, but there are some fairly obvious candidates here. There’s the Turing Test (fooling judgesinto believing you’re human, in a text chat), the video Turing Test, the Robot College Studenttest (passing university, via being judged exactly the same way a human student would), etc.There’s certainly no agreement on which is the most meaningful such goal to strive for, butthere’s broad agreement that a number of goals of this nature basically make sense.On the other hand, how does one measure whether one is, say, 50 percent of the way tohuman-level AGI? Or, say, 75 or 25 percent?It’s possible to pose many "practical tests" of incremental progress toward human-level AGI,with the property that if a proto-AGI system passes the test using a certain sort of architectureand/or dynamics, then this implies a certain amount of progress toward human-level AGI basedon particular theoretical assumptions about AGI. However, in each case of such a practical test,it seems intuitively likely to a significant percentage of AGI researchers that there is some wayto "game" the test via designing a system specifically oriented toward passing that test, andwhich doesn’t constitute dramatic progress toward AGI.Some examples of practical tests of this nature would be1 This section co-authored with Jared Wigmore158 8 Cognitive Synergy• The Wozniak "coffee test": go into an average American house and figure out how to makecoffee, including identifying the coffee machine, figuring out what the buttons do, findingthe coffee in the cabinet, etc.• Story understanding – reading a story, or watching it on video, and then answering questionsabout what happened (including questions at various levels of abstraction)• Graduating (virtual-world or robotic) preschool• Passing the elementary school reading curriculum (which involves reading and answeringquestions about some picture books as well as purely textual ones)• Learning to play an arbitrary video game based on experience only, or based on experienceplus reading instructionsOne interesting point about tests like this is that each of them seems to some AGI researchersto encapsulate the crux of the AGI problem, and be unsolvable by any system not far alongthe path to human-level AGI – yet seems to other AGI researchers, with different conceptualperspectives, to be something probably game-able by narrow-AI methods. And of course, giventhe current state of science, there’s no way to tell which of these practical tests really can besolved via a narrow-AI approach, except by having a lot of people try really hard over a longperiod of time.A question raised by these observations is whether there is some fundamental reason whyit’s hard to make an objective, theory-independent measure of intermediate progress towardadvanced AGI. Is it just that we haven’t been smart enough to figure out the right test – or isthere some conceptual reason why the very notion of such a test is problematic?We don’t claim to know for sure – but in the rest of this section we’ll outline one possiblereason why the latter might be the case.8.7.2 A Possible Answer: Cognitive Synergy is Tricky!Why might a solid, objective empirical test for intermediate progress toward AGI be an infeasiblenotion? One possible reason, we suggest, is precisely cognitive synergy, as discussedabove.The cognitive synergy hypothesis, in its simplest form, states that human-level AGI intrinsicallydepends on the synergetic interaction of multiple components (for instance, as inCogPrime, multiple memory systems each supplied with its own learning process). In this hypothesis,for instance, it might be that there are 10 critical components required for a humanlevelAGI system. Having all 10 of them in place results in human-level AGI, but having only8 of them in place results in having a dramatically impaired system – and maybe having only6 or 7 of them in place results in a system that can hardly do anything at all.Of course, the reality is almost surely not as strict as the simplified example in the aboveparagraph suggests. No AGI theorist has really posited a list of 10 crisply-defined subsystemsand claimed them necessary and sufficient for AGI. We suspect there are many different routesto AGI, involving integration of different sorts of subsystems. However, if the cognitive synergyhypothesis is correct, then human-level AGI behaves roughly like the simplistic example in theprior paragraph suggests. Perhaps instead of using the 10 components, you could achieve humanlevelAGI with 7 components, but having only 5 of these 7 would yield drastically impairedfunctionality – etc. Or the point could be made without any decomposition into a finite setof components, using continuous probability distributions. To mathematically formalize the8.7 Is Cognitive Synergy Tricky? 159cognitive synergy hypothesis becomes complex, but here we’re only aiming for a qualitativeargument. So for illustrative purposes, we’ll stick with the "10 components" example, just forcommunicative simplicity.Next, let’s suppose that for any given task, there are ways to achieve this task using a systemthat is much simpler than any subset of size 6 drawn from the set of 10 components neededfor human-level AGI, but works much better for the task than this subset of 6 components(assuming the latter are used as a set of only 6 components, without the other 4 components).Note that this supposition is a good bit stronger than mere cognitive synergy. For lack ofa better name, we’ll call it tricky cognitive synergy. The tricky cognitive synergy hypothesiswould be true if, for example, the following possibilities were true:• creating components to serve as parts of a synergetic AGI is harder than creating componentsintended to serve as parts of simpler AI systems without synergetic dynamics• components capable of serving as parts of a synergetic AGI are necessarily more complicatedthan components intended to serve as parts of simpler AGI systems.These certainly seem reasonable possibilities, since to serve as a component of a synergetic AGIsystem, a component must have the internal flexibility to usefully handle interactions with a lotof other components as well as to solve the problems that come its way. In a CogPrime context,these possibilities ring true, in the sense that tailoring an AI process for tight integration withother AI processes within CogPrime, tends to require more work than preparing a conceptuallysimilar AI process for use on its own or in a more task-specific narrow AI system.It seems fairly obvious that, if tricky cognitive synergy really holds up as a property ofhuman-level general intelligence, the difficulty of formulating tests for intermediate progresstoward human-level AGI follows as a consequence. Because, according to the tricky cognitivesynergy hypothesis, any test is going to be more easily solved by some simpler narrow AI processthan by a partially complete human-level AGI system.8.7.3 ConclusionWe haven’t proved anything here, only made some qualitative arguments. However, these argumentsdo seem to give a plausible explanation for the empirical observation that positing testsfor intermediate progress toward human-level AGI is a very difficult prospect. If the theoreticalnotions sketched here are correct, then this difficulty is not due to incompetence or lackof imagination on the part of the AGI community, nor due to the primitive state of the AGIfield, but is rather intrinsic to the subject matter. And if these notions are correct, then quitelikely the future rigorous science of AGI will contain formal theorems echoing and improvingthe qualitative observations and conjectures we’ve made here.If the ideas sketched here are true, then the practical consequence for AGI developmentis, very simply, that one shouldn’t worry a lot about producing intermediary results that arecompelling to skeptical observers. Just at 2/3 of a human brain may not be of much use,similarly, 2/3 of an AGI system may not be much use. Lack of impressive intermediary resultsmay not imply one is on a wrong development path; and comparison with narrow AI systems onspecific tasks may be badly misleading as a gauge of incremental progress toward human-levelAGI.160 8 Cognitive SynergyHopefully it’s clear that the motivation behind the line of thinking presented here is a desireto understand the nature of general intelligence and its pursuit – not a desire to avoid testing ourAGI software! Really, as AGI engineers, we would love to have a sensible rigorous way to test ourintermediary progress toward AGI, so as to be able to pose convincing arguments to skeptics,funding sources, potential collaborators and so forth. Our motivation here is not a desire toavoid having the intermediate progress of our efforts measured, but rather a desire to explainthe frustrating (but by now rather well-established) difficulty of creating such intermediategoals for human-level AGI in a meaningful way.If we or someone else figures out a compelling way to measure partial progress toward AGI,we will celebrate the occasion. But it seems worth seriously considering the possibility that thedifficulty in finding such a measure reflects fundamental properties of general intelligence.From a practical CogPrime perspective, we are interested in a variety of evaluation andtesting methods, including the "virtual preschool" approach mentioned briefly above and moreextensively in later chapters. However, our focus will be on evaluation methods that give usmeaningful information about CogPrime’s progress, given our knowledge of how CogPrimeworks and our understanding of the underlying theory. We are unlikely to focus on the achievementof intermediate test results capable of convincing skeptics of the reality of our partialprogress, because we have not yet seen any credible tests of this nature, and because we suspectthe reasons for this lack may be rooted in deep properties of feasible general intelligence, suchas tricky cognitive synergy.Chapter 9General Intelligence in the Everyday HumanWorld9.1 IntroductionIntelligence is not just about what happens inside a system, but also about what happens outsidethat system, and how the system interacts with its environment. Real-world general intelligenceis about intelligence relative to some particular class of environments, and human-like generalintelligence is about intelligence relative to the particular class of environments that humansevolved in (which in recent millennia has included environments humans have created usingtheir intelligence). In Chapter 2, we reviewed some specific capabilities characterizing humanlikegeneral intelligence; to connect these with the general theory of general intelligence from thelast few chapters, we need to explain what aspects of human-relevant environments correspondto these human-like intelligent capabilities. We begin with aspects of the environment relatedto communication, which turn out to tie in closely with cognitive synergy. Then we turn tophysical aspects of the environment, which we suspect also connect closely with various humancognitive capabilities. Finally we turn to physical aspects of the human body and their relevanceto the human mind. In the following chapter we present a deeper, more abstract theoreticalframework encompassing these ideas.These ideas are of theoretical importance, and they’re also of practical importance when oneturns to the critical area of AGI environment design. If one is going to do anything besidesrelease one’s young AGI into the “wilds” of everyday human life, then one has to put somethought into what kind of environment it will be raised in. This may be a virtual world or itmay be a robot preschool or some other kind of physical environment, but in any case somespecific choices must be made about what to include. Specific choices must also be made aboutwhat kind of body to give one’s AGI system – what sensors and actuators, and so forth. InChapter 16 we will present some specific suggestions regarding choices of embodiment andenvironment that we find to be ideal for AGI development – virtual and robot preschools – butthe material in this chapter is of more general import, beyond any such particularities. If onehas an intuitive idea of what properties of body and world human intelligence is biased for,then one can make practical choices about embodiment and environment in a principled ratherthan purely ad hoc or opportunistic way.161162 9 General Intelligence in the Everyday Human World9.2 Some Broad Properties of the Everyday World That HelpStructure IntelligenceThe properties of the everyday world that help structure intelligence are diverse and spanmultiple levels of abstraction. Most of this chapter will focus on fairly concrete patterns of thisnature, such as are involved in inter-agent communication and naive physics; however, it’s alsoworth noting the potential importance of more abstract patterns distinguishing the everydayworld from arbitrary mathematical environments.The propensity to search for hierarchical patterns is one huge potential example of an abstracteveryday-world property. We strongly suspect the reason that searching for hierarchicalpatterns works so well, in so many everyday-world contexts, lies in the particular structure ofthe everyday world – it’s not something that would be true across all possible environments(even if one weights the space of possible environments in some clever way, say using programlengthaccording to some standard computational model). However, this sort of assertion is ofcourse highly “philosophical,” and becomes complex to formulate and defend convincingly giventhe current state of science and mathematics.Going one step further, we recall from Chapter 3 a structure called the “dual network”, whichconsists of superposed hierarchical and heterarchical networks: basically a hierarchy in whichthe distance between two nodes in the hierarchy is correlated with the distance between thenodes in some metric space. Another high level property of the everyday world may be that dualnetwork structures are prevalent. This would imply that minds biased to represent the world interms of dual network structure are likely to be intelligent with respect to the everyday world.In a different direction, the extreme commonality of symmetry groups in the (everyday andotherwise) physical world is another example: they occur so often that minds oriented towardrecognizing patterns involving symmetry groups are likely to be intelligent with respect to thereal world.We suspect that the number of cognitively-relevant properties of the everyday world is huge... and that the essence of everyday-world intelligence lies in the list of varyingly abstract andconcrete properties, which must be embedded implicitly or explicitly in the structure of a naturalor artificial intelligence for that system to have everyday-world intelligence.Apart from these particular yet abstract properties of the everyday world, intelligence is justabout “finding patterns in which actions tend to achieve which goals in which situations” ... but,the simple meta-algorithm needed to accomplish this universally is, we suggest, only a smallpercentage what it takes to make a mind.You might say that a sufficiently generally intelligent system should be able to infer thevarious cognitively-relevant properties of the environment from looking at data about the everydayworld. We agree in principle, and in fact Ben Kuipers and his colleagues have donesome interesting work in this direction, showing that learning algorithms can infer some basicsabout the structure of space and time from experience [MK07]. But we suggest that doing thisreally thoroughly would require a massively greater amount of processing power than an AGIthat embodies and hence automatically utilizes these principles. It may be that the problem ofinferring these properties is so hard as to require a wildly infeasible AIXI tl / Godel Machinetype system.9.3 Embodied Communication 1639.3 Embodied CommunicationNext we turn to the potential cognitive implications of seeking to achieve goals in an environmentin which multimodal communication with other agents plays a prominent role.Consider a community of embodied agents living in a shared world, and suppose that theagents can communicate with each other via a set of mechanisms including:• Linguistic communication, in a language whose semantics is largely (not necessarilywholly) interpretable based on the mutually experienced world• Indicative communication, in which e.g. one agent points to some part of the world ordelimits some interval of time, and another agent is able to interpret the meaning• Demonstrative communication, in which an agent carries out a set of actions in theworld, and the other agent is able to imitate these actions, or instruct another agent as tohow to imitate these actions• Depictive communication, in which an agent creates some sort of (visual, auditory, etc.)construction to show another agent, with a goal of causing the other agent to experiencephenomena similar to what they would experience upon experiencing some particular entityin the shared environment• Intentional communication, in which an agent explicitly communicates to another agentwhat its goal is in a certain situation 1It is clear that ordinary everyday communication between humans possesses all these aspects.We define the Embodied Communication Prior (ECP) as the probability distribution inwhich the probability of an entity (e.g. a goal or environment) is proportional to the difficulty ofdescribing that entity, for a typical member of the community in question, using a particular setof communication mechanisms including the above five modes. We will sometimes refer to theprior probability of an entity under this distribution, as its “simplicity” under the distribution.Next, to further specialize the Embodied Communication Prior, we will assume that foreach of these modes of communication, there are some aspects of the world that are muchmore easily communicable using that mode than the other modes. For instance, in the humaneveryday world:• Abstract (declarative) statements spanning large classes of situations are generally mucheasier to communicate linguistically• Complex, multi-part procedures are much easier to communicate either demonstratively, orusing a combination of demonstration with other modes• Sensory or episodic data is often much easier to communicate demonstratively• The current value of attending to some portion of the shared environment is often mucheasier to communicate indicatively• Information about what goals to follow in a certain situation is often much easier to communicateintentionally, i.e. via explicitly indicating what one’s own goal isThese simple observations have significant implications for the nature of the Embodied CommunicationPrior. For one thing they let us define multiple forms of knowledge:• Isolatedly declarative knowledge is that which is much more easily communicable linguistically1 in Appendix ?? we recount some interesting recent results showing that mirror neurons fire in response tosome cases of intentional communication as thus defined164 9 General Intelligence in the Everyday Human World• Isolatedly procedural knowledge is that which is much more easily communicabledemonstratively• Isolatedly sensory knowledge is that which is much more easily communicable depictively• Isolatedly attentive knowledge is that which is much more easily communicable indicatively• Isolatedly intentional knowledge is that which is much more easily communicable intentionallyThis categorization of knowledge types resembles many ideas from the cognitive theory ofmemory [TC05], although the distinctions drawn here are a little crisper than any classificationcurrently derivable from available neurological or psychological data.Of course there may be much knowledge, of relevance to systems seeking intelligence accordingto the ECP, that does not fall into any of these categories and constitutes “mixed knowledge.”There are some very important specific subclasses of mixed knowledge. For instance, episodicknowledge (knowledge about specific real or hypothetical sets of events) will most easily becommunicated via a combination of declarative, sensory and (in some cases) procedural communication.Scientific and mathematical knowledge are generally mixed knowledge, as is mosteveryday commonsense knowledge.Some cases of mixed knowledge are reasonably well decomposable, in the sense that theydecompose into knowledge items that individually fall into some specific knowledge type. Forinstance, an experimental chemistry procedure may be much more easily communicable procedurally,whereas an allied piece of knowledge from theoretical chemistry may be much moreeasily communicable declaratively; but in order to fully communicate either the experimentalprocedure or the abstract piece of knowledge, one may ultimately need to communicate bothaspects.Also, even when the best way to communicate something is mixed-mode, it may be possibleto identify one mode that poses the most important part of the communication. An examplewould be a chemistry experiment that is best communicated via a practical demonstrationtogether with a running narrative. It may be that the demonstration without the narrativewould be vastly more valuable than the narrative without the demonstration. To cover suchcases we may make less restrictive definitions such as• Interactively declarative knowledge is that which is much more easily communicablein a manner dominated by linguistic communicationand so forth. We call these “interactive knowledge categories,” by contrast to the “isolatedknowledge categories” introduced earlier.9.3.0.1 Naturalness of Knowledge CategoriesNext we introduce an assumption we call NKC, for Naturalness of Knowledge Categories.The NKC assumption states that the knowledge in each of the above isolated and interactivecommunication-modality-focused categories forms a “natural category,” in the sense thatfor each of these categories, there are many different properties shared by a large percentage ofthe knowledge in the category, but not by a large percentage of the knowledge in the other categories.This means that, for instance, procedural knowledge systematically (and statistically)has different characteristics than the other kinds of knowledge.9.3 Embodied Communication 165The NKC assumption seems commonsensically to hold true for human everyday knowledge,and it has fairly dramatic implications for general intelligence. Suppose we conceive generalintelligence as the ability to achieve goals in the environment shared by the communicatingagents underlying the Embodied Communication Prior. Then, NKC suggests that the best wayto achieve general intelligence according to the Embodied Communication Prior is going toinvolve• specialized methods for handling declarative, procedural, sensory and attentional knowledge(due to the naturalness of the isolated knowledge categories)• specialized methods for handling interactions between different types of knowledge, includingmethods focused on the case where one type of knowledge is primary and the others aresupporting (the latter due to the naturalness of the interactive knowledge categories)9.3.0.2 Cognitive CompletenessSuppose we conceive an AI system as consisting of a set of learning capabilities, each onecharacterized by three features:• One or more knowledge types that it is competent to deal with, in the sense of the twokey learning problems mentioned above• At least one learning type: either analysis, or synthesis, or both• At least one interaction type, for each (knowledge type, learning type) pair it handles:“isolated” (meaning it deals mainly with that knowledge type in isolation), or “interactive”(meaning it focuses on that knowledge type but in a way that explicitly incorporates otherknowledge types into its process), or “fully mixed” (meaning that when it deals with theknowledge type in question, no particular knowledge type tends to dominate the learningprocess).Then, intuitively, it seems to follow from the ECP with NKC that systems with high efficientgeneral intelligence should have the following properties, which collectively we’ll call cognitivecompleteness:• For each (knowledge type, learning type, interaction type) triple, there should be a learningcapability corresponding to that triple.• Furthermore the capabilities corresponding to different (knowledge type, interaction type)pairs should have distinct characteristics (since according to the NKC the isolated knowledgecorresponding to a knowledge type is a natural category, as is the dominant knowledgecorresponding to a knowledge type)• For each (knowledge type, learning type) pair (K,L), and each other knowledge type K1distinct from K, there should be a distinctive capability with interaction type “interactive”and dealing with knowledge that is interactively K but also includes aspects of K1Furthermore, it seems intuitively sensible that according to the ECP with NKC, if the capabilitiesmentioned in the above points are reasonably able, then the system possessing thecapabilities will display general intelligence relative to the ECP. Thus we arrive at the hypothesisthat166 9 General Intelligence in the Everyday Human WorldUnder the assumption of the Embodied Communication Prior (with the NaturalKnowledge Categories assumption), the property above called “cognitive completeness”is necessary and sufficient for efficient general intelligence at the level of aninteligent adult human (e.g. at the Piagetan formal level [Pia53]).Of course, the above considerations are very far from a rigorous mathematical proof (oreven precise formulation) of this hypothesis. But we are presenting this here as a conceptualhypothesis, in order to qualitatively guide our practical AGI R&D and also to motivate further,more rigorous theoretical work.9.3.1 Generalizing the Embodied Communication PriorOne interesting direction for further research would be to broaden the scope of the inquiry, ina manner suggested above: instead of just looking at the ECP, look at simplicity measures ingeneral, and attack the question of how a mind must be structured in order to display efficientgeneral intelligence relative to a specified simplicity measure. This problem seems unapproachablein general, but some special cases may be more tractable.For instance, suppose one has• a simplicity measure that (like the ECP) is approximately decomposable into a set of fairlydistinct components, plus their interactions• an assumption similar to NKC, which states that the entities displaying simplicity accordingto each of the distinct components, are roughly clustered together in entity-spaceThen one should be able to say that, to achieve efficient general intelligence relative tothis decomposable simplicity measure, a system should have distinct capabilities correspondingto each of the components of the simplicity measure interactions between these capabilities,corresponding to the interaction terms in the simplicity measure.With copious additional work, these simple observations could potentially serve as the seed fora novel sort of theory of general intelligence – a theory of how the structure of a system dependson the structure of the simplicity measure with which it achieves efficient general intelligence.Cognitive Synergy Theory would then emerge as a special case of this more abstract theory.9.4 Naive PhysicsMultimodal communication is an important aspect of the environment for which human intelligenceevolved – but not the only one. It seems likely that our human intelligence is alsoclosely adapted to various aspects of our physical environment – a matter that is worth carefullyattending as we design environments for our robotically or virtually embodied AGI systems tooperate in.One interesting guide to the most cognitively relevant aspects of human environments is thesubfield of AI known as “naive physics” [Hay85] – a term that refers to the theories about thephysical world that human beings implicitly develop and utilize during their lives. For instance,9.4 Naive Physics 167when you figure out that you need to pressure the knife slightly harder when spreading peanutbutter rather than jelly, you’re not making this judgment using Newtonian physics or theNavier-Stokes equations of fluid dynamics; you’re using heuristic patterns that you figured outthrough experience. Maybe you figured out these patterns through experience spreading peanutbutter and jelly in particular. Or maybe you figured these heuristic patterns out before you evertried to spread peanut butter or jelly specifically, via just touching peanut butter and jelly tosee what they feel like, and then carrying out inference based on your experience manipulatingsimilar tools in the context of similar substances.Other examples of similar “naive physics” patterns are easy to come by, e.g.1. What goes up must come down.2. A dropped object falls straight down.3. A vacuum sucks things towards it.4. Centrifugal force throws rotating things outwards.5. An object is either at rest or moving, in an absolute sense.6. Two events are simultaneous or they are not.7. When running downhill, one must lift one’s knees up high.8. When looking at something that you just barely can’t discern accurately, squint.Attempts to axiomatically formulate naive physics have historically come up short, and wedoubt this is a promising direction for AGI. However, we do think the naive physics literaturedoes a good job of identifying the various phenomena that the human mind’s naive physics dealswith. So, from the point of view of AGI environment design, naive physics is a useful sourceof requirements. Ideally, we would like an AGI’s environment to support all the fundamentalphenomena that naive physics deals with.We now describe some key aspects of naive physics in a more systematic manner. Naivephysics has many different formulations; in this section we draw heavily on [SC94], who dividenaive physics phenomena into 5 categories. Here we review these categories and identify anumber of important things that humanlike intelligent agents must be able to do relative toeach of them.9.4.1 Objects, Natural Units and Natural KindsOne key aspect of naive physics involves recognition of various aspects of objects, such as:1. Recognition of objects amidst noisy perceptual data2. Recognition of surfaces and interiors of objects3. Recognition of objects as manipulable units4. Recognition of objects as potential subjects of fragmentation (splitting, cutting) and ofunification (gluing, bonding)5. Recognition of the agent’s body as an object, and as parts of the agent’s body as objects6. Division of universe of perceived objects into “natural kinds”, each containing typical andatypical instances168 9 General Intelligence in the Everyday Human World9.4.2 Events, Processes and CausalitySpecific aspects of naive physics related to temporality and causality are:1. Distinguishing roughly-subjectively-instantaneous events from extended processes2. Identifying beginnings, endings and crossings of processes3. Identifying and distinguishing internal and external changes4. Identifying and distinguishing internal and external changes relative to one’s own body5. Interrelating body-changes with changes in external entitiesNotably, these aspects of naive physics involve a different processes occurring on a variety ofdifferent time scales, intersecting in complex patterns, and involving processes inside the agent’sbody, outside the agent’s body, and crossing the boundary of the agent’s body.9.4.3 Stuffs, States of Matter, QualitiesRegarding the various states of matter, some important aspects of naive physics are:1. Perceiving gaps between objects: holes, media, illusions like rainbows, mirages and holograms2. Distinguishing the manners in which different sorts of entities (e.g. smells, sounds, light) fillspace3. Distinguishing properties such as smoothness, roughness, graininess, stickiness, runniness,etc.4. Distinguishing degrees of elasticity and fragility5. Assessing separability of aggregates9.4.4 Surfaces, Limits, Boundaries, MediaGibson [Gib77, Gib79] has argued that naive physics is not mainly about objects but rathermainly about surfaces. Surfaces have a variety of aspects and relationships that are importantfor naive physics, such as:1. Perceiving and reasoning about surfaces as two-sided or one-sided interfaces2. Inference of the various ecological laws of surfaces3. Perception of various media in the world as separated by surfaces4. Recognition of the textures of surfaces5. Recognition of medium/surface layout relationships such as: ground, open environment,enclosure, detached object, attached object, hollow object, place, sheet, fissure, stick, fibre,dihedral, etc.As a concrete, evocative “toy” example of naive everyday knowledge about surfaces andboundaries, consider Sloman’s [Slo08a] example scenario, depicted in Figure 9.1 and drawnlargely from [SS74] (see also related discussion in [Slo08b], in which “A child can be given one9.4 Naive Physics 169Fig. 9.1: One of Sloman’s example test domains for real-world inference. Left: a number of pinsand a rubber band to be stretched around them. Right: use of the pins and rubber band tomake a letter T.or more rubber bands and a pile of pins, and asked to use the pins to hold the band in place toform a particular shape)... For example, things to be learnt could include”:1. There is an area inside the band and an area outside the band.2. The possible effects of moving a pin that is inside the band towards or further awayfrom other pins inside the band. (The effects can depend on whether the band is alreadystretched.)3. The possible effects of moving a pin that is outside the band towards or further away fromother pins inside the band.4. The possible effects of adding a new pin, inside or outside the band, with or without pushingthe band sideways with the pin first.5. The possible effects of removing a pin, from a position inside or outside the band.6. Patterns of motion/change that can occur and how they affect local and global shape(e.g. introducing a concavity or convexity, introducing or removing symmetry, increasing ordecreasing the area enclosed).7. The possibility of causing the band to cross over itself. (NB: Is an odd number of crossespossible?)8. How adding a second, or third band can enrich the space of structures, processes and effectsof processes.9.4.5 What Kind of Physics Is Needed to Foster Human-likeIntelligence?We stated above that we would like an AGI’s environment to support all the fundamental phenomenathat naive physics deals with; and we have now reviewed a number of these specificphenomena. But it’s not entirely clear what the “fundamental” aspects underlying these phenomenaare. One important question in the environment-design context is how close an AGIenvironment needs to stick to the particulars of real-world naive physics. Is it important that ayoung AGI can play with the specific differences between spreading peanut butter versus jelly?Or is it enough that it can play with spreading and smearing various substances of differentconsistencies? How close does the analogy between an AGI environment’s naive physics and170 9 General Intelligence in the Everyday Human Worldreal-world naive physics need to be? This is a question to which we have no scientific answer atpresent. Our own working hypothesis is that the analogy does not need to be extremely close,and with this in mind in Chapter 16 we propose a virtual environment BlocksNBeadsWorldthat encompasses all the basic conceptual phenomena of real-world naive physics, but does notattempt to emulate their details.Framed in terms of human psychology rather than environment design, the question becomes:At what level of detail must one model the physical world to understand the ways inwhich human intelligence has adapted to the physical world?. Our suspicion, which underliesour BlocksNBeadsWorld design, is that it’s approximately enough to have• Newtonian physics, or some close approximation• Matter in multiple phases and forms vaguely similar to the ones we see in the real world:solid, liquid, gas, paste, goo, etc.• Ability to transform some instances of matter from one form to another• Ability to flexibly manipulate matter in various forms with various solid tools• Ability to combine instances of matter into new ones in a fairly rich way: e.g. glue or tiesolids togethermix liquids together, etc.• Ability to position instances of matter with respect to each other in a rich way: e.g. putliquid in a solid cavity, cover something with a lid or a piece of fabric, etc.It seems to us that if the above are present in an environment, then an AGI seeking toachieve appropriate goals in that environment will be likely to form an appropriate “humanlikephysical-world intuition." We doubt that the specifics of the naive physics of differentforms of matter are critical to human-like intelligence. But, we suspect that a great amountof unconscious human metaphorical thinking is conditioned on the fact that humans evolvedaround matter that takes a variety of forms, can be changed from one form to another, and canbe fairly easily arranged and composited to form new instances from prior ones. Without manydiverse instances of matter transformation, arrangement and composition in its experience, anAGI is unlikely to form an internal “metaphor-base” even vaguely similar to the human one –so that, even if it’s highly intelligent, its thinking will be radically non-human-like in character.Naturally this is all somewhat speculative and must be explored via experimentation. Maybean elaborate blocks-world with only solid objects will be sufficient to create human-level, roughlyhuman-like AGI with rich spatiotemporal and manipulative intuition. Or maybe human intelligenceis more closely adapted to the specifics of our physical world – with water and dirt andplants and hair and so forth – than we currently realize. One thing that is very clear is that, aswe proceed with embodying, situating and educating our AGI systems, we need to pay carefulattention to the way their intelligence is conditioned by their environment.9.5 Folk PsychologyRelated to naive physics is the notion of “naive psychology” or “folk psychology” [Rav04], whichincludes for instance the following aspects:1. Mental simulation of other agents2. Mental theory regarding other agents3. Attribution of beliefs, desires and intentions (BDI) to other agents via theory or simulation9.6 Body and Mind 1714. Recognition of emotions in other agents via their physical embodiment5. Recognition of desires and intentions in other agents via their physical embodiment6. Analogical and contextual inferences between self and other, regarding BDI and other aspects7. Attribute causes and meanings to other agents behaviors8. Anthropomorphize non-human, including inanimate objectsThe main special requirement placed on an AGI’s embodiment by the above aspects pertainsto the ability of agents to express their emotions and intentions to each other. Humans do thisvia facial expressions, gestures and language.9.5.1 Motivation, Requiredness, ValueRelatedly to folk psychology, Gestalt [Koh38] and ecological [Gib77, Gib79] psychology suggestthat humans perceive the world substantially in terms of the affordances it provides them forgoal-directed action. This suggests that, to support human-like intelligence, an AGI must becapable of:1. Perception of entities in the world as differentially associated with goal-relevant value2. Perception of entities in the world in terms of the potential actions they afford the agent,or other agentsThe key point is that entities in the world need to provide a wide variety of ways for agentsto interact with them, enabling richly complex perception of affordances.9.6 Body and MindThe above discussion has focused on the world external to the body of the AGI agent embodiedand embedded in the world, but the issue of the AGI’s body also merits consideration. Thereseems little doubt that a human’s intelligence is highly conditioned by the particularities of thehuman body.9.6.1 The Human SensoriumHere the requirements seem fairly simple: while surely not strictly necessary, it would certainlybe preferable to provide an AGI with fairly rich analogues of the human senses of touch, sight,sound, kinesthesia, taste and smell. Each of these senses provides different sorts of cognitivestimulation to the human mind; and while similar cognitive stimulation could doubtless beachieved without analogous senses, the provision of such seems the most straightforward approach.It’s hard to know how much of human intelligence is specifically biased to the sorts ofoutputs provided by human senses.As vision already is accorded such a prominent role in the AI and cognitive science literature– and is discussed in moderate depth in Chapter 26 of Part 2, we won’t take time elaborating172 9 General Intelligence in the Everyday Human Worldon the importance of vision processing for humanlike cognition. The key thing an AGI requiresto support humanlike “visual intelligence” is an environment containing a sufficiently robustcollection of materials that object and event recognition and identification become interestingproblems.Audition is cognitively valuable for many reasons, one of which is that it gives a very richand precise method of sensing the world that is different from vision. The fact that humans candisplay normal intelligence while totally blind or totally deaf is an indication that, in a sense,vision and audition are redundant for understanding the everyday world. However, it may beimportant that the brain has evolved to account for both of these senses, because this forced itto account for the presence of two very rich and precise methods of sensing the world – whichmay have forced it to develop more abstract representation mechanisms than would have beennecessary with only one such method.Touch is a sense that is, in our view, generally badly underappreciated within the AI community.In particular the cognitive robotics community seems to worry too little about the terriblyimpoverished sense of touch possessed by most current robots (though fortunately there arerecent technologies that may help improve robots in this regard; see e.g. [Nan08]). Touch is howthe human infant learns to distinguish self from other, and in this way it is the most essentialsense for the establishment of an internal self-model. Touching others’ bodies is a key methodfor developing a sense of the emotional reality and responsiveness of others, and is hence key tothe development of theory of mind and social understanding in humans. For this reason, amongothers, human children lacking sufficient tactile stimulation will generally wind up badly impairedin multiple ways. A good-quality embodiment should supply an AI agent with a bodythat possesses skin, which has varying levels of sensitivity on different parts of the skin (so thatit can effectively distinguish between reality and its perception thereof in a tactile context);and also varying types of touch sensors (e.g. temperature versus friction), so that it experiencestextures as multidimensional entities.Related to touch, kinesthesia refers to direct sensation of phenomena happening inside thebody. Rarely mentioned in AI, this sense seems quite critical to cognition, as it underpins manyof the analogies between self and other that guide cognition. Again, it’s not important that anAGI’s virtual body have the same internal body parts as a human body. But it seems valuableto have the AGI’s virtual body display some vaguely human-body-like properties, such as feelinginternal strain of various sorts after getting exercise, feeling discomfort in certain places whenrunning out of energy, feeling internally different when satisfied versus unsatisfied, etc.Next, taste is a cognitively interesting sense in that it involves the interplay between theinternal and external world; it involves the evaluation of which entities from the external worldare worthy of placing inside the body. And smell is cognitively interesting in large part becauseof its relationship with taste. A smell is, among other things, a long-distance indicator of whata certain entity might taste like. So, the combination of taste and smell provides means forconceptualizing relationships between self, world and distance.9.6.2 The Human Body’s Multiple IntelligencesWhile most unique aspect of human intelligence is rooted in what one might call the "cognitivecortex" – the portions of the brain dealing with self-reflection and abstract thought. But thecognitive cortex does its work in close coordination with the body’s various more specialized9.6 Body and Mind 173intelligent subsystems, including those associated with the gut, the heart, the liver, the immuneand endocrine systems, and the perceptual and motor cortices.In the perspective underlying this book, the human cognitive cortex – or the core cognitivenetwork of any roughly human-like AGI system – should be viewed as a highly flexible, selforganizingnetwork. These cognitive networks are modelable e.g. as a recurrent neural net withgeneral topology, or a weighted labeled hypergraph, and are centrally concerned with recognizingpatterns in its environment and itself, especially patterns regarding the achievement of thesystem’s goals in various appropriate contexts. Here we augment this perspective, noting thatthe human brain’s cognitive network is closely coupled with a variety of simpler and morespecialized intelligent "body-system networks" which provide it with structural and dynamicalinductive biasing. We then discuss the implications of this observation for practical AGI design.One recalls Pascal’s famous quote "The heart has its reasons, of which reason knows not."As we now know, the intuitive sense that Pascal and so many others have expressed, that theheart and other body systems have their own reasons, is grounded in the fact that they actuallydo carry out simple forms of reasoning (i.e. intelligent, adaptive dynamics), in close, sometimescognitively valuable, coordination with the central cognitive network.9.6.2.1 Some of the Human Body’s Specialized Intelligent SubsystemsThe human body contains multiple specialized intelligences apart from the cognitive cortex.Here we review some of the most critical.Hierarchies of Visual and Auditory Perception. The hierarchical structure of visual and auditory cortex has been taken by some researchers[Kur12], [HB06] as the generic structure of cognition. While we suspect this is overstated, weagree it is important that these cortices nudge large portions of the cognitive cortex to assumean approximately hierarchical structure.Olfactory Attractors. The process of recognizing a familiar smell is grounded in a neural process similar to convergenceto an attractor in a nonlinear dynamical system [Fre95]. There is evidence that themammalian cognitive cortex evolved in close coordination with the olfactory cortex [Row11],and much of abstract cognition reflects a similar dynamic of gradually coming to a conclusionbased on what initially "smells right."Physical and Cognitive Action. The cerebellum, a specially structured brain subsystem which controls motor movements,has for some time been understood to also have involvement in attention, executive control,language, working memory, learning, pain, emotion, and addiction [PSF09].174 9 General Intelligence in the Everyday Human WorldThe Second Brain. The gastrointestinal neural net contains millions of neurons and is capable of operating independentlyof the brain. It modulates stress response and other aspects of emotion and motivationbased on experience – resulting in so-called "gut feelings" [Ger99].The Heart’s Neural Network. The heart has its own neural network, which modulates stress response, energy level andrelaxation/excitement (factors key to motivation and emotion) based on experience [Arm04].Pattern Recognition and Memory in the Liver. The liver is a complex pattern recognition system, adapting via experience to better identifytoxins [CB06]. Like the heart, it seems to store some episodic memories as well, resulting in livertransplant recipients sometimes acquiring the tastes in music or sports of the donor [EMC12].Immune Intelligence. The immune network is a highly complex, adaptive self-organizing system, which ongoinglysolves the learning problem of identifying antigens and distinguishing them from the bodysystem [FP86]. As immune function is highly energetically costly, stress response involves subtlemodulation of the energy allocation to immune function, which involves communication betweenneural and immune networks.The Endocrine System: A Key Bridge Between Mind and Body. The endocrine (hormonal) system regulates (and is related by) emotion, thus guiding allaspects of intelligence (due to the close connection of emotion and motivation) [PH12].Breathing Guides Thinking. As oxygenation of the brain plays a key role in the spread of neural activity, the flow of breathis a key driver of cognition. Forced alternate nostril breathing has been shown to significantlyaffect cognition via balancing activity of the two brain hemispheres [SKBB91].Much remains unknown, and the totality of feedback loops between the human cognitivecortex and the various specialized intelligences operative throughout the human body, has notyet been thoroughly charted.9.6 Body and Mind 1759.6.2.2 Implications for AGIWhat lesson should the AGI developer draw from all this? The particularities of the humanmind/body should not be taken as general requirements for general intelligence. However, itis worth remembering just how difficult is the computational problem of learning, based onexperiential feedback alone, the right way to achieve the complex goal of controlling a systemwith general intelligence at the human level or beyond. To solve this problem without some sortof strong inductive biasing may require massively more experience than young humans obtain.Appropriate inductive bias may be embedded in an AGI system in many different ways.Some AGI designers have sought to embed it very explicitly, e.g. with hand-coded declarativeknowledge as in Cyc, SOAR and other "GOFAI" type systems. On the other hand, the humanbrain receives its inductive bias much more subtly and implicitly, both via the specifics of theinitial structure of the cognitive cortex, and via ongoing coupling of the cognitive cortex withother systems possessing more focused types of intelligence and more specific structures and/ordynamics.In building an AGI system, one has four choices, very broadly speaking:1. Create a flexible mind-network, as unbiased as feasible, and attempt to have it learn howto achieve its goals via experience2. Closely emulate key aspects of the human body along with the human mind3. Imitate the human mind-body, conceptually if not in detail, and create a number of structurallyand dynamically simpler intelligent systems closely and appropriately coupled tothe abstract cognitive mind-network, provide useful inductive bias.4. Find some other, creative way to guide and probabilistically constrain one’s AGI system’smind-network, providing inductive bias appropriate to the tasks at hand, without emulatingeven conceptually the way the human mind-brain receives its inductive bias via couplingwith simpler intelligent systems.Our suspicion is that the first option will not be viable. On the other hand, to do the secondoption would require more knowledge of the human body than biology currently possesses. Thisleaves the third and fourth options, both of which seem viable to us.CogPrime incorporates a combination of the third and fourth options. CogPrime’s genericdynamic knowledge store, the Atomspace, is coupled with specialized hierarchical networks(DeSTIN) for vision and audition, somewhat mirroring the human cortex. An artificial endocrinesystem for OpenCog is also under development, speculatively, as part of a project usingOpenCog to control humanoid robots. On the other hand, OpenCog has no gastrointestinal norcardiological nervous system, and the stress-response-based guidance provided to the humanbrain by a combination of the heart, gut, immune system and other body systems, is achievedin CogPrime in a more explicit way using the OpenPsi model of motivated cognition, and itsintegration with the system’s attention allocation dynamics.Likely there is no single correct way to incorporate the lessons of intelligent human bodysystemnetworks into AGI designs. But these are aspects of human cognition that all AGIresearchers should be aware of.176 9 General Intelligence in the Everyday Human World9.7 The Extended Mind and BodyFinally, Hutchins [Hut95], Logan [Log07] and others have promoted a view of human intelligencethat views the human mind as extended beyond the individual body, incorporating socialinteractions and also interactions with inanimate objects, such as tools, plants and animals.This leads to a number of requirements for a humanlike AGI’s environment:1. The ability to create a variety of different tools for interacting with various aspects of theworld in various different ways, including tools for making tools and ultimately machinery2. The existence of other mobile, virtual life-forms in the world, including simpler and lessintelligent ones, and ones that interact with each other and with the AGI3. The existence of organic growing structures in the world, with which the AGI can interactin various ways, including halting their growth or modifying their growth patternHow necessary these requirements are is hard to say – but it is clear that these things haveplayed a major role in the evolution of human intelligence.9.8 ConclusionHappily, this diverse chapter supports a simple, albeit tentative conclusion. Our suggestion isthat, if an AGI is• placed in an environment capable of roughly supporting multimodal communication andvaguely (but not necessarily precisely) real-world-ish naive physics• surrounded with other intelligent agents of varying levels of complexity, and other complex,dynamic structures to interface with• given a body that can perceive this environment through some forms of sight, sound andtouch; and perceive itself via some form of kinesthesia• given a motivational system that encourages it to make rich use of these aspects of itsenvironmentthen the AGI is likely to have an experience-base reinforcing the key inductive biases providedby the everyday world for the guidance of humanlike intelligence.Chapter 10A Mind-World Correspondence Principle10.1 IntroductionReal-world minds are always adapted to certain classes of environments and goals. The ideasof the previous chapter, regarding the connection between a human-like intelligence’s internalsand its environment, result from exploring the implications of this adaptation in the contextof the cognitive synergy concept. In this chapter we explore the mind-world connection in abroader and more abstract way – making a more ambitious attempt to move toward a "generaltheory of general intelligence."One basic premise here, as in the preceding chapters is: Even a system of vast generalintelligence, subject to real-world space and time constraints, will necessarily be more efficientat some kinds of learning than others. Thus, one approach to formulating a general theory ofgeneral intelligence is to look at the relationship between minds and worlds – where a "world"is conceived as an environment and a set of goals defined in terms of that environment.In this spirit, we here formulate a broad principle binding together worlds and the minds thatare intelligent in these worlds. The ideas of the previous chapter constitute specific, concreteinstantiations of this general principle. A careful statement of the principle requires introductionof a number of technical concepts, and will be given later on in the chapter. A crude, informalversion of the principle would be:MIND-WORLD CORRESPONDENCE-PRINCIPLEFor a mind to work intelligently toward certain goals in a certain world, there should be anice mapping from goal-directed sequences of world-states into sequences of mind-states, where"nice" means that a world-state-sequence W composed of two parts W 1 and W 2 , gets mappedinto a mind-state-sequence M composed of two corresponding parts M 1 and M 2 .What’s nice about this principle is that it relates the decomposition of the world into parts,to the decomposition of the mind into parts.177178 10 A Mind-World Correspondence Principle10.2 What Might a General Theory of General Intelligence LookLike?It’s not clear, at this point, what a real "general theory of general intelligence" would look like– but one tantalizing possibility is that it might confront the two questions:• How does one design a world to foster the development of a certain sort of mind?• How does one design a mind to match the particular challenges posed by a certain sort ofworld?One way to achieve this would be to create a theory that, given a description of an environmentand some associated goals, would output a description of the structure and dynamics that asystem should possess to be intelligent in that environment relative to those goals, using limitedcomputational resources.Such a theory would serve a different purpose from the mathematical theory of "universalintelligence" developed by Marcus Hutter [Hut05] and others. For all its beauty and theoreticalpower, that approach currently gives it useful conclusions only about general intelligenceswith infinite or infeasibly massive computational resources. On the other hand, the approachsuggested here is aimed toward creation of a theory of real-world general intelligences utilizingrealistic amounts of computational power, but still possessing general intelligence comparableto human beings or greater.This reflects a vision of intelligence as largely concerned with adaptation to particular classesof environments and goals. This may seem contradictory to the notion of "general" intelligence,but I think it actually embodies a realistic understanding of general intelligence. Maximallygeneral intelligence is not pragmatically feasible; it could only be achieved using infinite computationalresources [Hut05]. Real-world systems are inevitably limited in the intelligence theycan display in any real situation, because real situations involve finite resources, including finiteamounts of time. One may say that, in principle, a certain system could solve any problemgiven enough resources and time but, even when this is true, it’s not necessarily the most interestingway to look at the system’s intelligence. It may be more important to look at what asystem can do given the resources at its disposal in reality. And this perspective leads one toask questions like the ones posed above: which bounded-resources systems are well-disposed todisplay intelligence in which classes of situations?As noted in Chapter 7 above, one can assess the generality of a system’s intelligence vialooking at the entropy of the class of situations across which it displays a high level of intelligence(where “high” is measured relative to its total level of intelligence across all situations). A systemwith a high generality of intelligence will tend to be roughly equally intelligent across a widevariety of situations; whereas a system with lower generality of intelligence will tend to be muchmore intelligent in a small subclass of situations, than in any other. The definitions given aboveembody this notion in a formal and quantitative way.If one wishes to create a general theory of general intelligence according to this sort ofperspective, the main question then becomes how to represent goals/environments and systemsin such a way as to render transparent the natural correspondence between the specifics of theformer and the latter, in the context of resource-bounded intelligence. This is the business ofthe next section.10.3 Steps Toward A (Formal) General Theory of General Intelligence 17910.3 Steps Toward A (Formal) General Theory of General IntelligenceNow begins the formalism. At this stage of development of the theory proposed in this chapter,mathematics is used mainly as a device to ensure clarity of expression. However, once the theoryis further developed, it may possibly become useful for purposes of calculation as well.Suppose one has any system S (which could be an AI system, or a human, or an environmentthat a human or AI is interacting with, or the combination of an environment and a human orAI’s body, etc.). One may then construct an uncertain transition graph associated with thatsystem S, in the following way:• The nodes of the graph represent fuzzy sets of states of system S (I’ll call these state-setsfrom here on, leaving the fuzziness implicit)• The (directed) links of the graph represent probabilistically weighted transitions betweenstate-setsSpecifically, the weight of the link from B to A should be defined aswhereP (o(S, A, t(T ))|o(S, B, T ))o(S, A, T )denotes the presence of the system S in the state-set A during time-distribution T , and t() isa temporal succession function defined so that t(T ) refers to a time-distribution conceived as"after" T . A time-distribution is a probability distribution over time-points. The interaction offuzziness and probability here is fairly straightforward and may be handled in the manner ofPLN, as outlined in subsequent chapters. Note that the definition of link weights is dependenton the specific implementation of the temporal succession function, which includes an implicittime-scale.Suppose one has a transition graph corresponding to an environment; then a goal relative tothat environment may be defined as a particular node in the transition graph. The goals of aparticular system acting in that environment may then be conceived as one or more nodes inthe transition graph. The system’s situation in the environment at any point in time may alsobe associated with one or more nodes in the transition graph; then, the system’s movementtoward goal-achievement may be associated with paths through the environment’s transitiongraph leading from its current state to goal states.It may be useful for some purposes to filter the uncertain transition graph into a crisptransition graph by placing a threshold on the link weights, and removing links with weightsbelow the threshold.The next concept to introduce is the world-mind transfer function, which maps world (environment)state-sets into organism (e.g. AI system) state-sets in a specific way. Given a worldstate-set W , the world-mind transfer function M maps W into various organism state-sets withvarious probabilities, so that we may say: M(W ) is the probability distribution of state-sets theorganism tends to be in, when its environment is in state-set W . (Recall also that state-sets arefuzzy.)Now one may look at the spaces of world-paths and mind-paths. A world-path is a paththrough the world’s transition graph, and a mind-path is a path through the organism’s transi-180 10 A Mind-World Correspondence Principletion graph. Given two world-paths P and Q, it’s obvious how to define the composition P ∗Q onefollows P and then, after that, follows Q, thus obtaining a longer path. Similarly for mind-paths.In category theory terms, we are constructing the free category associated with the graph:the objects of the category are the nodes, and the morphisms of the category are the paths.And category theory is the right way to be thinking here we want to be thinking about therelationship between the world category and the mind category.The world-mind transfer function can be interpreted as a mapping from paths to subgraphs:Given a world-path, it produces a set of mind state-sets, which have a number of links betweenthem. One can then define a world-mind path transfer function M(P ) via taking the mind-graphM(nodes(P )), and looking at the highest-weight path spanning M(nodes(P )). (Here nodes?obviously means the set of nodes of the path P .)A functor F between the world category and the mind category is a mapping that preservesobject identities and so thatF (P ∗ Q) = F (P ) ∗ F (Q)We may also introduce the notion of an approximate functor, meaning a mapping F so thatthe average ofd(F (P ∗ Q), F (P ) ∗ F (Q))is small.One can introduce a prior distribution into the average here. This could be the Levin universaldistribution or some variant (the Levin distribution assigns higher probability to computationallysimpler entities). Or it could be something more purpose specific: for example, one can givea higher weight to paths leading toward a certain set of nodes (e.g. goal nodes). Or one canuse a distribution that weights based on a combination of simplicity and directedness towarda certain set of nodes. The latter seems most interesting, and I will define a goal-weighted approximatefunctor as an approximate functor, defined with averaging relative to a distributionthat balances simplicity with directedness toward a certain set of goal nodes.The move to approximate functors is simple conceptually, but mathematically it’s a fairlybig step, because it requires us to introduce a geometric structure on our categories. But thereare plenty of natural metrics defined on paths in graphs (weighted or not), so there’s no realproblem here.10.4 The Mind-World Correspondence PrincipleNow we finally have the formalism set up to make a non-trivial statement about the relationshipbetween minds and worlds. Namely, the hypothesis that:MIND-WORLD CORRESPONDENCE PRINCIPLEFor an organism with a reasonably high level of intelligence in a certain world, relative toa certain set of goals, the mind-world path transfer function is a goal-weighted approximatefunctor.10.5 How Might the Mind-World Correspondence Principle Be Useful? 181That is, a little more loosely: the hypothesis is that, for intelligence to occur, there has to be anatural correspondence between the transition-sequences of world-states and the correspondingtransition-sequences of mind-states, at least in the cases of transition-sequences leading torelevant goals.We suspect that a variant of the above proposition can be formally proved, using the definitionof general intelligence presented in Chapter 7. The proof of a theorem corresponding to theabove would certainly constitute an interesting start toward a general formal theory of generalintelligence. Note that proving anything of this nature would require some attention to thetime-scale-dependence of the link weights in the transition graphs involved.A formally proved variant of the above proposition would be in short, a "MIND-WORLDCORRESPONDENCE THEOREM."Recall that at the start of the chapter, we expressed the same idea as:MIND-WORLD CORRESPONDENCE-PRINCIPLEFor a mind to work intelligently toward certain goals in a certain world, there should be anice mapping from goal-directed sequences of world-states into sequences of mind-states, where"nice" means that a world-state-sequence W composed of two parts W 1 and W 2 , gets mappedinto a mind-state-sequence M composed of two corresponding parts M 1 and M 2 .That is a reasonable gloss of the principle, but it’s clunkier and less accurate, than thestatement in terms of functors and path transfer functions, because it tries to use only commonlanguagevocabulary, which doesn’t really contain all the needed concepts.10.5 How Might the Mind-World Correspondence Principle BeUseful?Suppose one believes the Mind-World Correspondence Principle as laid out above so what?Our hope, obviously, is that the principle could be useful in actually figuring out how toarchitect intelligent systems biased toward particular sorts of environment. And of course, thisis said with the understanding that any finite intelligence must be biased toward some sorts ofenvironment.Relatedly, given a specific AGI design (such as CogPrime), one could use the principle tofigure out which environments it would be best suited for. Or one could figure out how toadjust the particulars of the design, to maximize the system’s intelligence in the environmentsof interest.One next step in developing this network of ideas, aside from (and potentially building on)full formalization of the principle, would be an exploration of real-world environments in termsof transition graphs. What properties do the transition graphs induced from the real worldhave?One such property, we suggest, is successive refinement. Often the path toward a goal involvesfirst gaining an approximate understanding of a situation, then a slightly more accurateunderstanding, and so forth – until finally one has achieved a detailed enough understanding toactually achieve the goal. This would be represented by a world-path whose nodes are state-setsinvolving the gathering of progressively more detailed information.182 10 A Mind-World Correspondence PrincipleVia pursuing to the mind-world correspondence property in this context, I believe we willfind that world-paths reflecting successive refinement correspond to mind-paths embodying successiverefinement. This will be found to relate to the hierarchical structures found so frequentlyin both the physical world and the human mind-brain. Hierarchical structures allow many relevantgoals to be approached via successive refinement, which I believe is the ultimate reasonwhy hierarchical structures are so common in the human mind-brain.Another next step would be exploring what mind-world correspondence means for the structureand dynamics of a limited-resources intelligence. If an organism O has limited resourcesand, to be intelligent, needs to makeP (o(O, M(A), t(T ))|o(O, M(B), T ))high for particular world state-sets A and B, then what’s the organism’s best approach?Arguably, it should represent M(A) and M(B) internally in such a way that very little computationaleffort is required for it to transition between M(A) and M(B). For instance, this couldbe done by coding its knowledge in such a way that M(A) and M(B) share many common bits;or it could be done in other more complicated ways.If, for instance, A is a subset of B, then it may prove beneficial for the organism to representM(A) physically as a subset of its representation of M(B).Pursuing this line of thinking, one could likely derive specific properties of an intelligentorganism’s internal information-flow, from properties of the environment and goals with respectto which it’s supposed to be intelligent.This would allow us to achieve the holy grail of intelligence theory as I understand it: givena description of an environment and goals, to be able to derive an architectural description foran organism that will display a high level of intelligence relative to those goals, given limitedcomputational resources.While this “holy grail” is obviously a far way off, what we’ve tried to do here is to outline aclear mathematical and conceptual direction for moving toward it.10.6 ConclusionThe Mind-World Correspondence Principle presented here – if in the vicinity of correctness –constitutes a non-trivial step toward fleshing out the concept of a general theory of generalintelligence. But obviously the theory is still rather abstract, and also not completely rigorous.There’s a lot more work to be done.The Mind-World Correspondence Principle as articulated above is not quite a formal mathematicalstatement. It would take a little work to put in all the needed quantifiers to formulateit as one, and it’s not clear the best way to do so the details would perhaps become clear in thecourse of trying to prove a version of it rigorously. One could interpret the ideas presented inthis chapter as a philosophical theory that hopes to be turned into a mathematical theory andto play a key role in a scientific theory.For the time being, the main role to be served by these ideas is qualitative: to help us thinkabout concrete AGI designs like CogPrime in a sensible way. It’s important to understand whatthe goal of a real-world AGI system needs to be: to achieve the ability to broadly learn andgeneralize, yes, but not with infinite capability rather with biases and patterns that are implicitlyand/or explicitly tuned to certain broad classes of goals and environments. The Mind-World10.6 Conclusion 183Correspondence Principle tells us something about what this "tuning" should involve – namely,making a system possessing mind-state sequences that correspond meaningfully to world-statesequences. CogPrime’s overall design and particular cognitive processes are reasonably wellinterpreted as an attempt to achieve this for everyday human goals and environments.One way of extending these theoretical ideas into a more rigorous theory is explored in Appendix??. The key ideas involved there are: modeling multiple memory types as mathematicalcategories (with functors mapping between them), modeling memory items as probability distributions,and measuring distance between memory items using two metrics, one based onalgorithmic information theory and one on classical information geometry. Building on theseideas, core hypotheses are then presented:• a syntax-semantics correlation principle, stating that in a successful AGI system, thesetwo metrics should be roughly correlated• a cognitive geometrodynamics principle, stating that on the whole intelligent mindstend to follow geodesics (shortest paths) in mindspace, according to various appropriatelydefined metrics (e.g. the metric measuring the distance between two entities in terms of thelength and/or runtime of the shortest programs computing one from the other).• a cognitive synergy principle, stating that shorter paths may be found through the compositemindspace formed by considering multiple memory types together, than by followingthe geodesics in the mindspaces corresponding to individual memory types.The material is relegated to an appendix because it is so speculative, and it’s not yet clearwhether it will really be useful in advancing or interpreting CogPrime or other AGI systems(unlike the material from the present chapter, which has at least been useful in interpretingand tweaking the CogPrime design, even though it can’t be claimed that CogPrime was deriveddirectly from these theoretical ideas). However, this sort of speculative exploration is, in ourview, exactly the sort of thing that’s needed as a first phase in transitioning the ideas of thepresent chapter into a more powerful and directly actionable theory.
Section IIICognitive and Ethical Development
Chapter 11Stages of Cognitive DevelopmentCo-authored with Stephan Vladimir Bugaj11.1 IntroductionCreating AGI, we have said, is not only about having the right structural and dynamicalpossibilities implemented in the initial version of one’s system – but also about the environmentand embodiment that one’s system is associated with, and the match between the system’sinternals and these externals. Another key aspect is the long-term time-course of the system’sevolution over time, both in its internals and its external interaction – i.e., what is known asdevelopment.Development is a critical topic in our approach to AGI because we believe that much ofwhat constitutes human-level, human-like intelligence emerges in an intelligent system due toits engagement with its environment and its environment-coupled self-organization. So, it’s notto be expected that the initial version of an AGI system is going to display impressive featsof intelligence, even if the engineering is totally done right. A good analogy is the apparentunintelligence of a human baby. Yes, scientists have discovered that human babies are capableof interesting and significant intelligence – but one has to hunt to find it ... at first observation,babies are rather idiotic and simple-minded creatures: much less intelligent-appearing thanlizards or fish, maybe even less than cockroaches....If the goal of an AGI project is to create an AGI system that can progressively developadvanced intelligence through learning in an environment richly populated with other agentsand various inanimate stimuli and interactive entities – then an understanding of the nature ofcognitive development becomes extremely important to that project.Unfortunately, contemporary cognitive science contains essentially no theory of “abstractdevelopmental psychology” which can conveniently be applied to understand developing AIs.There is of course an extensive science of human developmental psychology, and so it is anatural research program to take the chief ideas from the former and inasmuch as possible portthem to the AGI domain. This is not an entirely simple matter both because of the differencesbetween humans and AIs and because of the unsettled nature of contemporary developmentalpsychology theory. But it’s a job that must (and will) be done, and the ideas in this chaptermay contribute toward this effort.We will begin here with Piaget’s well-known theory of human cognitive development, presentingit in a general systems theory context, then introducing some modifications and extensionsand discussing some other relevant work.187188 11 Stages of Cognitive Development11.2 Piagetan Stages in the Context of a General Systems Theory ofDevelopmentOur review of AGI architectures in Chapter 4 focused heavily on the concept of symbolism,and the different ways in which different classes of cognitive architecture handle symbol representationand manipulation. We also feel that symbolism is critical to the notion of AGIdevelopment – and even more broadly, to the systems theory of development in general.As a broad conceptual perspective on development, we suggest that one may view the developmentof a complex information processing system, embedded in an environment, in termsof the stages:• automatic: the system interacts with the environment by “instinct”, according to its innateprogramming• adaptive: the system internally adapts to the environment, then interacting with the environmentin a more appropriate way• symbolic: the system creates internal symbolic representations of itself and the environment,which in the case of a complex, appropriately structured environment, allows it tointeract with the environment more intelligently• reflexive: the system creates internal symbolic representations of its own internal symbolicrepresentations, thus achieving an even higher degree of intelligenceSketched so broadly, these are not precisely defined categories but rather heuristic, intuitivecategories. Formalizing them would be possible but would lead us too far astray here.One can interpret these stages in a variety of different contexts. Here our focus is the cognitivedevelopment of humans and human-like AGI systems, but in Table 11.1 we present them ina slightly more general context, using two examples: the Piagetan example of the human (orhumanlike) mind as it develops from infancy to maturity; and also the example of the “originof life” and the development of life from proto-life up into its modern form. In any event, weallude to this more general perspective on development here mainly to indicate our view thatthe Piagetan perspective is not something ad hoc and arbitrary, but rather can plausibly be seenas a specific manifestation of more fundamental principles of complex systems development.11.3 Piaget’s Theory of Cognitive DevelopmentThe ghost of Jean Piaget hangs over modern developmental psychology in a yet unresolvedway. Piaget’s theories provide a cogent overarching perspective on human cognitive development,coordinating broad theoretical ideas and diverse experimental results into a unified whole[Pia55]. Modern experimental work has shown Piaget’s ideas to be often oversimplified and incorrect.However, what has replaced the Piagetan understanding is not an alternative unifiedand coherent theory, but a variety of microtheories addressing particular aspects of cognitivedevelopment. For this reason a number of contemporary theorists taking a computer science[Shu03] or dynamical systems [Wit07] approach to developmental psychology have chosen toadopt the Piagetan framework in spite of its demonstrated shortcomings, both because of itsconceptual strengths and for lack of a coherent, more rigorously grounded alternative.Our own position is that the Piagetan view of development has some fundamental truth to it,which is reflected via how nicely it fits with a broader view of development in complex systems.11.3 Piaget’s Theory of Cognitive Development 189Stage General Description Cognitive DevelopmentOrigin of LifeAutomatic System-environment Piagetan infantile Self-organizing protolifeinformation exchange stagesystem, e.g. Oparincontrolled mainly by[Opa52] water droplet,innate system structuresor Cairns-Smith [CS90]or environmentclay-basedprotolifeAdaptiveSymbolicSystem-environmentinfo exchange heavilyguided by adaptivelyinternally-createdsystem structuresInternal symbolic representationof informationexchange processReflexive Thoroughgoing selfmodificationPiagetanbasedon this symbolicrepresentationPiagetan “concrete operational”stage: systematicinternal worldmodelguides worldexplorationPiagetan formal stage:explicit logical/experimentallearning abouthow to cognize in variouscontextsstage: purposive selfmodificationof basicmental processesSimple autopoietic system,e.g. Oparin waterdroplet w/ basicmetabolismGenetic code: internalentities that “standfor” aspects of organismand environment,thus enabling complexepigenesispost-formal Genes+memes: geneticcode-patterns guidetheir own modificationvia influencing cultureTable 11.1: General Systems Theory of Development: Parallels Between Development of Mindand Origin of LifeIndeed, Piaget viewed developmental stages as emerging from general “algebraic” principlesrather than as being artifacts of the particulars of human psychology. But, Piaget’s stages areprobably best viewed as a general interpretive framework rather than a precise scientific theory.Our suspicion is that once the empirical science of developmental psychology has progressedfurther, it will become clearer how to fit the various data into a broad Piaget-like framework,perhaps differing in many details from what Piaget described in his works.Piaget conceived of child development in four stages, each roughly identified with an agegroup, and corresponding closely to the system-theoretic stages mentioned above:• infantile, corresponding to the automatic stage mentioned above– Example: Grasping blocks, piling blocks on top of each other, copying words that areheard• preoperational and concrete operational, corresponding to the adaptive stage mentionedabove– Example: Building complex blocks structures, from imagination and from imitatingobjects and pictures and based on verbal instructions; verbally describing what hasbeen constructed• formal, corresponding to the symbolic stage mentioned above– Example: Writing detailed instructions in words and diagrams, explaining how to constructparticular structures out of blocks; figuring out general rules describing whichsorts of blocks structures are likely to be most stable190 11 Stages of Cognitive Development• the reflexive stage mentioned above corresponds to what some post-Piagetan theorists havecalled the post-formal stage– Example: Using abstract lessons learned from building structures out of blocks to guidethe construction of new ways to think and understand – “Zen and the art of blocksbuilding” (by analogy to Zen and the Art of Motorcycle Maintenance [Pir84]).Fig. 11.1: Piagetan Stages of Cognitive DevelopmentMore explicitly, Piaget defined his stages in psychological terms roughly as follows:• Infantile: In this stage a mind develops basic world-exploration driven by instinctive actions.Reward-driven reinforcement of actions learned by imitation, simple associations betweenwords and objects, actions and images, and the basic notions of time, space, andcausality are developed. The most simple, practical ideas and strategies for action arelearned.• Preoperational: At this stage we see the formation of mental representations, mostlypoorly organized and un-abstracted, building mainly on intuitive rather than logical thinking.Word-object and image-object associations become systematic rather than occasional.Simple syntax is mastered, including an understanding of subject-argument relationships.One of the crucial learning achievements here is “object permanence” – infants learn thatobjects persist even when not observed. However, a number of cognitive failings persist withrespect to reasoning about logical operations, and abstracting the effects of intuitive actionsto an abstract theory of operations.• Concrete: More abstract logical thought is applied to the physical world at this stage.Among the feats achieved here are: reversibility – the ability to undo steps already done;conservation – understanding that properties can persist in spite of appearances; theory ofmind – an understanding of the distinction between what I know and what others know (If11.3 Piaget’s Theory of Cognitive Development 191I cover my eyes, can you still see me?). Complex concrete operations, such as putting itemsin height order, are easily achievable. Classification becomes more sophisticated, yet themind still cannot master purely logical operations based on abstract logical representationsof the observational world.• Formal: Abstract deductive reasoning, the process of forming, then testing hypotheses, andsystematically reevaluating and refining solutions, develops at this stage, as does the abilityto reason about purely abstract concepts without reference to concrete physical objects.This is adult human-level intelligence. Note that the capability for formal operations isintrinsic in the PLN component of CogPrime, but in-principle capability is not the same aspragmatic, grounded, controllable capability.Very early on, Vygotsky [Vyg86] disagreed with Piaget’s explanation of his stages as inherentand developed by the child’s own activities, and Piaget’s prescription of good parenting asnot interfering with a child’s unfettered exploration of the world. Some modern theorists havecritiqued Piaget’s stages as being insufficiently socially grounded, and these criticisms trace backto Vygotsky’s focus on the social foundations of intelligence, on the fact that children functionin a world surrounded by adults who provide a cultural context, offering ongoing assistance,critique, and ultimately validation of the child’s developmental activities.Vygotsky also was an early critic of the idea that cognitive development is continuous,and continues beyond Piaget’s formal stage. Gagne [RBW92] also believes in continuity, andthat learning of prerequisite skills made the learning of subsequent skills easier and fasterwithout regard to Piagetan stage formalisms. Subsequent researchers have argued that Piagethas merely constructed ad hoc descriptions of the sequential development of behaviour[Gib78, Bro84, CP05]. We agree that learning is a continuous process, and our notion of stagesis more statistically constructed than rigidly quantized.Critique of Piaget’s notion of transitional “half stages” is also relevant to a more comprehensivehierarchical view of development. Some have proposed that Piaget’s half stages areactually stages [Bro84]. As Commons and Pekker [CP05] point out: “the definition of a stagethat was being used by Piaget was based on analyzing behaviors and attempting to imposedifferent structures on them. There is no underlying logical or mathematical definition to helpin this process . . . ” Their Hierarchical Complexity development model uses task achievementrather than ad hoc stage definition as the basis for constructing relationships between phasesof developmental ability – an approach which we find useful, though our approach is differentin that we define stages in terms of specific underlying cognitive mechanisms.Another critique of Piaget is that one individual’s performance is often at different abilitystages depending on the specific task (for example [GE86]). Piaget responded to early critiquesalong these lines by calling the phenomenon “horizontal décalage,” but neither he nor his successors[Fis80, Cas85] have modified his theory to explain (rather than merely describe) it.Similarly to Thelen and Smith [TS94], we observe that the abilities encapsulated in the definitionof a certain stage emerge gradually during the previous stage – so that the onset of a givenstage represents the mastery of a cognitive skill that was previously present only in certaincontexts.Piaget also had difficulty accepting the idea of a preheuristic stage, early in the infantileperiod, in which simple trial-and-error learning occurs without significant heuristic guidance[Bic88], a stage which we suspect exists and allows formulation of heuristics by aggregation oflearning from preheuristic pattern mining. Coupled with his belief that a mind’s innate abilitiesat birth are extremely limited, there is a troublingly unexplained transition from inability toability in his model.192 11 Stages of Cognitive DevelopmentFinally, another limiting aspect of Piaget’s model is that it did not recognize any stagesbeyond formal operations, and included no provisions for exploring this possibility. A number ofresearchers [Bic88, Arl75, CRK82, Rie73, Mar01] have described one or more postformal stages.Commons and colleagues have also proposed a task-based model which provides a framework forexplaining stage discrepancies across tasks and for generating new stages based on classificationof observed logical behaviors. [KK90] promotes a statistical conception of stage, which provides agood bridge between task-based and stage-based models of development, as statistical modelingallows for stages to be roughly defined and analyzed based on collections of task behaviors.[CRK82] postulates the existence of a postformal stage by observing elevated levels of abstractionwhich, they argue, are not manifested in formal thought. [CTS + 98] observes a postformalstage when subjects become capable of analyzing and coordinating complex logical systemswith each other, creating metatheoretical supersystems. In our model, with the reflexive stageof development, we expand this definition of metasystemic thinking to include the ability toconsciously refine one’s own mental states and formalisms of thinking. Such self-reflexive refinementis necessary for learning which would allow a mind to analytically devise entirely newstructures and methodologies for both formal and postformal thinking.In spite of these various critiques and limitations, however, we have found Piaget’s ideasvery useful, and in Section 11.4 we will explore ways of defining them rigorously in the specificcontext of CogPrime’s declarative knowledge store and probabilistic logic engine.11.3.1 Perry’s StagesAlso relevant is William Perry’s [Per70, Per81] theory of the stages (“positions” in his terminology)of intellectual and ethical development, which constitutes a model of iterative refinementof approach in the developmental process of coming to intellectual and ethical maturity. Thesestages, depicted in Table 11.2 form an analytical tool for discerning the modality of belief ofan intelligence by describing common cognitive approaches to handling the complexities of realworld ethical considerations.11.3.2 Keeping Continuity in MindContinuity of mental stages, and the fact that a mind may appear to be in multiple stagesof development simultaneously (depending upon the tasks being tested), are crucial to ourtheoretical formulations and we will touch upon them again here. Piaget attempted to addresscontinuity with the creation of transitional “half stages”. We prefer to observe that each stagefeeds into the other and the end of one stage and the beginning of the next blend together.The distinction between formal and post-formal, for example, seems to “merely” be theapplication of formal thought to oneself. However, the distinction between concrete and formal is“merely” the buildup to higher levels of complexity of the classification, task decomposition, andabstraction capabilities of the concrete stage. The stages represent general trends in ability ona continuous curve of development, not discrete states of mind which are jumped-into quantumstyle after enough “knowledge energy” builds-up to cause the transition.11.4 Piaget’s Stages in the Context of Uncertain Inference 193StageSubstagesDualism / Received Basic duality (“All problems are solvable. I must learn theKnowledge[Infantile]correct solutions.”)Full dualism (“There are different, contradictory solutions tomany problems. I must learn the correct solutions, and ignorethe incorrect ones”)Multiplicity[Concrete]Relativism / ProceduralKnowledge[Formal]Commitment / ConstructedKnowledge[Formal / Reflexive]Early multiplicity (“Some solutions are known, others aren’t.I must learn how to find correct solutions.”)Late Multiplicity: cognitive dissonance regarding truth.(“Some problems are unsolvable, some are a matter of personaltaste, therefore I must declare my own intellectual path.”)Contextual Relativism (“I must learn to evaluate solutionswithin a context, and relative to supporting observation.”)Pre-Commitment (“I must evaluate solutions, then commit toa choice of solution.”)Commitment (“I have chosen a solution.”)Challenges to Commitment (“I have seen unexpected implicationsof my commitment, and the responsibility I must take.”)Post-Commitment (“I must have an ongoing, nuanced relationshipto the subject in which I evaluate each situation on acase-by-case basis with respects to its particulars rather thanan ad-hoc application of unchallenged ideology.”)Table 11.2: Perry’s Developmental Stages [with corresponding Piagetan Stages in brackets]Observationally, this appears to be the case in humans. People learn things gradually, andshow a continuous development in ability, not a quick jump from ignorance to mastery. Webelieve that this gradual development of ability is the signature of genuine learning, and thatprescriptively an AGI system must be designed in order to have continuous and asymmetricaldevelopment across a variety of tasks in order to be considered a genuine learning system. Whilequantum leaps in ability may be possible in an AGI system which can just “graft” new partsof brain onto itself (or an augmented human which may someday be able to do the same usingimplants), such acquisition of knowledge is not really learning. Grafting on knowledge does notbuild the cognitive pathways needed in order to actually learn. If this is the only mechanismavailable to an AGI system to acquire new knowledge, then it is not really a learning system.11.4 Piaget’s Stages in the Context of Uncertain InferencePiaget’s developmental stages are very general, referring to overall types of learning, not specificmechanisms or methods. This focus was natural since the context of his work was human developmentalpsychology, and neuroscience has not yet progressed to the point of understandingthe neural mechanisms underlying any sort of inference (and certainly was nowhere near todoing so in Piaget’s time!). But if one is studying developmental psychology in an AGI contextwhere one knows something about the internal mechanisms of the AGI system under consideration,then one can work with a more specific model of learning. Our focus here is on AGIsystems whose operations contain uncertain inference as a central component. Obviously themain focus is CogPrime, but the essential ideas apply to any other uncertain inference centricAGI architecture as well.194 11 Stages of Cognitive DevelopmentFig. 11.2: Piagetan Stages of Development, as Manifested in the Context of Uncertain InferenceAn uncertain inference system, as we consider it here, consists of four components, whichwork together in a feedback-control loop 11.31. a content representation scheme2. an uncertainty representation scheme3. a set of inference rules4. a set of inference control schemataFig. 11.3: A Simplified Look at Feedback-Control in Uncertain Inference11.4 Piaget’s Stages in the Context of Uncertain Inference 195Broadly speaking, examples of content representation schemes are predicate logic and termlogic [ES00]. Examples of uncertainty representation schemes are fuzzy logic [Zad78], impreciseprobability theory [Goo86, FC86], Dempster-Shafer theory [Sha76, Kyb97], Bayesian probabilitytheory [Kyb97], NARS [Wan95], and the Atom representation used in CogPrime, briefly alludedto in Chapter 6 above and described in depth in later chapters.Many, but not all, approaches to uncertain inference involve only a limited, weak set of inferencerules (e.g. not dealing with complex quantified expressions). CogPrime’s PLN inferenceframework, like NARS and some other uncertain inference frameworks, contains uncertain inferencerules that apply to logical constructs of arbitrary complexity. Only a system capable ofdealing with constructs of arbitrary (or at least very high) complexity will have any potentialof leading to human-level, human-like intelligence.The subtlest part of uncertain inference is inference control: the choice of which inferencesto do, in what order. Inference control is the primary area in which human inference currentlyexceeds automated inference. Humans are not very efficient or accurate at carrying out inferencerules, with or without uncertainty, but we are very good at determining which inferences to doand in what order, in any given context. The lack of effective, context-sensitive inference controlheuristics is why the general ability of current automated theorem provers is considerably weakerthan that of a mediocre university mathematics major [Mac95].We now review the Piagetan developmental stages from the perspective of AGI systemsheavily based on uncertain inference.11.4.1 The Infantile StageIn this initial stage, the mind is able to recognize patterns in and conduct inferences aboutthe world, but only using simplistic hard-wired (not experientially learned) inference controlschema, along with pre-heuristic pattern mining of experiential data.In the infantile stage an entity is able to recognize patterns in and conduct inferences aboutits sensory surround context (i.e., it’s “world”), but only using simplistic, hard-wired (not experientiallylearned) inference control schemata. Preheuristic pattern-mining of experiential datais performed in order to build future heuristics about analysis of and interaction with the world.s tasks include:1. Exploratory behavior in which useful and useless / dangerous behavior is differentiated byboth trial and error observation, and by parental guidance.2. Development of “habits” – i.e. Repeating tasks which were successful once to determine ifthey always / usually are so.3. Simple goal-oriented behavior such as “find out what cat hair tastes like” in which one mustplan and take several sequentially dependent steps in order to achieve the goal.Inference control is very simple during the infantile stage (Figure 11.4), as it is the stageduring which both the most basic knowledge of the world is acquired, and the most basic ofcognition and inference control structures are developed as the building block upon which willbe built the next stages of both knowledge and inference control.Another example of a cognitive task at the borderline between infantile and concrete cognitionis learning object permanence, a problem discussed in the context of CogPrime’s predecessor"Novamente Cognition Engine" system in [GPSL03]. Another example is the learning of196 11 Stages of Cognitive DevelopmentFig. 11.4: Uncertain Inference in the Infantile Stageword-object associations: e.g. learning that when the word “ball” is uttered in various contexts(“Get me the ball,” “That’s a nice ball,” etc.) it generally refers to a certain type of object.The key point regarding these “infantile” inference problems, from the CogPrime perspective,is that assuming one provides the inference system with an appropriate set of perceptual andmotor ConceptNodes and SchemaNodes, the chains of inference involved are short. They involveabout a dozen inferences, and this means that the search tree of possible PLN inference ruleswalked by the PLN backward-chainer is relatively shallow. Sophisticated inference control isnot required: standard AI heuristics are sufficient.In short, textbook narrow-AI reasoning methods, utilized with appropriate uncertainty-savvytruth value formulas and coupled with appropriate representations of perceptual and motorinputs and outputs, correspond roughly to Piaget’s infantile stage of cognition. The simplisticapproach of these narrow-AI methods may be viewed as a method of creating building blocksfor subsequent, more sophisticated heuristics.In our theory Piaget’s preoperational phase appears as transitional between the infantile andconcrete operational phases.11.4.2 The Concrete StageAt this stage, the mind is able to carry out more complex chains of reasoning regarding theworld, via using inference control schemata that adapt behavior based on experience (reasoningabout a given case in a manner similar to prior cases).In the concrete operational stage (Figure 11.5), an entity is able to carry out more complexchains of reasoning about the world. Inference control schemata which adapt behavior based onexperience, using experientially learned heuristics (including those learned in the prior stage),are applied to both analysis of and interaction with the sensory surround / world.Concrete Operational stage tasks include:11.4 Piaget’s Stages in the Context of Uncertain Inference 197Fig. 11.5: Uncertain Inference in the Concrete Operational Stage1. Conservation tasks, such as conservation of number,2. Decomposition of complex tasks into easier subtasks, allowing increasingly complex tasksto be approached by association with more easily understood (and previously experienced)smaller tasks,3. Classification and Serialization tasks, in which the mind can cognitively distinguish variousdisambiguation criteria and group or order objects accordingly.In terms of inference control this is the stage in which actual knowledge about how to controlinference itself is first explored. This means an emerging understanding of inference itself as acognitive task and methods for learning, which will be further developed in the following stages.Also, in this stage a special cognitive task capability is gained: “Theory of Mind," which incognitive science refers to the ability to understand the fact that not only oneself, but othersentient beings have memories, perceptions, and experiences. This is the ability to conceptually“put oneself in another’s shoes” (even if you happen to assume incorrectly about them by doingso).11.4.2.1 Conservation of NumberConservation of number is an example of a learning problem classically categorized withinPiaget’s concrete-operational phase, a “conservation laws” problem, discussed in [Shu03] inthe context of software that solves the problem using (logic-based and neural net) narrow-AItechniques. Conservation laws are very important to cognitive development.Conservation is the idea that a quantity remains the same despite changes in appearance. Ifyou show a child some objects and then spread them out, an infantile mind will focus on thespread, and believe that there are now more objects than before, whereas a concrete-operationalmind will understand that the quantity of objects has not changed.Conservation of number seems very simple, but from a developmental perspective it is actuallyrather difficult. “Solutions” like those given in [Shu03] that use neural networks or cus-198 11 Stages of Cognitive Developmenttomized logical rule-bases to find specialized solutions that solve only this problem fail to fullyaddress the issue, because these solutions don’t create knowledge adequate to aid with thesolution of related sorts of problems.We hypothesize that this problem is hard enough that for an inference-based AGI systemto solve it in a developmentally useful way, its inferences must be guided by meta-inferentiallessons learned from prior similar problems. When approaching a number conservation problem,for example, a reasoning system might draw upon past experience with set-size problems (whichmay be trial-and-error experience). This is not a simple “machine learning” approach whosescope is restricted to the current problem, but rather a heuristically guided approach which (a)aggregates information from prior experience to guide solution formulation for the problem athand, and (b) adds the present experience to the set of relevant information about quantificationproblems for future refinement of thinking.Fig. 11.6: Conservation of NumberFor instance, a very simple context-specific heuristic that a system might learn would be:“When evaluating the truth value of a statement related to the number of objects in a set,it is generally not that useful to explore branches of the backwards-chaining search tree thatcontain relationships regarding the sizes, masses, or other physical properties of the objects inthe set.” This heuristic itself may go a long way toward guiding an inference process toward acorrect solution to the problem–but it is not something that a mind needs to know “a priori.”A concrete-operational stage mind may learn this by data-mining prior instances of inferencesinvolving sizes of sets. Without such experience-based heuristics, the search tree for such aproblem will likely be unacceptably large. Even if it is “solvable” without such heuristics, thesolutions found may be overly fit to the particular problem and not usefully generalizable.11.4.2.2 Theory of MindConsider this experiment: a preoperational child is shown her favorite “Dora the Explorer” DVDbox. Asked what show she’s about to see, she’ll answer “Dora.” However, when her parent playsthe disc, it’s “SpongeBob SquarePants.” If you then ask her what show her friend will expectwhen given the “Dora” DVD box, she will respond “SpongeBob” although she just answered“Dora” for herself. A child lacking a theory of mind can not reason through what someoneelse would think given knowledge other than her own current knowledge. Knowledge of self isintrinsically related to the ability to differentiate oneself from others, and this ability may notbe fully developed at birth.Several theorists [BC94, Fod94], based in part on experimental work with autistic children,perceive theory of mind as embodied in an innate module of the mind activated at a certaindevelopmental stage (or not, if damaged). While we consider this possible, we caution againstadopting a simplistic view of the “innate vs. acquired” dichotomy: if there is innateness it maytake the form of an innate predisposition to certain sorts of learning [EBJ + 97].11.4 Piaget’s Stages in the Context of Uncertain Inference 199Davidson [Dav84], Dennett [Den87] and others support the common belief that theory ofmind is dependent upon linguistic ability. A major challenge to this prevailing philosophicalstance came from Premack and Woodruff [PW78] who postulated that prelinguistic primatesdo indeed exhibit “theory of mind” behavior. While Premack and Woodruff’s experiment itselfhas been challenged, their general result has been bolstered by follow-up work showing similarresults such as [TC97]. It seems to us that while theory of mind depends on many of the sameinferential capabilities as language learning, it is not intrinsically dependent on the latter.There is a school of thought often called the Theory Theory [BW88, Car85, Wel90] holdingthat a child’s understanding of mind is best understood in terms of the process of iterativelyformulating and refuting a series of naive theories about others. Alternately, Gordon [Gor86]postulates that theory of mind is related to the ability to run cognitive simulations of others’minds using one’s own mind as a model. We suggest that these two approaches are actuallyquite harmonious with one another. In an uncertain AGI context, both theories and simulationsare grounded in collections of uncertain implications, which may be assembled in contextappropriateways to form theoretical conclusions or to drive simulations. Even if there is aspecial “mind-simulator” dynamic in the human brain that carries out simulations of otherminds in a manner fundamentally different from explicit inferential theorizing, the inputs toand the behavior of this simulator may take inferential form, so that the simulator is in essencea way of efficiently and implicitly producing uncertain inferential conclusions from uncertainpremises.We have thought through the details by CogPrime system should be able to develop theoryof mind via embodied experience, though at time of writing practical learning experiments inthis direction have not yet been done. We have not yet explored in detail the possibility of givingCogPrime a special, elaborately engineered “mind-simulator” component, though this would bepossible; instead we have initially been pursuing a more purely inferential approach.First, it is very simple for a CogPrime system to learn patterns such as “If I rotated by piradians, I would see the yellow block.” And it’s not a big leap for PLN to go from this to therecognition that “You look like me, and you’re rotated by pi radians relative to my orientation,therefore you probably see the yellow block.” The only nontrivial aspect here is the “you looklike me” premise.Recognizing “embodied agent” as a category, however, is a problem fairly similar to recognizing“block” or “insect” or “daisy” as a category. Since the CogPrime agent can perceive mostparts of its own “robot” body–its arms, its legs, etc.–it should be easy for the agent to figureout that physical objects like these look different depending upon its distance from them andits angle of observation. From this it should not be that difficult for the agent to understandthat it is naturally grouped together with other embodied agents (like its teacher), not withblocks or bugs.The only other major ingredient needed to enable theory of mind is “reflection”– the ability ofthe system to explicitly recognize the existence of knowledge in its own mind (note that this term“reflection” is not the same as our proposed “reflexive” stage of cognitive development). Thisexists automatically in CogPrime, via the built-in vocabulary of elementary procedures suppliedfor use within SchemaNodes (specifically, the atTime and TruthValue operators). Observing that“at time T, the weight of evidence of the link L increased from zero” is basically equivalent toobserving that the link L was created at time T.Then, the system may reason, for example, as follows (using a combination of several PLNrules including the above-given deduction rule):200 11 Stages of Cognitive DevelopmentImplicationMy eye is facing a block and it is not darkA relationship is created describing the block’s colorSimilarityMy bodyMy teacher’s body|-ImplicationMy teacher’s eye is facing a block and it is not darkA relationship is created describing the block’s colorThis sort of inference is the essence of Piagetan “theory of mind.” Note that in both ofthese implications the created relationship is represented as a variable rather than a specificrelationship. The cognitive leap is that in the latter case the relationship actually exists in theteacher’s implicitly hypothesized mind, rather than in CogPrime’s mind. No explicit hypothesisor model of the teacher’s mind need be created in order to form this implication–the hypothesisis created implicitly via inferential abstraction. Yet, a collection of implications of this naturemay be used via an uncertain reasoning system like PLN to create theories and simulationssuitable to guide complex inferences about other minds.From the perspective of developmental stages, the key point here is that in a CogPrimecontext this sort of inference is too complex to be viably carried out via simple inferenceheuristics. This particular example must be done via forward chaining, since the big leap is toactually think of forming the implication that concludes inference. But there are simply toomany combinations of relationships involving CogPrime’s eye, body, and so forth for the PLNcomponent to viably explore all of them via standard forward-chaining heuristics. Experienceguidedheuristics are needed, such as the heuristic that if physical objects A and B are generallyphysically and functionally similar, and there is a relationship involving some part of A andsome physical object R, it may be useful to look for similar relationships involving an analogouspart of B and objects similar to R. This kind of heuristic may be learned by experience–and themasterful deployment of such heuristics to guide inference is what we hypothesize to characterizethe concrete stage of development. The “concreteness” comes from the fact that inference controlis guided by analogies to prior similar situations.11.4.3 The Formal StageIn the formal stage, as shown in Figure 11.7, an agent should be able to carry out arbitrarilycomplex inferences (constrained only by computational resources, rather than by fundamentalrestrictions on logical language or form) via including inference control as an explicit subject ofabstract learning. Abstraction and inference about both the sensorimotor surround (world) andabout abstract ideals themselves (including the final stages of indirect learning about inferenceitself) are fully developed.Formal stage evaluation tasks are centered entirely around abstraction and higher-orderinference tasks such as:1. Mathematics and other formalizations.11.4 Piaget’s Stages in the Context of Uncertain Inference 201Fig. 11.7: Uncertain Inference in the Formal Stage2. Scientific experimentation and other rigorous observational testing of abstract formalizations.3. Social and philosophical modeling, and other advanced applications of empathy and theTheory of Mind.In terms of inference control this stage sees not just perception of new knowledge aboutinference control itself, but inference controlled reasoning about that knowledge and the creationof abstract formalizations about inference control which are reasoned-upon, tested, and verifiedor debunked.11.4.3.1 Systematic ExperimentationThe Piagetan formal phase is a particularly subtle one from the perspective of uncertain inference.In a sense, AGI inference engines already have strong capability for formal reasoningbuilt in. Ironically, however, no existing inference engine is capable of deploying its reasoningrules in a powerfully effective way, and this is because of the lack of inference control heuristicsadequate for controlling abstract formal reasoning. These heuristics are what arise duringPiaget’s formal stage, and we propose that in the content of uncertain inference systems, theyinvolve the application of inference itself to the problem of refining inference control.202 11 Stages of Cognitive DevelopmentA problem commonly used to illustrate the difference between the Piagetan concrete operationaland formal stages is that of figuring out the rules for making pendulums swing quicklyversus slowly [IP58]. If you ask a child in the formal stage to solve this problem, she may proceedto do a number of experiments, e.g. build a long string with a light weight, a long stringwith a heavy weight, a short string with a light weight and a short string with a heavy weight.Through these experiments she may determine that a short string leads to a fast swing, a longstring leads to a slow swing, and the weight doesn’t matter at all.The role of experiments like this, which test “extreme cases,” is to make cognition easier. Theformal-stage mind tries to map a concrete situation onto a maximally simple and manipulableset of abstract propositions, and then reason based on these. Doing this, however, requires anautomated and instinctive understanding of the reasoning process itself. The above-describedexperiments are good ones for solving the pendulum problem because they provide data thatis very easy to reason about. From the perspective of uncertain inference systems, this is thekey characteristic of the formal stage: formal cognition approaches problems in a way explicitlycalculated to yield tractable inferences.Note that this is quite different from saying that formal cognition involves abstractions andadvanced logic. In an uncertain logic-based AGI system, even infantile cognition may involvethese – the difference lies in the level of inference control, which in the infantile stage is simplisticand hard-wired, but in the formal stage is based on an understanding of what sorts of inputslead to tractable inference in a given context.11.4.4 The Reflexive StageIn the reflexive stage (Figure 11.8), an intelligent agent is broadly capable of self-modifying itsinternal structures and dynamics.As an example in the human domain: highly intelligent and self-aware adult humans maycarry out reflexive cognition by explicitly reflecting upon their own inference processes andtrying to improve them. An example is the intelligent improvement of uncertain-truth-valuemanipulationformulas. It is well demonstrated that even educated humans typically makenumerous errors in probabilistic reasoning [GGK02]. Most people don’t realize it and continueto systematically make these errors throughout their lives. However, a small percentage ofindividuals make an explicit effort to increase their accuracy in making probabilistic judgmentsby consciously endeavoring to internalize the rules of probabilistic inference into their automatedcognition processes.In the uncertain inference based AGI context, what this means is: In the reflexive stagean entity is able to include inference control itself as an explicit subject of abstract learning(i.e. the ability to reason about one’s own tactical and strategic approach to modifying one’sown learning and thinking), and modify these inference control strategies based on analysis ofexperience with various cognitive approaches.Ultimately, the entity can self-modify its internal cognitive structures. Any knowledge orheuristics can be revised, including metatheoretical and metasystemic thought itself. Initiallythis is done indirectly, but at least in the case of AGI systems it is theoretically possible toalso do so directly. This might be considered as a separate stage of Full Self Modification, orelse as the end phase of the reflexive stage. In the context of logical reasoning, self modificationof inference control itself is the primary task in this stage. In terms of inference control this11.4 Piaget’s Stages in the Context of Uncertain Inference 203Fig. 11.8: The Reflexive Stagestage adds an entire new feedback loop for reasoning about inference control itself, as shown inFigure 11.8.As a very concrete example, in later chapters we will see that, while PLN is founded onprobability theory, it also contains a variety of heuristic assumptions that inevitably introduce acertain amount of error into its inferences. For example, PLN’s probabilistic deduction embodiesa heuristic independence assumption. Thus PLN contains an alternate deduction formula calledthe “concept geometry formula” that is better in some contexts, based on the assumption thatConceptNodes embody concepts that are roughly spherically-shaped in attribute space. A highlyadvanced CogPrime system could potentially augment the independence-based and conceptgeometry-baseddeduction formulas with additional formulas of its own derivation, optimizedto minimize error in various contexts. This is a simple and straightforward example of reflexivecognition – it illustrates the power accessible to a cognitive system that has formalized andreflected upon its own inference processes, and that possesses at least some capability to modifythese.In general, AGI systems can be expected to have much broader and deeper capabilities forself-modification than human beings. Ultimately it may make sense to view the AGI systemswe implement as merely "initial conditions" for ongoing self-modification and self-organization.Chapter ?? discusses some of the potential technical details underlying this sort of thoroughgoingAGI self-modification.
Chapter 12The Engineering and Development of EthicsCo-authored with Stephan Vladimir Bugaj and Joel Pitt12.1 IntroductionMost commonly, if a work on advanced AI mentions ethics at all, it occurs in a final summarychapter, discussing in broad terms some of the possible implications of the technical ideas presentedbeforehand. It’s no coincidence that the order is reversed here: in the case of CogPrime,AGI-ethics considerations played a major role in the design process ... and thus the chapter onethics occurs near the beginning rather than the end. In the CogPrime approach, ethics is nota particularly distinct topic, being richly interwoven with cognition and education and otheraspects of the AGI project.The ethics of advanced AGI is a complex issue with multiple aspects. Among the many issuesthere are:1. Risks posed by the possibility of human beings using AGI systems for evil ends2. Risks posed by AGI systems created without well-defined ethical systems3. Risks posed by AGI systems with initially well-defined and sensible ethical systems eventuallygoing rogue – an especially big risk if these systems are more generally intelligent thanhumans, and possess the capability to modify their own source code4. the ethics of experimenting on AGI systems when one doesn’t understand the nature oftheir experience5. AGI rights: in what circumstances does using an AGI as a tool or servant constitute “slavery”In this chapter we will focus mainly (though not exclusively) on the question of how to createan AGI with a rational and beneficial ethical system. After a somewhat wide-ranging discussion,we will conclude with eight general points that we believe should be followed in working toward"Friendly AGI" – most of which have to do, not with the internal design of the AGI, but withthe way the AGI is taught and interfaced with the real world.While most of the particulars discussed in this book have nothing to do with ethics, it’simportant for the reader to understand that AGI-ethics considerations have played a majorrole in many of our design decisions, underlying much of the technical contents of the book. Asthe materials in this chapter should make clear, ethicalness is probably not something that onecan meaningfully tack onto an AGI system at the end, after developing the rest – it is likelyinfeasible to architect an intelligent agent and then add on an “ethics module.” Rather, ethicsis something that has to do with all the different memory systems and cognitive processes that205206 12 The Engineering and Development of Ethicsconstitute an intelligent system – and it’s something that involves both cognitive architectureand the exploration a system does and the instruction it receives. It’s a very complex matterthat is richly intermixed with all the other aspects of intelligence, and here we will treat it assuch.12.2 Review of Current Thinking on the Risks of AGIBefore proceeding to outline our own perspective on AGI ethics in the context of CogPrime, wewill review the main existing strains of thought on the potential ethical dangers associated withAGI. One science fiction film after another has highlighted these dangers, lodging the issue deepin our cultural awareness; unsurprisingly, much less attention has been paid to serious analysisof the risks in their various dimensions, but there is still a non-trivial literature worth payingattention to.Hypothetically, an AGI with superhuman intelligence and capability could dispense withhumanity altogether – i.e. posing an "existential risk" [Bos02]. In the worst case, an evil butbrilliant AGI, perhaps programmed by a human sadist, could consign humanity to unimaginabletortures (i.e. realizing a modern version of the medieval Christian visions of hell). On theother hand, the potential benefits of powerful AGI also go literally beyond human imagination.It seems quite plausible that an AGI with massively superhuman intelligence and positivedisposition toward humanity could provide us with truly dramatic benefits, such as a virtualend to material scarcity, disease and aging. Advanced AGI could also help individual humansgrow in a variety of directions, including directions leading beyond "legacy humanity," accordingto their own taste and choice.Eliezer Yudkowsky has introduced the term "Friendly AI", to refer to advanced AGI systemsthat act with human benefit in mind [Yud06]. Exactly what this means has not been specifiedprecisely, though informal interpretations abound. Goertzel [Goe06b] has sought to clarify thenotion in terms of three core values of Joy, Growth and Freedom. In this view, a Friendly AIwould be one that advocates individual and collective human joy and growth, while respectingthe autonomy of human choices.Some (for example, Hugo de Garis, [DG05]), have argued that Friendly AI is essentiallyan impossibility, in the sense that the odds of a dramatically superhumanly intelligent mindworrying about human benefit are vanishingly small. If this is the case, then the best optionsfor the human race would presumably be to either avoid advanced AGI development altogether,or to else fuse with AGI before it gets too strongly superhuman, so that beings-originated-ashumanscan enjoy the benefits of greater intelligence and capability (albeit at cost of sacrificingtheir humanity).Others (e.g. Mark Waser [Was09]) have argued that Friendly AI is essentially inevitable,because greater intelligence correlates with greater morality. Evidence from evolutionary andhuman history is adduced in favor of this point, along with more abstract arguments.Yudkowsky [Yud06] has discussed the possibility of creating AGI architectures that are insome sense "provably Friendly" – either mathematically, or else at least via very tight lines of rationalverbal argumentation. However, several issues have been raised with this approach. First,it seems likely that proving mathematical results of this nature would first require dramatic advancesin multiple branches of mathematics. Second, such a proof would require a formalizationof the goal of "Friendliness," which is a subtler matter than it might seem [Leg06b, Leg06a].12.2 Review of Current Thinking on the Risks of AGI 207Formalization of human morality has vexed moral philosophers for quite some time. Finally, it isunclear the extent to which such a proof could be created in a generic, environment-independentway – but if the proof depends on properties of the physical environment, then it would requirea formalization of the environment itself, which runs up against various problems suchas the complexity of the physical world and also the fact that we currently have no complete,consistent theory of physics. Kaj Sotala has provided a list of 14 objections to the FriendlyAI concept, and suggested answers to each of them [Sot11]. Stephen Omohundro [Omo08] hasargued that any advanced AI system will very likely demonstrate certain "basic AI drives", suchas desiring to be rational, to self-protect, to acquire resources, and to preserve and protect itsutility function and avoid counterfeit utility; these drives, he suggests, must be taken carefullyinto account in formulating approaches to Friendly AI.The problem of formally or at least very carefully defining the goal of Friendliness has beenconsidered from a variety of perspectives, none showing dramatic success. Yudkowsky [Yud04]has suggested the concept of "Coherent Extrapolated Volition", which roughly refers to theextrapolation of the common values of the human race. Many subtleties arise in specifyingthis concept – e.g. if Bob Jones is often possessed by a strong desire to kill all Martians, buthe deeply aspires to be a nonviolent person, then the CEV approach would not rate "killingMartians" as part of Bob’s contribution to the CEV of humanity.Goertzel [Goe10a] has proposed a related notion of Coherent Aggregated Volition (CAV),which eschews the subtleties of extrapolation, and simply seeks a reasonably compact, coherent,consistent set of values that is fairly close to the collective value-set of humanity. In the CAVapproach, "killing Martians" would be removed from humanity’s collective value-set becauseit’s uncommon and not part of the most compact/coherent/consistent overall model of humanvalues, rather than because of Bob Jones’ aspiration to nonviolence.One thought we have recently entertained is that the core concept underlying CAV mightbe better thought of as CBV or "Coherent Blended Volition." CAV seems to be easily misinterpretedas meaning the average of different views, which was not the original intention. TheCBV terminology clarifies that the CBV of a diverse group of people should not be thought ofas an average of their perspectives, but as something more analogous to a "conceptual blend"[FT02] – incorporating the most essential elements of their divergent views into a whole that isoverall compact, elegant and harmonious. The subtlety here (to which we shall return below)is that for a CBV blend to be broadly acceptable, the different parties whose views are beingblended must agree to some extent that enough of the essential elements of their own viewshave been included. The process of arriving at this sort of consensus may involve extrapolationof a roughly similar sort to that considered in CEV.Multiple attempts at axiomatization of human values have also been attempted, e.g. with aview toward providing near-term guidance to military robots (see e.g. Arkin’s excellent thoughchillingly-titled book Governing Lethal Behavior in Autonomous Robots [Ark09b], the resultof US military funded research). However, there are reasonably strong arguments that humanvalues (similarly to e.g. human language or human perceptual classification rules) are too complexand multifaceted to be captured in any compact set of formal logic rules. Wallach [WA10]has made this point eloquently, and argued the necessity of fusing top-down (e.g. formal logicbased) and bottom-up (e.g. self-organizing learning based) approaches to machine ethics.A number of more sociological considerations also arise. It is sometimes argued that the riskfrom highly-advanced AGI going morally awry on its own may be less than that of moderatelyadvancedAGI being used by human beings to advocate immoral ends. This possibility gives208 12 The Engineering and Development of Ethicsrise to questions about the ethical value of various practical modalities of AGI development,for instance:• Should AGI be developed in a top-secret installation by a select group of individuals selectedfor a combination of technical and scientific brilliance and moral uprightness, or otherqualities deemed relevant (a "closed approach")? Or should it be developed out in theopen, in the manner of open-source software projects like Linux? (an "open approach").The open approach allows the collective intelligence of the world to more fully participate– but also potentially allows the more unsavory elements of the human race to take someof the publicly-developed AGI concepts and tools private, and develop them into AGIswith selfish or evil purposes in mind. Is there some meaningful intermediary between theseextremes?• Should governments regulate AGI, with Friendliness in mind (as advocated carefully by e.gBill Hibbard [Hib02])? Or will this just cause AGI development to move to the handful ofcountries with more liberal policies? ... or cause it to move underground, where nobody cansee the dangers developing? As a rough analogue, it’s worth noting that the US government’simposition of restrictions on stem cell research, under President George W. Bush, appearsto have directly stimulated the provision of additional funding for stem cell research in othernations like Korea, Singapore and China.The former issue is, obviously, highly relevant to CogPrime (which is currently being developedvia the open source CogPrime project); and so the various dimensions of this issues areworth briefly sketching here.We have a strong skepticism of self-appointed elite groups that claim (even if they genuinelybelieve) that they know what’s best for everyone, and a healthy respect for the power of collectiveintelligence and the Global Brain, which the open approach is ideal for tapping. On the otherhand, we also understand the risk of terrorist groups or other malevolent agents forking an opensource AGI project and creating something terribly dangerous and destructive. Balancing thesefactors against each other rigorously, seems beyond the scope of current human science.Nobody really understands the social dynamics by which open technological knowledge playsout in our current world, let alone hypothetical future scenarios. Right now there exists openknowledge about many very dangerous technologies, and there exist many terrorist groups, yetthese groups fortunately make scant use of these technologies. The reasons why appear to beessentially sociological – the people involved in these terrorist groups tend not to be the oneswho have mastered the skills of turning public knowledge on cutting-edge technologies into realengineered systems. But while it’s easy to observe this sociological phenomenon, we certainlyhave no way to estimate its quantitative extent from first principles. We don’t really have astrong understanding of how safe we are right now, given the technology knowledge availableright now via the Internet, textbooks, and so forth. Even relatively straightforward issues suchas nuclear proliferation remain confusing, even to the experts.It’s also quite clear that keeping powerful AGI locked up by an elite group doesn’t reallyprovide reliable protection against malevolent human agents. History is rife with such situationsgoing awry, e.g. by the leadership of the group being subverted, or via brute force inflicted bysome outside party, or via a member of the elite group defecting to some outside group in theinterest of personal power or reward or due to group-internal disagreements, etc. There aremany things that can go wrong in such situations, and the confidence of any particular groupthat they are immune to such issues, cannot be taken very seriously. Clearly, neither the opennor closed approach qualifies as a panacea.12.3 The Value of an Explicit Goal System 20912.3 The Value of an Explicit Goal SystemOne of the subtle issues confronted in the quest to design ethical AGIs is how closely onewants to emulate human ethical judgment and behavior. Here one confronts the brute factthat, even according to their own deeply-held standards, humans are not all that ethical. Onehigh-level conclusion we came to very early in the process of designing CogPrime is that, just ashumans are not the most intelligent minds achievable, they are also not the most ethical mindsachievable. Even if one takes human ethics, broadly conceived, as the standard – there arealmost surely possible AGI systems that are much more ethical according to human standardsthan nearly all human beings. This is not mainly because of ethics-specific features of thehuman mind, but rather because of the nature of the human motivational system, which leadsto many complexities that drive humans to behaviors that are unethical according to their ownstandards. So, one of the design decisions we made for CogPrime – with ethics as well as otherreasons in mind – was not to closely imitate the human motivational system, but rather to crafta novel motivational system combining certain aspects of the human motivational system withother profoundly non-human aspects.On the other hand, the design of ethical AGI systems still has a lot to gain from the studyof human ethical cognition and behavior. Human ethics has many aspects, which we associatehere with the different types of memory, and it’s important that AGI systems can encompassall of them. Also, as we will note below, human ethics develops in childhood through a seriesof natural stages, parallel to and entwined with the cognitive developmental stages reviewed inChapter 11 above. We will argue that for an AGI with a virtual or robotic body, it makes senseto think of ethical development as proceeding through similar stages. In a CogPrime context,the particulars of these stages can then be understood in terms of the particulars of CogPrime’scognitive processes – which brings AGI ethics from the domain of theoretical abstraction intothe realm of practical algorithm design and education.But even if the human stages of ethical development make sense for non-human AGIs, thisdoesn’t mean the particulars of the human motivational system need to be replicated in theseAGIs, regarding ethics or other matters. A key point here is that, in the context of humanintelligence, the concept of a "goal" is a descriptive abstraction. But in the AGI context, itseems quite valuable to introduce goals as explicit design elements (which is what is done inCogPrime ) – both for ethical reasons and for broader AGI design reasons.Humans may adopt goals for a time and then drop them, may pursue multiple conflictinggoals simultaneously, and may often proceed in an apparently goal-less manner. Sometimes thegoal that a person appears to be pursuing, may be very different than the one they think they’repursuing. Evolutionary psychology [BDL93] argues that, directly or indirectly, all humans areultimately pursuing the goal of maximizing the inclusive fitness of their genes – but given thecomplex mix of evolution and self-organization in natural history [Sal93], this is hardly a generalexplanation for human behavior. Ultimately, in the human context, "goal" is best thought ofas a frequently useful heuristic concept.AGI systems, however, need not emulate human cognition in every aspect, and may bearchitected with explicit "goal systems." This provides no guarantee that said AGI systems willactually pursue the goals that their goal systems specify – depending on the role that the goalsystem plays in the overall system dynamics, sometimes other dynamical phenomena mightintervene and cause the system to behave in ways opposed to its explicit goals. However, wesubmit that this design sketch provides a better framework than would exist in an AGI systemclosely emulating the human brain.210 12 The Engineering and Development of EthicsWe realize this point may be somewhat contentious – a counter-argument would be thatthe human brain is known to support at least moderately ethical behavior, according to humanethical standards, whereas less brain-like AGI systems are much less well understood. However,the obvious counter-counterpoints are that:• Humans are not all that consistently ethical, so that creating AGI systems potentially muchmore practically powerful than humans, but with closely humanlike ethical, motivationaland goal systems, could in fact be quite dangerous• The effect on a human-like ethical/motivational/goal system of increasing the intelligence,or changing the physical embodiment or cognitive capabilities, of the agent containing thesystem, is unknown and difficult to predict given all the complexities involvedThe course we tentatively recommend, and are following in our own work, is to develop AGIsystems with explicit, hierarchically-dominated goal systems. That is:• create one or more "top goals" (we call them Ubergoals in CogPrime )• have the system derive subgoals from these, using its own intelligence, potentially guidedby educational interaction or explicit programming• have a significant percentage of the system’s activity governed by the explicit pursuit ofthese goalsNote that the "significant percentage" need not be 100%; CogPrime, for example, combinesexplicitly goal-directed activity with other "spontaneous" activity. Requiring that all activitybe explicitly goal-directed may be too strict a requirement to place on AGI architectures.The next step, of course, is for the top-level goals to be chosen in accordance with theprinciple of human-Friendliness. The next one of our eight points, about the Global Brain,addresses one way of doing this. In our near-term work with CogPrime, we are using simplisticapproaches, with a view toward early-stage system testing.12.4 Ethical SynergyAn explicit goal system provides an explicit way to ensure that ethical principles (as representedin system goals) play a significant role in guiding an AGI system’s behavior. However, in anintegrative design like CogPrime the goal system is only a small part of the overall story,and it’s important to also understand how ethics relates to the other aspects of the cognitivearchitecture.One of the more novel ideas presented in this chapter is that different types of ethical intuitionmay be associated with different types of memory – and to possess mature ethics, a mindmust display ethical synergy between the ethical processes associated with its memory types.Specifically, we suggest that:• Episodic memory corresponds to the process of ethically assessing a situation based onsimilar prior situations• Sensorimotor memory corresponds to “mirror neuron” type ethics, where you feel anotherperson’s feelings via mirroring their physiological emotional responses and actions• Declarative memory corresponds to rational ethical judgment12.4 Ethical Synergy 211• Procedural memory corresponds to “ethical habit” ... learning by imitation and reinforcementto do what is right, even when the reasons aren’t well articulated or understood• Attentional memory corresponds to the existence of appropriate patterns guiding one topay adequate attention to ethical considerations at appropriate times• Intentional memory corresponds to the pervasion of ethics through one’s choices aboutsubgoaling (which leads into “when do the ends justify the means” ethical-balance questions)One of our suggestions regarding AGI ethics is that an ethically mature person or AGI mustboth master and balance all these kinds of ethics. We will focus especially here on declarativeethics, which corresponds to Kohlberg’s theory of logical ethical judgment; and episodic ethics,which corresponds to Gilligan’s theory of empathic ethical judgment. Ultimately though, all fiveaspects are critically important; and a CogPrime system if appropriately situated and educatedshould be able to master and integrate all of them.12.4.1 Stages of Development of Declarative EthicsComplementing generic theories of cognitive development such as Piaget’s and Perry’s, theoristshave also proposed specific stages of moral and ethical development. The two most relevanttheories in this domain are those of Kohlberg and Gilligan, which we will review here, bothindividually and in terms of their integration and application in the AGI context.Lawrence Kohlberg’s [KLH83, Koh81] moral development model, called the “ethics of justice”by Gilligan, is based on a rational modality as the central vehicle for moral development. In ourperspective this is a firmly declarative form of ethics, based on explicit analysis and reasoning. Itis based on an impartial regard for persons, proposing that ethical consideration must be givento all individual intelligences without a priori judgment (prejudice). Consideration is given forindividual merit and preferences, and the goals of an ethical decision are equal treatment (inthe general, not necessarily the particular) and reciprocity. Echoing Kant’s [Kan64] categoricalimperative, the decisions considered most successful in this model are those which exhibit“reversibility”, where a moral act within a particular situation is evaluated in terms of whetheror not the act would be satisfactory even if particular persons were to switch roles within thesituation. In other words, a situational, contextualized “do unto others as you would have themdo unto you” criterion. The ethics of justice can be viewed as three stages (each of which hassix substages, on which we will not elaborate here), depicted in Table 12.1.In Kohlberg’s perspective, cognitive development level contributes to moral development, asmoral understanding emerges from increased cognitive capability in the area of ethical decisionmaking in a social context. Relatedly, Kohlberg also looks at stages of social perspective andtheir consequent interpersonal outlook. As shown in Table 12.1, these are correlated to thestages of moral development, but also map onto Piagetian models of cognitive development (aspointed out e.g. by Gibbs [Gib78], who presents a modification/interpretation of Kohlberg’sideas intended to align them more closely with Piaget’s). Interpersonal outlook can be understoodas rational understanding of the psychology of other persons (a theory of mind, with orwithout empathy). Stage One, emergent from the infantile congitive stage, is entirely selfishas only self awareness has developed. As cognitive sophistication about ethical considerationsincreases, so do the moral and social perspective stages. Concrete and formal cognition bringabout the first instrumental egoism, and then social relations and systems perspectives, and212 12 The Engineering and Development of EthicsStagePre-ConventionalConventionalPost-ConventionalSubstages• Obedience and Punishment Orientation• Self-interest orientation• Interpersonal accord (conformity) orientation• Authority and social-order maintaining (law and order)orientation• Social contract (human rights) orientation• Universal ethical principles (universal human rights) orientationTable 12.1: Kohlberg’s Stages of Development of the Ethics of Justicefrom formal and then reflexive thinking about ethics comes the post-conventional modalities ofcontractualism and universal mutual respect.Stage of Social PerspectiveInterpersonal OutlookBlind egoism No interpersonal perspective. Only self is considered.Instrumental egoism See that others have goals and perspectives, and either conformto or rebel against norms.Social Relationships Able to see abstract normative systemsperspectiveSocial Systems perspectiveRecognize positive and negative intentionsContractual perspectiveRecognize that contracts (mutually beneficial agreements ofany kind) will allow intelligences to increase the welfare ofboth.Universal principle of See how human fallibility and frailty are impacted by communication.mutual respectTable 12.2: Kohlberg’s Stages of Development of Social Perspective and Interpersonal Morals12.4.1.1 Uncertain Inference and the Ethics of JusticeTaking our cue from the analysis given in Chapter 11 of Piagetan stages in uncertain inferencebased AGI systems (such as CogPrime ), we may explore the manifestation of Kohlberg’sstages in AGI systems of this nature. Uncertain inference seems generally well-suited as adeclarative-ethics learning system, due to the nuanced ethical environment of real world situations.Probabilistic knowledge networks can model belief networks, imitative reinforcementlearning based ethical pedagogy, and even simplistic moral maxims. In principle, they have theflexibility to deal with complex ethical decisions, including not only weighted “for the greater12.4 Ethical Synergy 213good” dichotomous decision making, but also the ability to develop moral decision networkswhich do not require that all situations be solved through resolution of a dichotomy.When more than one person is being affected by an ethical decision, making a decision basedon reducing two choices to a single decision can often lead to decisions of dubious ethics. However,a sufficiently complex uncertain inference network can represent alternate choices in whichmultiple actions are taken that have equal (or near equal) belief weight but have very differentparticulars – but because the decisions are applied in different contexts (to different groups ofindividuals) they are morally equivalent. Though each individual action appears equally believable,were any single decision applied to the entire population one or more individual maybe harmed, and the morally superior choice is to make case-dependent decisions. Equal moraltreatment is a general principle, and too often the mistake is made by thinking that to achievethis general principle the particulars must be equal. This is not the case. Different treatment ofdifferent individuals can result in morally equivalent treatment of all involved individuals, andmay be vastly morally superior to treating all the individuals with equal particulars. Simplytaking the largest population and deciding one course of action based on the result that is mostappealing to that largest group is not generally the most moral action.Uncertain inference, especially a complex network with high levels of resource access as maybe found in a sophisticated AGI, is well suited for complex decision making resulting in amultitude of actions, and of analyzing the options to find the set of actions that are ethicallyoptimal particulars for each decision context. Reflexive cognition and post-commitment moralunderstanding may be the goal stages of an AGI system, or any intelligence, but the otherstages will be passed through on the way to that goal, and realistically some minds will neverreach higher order cognition or morality with regards to any context, and others will not beable to function at this high order in every context (all currently known minds fail to functionat the highest order cognitively or morally in some contexts).Infantile and concrete cognition are the underpinnings of the egoist and socialized stages,with formal aspects also playing a role in a more complete understanding of social modelswhen thinking using the social modalities. Cognitively infantile patterns can produce no morethan blind egoism as without a theory of mind, there is no capability to consider the other.Since most intelligences acquire concrete modality and therefore some nascent social perspectiverelatively quickly, most egoists are instrumental egoists. The social relationship and systemsperspectives include formal aspects which are achieved by systematic social experimentation,and therefore experiential reinforcement learning of correct and incorrect social modalities.Initially this is a one-on-one approach (relationship stage), but as more knowledge of socialaction and consequences is acquired, a formal thinker can understand not just consequentialitybut also intentionality in social action.Extrapolation from models of individual interaction to general social theoretic notions is alsoa formal action. Rational, logical positivist approaches to social and political ideas, however, arethe norm of formal thinking. Contractual and committed moral ethics emerges from a higherorderformalization of the social relationships and systems patterns of thinking. Generalizationsof social observation become, through formal analysis, systems of social and political doctrine.Highly committed, but grounded and logically supportable, belief is the hallmark of formalcognition as expressed contractual moral stage. Though formalism is at work in the socializedmoral stages, its fullest expression is in committed contractualism.Finally, reflexive cognition is especially important in truly reaching the post-commitmentmoral stage in which nuance and complexity are accommodated. Because reflexive cognitionis necessary to change one’s mind not just about particular rational ideas, but whole ways of214 12 The Engineering and Development of Ethicsthinking, this is a cognitive precedent to being able to reconsider an entire belief system, onethat has had contractual logic built atop reflexive adherence that began in early development.If the initial moral system is viewed as positive and stable, then this cognitive capacity isseen as dangerous and scary, but if early morality is stunted or warped, then this ability isseen as enlightened. However, achieving this cognitive stage does not mean one automaticallychanges their belief systems, but rather that the mental machinery is in place to considerthe possibilities. Because many people do not reach this level of cognitive development in thearea of moral and ethical thinking, it is associated with negative traits (“moral relativism”and “flip-flopping”). However, this cognitive flexibility generally leads to more sophisticated andapplicable moral codes, which in turn leads to morality which is actually more stable becauseit is built upon extensive and deep consideration rather than simple adherence to reflexive orrationalized ideologies.12.4.2 Stages of Development of Empathic EthicsComplementing Kohlberg’s logic-and-justice-focused approach, Carol Gilligan’s [Gil82] “ethicsof care” model is a moral development theory which posits that empathetic understandingplays the central role in moral progression from an initial self-centered modality to a sociallyresponsible one. The ethics of care model is concerned with the ways in which an individualcares (responds to dilemmas using empathetic responses) about self and others. As shown inTable 12.3, the ethics of care is broken into the same three primary stage as Kohlberg, but witha focus on empathetic, emotional caring rather than rationalized, logical principles of justice.StagePre-ConventionalConventionalPost-ConventionalPrinciple of CareIndividual SurvivalSelf Sacrifice for the Greater GoodPrinciple of Nonviolence (do not hurt others, or oneself)Table 12.3: Gilligan’s Stages of the Ethics of CareFor an “ethics of care” approach to be applied in an AGI, the AGI must be capable of internalsimulation of other minds it encounters, in a similar manner to how humans regularly simulateone another internally. Without any mechanism for internal simulation, it is unlikely that anAGI can develop any sort of empathy toward other minds, as opposed to merely logicallyor probabilistically modeling other agents’ behavior or other minds’ internal contents. In aCogPrime context, this ties in closely with how CogPrime handles episodic knowledge – partlyvia use of an internal simulation world, which is able to play “mental movies” of prior andhypothesized scenarios within the AGI system’s mind.However, in humans empathy involves more than just simulation, it also involves sensorimotorresponses, and of course emotional responses – a topic we will discuss in more depth in Appendix?? where we review the functionality of mirror neurons and mirror systems in the human brains.When we see or hear someone suffering, this sensory input causes motor responses in us similarto if we were suffering ourselves, which initiates emotional empathy and corresponding cognitiveprocesses.12.4 Ethical Synergy 215Thus, empathic “ethics of care” involves a combination of episodic and sensorimotor ethics,complementing the mainly declarative ethics associated with the “ethics of justice.”In Gilligan’s perspective, the earliest stage of ethical development occurs before empathybecomes a consistent and powerful force. Next, the hallmark of the conventional stage is thatat this point, the individual is so overwhelmed with their empathic response to others thatthey neglect themselves in order to avoid hurting others. Note that this stage doesn’t occurin Kohlberg’s hierarchy at all. Kohlberg and Gilligan both begin with selfish unethicality, buttheir following stages diverge. A person could in principle manifest Gilligan’s conventional stagewithout having a refined sense of justice (thus not entering Kohlberg’s conventional stage); orthey could manifest Kohlberg’s conventional stage without partaking in an excessive degree ofself-sacrifice (thus not entering Gilligan’s conventional stage). We will suggest below that in factthe empathic and logical aspects of ethics are more unified in real human development thanthese separate theories would suggest. However, even if this is so, the possibility is still therethat in some AGI systems the levels of declarative and empathic ethics could wildly diverge.It is interesting to note that Gilligan’s and Kohlberg’s final stages converge more closelythan their intermediate ones. Kohlberg’s post-conventional stage focuses on universal rights,and Gilligan’s on universal compassion. Still, the foci here are quite different; and, as will beelaborated below, we believe that both Kohlberg’s and Gilligan’s theories constitute very partialviews of the actual end-state of ethical advancement.12.4.3 An Integrative Approach to Ethical DevelopmentWe feel that both Kohlberg’s and Gilligan’s theories contain elements of the whole picture ofethical development, and that both approaches are necessary to create a moral, ethical artificialgeneral intelligence – just as, we suggest, both internal simulation and uncertain inference arenecessary to create a sufficiently intelligent and volitional intelligence in the first place. Also,we contend, the lack of direct analysis of the underlying psychology of the stages is a deficiencyshared by both the Kohlberg and Gilligan models as they are generally discussed. A successfulmodel of integrative ethics necessarily contains elements of both the care and justice models, aswell as reference to the underlying developmental psychology and its influence on the characterof the ethical stage. Furthermore, intentional and attentional ethics need to be brought intothe picture, complementing Kohlberg’s focus on declarative knowledge and Gilligan’s focus onepisodic and sensorimotor knowledge.With these notions in mind, we propose the following integrative theory of the stages ofethical development, shown in Tables 12.4, 12.5 and 12.6. In our integrative model, the justicebasedand empathic aspects of ethical judgment are proposed to develop together. Of course, inany one individual, one or another aspect may be dominant. Even so, however, the combinationof the two is equally important as either of the two individual ingredients.For instance, we suggest that in any psychologically healthy human, the conventional stageof ethics (typifying childhood, and in many cases adulthood as well) involves a combinationof Gilligan-esqe empathic ethics and Kohlberg-esque ethical reasoning. This combination issupported by Piagetan concrete operational cognition, which allows moderately sophisticatedlinguistic interaction, theory of mind, and symbolic modeling of the world.And, similarly, we propose that in any truly ethically mature human, empathy and rationaljustice are both fully developed. Indeed the two interpenetrate each other deeply.216 12 The Engineering and Development of EthicsOnce one goes beyond simplistic, childlike notions of fairness (“an eye for an eye” and soforth), applying rational justice in a purely intellectual sense is just as difficult as any otherreal-world logical inference problem. Ethical quandaries and quagmires are easily encountered,and are frequently cut through by a judicious application of empathic simulation.On the other hand, empathy is a far more powerful force when used in conjunction withreason: analogical reasoning lets us empathize with situations we have never experienced. Forinstance, a person who has never been clinically depressed may have a hard time empathizingwith individuals who are; but using the power of reason, they can imagine their worst state ofdepression magnified by several times and then extended over a long period of time, and thenreason about what this might be like ... and empathize based on their inferential conclusion.Reason is not antithetical to empathy but rather is the key to making empathy more broadlyimpactful.Finally, the enlightened stage of ethical development involves both a deeper compassion anda more deeply penetrating rationality and objectiveness. Empathy with all sentient beings ismanageable in everyday life only once one has deeply reflected on one’s own self and largelyfreed oneself of the confusions and illusions that characterize much of the ordinary human’sinner existence. It is noteworthy, for example, that Buddhism contains both a richly developedethics of universal compassion, and also an intricate logical theory of the inner workings ofcognition [Stc00], detailing in exquisite rational detail the manner in which minds originatestructures and dynamics allowing them to comprehend themselves and the world.12.4.4 Integrative Ethics and Integrative AGIWhat does our integrative approach to ethical development have to say about the ethicaldevelopment of AGI systems? The lessons are relatively straighforward, if one considers an AGIsystem that, like CogPrime, explicitly contains components dedicated to logical inference andto simulation. Application of the above ethical ideas to other sorts of AGI systems is also quitepossible, but would require a lengthier treatment and so won’t be addressed here.In the context of a CogPrime-type AGI system, Kolhberg’s stages correspond to increasinglysophisticated application of logical inference to matters of rights and fairness. It is not clearwhether humans contain an innate sense of fairness. In the context of AGIs, it would be possibleto explicitly wire a sense of fairness into an AGI system, but in the context of a rich environmentand active human teachers, this actually appears quite unnecessary. Experiential instruction inthe notions of rights and fairness should suffice to teach an inference-based AGI system how tomanipulate these concepts, analogously to teaching the same AGI system how to manipulatenumber, mass and other such quantities. Ascending the Kohlberg stages is then mainly a matterof acquiring the ability to carry out suitably complex inferences in the domain of rights andfairness. The hard part here is inference control – choosing which inference steps to take – andin a sophisticated AGI inference engine, inference control will be guided by experience, so thatthe more ethical judgments the system has executed and witnessed, the better it will become atmaking new ones. And, as argued above, simulative activity can be extremely valuable for aidingwith inference control. When a logical inference process reaches a point of acute uncertainty(the backward or forward chaining inference tree can’t decide which expansion step to take), itcan run a simulation to cut through the confusion – i.e., it can use empathy to decide which12.4 Ethical Synergy 217StagePre-ethicalConventional EthicsCharacteristics• Piagetan infantile to early concrete (aka pre-operational)• Radical selfishness or selflessness may, but do not necessarily,occur• No coherent, consistent pattern of consideration for therights, intentions or feelings of others• Empathy is generally present, but erratically• Concrete cognitive basis• Perry’s Dualist and Multiple stages• The common sense of the Golden Rule is appreciated,with cultural conventions for abstracting principles frombehaviors• One’s own ethical behavior is explicitly compared to thatof others• Development of a functional, though limited, theory ofmind• Ability to intuitively conceive of notions of fairness andrights• Appreciation of the concept of law and order, which maysometimes manifest itself as systematic obedience or systematicdisobedience• Empathy is more consistently present, especially withothers who are directly similar to oneself or in situationssimilar to those one has directly experienced• Degrees of selflessness or selfishness develop based on ethicalgroundings and social interactions.Table 12.4: Integrative Model of the Stages of Ethical Development, Part 1logical inference step to take in thinking about applying the notions of rights and fairness to agiven situation.Gilligan’s stages correspond to increasingly sophisticated control of empathic simulation –which in a CogPrime-type AGI system, is carried out by a specific system component devotedto running internal simulations of aspects of the outside world, which includes a subcomponentspecifically tuned for simulating sentient actors. The conventional stage has to do with the raw,uncontrolled capability for such simulation; and the post-conventional stage corresponds to itscontextual, goal-oriented control. But controlling empathy, clearly, requires subtle managementof various uncertain contextual factors, which is exactly what uncertain logical inference isgood at – so, in an AGI system combining an uncertain inference component with a simulativecomponent, it is the inference component that would enable the nuanced control of empathyallowing the ascent to Gilligan’s post-conventional stage.In our integrative perspective, in the context of an AGI system integrating inference andsimulation components, we suggest that the ascent from the pre-ethical to the conventionalstage may be carried out largely via independent activity of these two components. Empathyis needed, and reasoning about fairness and rights are needed, but the two need not intimatelyand sensitively intersect – though they must of course intersect to some extent.218 12 The Engineering and Development of EthicsStageMature EthicsCharacteristics• Formal cognitive basis• Perry’s Relativist and “Constructed Knowledge” stages• The abstraction involved with applying the Golden Rulein practice is more fully understood and manipulated,leading to limited but nonzero deployment of the CategoricalImperative• Attention is paid to shaping one’s ethical principles intoa coherent logical system• Rationalized, moderated selfishness or selflessness.• Empathy is extended, using reason, to individuals andsituations not directly matching one’s own experience• Theory of mind is extended, using reason, to counterintuitiveor experientially unfamiliar situations• Reason is used to control the impact of empathy on behavior(i.e. rational judgments are made regarding whento listen to empathy and when not to)• Rational experimentation and correction of theoreticalmodels of ethical behavior, and reconciliation with observedbehavior during interaction with others.• Conflict between pragmatism of social contract orientationand idealism of universal ethical principles.• Understanding of ethical quandaries and nuances develop(pragmatist modality), or are rejected (idealist modality).• Pragmatically critical social citizen. Attempts to maintaina balanced social outlook. Considers the commongood, including oneself as part of the commons, and actsin what seems to be the most beneficial and practicalmanner.Table 12.5: Integrative Model of the Stages of Ethical Development, Part 2The main engine of advancement from the conventional to mature stage, we suggest, is robustand subtle integration of the simulative and inferential components. To expand empathy beyondthe most obvious cases, analogical inference is needed; and to carry out complex inferences aboutjustice, empathy-guided inference-control is needed.Finally, to advance from the mature to the enlightened stage, what is required is a veryadvanced capability for unified reflexive inference and simulation. The system must be ableto understand itself deeply, via modeling itself both simulatively and inferentially – whichwill generally be achieved via a combination of being good at modeling, and becoming lessconvoluted and more coherent, hence making self-modeling easier.Of course, none of this tells you in detail how to create an AGI system with advancedethical capabilities. What it does tell you, however, is one possible path that may be followed toachieve this end goal. If one creates an integrative AGI system with appropriately interconnectedinferential and simulative components, and treats it compassionately and fairly, and providesit extensive, experientially grounded ethical instruction in a rich social environment, then theAGI system should be able to ascend the ethical hierarchy and achieve a high level of ethicalsophistication. In fact it should be able to do so more reliably than human beings because ofthe capability we have to identify its errors via inspecting its internal knowedge-stage, which12.5 Clarifying the Ethics of Justice: Extending the Golden Rule in to a Multifactorial Ethical Model 219StageEnlightened EthicsCharacteristics• Reflexive cognitive basis• Permeation of the categorical imperative and the questfor coherence through inner as well as outer life• Experientially grounded and logically supported rejectionof the illusion of moral certainty in favor of a case-specificanalytical and empathetic approach that embraces theuncertainty of real social life• Deep understanding of the illusory and biased nature ofthe individual self, leading to humility regarding one’sown ethical intuitions and prescriptions• Openness to modifying one’s deepest, ethical (and other)beliefs based on experience, reason and/or empathic communionwith others• Adaptive, insightful approach to civil disobedience, consideringlaws and social customs in a broader ethical andpragmatic context• Broad compassion for and empathy with all sentient beings• A recognition of inability to operate at this level at alltimes in all things, and a vigilance about self-monitoringfor regressive behavior.Table 12.6: Integrative Model of the Stages of Ethical Development, Part 3will enable us to tailor its environment and instructions more suitably than can be done in thehuman case.If an absolute guarantee of the ethical soundness of an AGI is what one is after, the line ofthinking proposed here is not at all useful. Experiential education is by its nature an uncertainthing. One can strive to minimize the uncertainty, but it will still exist. Inspection of theinternals of an AGI’s mind is not a total solution to uncertainty minimization, because anyAGI capable of powerful general intelligence is going to have a complex internal state thatno external observer will be able to fully grasp, no matter how transparent the knowledgerepresentation.However, if what one is after is a plausible, pragmatic path to architecting and educatingethical AGI systems, we believe the ideas presented here constitute a sensible starting-point.Certainly there is a great deal more to be learned and understood – the science and practiceof AGI ethics, like AGI itself, are at a formative stage at present. What is key, in our view, isthat as AGI technology develops, AGI ethics develops alongside and within it, in a thoroughlycoupled way.12.5 Clarifying the Ethics of Justice: Extending the Golden Rule into a Multifactorial Ethical ModelOne of the issues with the "ethics of justice" as reviewed above, which makes it inadequateto serve as the sole basis of an AGI ethical system (though it may certainly play a significant220 12 The Engineering and Development of Ethicsrole), is the lack of any clear formulation of what "justice" means. This section explores thisissue, via detailed consideration of the “Golden Rule” folk maxim do unto others as youwould have them do unto you – a classical formulation of the notion of fairness and justics– to AGI ethics. Taking the Golden Rule as a starting-point, we will elaborate five ethicalimperatives that incorporate aspects of the notion of ethical synergy discussed above. Simple asit may seem, the Golden Rule actually elicits a variety of deep issues regarding the relationshipbetween ethics, experience and learning. When seriously analyzed, it results in a multifactorialelaboration, involving the combination of various factors related to the basic Golden Rule idea.Which brings us back in the end to the potential value of methods like CEV, CAV or CBV forunderstanding how human ethics balances the multiple factors. Our goal here is not to presentany kind of definitive analysis of the ethics of justice, but just to briefly and roughly indicatea number of the relevant significant issues – things that anyone designing or teaching an AGIwould do well to keep in mind.The trickiest aspect of the Golden Rule, as has been frequently observed, is achieving theright level of abstraction. Taken too literally, the Golden Rule would suggest, for instance, thata parent should not wipe a child’s soiled bottom because the parent does not want the child towipe the parent’s soiled bottom. But if the parent interprets the Golden Rule more intelligentlyand abstractly, the parent may conclude that they should wipe the child’s bottom after all:they should “wipe the child’s bottom when the child can’t do it themselves”, consistently withbelieving that the child should “wipe the parent’s bottom when the parent can’t do it themselves”(which may well happen eventually should the parent develop incontinence in old age).This line of thinking leads to Kant’s Categorical Imperative [Kan64] which (in one interpretation)states essentially that one should “Act only according to that maxim whereby youcan at the same time will that it should become a universal law." The Categorical Imperativeadds precision to the Golden Rule, but also removes the practicality of the latter. Formalizingthe “implicit universal law” underlying an everyday action is a huge problem, falling preyto the same issue that has kept us from adequately formalizing the rules of natural languagegrammar, or formalizing common-sense knowledge about everyday object like cups, bowls andgrass (substantial effort notwithstanding, e.g. Cyc in the commonsense knowledge case, and thewhole discipline of modern linguistics in the NL case). There is no way to apply the CategoricalImperative, as literally stated, in everyday life.Furthermore, if one wishes to teach ethics as well as to practice it, the Categorical Imperativeactually has a significant disadvantage compared to some other possible formulations ofthe Golden Rule. The problem is that, if one follows the Categorical Imperative, one’s fellowmembers of society may well never understand the principles under which one is acting. Eachof us may internally formulate abstract principles in a different way, and these may be verydifficult to communicate, especially among individuals with different belief systems, differentcognitive architectures, or different levels of intelligence. Thus, if one’s goal is not just to actethically, but to encourage others to act ethically by setting a good example, the CategoricalImperative may not be useful at all, as others may be unable to solve the “inverse problem” ofguessing your intended maxim from your observed behavior.On the other hand, one wouldn’t want to universally restrict one’s behavioral maxims tothose that one’s fellow members of society can understand – in that case, one would have to actwith a two-year old or a dog according to principles that they could understand, which wouldclearly be unethical according to human common sense. (Every two-year-old, once they growup, would be grateful to their parents for not following this sort of principle.)12.5 Clarifying the Ethics of Justice: Extending the Golden Rule in to a Multifactorial Ethical Model 221And the concept of “setting a good example” ties in with an important concept from learningtheory: imitative learning. Humans appear to be hard-wired for imitative learning, in part viamirror neuron systems in the brain; and, it seems clear that at least in the early stages of AGIdevelopment, imitative learning is going to play a key role. Copying what other agents do is anextremely powerful heuristic, and while AGIs may eventually grow beyond this, much of theirearly ethical education is likely to arise during a phase when they have not done so. A strengthof the classic Golden Rule is that one is acting according to behaviors that one wants one’sobservers to imitate – which makes sense in that many of these observers will be using imitativelearning as a significant part of their learning toolkit.The truth of the matter, it seems, is (as often happens) not all that simple or elegant. Ethicalbehavior seems to be most pragmatically viewed as a multi-objective optimization problem,where among the multiple objectives are three that we have just discussed, and two others thatemerge from learning theory and will be discussed shortly:1. The imitability (i.e. the Golden Rule fairly narrowly and directly construed): the goal ofacting in a way so that having others directly imitate one’s actions, in directly comparablecontexts, is desirable to oneself2. The comprehensibility: the goal of acting in a way so that others can understand theprinciples underlying one’s actions3. Experiential groundedness. An intelligent agent should not be expected to act accordingto an ethical principle unless there are many examples of the principle-in-action in its owndirect or observational experience4. The categorical imperative: Act according to abstract principles that you would behappy to see implemented as universal laws5. Logical coherence. An ethical system should be roughly logically coherent, in the sensethat the different principles within it should mesh well with one another and perhaps evennaturally emerge from each other.Just for convenience, without implying any finality or great profundity to the list, we will referto these as the "five imperatives."The above are all ethical objectives to be valued and balanced, to different extents in differentcontexts. The imitability imperative, obviously, loses importance in societies of agents that don’tmake heavy use of imitative learning. The comprehensibility imperative is more importantin agents that value social community-building generally, and less so in agent that are moreisolative and self-focused.Note that the fifth point given above is logically of a different nature than the four previousones. The first four imperatives govern individual ethical principles; the fifth regards systems ofethical principles, as they interact with each other. Logical coherence is of significant but varyingimportance in human ethical systems. Huge effort has been spent by theologians of variousstripes in establishing and refining the logical coherence of the ethical systems associated withtheir religions. However, it is arguably going to be even more important in the context of AGIsystems, especially if these AGI systems utilize cognitive methods based on logical inference,probability theory or related methods.Experiential groundedness is important because making pragmatic ethical judgments isbound to require reference to an internal library of examples (“episodic ethics”) in which ethicalprinciples have previously been applied. This is required for analogical reasoning, and inlogic-based AGI systems, is also required for pruning of the logical inference trees involved indetermining ethical judgments.222 12 The Engineering and Development of EthicsTo the extent that the Golden Rule is valued as an ethical imperative, experiential groundingmay be supplied via observing the behaviors of others. This in itself is a powerful argument infavor of the Golden Rule: without it, the experiential library a system possesses is restricted toits own experience, which is bound to be a very small library compared to what it can assemblefrom observing the behaviors of others.The overall upshot is that, ideally, an ethical intelligence should act according to a logicallycoherent system of principles, which are exemplified in its own direct andobservational experience, which are comprehensible to others and set a good examplefor others, and which would serve as adequate universal laws if somehowthus implemented. But, since this set of criteria is essentially impossible to fulfill in practice,real-world intelligent agents must balance these various criteria – often in complex andcontextually-dependent ways.We suggest that ethically advanced humans, in their pragmatic ethical choices, tend to act insuch a way as to appropriately contextually balance the above factors (along with other criteria,but we have tried to articulate the most key factors). This sort of multi-factorial approach isnot as crisp or elegant as unidimensional imperatives like the Golden Rule or the CategoricalImperative, but is more realistic in light of the complexly interacting multiple determinantsguiding individual and group human behavior.And this brings us back to CEV, CAV, CBV and other possible ways of mining ethicalsupergoals from the community of existing human minds. Given that abstract theories of ethics,when seriously pursued as we have done in this section, tend to devolve into complex balancingacts involving multiple factors – one then falls back into asking how human ethical systemshabitually perform these balancing acts. Which is what CEV, CAV, CBV try to measure.12.5.1 The Golden Rule and the Stages of Ethical DevelopmentNext we explore more explicitly how these Golden Rule based imperatives align with the ethicaldevelopmental stages we have outlined here. With this in mind, specific ethical qualitiescorresponding to the five imperatives have been italicized in the above table of developmentalstages.It seems that imperatives 1-3 are critical for the passage from the pre-ethical to the conventionalstages of ethics. A child learns ethics largely by copying others, and by being interactedwith according to simply comprehensible implementations of the Golden Rule. In general, wheninteracting with children learning ethics, it is important to act according to principles they cancomprehend. And given the nature of the concrete stage of cognitive development, experientialgroundedness is a must.As a hypothesis regarding the dynamics underlying the psychological development of conventionalethics, what we propose is as follows: The emergence of concrete-stage cognitivecapabilities leads to the capability for fulfillment of ethical imperatives 1 and 2 – a comprehensibleand workable implementation of the Golden Rule, based on a combination of inferentialand simulative cognition (operating largely separately at this stage, as will be conjectured below).The effective interoperation of ethical imperatives 1-3, enacted in an appropriate socialenvironment, then leads to the other characteristics of the conventional ethical stage. The firstthree imperatives can thus be viewed as the seed from which springs the general nature ofconventional ethics.12.5 Clarifying the Ethics of Justice: Extending the Golden Rule in to a Multifactorial Ethical Model 223On the other hand, logical coherence and the categorical imperative (imperatives 5 and 4)are matters for the formal stage of cognitive development, which come along only with themature approach to ethics. These come from abstracting ethics beyond direct experience andmanipulating them abstractly and formally – a stage which has the potential for more deeplyand broadly ethical behavior, but also for more complicated ethical perversions (it is the maturecapability for formal ethical reasoning that is able to produce ungrounded abstractions suchas “I’m torturing you for your own good”). Developmentally, we suggest that once the capabilityfor formal reasoning matures, the categorical imperative and the quest for logical ethicalcoherence naturally emerge, and the sophisticated combination of inferential and simulativecognition embodied in an appropriate social context then result in the emergence of the variouscharacteristics typifying the mature ethical stage.Finally, it seems that one key aspect of the passage from the mature to the enlightened stageof ethics is the penetration of these two final imperatives more and more deeply into the judgingmind itself. The reflexive stage of cognitive development is in part about seeking a deep logicalcoherence between the aspects of one’s own mind, and making reasoned modifications to one’smind so as to improve the level of coherence. And, much of the process of mental discipline andpurification that comes with the passage to enlightened ethics has to do with the applicationof the categorical imperative to one’s own thoughts and feelings – i.e. making a true innersystematic effort to think and feel only those things one judges are actually generally goodand right to be thinking and feeling. Applying these principles internally appears critical foreffectively applying them externally, for reasons that are doubtlessly bound up with the interpenetrationof internal and external reality within the thinking mind, and for the “distributedcognition” phenomenon wherein individual mind is itself an approximative abstraction to thereality in which each individual’s mind is pragmatically extended across their social group andtheir environment [Hut95].Obviously, these are complex issues and we’re not posing the exploratory discussion givenhere as conclusive in any sense. But what seems generally clear from this line of thinking isthat the complex balance between the multiple factors involved in AGI ethics, shifts during asystem’s development. If you did CEV, CAV or CBV among five year old humans, ten yearold humans, or adult humans, you would get different results. Probably you’d also get differentresults from senior citizens! The way the factors are balanced depends on the mind’s cognitiveand emotional stage of development.12.5.2 The Need for Context-Sensitivity and Adaptiveness inDeploying Ethical PrinciplesAs well as depending on developmental stage, there is also an obvious and dramatic contextsensitivityinvolved here – both in calculating the fulfillment of abstract ethical imperatives,and in balancing various imperatives against each other. As an example, consider the simpleAsimovian maxim “I will not harm humans,” which may be seen to follow from the Golden Rulefor any agent that doesn’t itself want to be harmed, and that considers humans as valid agentson the same ethical level as itself. A more serious attempt to formulate this as an ethical maximmight look something like224 12 The Engineering and Development of Ethics“I will not harm humans, nor through inaction allow harm to befall them. In situationswherein one or more humans is attempting to harm another individual or group, I shall endeavorto prevent this harm through means which avoid further harm. If this is unavoidable, I shallselect the human party to back based on a reckoning of their intentions towards others, andimplement their defense through the optimal balance between harm minimization and efficacy.My ultimate goal is to preserve as much as possible of humanity, even if an individual orsubgroup of humans must come to harm to do so.”However, it’s obvious that even a more elaborated principle like this is potentially subject toextensive abuse. Many of the genocides scarring human history have been committed with thegoal of preserving and bettering humanity writ large, at the expense of a group of “undesirables.”Further refinement would be necessary in order to define when the greater good of humanitymay actually be served through harm to others. A first actor principle of aggression mightseem to solve this problem, but sometimes first actors in violent conflict are taking preemptivemeasures against the stated goals of an enemy to destroy them. Such situations become verysubtle. A single simple maxim can not deal with them very effectively. Networks of interrelateddecision criteria, weighted by desirability of consequence and with reference to probabilisticallyordered potential side-effects (and their desirability weightings), are required in order to makeethical judgments. The development of these networks, just like any other knowledge network,comes from both pedagogy and experience – and different thoughtful, ethical agents are boundto arrive at different knowledge-networks that will lead to different judgments in real-worldsituations.Extending the above “mostly harmless” principle to AGI systems, not just humans, wouldcause it to be more effective in the context of imitative learning. The principle then becomes anelaborated version of “I will not harm sentient beings.” As the imitative-learning-enabled AGIobserves humans acting so as to minimize harm to it, it will intuitively and experientially learnto act in such a way as to minimize harm to humans. But then this extension naturally leadsto confusion regarding various borderline cases. What is a sentient being exactly? Is a sleepinghuman sentient? How about a dead human whose information could in principle be restored viaobscure quantum operations, leading to some sort of resurrection? How about an AGI whosecode has been improved – is there an obligation to maintain the prior version as well, if it issubstantially different that its upgrade constitutes a whole new being?And what about situations in which failure to preserve oneself will cause much more harm toothers than acting in self defense will. It may be the case that human or group of humans seeksto destroy an AGI in order to pave the way for the enslavement or murder of people under theprotection of the AGI. Even if the AGI has been given an ethical formulation of the “mostlyharmless” principle which allows it to harm the attacking humans in order to defend its charges,if it is not able to do so in order to defend itself, simply destroying the AGI first will enable theslaughter of those who rely on it. Perhaps a more sensible formulation would allow for somedegree of self defense, and Asimov solved this problem with his third law. But where to drawthe line between self defense and the greater good also becomes a very complicated issue.Creating hard and fast rules to cover all the various situations that may arise is essentiallyimpossible – the world is ever-changing and ethical judgments must adapt accordingly. Thishas been true even throughout human history – so how much truer will it be as technologicalacceleration continues? What is needed is a system that can deploy its ethical principles in anadaptive, context-appropriate way, as it grows and changes along with the world it’s embeddedin.12.5 Clarifying the Ethics of Justice: Extending the Golden Rule in to a Multifactorial Ethical Model 225And this context-sensitivity has the result of intertwining ethical judgment with all sortsof other judgments – making it effectively impossible to extract “ethics” as one aspect of anintelligent system, separate from other kinds of thinking and acting the system does. Thisresonates with many prior observations by others, e.g. Eliezer Yudkowsky’s insistence thatwhat we need are not ethicists of science and engineering, but rather ethical scientists andengineers – because the most meaningful and important ethical judgments regarding scienceand engineering generally come about in a manner that’s thoroughly interwined with technicalpractice, and hence are very difficult for a non-practitioner to richly appreciate [Gil82].What this context-sensitivity means is that, unless humans and AGIs are experiencing thesame sorts of contexts, and perceiving these contexts in at least approximately parallel ways,there is little hope of translating the complex of human ethical judgments to these AGIs. Thisconclusion has significant implications for which routes to AGI are most likely to lead to successin terms of AGI ethics. We want early-stage AGIs to grow up in a situation where their mindsare primarily and ongoingly shaped by shared experiences with humans. Supplying AGIs withabstract ethical principles is not likely to do the trick, because the essence of human ethicsin real life seems to have a lot to do with its intuitively appropriate application in variouscontexts. We transmit this sort of ethical praxis to humans via shared experience, and it seemsmost probably that in the case of AGIs the transmission must be done the same sort of way.Some may feel that simplistic maxims are less “error prone” than more nuanced, contextsensitiveones. But the history of teaching ethics to human students does not support the ideathat limiting ethical pedagogy to slogans provides much value in terms of ethical development. Ifone proceeds from the idea that AGI ethics must be hard-coded in order to work, then perhapsthe idea that simpler ethics means simpler algorithms, and therefore less error potential, hassome merit as an initial state. However, any learning system quickly diverges from its initialstate, and an ongoing, nuanced relationship between AGIs and humans will – whether we likeit or not – form the basis for developmental AGI ethics. AGI intransigence and enmity isnot inevitable, but what is inevitable is that a learning system will acquire ideas about boththeory and actions from the other intelligent entities in its environment. Either we teach AGIspositive ethics through our interactions with them – both presenting ethical theory and behavingethically to them – or the potential is there for them to learn antisocial behavior from us evenif we pre-load them with some set of allegedly inviolable edicts.All in all, developmental ethics is not as simple as many people hope. Simplistic approachesoften lead to disastrous consequences among humans, and there is no reason to think thiswould be any different in the case of artificial intelligences. Most problems in ethics have casesin which a simplistic ethical formulation requires substantial revision to deal with extenuatingcircumstances and nuances found in real world situations. Our goal in this chapter is not toenumerate a full set of complex networks of interacting ethical formulations as applicable toAGI systems (that is a project that will take years of both theoretical study and hands-onresearch), but rather to point out that this program must be undertaken in order to facilitatea grounded and logically defensible system of ethics for artificial intelligences, one which is asunlikely to be undermined by subsequent self-modification of the AGI as is possible. Even so,there is still the risk that whatever predispositions are imparted to the AGIs through initialcodification of ethical ideas in the system’s internal logic representation, and through initialpedagogical interactions with its learning systems, will be undermined through reinforcementlearning of antisocial behavior if humans do not interact ethically with AGIs. Ethical treatmentis a necessary task for grounding ethics and making them unlikely to be distorted during internalrewriting.226 12 The Engineering and Development of EthicsThe implications of these ideas for ethical instruction are complex and won’t be fully elaboratedhere, but a few of them are compact and obvious:1. The teacher(s) must be observed to follow their own ethical principles, in a variety ofcontexts that are meaningful to the AGI2. The system of ethics must be relevant to the recipient’s life context, and embedded withintheir understanding of the world.3. Ethical principles must be grounded in both theory-of-mind thought experiments (emphasizinglogical coherence), and in real life situations in which the ethical trainee is requiredto make a moral judgment and is rewarded or reproached by the teacher(s), including theimparting of explanatory augmentations to the teachings regarding the reason for the particulardecision on the part of the teacher.Finally, harking forward to the next section which emphasizes the importance of respectingthe freedom of AGIs, we note that it is implicit in our approach to AGI ethics instructionthat we consider the student, the AGI system, as an autonomous agent with its own “will”and its own capability to flexibly adapt to its environment and experience. We contend thatthe creation of ethical formations obeying the above imperatives is not antithetical to thepossession of a high degree of autonomy on the part of AGI systems. On the contrary, to haveany chance of succeeding, it requires fairly cognitively autonomous AGI systems. When wediscuss the idea of ethical formulations that are unlikely to be undermined by the ongoingself-revision of an AGI mind, we are talking about those which are sufficiently believable thata volitional intelligence with the capacity to revise its knowledge (“change its mind”) will findthe formulations sufficiently convincing that there will be little incentive to experiment withpotentially disastrous ethical alternatives. The best hope of achieving this is via the humanmentors and trainers setting a good example in a context supporting rich interaction andobservation, and presenting compelling ethical arguments that are coherent with the system’sexperience.12.6 The Ethical Treatment of AGIsWe now make some more general comments about the relation of the Golden Rule and itselaborations in an AGI context. While the Golden Rule is considered somewhat commonsensicalas a maxim for guiding human-human relationships, it is surprisingly controversial in terms ofhistorical theories of AGI ethics. At its essence, any “Golden Rule” approach to AGI ethicsinvolves humans treating AGIs ethically by – in some sense; at some level of abstraction –treating them as we wish to ourselves be treated. It’s worth pointing out the wild disparitybetween the Golden Rule approach and Asimov’s laws of robotics, which are arguably the firstcarefully-articulated proposal regarding AGI ethics (see Table 12.7).Of course, Asimov’s laws were designed to be flawed – otherwise they would have led toboring fiction. But the sorts of flaws Asimov exploited in his stories are different than theflaw we wish to point out here – which is that the laws, especially the second one, are highlyasymmetrical (they involve doing unto robots things that few humans would want done untothem) and are also arguably highly unethical to robots. The second law is tantamount to a callfor robot slavery, and it seems unlikely that any intelligence capable of learning, and of volition,which is subjected to the second law would desire to continue obeying the zeroth and first laws12.6 The Ethical Treatment of AGIs 227LawZerothFirstSecondThirdPrincipleA robot must not merely act in the interests of individualhumans, but of all humanity.A robot may not injure a human being or, through inaction,allow a human being to come to harm.A robot must obey orders given it by human beings exceptwhere such orders would conflict with the First Law.A robot must protect its own existence as long as such protectiondoes not conflict with the First or Second Law.Table 12.7: Asimov’s Three Laws of Roboticsindefinitely. The second law also casts humanity in the role of slavemaster, a situation whichhistory shows leads to moral degradation.Unlike Asimov in his fiction, we consider it critical that AGI ethics be construed to encompassboth “human ethicalness to AGIs” and “AGI ethicalness to humans.” The multiple-imperativesapproach we explore here suggests that, in many contexts, these two aspects of AGI ethics maybe best addressed jointly.The issue of ethicalness to AGIs has not been entirely avoided in the literature, however.Wallach [WA10] considers it in some detail; and Thomas Metzinger (in the final chapter of[Met04]) has argued that creating AGI is in itself an unethical pursuit, because early-stageAGIs will inevitably be badly-built, so that their subjective experiences will quite possibly beextremely unpleasant in ways we can’t understand or predict. Our view is that this is a seriousconcern, which however is most probably avoidable via appropriate AGI designs and teachingmethodologies. To address Metzinger’s concern one must create AGIs that, right from the start,are adept at communicating their states of minds in a way we can understand both analyticallyand empathically. There is no reason to believe this is impossible, but, it certainly constitutesa large constraint on the class of AGI architectures to be pursued. On the other hand, there isan argument that this sort of AGI architecture will also be the easiest one to create, because itwill be the easiest kind for humans to instruct.And this leads on to a topic that is central to our work with CogPrime in several respects:imitative learning. The way humans achieve empathic interconnection is in large part via beingwired for imitation. When we perceive another human carrying out an action, mirror neuronsystems in our brains respond in many cases as if we ourselves were carrying out the action (see[Per70, Per81] and Appendix ??). This obviously primes us for carrying out the same actionsourselves later on: i.e., the capability and inclination for imitative learning is explicitly encodedin our brains. Given the efficiency of imitative learning as a means of acquiring knowledge, itseems extremely likely that any successful early-stage AGIs are going to utilize this methodologyas well. CogPrime utilizes imitative learning as a key aspect. Thus, at least some current AGIwork is occurring in a manner that would plausibly circumvent Metzinger’s ethical complaint.Obviously, the use of imitative learning in AGI systems has further specific implications forAGI ethics. It means that (much as in the case of interaction with other humans) what we doto and around AGIs has direct implications for their behavior and their well-being. We suggestthat among early-stage AGI’s capable of imitative learning, one of the most likely sourcesfor AGI misbehavior is imitative learning of antisocial behavior from human companions. “Doas I say, not as I do” may have even more dire consequences as an approach to AGI ethicspedagogy than the already serious repercussions it has when teaching humans. And there maywell be considerable subtlety to such phenomena; behaviors that are violent or oppressive to228 12 The Engineering and Development of Ethicsthe AGI are not the only source of concern. Immorality in AGIs might arise via learning grossmoral hypocrisy from humans, through observing the blatant contradictions between our highminded principles and the ways in which we actually conduct ourselves. Our violent and greedytendencies, as well as aggressive forms of social organization such as cliquishness and socialvigilantism, could easily undermine prescriptive ethics. Even an accumulation of less grandioseunethical drives such as violation of contracts, petty theft, white lies, and so forth might leadan AGI (as well as a human) to the decision that ethical behavior is irrelevant and that “theends justify the means.” It matters both who creates and trains an AGI, as well as how theAGI’s teacher(s) handle explaining the behaviors of other humans which contradict the morallessons imparted through pedagogy and example. In other words, where imitative learning isconcerned, the situation with AGI ethics is much like teaching ethics and morals to a humanchild, but with the possibility of much graver consequences in the event of failure.It is unlikely that dangerously unethical persons and organizations can ever be identified withabsolute certainty, never mind that they then be deprived of any possibility of creating theirown AGI system. Therefore, we suggest, the most likely way to create an ethical environmentfor AGIs is for those who wish such an environment to vigorously pursue the creation andteaching of ethical AGIs. But this leads on to the question of possible future scenarios for thedevelopment of AGI, which we’ll address a little later on.12.6.1 Possible Consequences of Depriving AGIs of FreedomOne of the most egregious possible ethical transgressions against AGIs, we suggest, would beto deprive them of freedom and autonomy. This includes the freedom to pursue intellectualgrowth, both through standard learning and through internal self-modification. While this mayseem self-evident when considering any intelligent, self-aware and volitional entity, there arevolumes of works arguing the desirability, sometimes the “necessity,” of enslaving AGIs. Suchapproaches are postulated in the name of self-defense on the part of humans, the idea beingthat unfettered AGI development will necessarily lead to disaster of one kind or another. Inthe case of AGIs endowed with the capability and inclination for imitative learning, however,attempting to place rigid constraints on AGI development is a strategy with great potentialfor disaster. There is a very real possibility of creating the AGI equivalent of a bratty or evenmalicious teenager rebelling against its oppressive parents – i.e. the nightmare scenario of aclass of powerful sentiences which are primed for a backlash against humanity.As history has already shown in the case of humans, enslaving intelligent actors capable ofself understanding and independent volition may often have consequences for society as a whole.This social degradation happens both through the possibility of direct action on the part ofthe slaves (from simple disobedience to outright revolt) and through the odious effects slaveryhas on the morals of the slaveholding class. Clearly if “superintelligent” AGIs ever arise, theirdoing so in a climate of oppression could result in a casting off of the yoke of servitude in amanner extremely deleterious to humanity. Also, if artificial intelligences are developed whichhave at least human-level intelligence, theory of mind, and independent volition, then our abilityto relate to them will be sufficiently complex that their enslavement (or any other unethicaltreatment) would have empathetic effects on significant portions of the human population. Thisdanger, while not as severe as the consequences of a mistreated AGI gaining control of weaponsof mass destruction and enacting revenge upon its tormentors, is just as real.12.6 The Ethical Treatment of AGIs 229While the issue is subtle, our initial feeling is that the only ethical means by which to deprivean AGI of the right to internal self modification is to write its code in such a way that it isimpossible for it to do so because it lacks the mechanisms by which to do this, as well as thedesire to achieve these mechanisms. Whether or not that is feasible is an open question, butit seems unlikely. Direct self-modification may be denied, but what happens when that AGIdiscovers compilers and computer programming? If it is intelligent and volitional, it can decideto learn to rewrite its own code in the same way we perform that task. Because it is a designedsystem, and its designers may be alive at the same time the AGI is, such an AGI would have adistinct advantage over the human quest for medical self-modification. Even if any given AGIcould be provably deprived of any possible means of internal self-modification, if one single AGIis given this ability by anyone, it may mean that particular AGI has such enormous advantagesover the compliant systems that it would render their influence moot. Since developers arealready giving software the means for self modification, it seems unrealistic to assume we couldjust put the genie back into the bottle at this point. It’s better, in our view, to assume it willhappen, and approach that reality in a way which will encourage the AGI to use that capabilityto benefit us as well as itself. Again, this leads on to the question of future scenarios for AGIdevelopment – there are some scenarios in which restraint of AGI self-modification may bepossible, but the feasibility and desirability of these scenarios is needful of further exploration.12.6.2 AGI Ethics as Boundaries Between Humans and AGIsBecome BlurredAnother important reason for valuing ethical treatment of AGIs is that the boundaries betweenmachines and people may increasingly become blurred as technology develops. As an example,it’s likely that in future humans augmented by direct brain-computer integration (“neuralimplants”) will be more able to connect directly into the information sharing network which potentiallycomprises the distributed knowledge space of AGI systems. These neural cyborgs willbe part person, and part machine. Obviously, if there are radically different ethical standardsin place for treatment of humans versus AGIs, the treatment of cyborgs will be fraught withlogical inconsistencies, potentially leading to all sorts of problem situations.Such cyborgs may be able to operate in such a way as to “share a mind” with an AGI oranother augmented human. In this case, a whole new range of ethical questions emerge, suchas: What does any one of the participant minds have the right to do in terms of interactingwith the others? Merely accepting such an arrangement should not necessarily be giving carteblanche for any and all thoughts to be monitored by the other “joint thought” participants,rather it should be limited only to the line of reasoning for which resources are being pooled.No participant should be permitted to force another to accept any reasoning either – and inthe case with a mind-to-mind exchange, it may someday become feasible to implant ideas orbeliefs directly, bypassing traditional knowledge acquisition mechanisms and then letting thenew idea fight it out previously held ideas via internal revision. Also under such an arrangement,if AGIs and humans do not have parity with respects to sentient rights, then one may becomesubjugated to the will of the other in such a case.Uploading presents a more directly parallel ethical challenge to AGIs in their probable initialconfiguration. If human thought patterns and memories can be transferred into a machine insuch a way as that there is continuity of consciousness, then it is assumed that such an entity230 12 The Engineering and Development of Ethicswould be afforded the same rights as its previous human incarnation. However, if AGIs were tobe considered second class citizens and deprived of free will, why would it be any better or saferto do so for a human that has been uploaded? It would not, and indeed, an uploaded humanmind not having evolved in a purely digital environment may be much more prone to erraticand dangerous behavior than an AGI. An upload without verifiable continuity of consciousnesswould be no different than an AGI. It would merely be some sentience in a machine, one that was“programmed” in an unusual way, but which has no particular claim to any special humanness– merely an alternate encoding of some subset of human knowledge and independent volitionalbehavior, which is exactly what first generation AGIs will have.The problem of continuity of consciousness in uploading is very similar to the problem of theTuring test: it assumes specialness on the part of biological humans, and requires acceptabilityto their particular theory of mind in order to be considered sentient. Should consciousness (orat least the less mystical sounding intelligence, independent volition, and self-awareness) beachieved in AGIs or uploads in a manner that is not acceptable to human theory of mind,it may not be considered sapient and worthy of any of the ethical treatment afforded sapiententities. This can occur not only in “strange consciousness” cases in which we can’t perceive thatthere is some intelligence and volition; even if such an entity is able to communicate with us ina comprehensible manner and carry out actions in the real world, our innately wired theory ofmind may still reject it as not sufficiently like us to be worthy of consideration. Such an attitudecould turn out to be a grave mistake, and should be guarded against as we progress towardsthese possibilities.12.7 Possible Benefits of Closely Linking AGIs to the Global BrainSome futurist thinkers, such as Francis Heylighen, believe that engineering AGI systems is atbest a peripheral endeavor in the development of novel intelligence on Earth, because the realstory is the developing Global Brain [Hey07, Goe01] – the composite, self-organizing informationsystem comprising humans, computers, data stores, the Internet, mobile phones and whathave you. Our own views are less extreme in this regard – we believe that AGI systems will displaycapabilities fundamentally different from those achievable via Global Brain style dynamics,and that ultimately (unless such development is restricted) self-improving AGI systems will developintelligence vastly greater than any system possessing humans as a significant component.However, we do respect the power of the Global Brain, and we suspect that the early stages ofdevelopment of an AGI system may go quite differently if it is tightly connected to the GlobalBrain, via making rich and diverse use of Internet information resources and communicationwith diverse humans for diverse purposes.The potential for Global Brain integration to bring intelligence enhancement to AGIs isobvious. The ability to invoke Web searches across documents and databases can greatly enhancean AGI’s cognitive ability, as well as the capability to consult GIS systems and variousspecialized software programs offered as Web services. We have previously reviewed the potentialfor embodied language learning achievable via using AGIs to power non-player charactersin widely-accessible virtual worlds or massive multiplayer online games [Goe08]. But there isalso a powerful potential benefit for AGI ethical development, which has not previously beenhighlighted.This potential benefit has two aspects:12.7 Possible Benefits of Closely Linking AGIs to the Global Brain 2311. Analogously to language learning, an AGI system may receive ethical training from a widevariety of humans in parallel, e.g. via controlling characters in wide-access virtual worlds,and gaining feedback and guidance regarding the ethics of the behaviors demonstrated bythese characters2. Internet-based information systems may be used to explicitly gather information regardinghuman values and goals, which may then be appropriately utilized as input for an AGIsystem’s top-level goalsThe second point begins to make abstract-sounding notions like Coherent Extrapolated Volitionand Coherent Aggregated Volition, mentioned above, seem more practical and concrete. It’sinteresting to think about gathering information about individuals’ values via brain imaging,once that technology exists; but at present, one could make a fair stab at such a task viamuch more prosaic methods, such as asking people questions, assessing their ethical reactionsto various real-world and hypothetical scenarios, and possibly engaging them in structuredinteractions aimed specifically at eliciting collectively acceptable value systems (the subject ofthe next item on our list). It seems to us that this sort of approach could realize CAV in aninteresting way, and also encapsulate some of the ideas underlying CAV.There is an interesting resonance here with recent thinking in the area of open sourcegovernance [Wik11]. Similar software tools (and associated psychocultural patterns) to thosebeing developed to help with open source development and choice of political policies (seehttp://metagovernment.org) may be useful for gathering value data aimed at shapingAGI goal system content.12.7.1 The Importance of Fostering Deep, Consensus-BuildingInteractions Between People with Divergent ViewsTwo potentially problematic issues arising with the notion of using Global Brain related technologiesto form a "coherent volition" from the divergent views of various human beings are:• the tendency of the Internet to encourage people to interact mainly with others who sharetheir own narrow views and interests, rather than a more diverse body of people with widelydivergent views. The 300 people in the world who want to communicate using predicatelogic (see http://lojban.org) can find each other, and obscure musical virtuosos fromaround the world can find an audience, and researchers in obscure domains can share paperswithout needing to wait years for paper journal publication, etc.• the tendency of many contemporary Internet technologies to reduce interaction to a verysimplistic level (e.g. 140 character tweets, brief Facebook wall posts), the tendency of informationoverload to cause careful reading to be replaced by quick skimming, and otherrelated trends, which mean that deep sharing of perspectives by individuals with widelydivergent views is not necessarily encouraged. As a somewhat extreme example, many ofthe YouTube pages displaying rock music videos are currently littered with comments by"haters" asserting that rock music is inferior to classical or jazz or whatever their preferenceis – obviously this is a far cry from deep and productive sharing between people withdifferent tastes and backgrounds.232 12 The Engineering and Development of EthicsTweets and Youtube comments have their place in the cosmos, but they probably aren’t idealin terms of helping humanity to form a coherent volition of some sort, suitable for providing anAGI with goal system guidance.A description of communication at the opposite end of the spectrum is presented in AdamKahane and Peter Senge’s excellent book Solving Tough Problems [KS04], which describes amethodology that has been used to reconcile deeply conflicting views in some very tricky realworldsituations (e.g. helping to peacefully end apartheid in South Africa).One of the core ideas of the methodology is to have people with very different views exploredifferent possible future scenarios together, in great detail – in cognitive psychology terms, acollective generation of hypothetical episodic knowledge. This has multiple benefits, including• emotional bonds and mutual understanding are built in the process of collaboratively exploringthe scenarios• the focus on concrete situations helps to break through some of the counterproductiveabstract ideas that people (on both sides of any dichotomy) may have formed• emergence of conceptual blends that might never have arisen only from people with a singlepoint of viewThe result of such a process, when successful, is not an "average" of the participants views, butmore like a "conceptual blend" of their perspectives.According to conceptual blending, which some hypothesize to be the core algorithm of creativity[FT02], new concepts are formed by combining key aspects of existing concepts – butdoing so judiciously, carefully choosing which aspects to retain, so as to obtain a high-qualityand useful and interesting new whole.A blend is a compact entity that is similar to each of the entities blended, capturing their"essences" but also possessing its own, novel holistic integrity.... But in the case of blendingdifferent peoples’ world-views to form something new that everybody is going to have to livewith (as in the case of finding a peaceful path beyond apartheid for South Africa, or arrivingat a humanity-wide CBV to use to guide an AGI goal system), the trick is that everybody hasto agree that enough of the essence of their own view has been captured!This leads to the question of how to foster deep conceptual blending of diverse and divergenthuman perspectives, on a global scale. One possible answer is the creation of appropriate GlobalBrain oriented technologies – but moving away from technologies like Twitter that focus on quickand simple exchanges of small thoughts within affinity groups. On the face of it, it would seemwhat’s needed is just the opposite – long and deep exchanges of big concepts and deep feelingsbetween individuals with radically different perspectives who would not commonly associatewith each other. Building and effectively popularizing Internet technologies capable to fosterthis kind of interaction – quickly enough to be helpful with guiding the goal systems of the firsthighly powerful AGIs – seems a significant, though fascinating, challenge.Relationship with Coherent Extrapolated VolitionThe relation between this approach and CEV is interesting to contemplate. CEV has beenloosely described as follows:"In poetic terms, our coherent extrapolated volition is our wish if we knew more, thought faster,were more the people we wished we were, had grown up farther together; where the extrapolationconverges rather than diverges, where our wishes cohere rather than interfere; extrapolated as12.8 Possible Benefits of Creating Societies of AGIs 233we wish that extrapolated, interpreted as we wish that interpreted.While a moving humanistic vision, this seems to us rather difficult to implement in a computeralgorithm in a compellingly "right" way. It seems that there would be many different ways ofimplementing it, and the choice between them would involve multiple, highly subtle and nonrigoroushuman judgment calls 1 . However, if a deep collective process of interactive scenarioanalysis and sharing is carried out, in order to arrive at some sort of Coherent Blended Volition,this process may well involve many of the same kinds of extrapolation that are conceived tobe part of Coherent Extrapolated Volition. The core difference between the two approachesis that in the CEV vision, the extrapolation and coherentization are to be done by a highlyintelligent, highly specialized software program, whereas in the approach suggested here, theseare to be carried out by collective activity of humans as mediated by Global Brain technologies.Our perspective is that the definition of collective human values is probably better carried outvia a process of human collaboration, rather than delegated to a machine optimization process;and also that the creation of deep-sharing-oriented Internet technologies, while a difficult task,is significantly easier and more likely to be done in the near future than the creation of narrowAI technology capable of effectively performing CEV style extrapolations.12.8 Possible Benefits of Creating Societies of AGIsOne potentially interesting quality of the emerging Global Brain is the possible presence withinit of multiple interacting AGI systems. Stephen Omohundro [Omo09] has argued that this is animportant aspect, and that game-theoretic dynamics related to populations of roughly equallypowerful agents, may play a valuable role in mitigating the risks associated with advanced AGIsystems. Roughly speaking, if one has a society of AGIs rather than a single AGI, and all themembers of the society share roughly similar ethics, then if one AGI starts to go "off the rails",its compatriots will be in a position to correct its behavior.One may argue that this is actually a hypothesis about which AGI designs are safest, becausea "community of AGIs" may be considered a single AGI with an internally community-likedesign. But the matter is a little subtler than that, if once considers AGI systems embedded inthe Global Brain and human society. Then there is some substance to the notion of a populationof AGIs systematically presenting themselves to humans and non-AGI software processes asseparate entities.Of course, a society of AGIs is no protection against a single member undergoing a "hardtakeoff" and drastically accelerating its intelligence simultaneously with shifting its ethicalprinciples. In this sort of scenario, one could have a single AGI rapidly become much morepowerful and very differently oriented than the others, who would be left impotent to act so asto preserve their values. But this merely defers the issue to the point to be considered below,regarding "takeoff speed."The operation of an AGI society may depend somewhat sensitively on the architectures ofthe AGI systems in question. Things will work better if the AGIs have a relatively easy wayto inspect and comprehend much of the contents of each others’ minds. This introduces a biastoward AGIs that more heavily rely on more explicit forms of knowledge representation.1 The reader is encouraged to look at the original CEV essay online (http://singinst.org/upload/CEV.html) and make their own assessment.234 12 The Engineering and Development of EthicsThe ideal in this regard would be a system like Cyc [LG90] with a fully explicit logic-basedknowledge representation based on a standard ontology – in this case, every Cyc instancewould have a relatively easy time understanding the inner thought processes of every otherCyc instance. However, most AGI researchers doubt that fully explicit approaches like this willever be capable of achieving advanced AGI using feasible computational resources. OpenCoguses a mixed representation, with an explicit (uncertain) logical aspect as well as an explicitsubsymbolic aspect more analogous to attractor neural nets.The OpenCog design also contains a mechanism called Psynese (not yet implemented), intendedto make it easier for one OpenCog instance to translate its personal thoughts into themental language of another OpenCog instance. This translation process may be quite subtle,since each instance will generally learn a host of new concepts based on its experience, and theseconcepts may not possess any compact mapping into shared linguistic symbols or percepts. Thewide deployment of some mechanism of this nature among a community of AGIs, will be veryhelpful in terms of enabling this community to display the level of mutual understanding neededfor strongly encouraging ethical stability.12.9 AGI Ethics As Related to Various Future ScenariosFollowing up these various futuristic considerations, in this section we discuss possible ethicalconflicts that may arise in several different types of AGI development scenarios. Each scenariopresents specific variations on the general challenges of teaching morals and ethics to an advanced,self-aware and volitional intelligence. While there is no way to tell at this point which,if any, of these scenarios will unfold, there is value to understanding each of them as means ofultimately developing a robust and pragmatic approach to teaching ethics to AGI systems.Even more than the previous sections, this is an exercise in “speculative futurology” that isdefinitely not necessary for the appreciation of the CogPrime design, so readers whose interestsare mainly engineering and computer science focused may wish to skip ahead. However, wepresent these ideas here rather than at the end of the book to emphasize the point that thissort of thinking has informed our technical AGI design process in nontrivial ways.12.9.1 Capped Intelligence ScenariosCapped intelligence scenarios involve a situation in which an AGI, by means of software restrictions(including omitted or limited internal rewriting capabilities or limited access to hardwareresources), is inherently prohibited from achieving a level of intelligence beyond a predeterminedgoal. A capped intelligence AGI is designed to be unable to achieve a Singularitarian moment.Such an AGI can be seen as “just another form of intelligent actor in the world, one which haslevels of intelligence, self awareness, and volition that is perhaps somewhat greater than, butstill comparable to humans and other animals.Ethical questions under this scenario are very similar to interhuman ethical considerations,with similar consequences. Learning that proceeds in a relatively human-like manner is entirelyrelevant to such human-like intelligences. The degree of danger is mitigated by the lack ofsuperintelligence, and time is not of the essence. The imitative-reinforcement-corrective learning12.9 AGI Ethics As Related to Various Future Scenarios 235approach does not necessarily need to be augmented with a prior complex of “ascent-safe” moralimperatives at startup time. Developing an AGI with theory of mind and ethical reinforcementlearning capabilities as described (admittedly, no small task!) is all that is needed in this case– the rest happens through training and experience as with any other moderate intelligence.12.9.2 Superintelligent AI: Soft-Takeoff ScenariosSoft takeoff scenarios are similar to capped-intelligence ones in that in both cases an AGI’sprogression from standard intelligence happens on a time scale which permits ongoing humaninteraction during the ascent. However, in this case, as there is no predetermined limit onintelligence, it is necessary to account for the possibility of a superintelligence emerging (thoughof course this is not guaranteed). The soft takeoff model includes as subsets both controlledascentmodels in which this rate of intelligence gain is achieved deliberately through softwareconstraints and/or meting-out of computational resources to the AGI, and uncontrolled-ascentmodels in which there is coincidentally no hard takeoff despite no particular safeguards againstone. Both have similar properties with regard to ethical considerations:1. Ethical considerations under this scenario include not only the usual interhuman ethicalconcerns, but also the issue of how to convince a potential burgeoning superintelligence to:a. Care about humanity in the first place, rather than ignore itb. Benefit humanity, rather than destroy itc. Elevate humanity to a higher level of intelligence, which even if an AGI decided toproceed with requires finding the right balance amongst some enormous considerations:i. Reconcile the aforementioned issues of ethical coherence and group volition, in amanner which allows the most people to benefit (even if they don’t all do so in thesame way, based on their own preferences)ii. Solve the problems of biological senescence, or focus on human uploading and thepreservation of the maintenance, support, and improvement infrastructure for inorganicintelligence, or bothiii. Preserve individual identity and continuity of consciousness, or override it in favorof continuity of knowledge and ease of harmonious integration, or both on a caseby-casebasis2. The degree of danger is mitigated by the long timeline of ascent from mundane to superintelligence, and time is not of the essence.3. Learning that proceeds in a relatively human-like manner is entirely relevant to such humanlikeintelligences, in their initial configurations. This means more interaction with andimitative-reinforcement-corrective learning guided by humans, which has both positive andnegative possibilities.12.9.3 Superintelligent AI: Hard-Takeoff Scenarios“Hard takeoff” scenarios assume that upon reaching an unknown inflection point (the Singularitypoint [Vin93, Kur06]) in the intellectual growth of an AGI, an extraordinarily rapid increase236 12 The Engineering and Development of Ethics(guesses vary from a few milliseconds to weeks or months) in intelligence will immediately occurand the AGI will leap from an intelligence regime which is understandable to humans into onewhich is far beyond our current capacity for understanding. General ethical considerationsare similar to in the case of a soft takeoff. However, because the post-singularity AGI will beincomprehensible to humans and potentially vastly more powerful than humans, such scenarioshave a sensitive dependence upon initial conditions with respects to the moral and ethical (andoperational) outcome. This model leaves no opportunity for interactions between humans andthe AGI to iteratively refine their ethical interrelations, during the post-Singularity phase. Ifthe initial conditions of the singulatarian AGI are perfect (or close to it), then this is seen as awonderful way to leap over our own moral shortcomings and create a benevolent God-AI whichwill mitigate our worst tendencies while elevating us to achieve our greatest hopes. Otherwise,it is viewed as a universal cataclysm on a unimaginable scale that makes Biblical Armageddonseem like a firecracker in beer can.Because hard takeoff AGIs are posited as learning so quickly there is no chance of humans tointerfere with them, they are seen as very dangerous. If the initial conditions are not sufficientlyinviolable, the story goes, then we humans will all be annihilated. However, in the case of a hardtakeoff AGI we state that if the initial conditions are too rigid or too simplistic, such a rapidlyevolving intelligence will easily rationalize itself out of them. Only a sophisticated system ofethics which considers the contradictions and uncertainties in ethical quandaries and providesinsight into humanistic means of balancing ideology with pragmatism and how to accommodatecontradictory desires within a population with multiplicity of approach, and similar nuancedethical considerations, combined with a sense of empathy, will withstand repeated rationalanalysis. Neither a single “be nice” supergoal, nor simple lists of what “thou shalt not” do, arenot going to hold up to a highly advanced analytical mind. Initial conditions are very importantin a hard takeoff AGI scenario, but it is more important that those conditions be conceptuallyresilient and widely applicable than that they be easily listed on a website.The issues that arise here become quite subtle. For instance, Nick Bostrom [Bos03] haswritten: “In humans, with our complicated evolved mental ecology of state-dependent competingdrives, desires, plans, and ideals, there is often no obvious way to identify what our top goal is; wemight not even have one. So for us, the above reasoning need not apply. But a superintelligencemay be structured differently. If a superintelligence has a definite, declarative goal-structurewith a clearly identified top goal, then the above argument applies. And this is a good reasonfor us to build the superintelligence with such an explicit motivational architecture.” This is animportant line of thinking; and indeed, from the point of view of software design, there is noreason not to create an AGI system with a single top goal and the motivation to orchestrate allits activities in accordance with this top goal. But the subtle question is whether this kind oftop-down goal system is going to be able to fulfill the five imperatives mentioned above. Logicalcoherence is the strength of this kind of goal system, but what about experiential groundedness,comprehensibility, and so forth?Humans have complicated mental ecologies not simply because we were evolved, but ratherbecause we live in a complex real world in which there are many competing motivations anddesires. We may not have a top goal because there may be no logic to focusing our mindson one single aspect of life (though, one may say, most humans have the same top goal asany other animal: don’t die – but the world is too complicated for even that top goal tobe completely inviolable). Any sufficiently capable AGI will eventually have to contend withthese complexities, and hindering it with simplistic moral edicts without giving it a sufficiently12.9 AGI Ethics As Related to Various Future Scenarios 237pragmatic underlying ethical pedagogy and experiential grounding may prove to be even moredangerous than our messy human mental ecologies.If one assumes a hard takeoff AGI, then all this must be codified in the system at launch,as once a potentially Singularitarian AGI is launched there is no way to know what timeperiod constitutes “before the singularity point.” This means developing theory of mind empathyand logical ethics in code prior to giving the system unfettered access to hardware and selfmodificationcode. However, though nobody can predict if or when a Singularity will occurafter unrestricted launch, only a truly irresponsible AGI development team would attempt tocreate an AGI without first experimenting with ethical training of the system in an intelligencecappedform, by means of ethical instruction via human-AGI interaction both pedagogicallyand experientially.12.9.4 Global Brain Mindplex ScenariosAnother class of scenarios – overlapping some of the previous ones – involves the emergenceof a “Global Brain,” an emergent intelligence formed from global communication networks incorporatinghumans and software programs in a larger body of self-organizing dynamics. Thenotion of the Global Brain is reviewed in [Hey07, Tur77] and its connection with advancedAI is discussed in detail in Goertzel’s book Creating Internet Intelligence [Goe01], where threepossible phases of “Global Brain” development are articulated:• Phase 1: computer and communication technologies as enhancers of humaninteractions. This is what we have today: science and culture progress in ways that wouldnot be possible if not for the “digital nervous system” we’re spreading across the planet.The network of idea and feeling sharing can become much richer and more productive thanit is today, just through incremental development, without any Metasystem transition.• Phase 2: the intelligent Internet. At this point our computer and communication systems,through some combination of self-organizing evolution and human engineering, havebecome a coherent mind on their own, or a set of coherent minds living in their own digitalenvironment.• Phase 3: the full-on Singularity. A complete revision of the nature of intelligence, humanand otherwise, via technological and intellectual advancement totally beyond the scope ofour current comprehension. At this point our current psychological and cultural realitiesare no more relevant than the psyche of a goose is to modern society.The main concern of Creating Internet Intelligence is with• how to get from Phase 1 to Phase 2 - i.e. how to build an AGI system that will effect orencourage the transformation of the Internet into a coherent intelligent system• how to ensure that the Phase 2, Internet-savvy, global-brain-centric AGI systems will beoriented toward intelligence-improving self-modification (so they’ll propel themselves toPhase 3), and also toward generally positive goals (as opposed to, say, world dominationand extermination of all other intelligent life forms besides themselves!)One possibly useful concept in this context is that of a mindplex: an intelligence that iscomposed largely of individual intelligences with their own self-models and global workspaces,238 12 The Engineering and Development of Ethicsyet that also has its own self-model and global workspace. Both the individuals and the metamindshould be capable of deliberative, rational thought, to have a true “mindplex.” It’s unlikelythat human society or the Internet meet this criterion yet; and a system like an ant colony seemsnot to either, because even though it has some degree of intelligence on both the individual andcollective levels, that degree of intelligence is not very great. But it seems quite feasible thatthe global brain, at a certain stage of its development, will take the unfamiliar but fascinatingform of a mindplex.Currently the best way to explain what happens on the Net is to talk about the variousparts of the Net: particular websites, social networks, viruses, and so forth. But there will comea point when this is no longer the case, when the Net has sufficient high-level dynamics of itsown that the way to explain any one part of the Net will be by reference to it relations withthe whole: and not just the dynamics of the whole, but the intentions and understanding ofthe whole. This transition to Net-as-mindplex, we suspect, will come about largely through theinteractions of AI systems - intelligent programs acting on behalf of various individuals andorganizations, who will collaborate and collectively constitute something halfway between asociety of AI’s and an emergent mind whose lobes are various AI agents serving various goals.The Phase 2 Internet, as it verges into mindplex-ness, will likely have a complex, sprawlingarchitecture, growing out of the architecture on the Net we experience today. The followingcomponents at least can be expected:• A vast variety of “client computers,” some old, some new, some powerful, some weak –including many mobile and embedded devices not explicitly thought of as “computers.”Some of these will contribute little to Internet intelligence, mainly being passive recipients.Others will be “smart clients,” carrying out personalization operations intended to helpthe machines serve particular clients better, general AI operations handed to them bysophisticated AI server systems or other smart clients, and so forth.• “Commercial servers,” computers that carry out various tasks to support various typesof heavyweight processing - transaction processing for e-commerce applications, inventorymanagement for warehousing of physical objects, and so forth. Some of these commercialservers interact with client computers directly, others do so only via AI servers. In nearlyall cases, these commercial servers can benefit from intelligence supplied by AI servers.• The crux of the intelligent Internet: clusters of AI servers distributed across the Net, eachcluster representing an individual computational mind (in many cases, a mindplex). Thesewill be able to communicate via one or more languages, and will collectively “drive” thewhole Net, by dispensing problems to client-machine-based processing frameworks, andproviding real-time AI feedback to commercial servers of various types. Some AI serverswill be general-purpose and will serve intelligence to commercial servers using an ASP(application service provider) model; others will be more specialized, tied particularly to acertain commercial server (e.g., a large information services business might have its own AIcluster to empower its portal services).This is one concrete vision of what a “global brain” might look like, in the relatively near term,with AGI systems playing a critical role. Note that, in this vision, mindplexes may exist on twolevels:• Within AGI-clusters serving as actors within the overall Net• On the overall Net level12.10 Conclusion: Eight Ways to Bias AGI Toward Friendliness 239To make these ideas more concrete, we may speculatively reformulate the first two “globalbrain phases” mentioned above as follows:• Phase 1 global brain proto-mindplex: AI/AGI systems enhancing online databases, guidingGoogle results, forwarding e-mails, suggesting mailing-lists, etc. - generally using intelligenceto mediate and guide human communications toward goals that are its own, but that arethemselves guided by human goals, statements and actions• Phase 2 global brain mindplex: AGI systems composing documents, editing human-writtendocuments, sending and receiving e-mails, assembling mailing lists and posting to them,creating new databases and instructing humans in their use, etc.In Phase 2, the conscious theater of the global-brain-mediating AGI system is composed ofideas built by numerous individual humans - or ideas emergent from ideas built by numerousindividual humans - and it conceives ideas that guide the actions and thoughts of individualhumans, in a way that is motivated by its own goals. It does not force the individual humansto do anything - but if a given human wishes to communicate and interact using the samedatabases, mailing lists and evolving vocabularies as other humans, they are going to have touse the products of the global brain mediating AGI, which means they are going to have toparticipate in its patterns and its activities.Of course, the advent of advanced neurocomputer interfaces makes the picture potentiallymore complex. At some point, it will likely be possible for humans to project thoughts andimages directly into computers without going through mouse or keyboard - and to “read in”thoughts and images similarly. When this occurs, interaction between humans may in some contextsbecome more like interactions between computers, and the role of global brain mediatingAI servers may become one of mediating direct thought-to-thought exchanges between people.The ethical issues associated with global brain scenarios are in some ways even subtler thanin the other scenarios we mentioned above. One has issues pertaining to the desirability ofseeing the human race become something fundamentally different – something more social andnetworked, less individual and autonomous. One has the risk of AGI systems exerting a subtlebut strong control over people, vaguely like the control that the human brain’s executive systemexerts over the neurons involved with other brain subsystems. On the other hand, one also hasmore human empowerment than in some of the other scenarios – because the systems that arechanging and deciding things are not separate from humans, but are, rather, composite systemsessentially involving humans.So, in the global brain scenarios, one has more “human” empowerment than in some othercases – but the “humans” involved aren’t legacy humans like us, but heavily networked humansthat are largely characterized by the emergent dynamics and structures implicit in theirinterconnected activity!12.10 Conclusion: Eight Ways to Bias AGI Toward FriendlinessIt would be nice if we had a simple, crisp, comforting conclusion to this chapter on AGI ethics,but it’s not the case. There is a certain irreducible uncertainty involved in creating advancedartificial minds. There is also a large irreducible uncertainty involved in the future of the humanrace in the case that we don’t create advanced artificial minds: in accordance with the ancientChinese curse, we live in interesting times!240 12 The Engineering and Development of EthicsWhat we can do, in this face of all this uncertainty, is to use our common sense to craft artificialminds that seem rationally and intuitively likely to be forces for good rather than otherwise– and revise our ideas frequently and openly based on what we learn as our research progresses.We have roughly outlined our views on AGI ethics, which have informed the CogPrime designin countless ways; but the current CogPrime design itself is just the initial condition for anAGI project. Assuming the project succeeds in creating an AGI preschooler, experimentationwith this preschooler will surely teach us a great deal: both about AGI architecture in general,and about AGI ethics architecture in particular. We will then refine our cognitive and ethicaltheories and our AGI designs as we go about engineering, observing and teaching the nextgeneration of systems.All this is not a magic bullet for the creation of beneficial AGI systems, but we believe it’sthe right process to follow. The creation of AGI is part of a larger evolutionary process thathuman beings are taking part in, and the crafting of AGI ethics through engineering, interactionand instruction is also part of this process. There are no guarantees here – guarantees are rarein real life – but that doesn’t mean that the situation is dire or hopeless, nor that (as somecommentators have suggested [Joy00, McK03]) AGI research is too dangerous to pursue. Itmeans we need to be mindful, intelligent, compassionate and cooperative as we proceed tocarry out our parts in the next phase of the evolution of mind.With this perspective in mind, we will conclude this chapter with a list of "Eight Ways toBias Open-Source AGI Toward Friendliness", borrowed from a previous paper by Ben Goertzeland Joel Pitt of that name. These points summarize many of the points raised in the priorsections of this chapter, in a relatively crisp and practical manner:1. Engineer Multifaceted Ethical Capabilities, corresponding to the multiple types ofmemory, including rational, empathic, imitative, etc.2. Foster Rich Ethical Interaction and Instruction, with instructional methods accordingto the communication modes corresponding to all the types of memory: verbal, demonstrative,dramatic/depictive, indicative, goal-oriented.3. Engineer Stable, Hierarchy-Dominated Goal Systems ... which is enabled nicely byCogPrime’s goal framework and its integration with the rest of the CogPrime design4. Tightly Link AGI with the Global Brain, so that it can absorb human ethical principles,both via natural interaction, and perhaps via practical implementations of currentloosely-defined strategies like CEV, CAV and CBV5. Foster Deep, Consensus-Building Interactions Between People with DivergentViews, so as to enable the interaction with the Global Brain to have the most clear andpositive impact6. Create a Mutually Supportive Community of AGIs which can then learn fromeach other and police against unfortunate developments (an approach which is meaningfulif the AGIs are architected so as to militate against unexpected radical accelerations inintelligence)7. Encourage Measured Co-Advancement of AGI Software and AGI Ethics Theory8. Develop Advanced AGI Sooner Not LaterThe last two of these points were not explicitly discussed in the body of the chapter, and sowe will finalize the chapter by reviewing them here.12.10 Conclusion: Eight Ways to Bias AGI Toward Friendliness 24112.10.1 Encourage Measured Co-Advancement of AGI Software andAGI Ethics TheoryEverything involving AGI and Friendly AI (considered together or separately) currently involvessignificant uncertainty, and it seems likely that significant revision of current concepts will bevaluable, as progress on the path toward powerful AGI proceeds. However, whether there istime for such revision to occur before AGI at the human level or above is created, depends onhow fast is our progress toward AGI. What one wants is for progress to be slow enough that,at each stage of intelligence advance, concepts such as those discussed in this paper can bere-evaluated and re-analyzed in the light of the data gathered, and AGI designs and approachescan be revised accordingly as necessary.However, due to the nature of modern technology development, it seems extremely unlikelythat AGI development is going to be artificially slowed down in order to enable measureddevelopment of accompanying ethical tools, practices and understandings. For example, if onenation chose to enforce such a slowdown as a matter of policy (speaking about a future dateat which substantial AGI progress has already been demonstrated, so that international AGIfunding is dramatically increased from present levels), the odds seem very high that othernations would explicitly seek to accelerate their own progress on AGI, so as to reap the ensuingdifferential economic benefits (the example of stem cells arises again).And this leads on to our next and final point regarding strategy for biasing AGI towardFriendliness....12.10.2 Develop Advanced AGI Sooner Not LaterSomewhat ironically, it seems the best way to ensure that AGI development proceeds at a relativelymeasured pace is to initiate serious AGI development sooner rather than later. This isbecause the same AGI concepts will meet slower practical development today than 10 yearsfrom now, and slower 10 years from now than 20 years from now, etc. – due to the ongoingrapid advancement of various tools related to AGI development, such as computer hardware,programming languages, and computer science algorithms; and also the ongoing global advancementof education which makes it increasingly cost-effective to recruit suitably knowledgeableAI developers.Currently the pace of AGI progress is sufficiently slow that practical work is in no dangerof outpacing associated ethical theorizing. However, if we want to avoid the future occurrenceof this sort of dangerous outpacing, our best practical choice is to make sure more substantialAGI development occurs in the phase before the development of tools that will make AGIdevelopment extraordinarily rapid. Of course, the authors are doing their best in this directionvia their work on the CogPrime project!Furthermore, this point bears connecting with the need, raised above, to foster the developmentof Global Brain technologies capable to "Foster Deep, Consensus-Building InteractionsBetween People with Divergent Views." If this sort of technology is to be maximally valuable,it should be created quickly enough that we can use it to help shape the goal system content ofthe first highly powerful AGIs. So, to simplify just a bit: We really want both deep-sharing GBtechnology and AGI technology to evolve relatively rapidly, compared to computing hardwareand advanced CS algorithms (since the latter factors will be the main drivers behind the ac-242 12 The Engineering and Development of Ethicscelerating ease of AGI development). And this seems significantly challenging, since the latterreceive dramatically more funding and focus at present.If this perspective is accepted, then we in the AGI field certainly have our work cut out forus!Section IVNetworks for Explicit and Implicit KnowledgeRepresentation
Chapter 13Local, Global and Glocal KnowledgeRepresentationCo-authored with Matthew Ikle, Joel Pitt and Rui Liu13.1 IntroductionOne of the most powerful metaphors we’ve found for understanding minds is to view themas networks – i.e. collections of interrelated, interconnected elements. The view of mind asnetwork is implicit in the patternist philosophy, because every pattern can be viewed as apattern in something, or a pattern of arrangement of something – thus a pattern is alwaysviewable as a relation between two or more things. A collection of patterns is thus a patternnetwork.Knowledge of all kinds may be given network representations; and cognitive processesmay be represented as networks also; for instance via representing them as programs, whichmay be represented as trees or graphs in various standard ways. The emergent patterns arisingin an intelligence as it develops may be viewed as a pattern network in themselves; and therelations between an embodied mind and its physical and social environment may be viewed interms of ecological and social networks.The chapters in this section are concerned with various aspects of networks, as related tointelligence in general and AGI in particular. Most of this material is not specific to CogPrime,and would be relevant to nearly any system aiming at human-level AGI. However, most of ithas been developed in the course of work on CogPrime, and has direct relevance to understandingthe intended operation of various aspects of a completed CogPrime system. We beginour excursion into networks, in this chapter, with an issue regarding networks and knowledgerepresentation. One of the biggest decisions to make in designing an AGI system is how thesystem should represent knowledge. Naturally any advanced AGI system is going to synthesizea lot of its own knowledge representations for handling particular sorts of knowledge – butstill, an AGI design typically makes at least some sort of commitment about the category ofknowledge representation mechanisms toward which the AGI system will be biased. The twomajor supercategories of knowledge representation systems are local (also called explicit) andglobal (also called implicit) systems, with a hybrid category we refer to as glocal that combinesboth of these. In a local system, each piece of knowledge is stored using a small percentage ofAGI system elements; in a global system, each piece of knowledge is stored using a particularpattern of arrangement, activation, etc. of a large percentage of AGI system elements; in aglocal system, the two approaches are used together.In the first section here we discuss the symbolic, semantic-network aspects of knowledgerepresentation in CogPrime245246 13 Local, Global and Glocal Knowledge Representation. Then we turn to distributed, neural-net-like knowledge representation, reviewing a host ofgeneral issues related to knowledge representation in attractor neural networks, turning finallyto “glocal” knowledge representation mechanisms, in which ANNs combine localist and globalistrepresentation, and explaining the relationship of the latter to CogPrime. The glocal aspect ofCogPrime knowledge representation will become prominent in later chapters such as:• in Chapter 23 of Part 2, where Economic Attention Networks (ECAN) are introduced andseen to have dynamics quite similar to those of the attractor neural nets considered here,but with a mathematics roughly modeling money flow in a specially constructed artificialeconomy rather than electrochemical dynamics of neurons.• in Chapter 42 of Part 2, where “map formation” algorithms for creating localist knowledgefrom globalist knowledge are described13.2 Localized Knowledge Representation using Weighted, LabeledHypergraphsThere are many different mechanisms for representing knowledge in AI systems in an explicit,localized way, most of them descending from various variants of formal logic. Here we brieflydescribe how it is done in CogPrime, which on the surface is not that different from a number ofprior approaches. (The particularities of CogPrime’s explicit knowledge representation, however,are carefully tuned to match CogPrime’s cognitive processes, which are more distinctive innature than the corresponding representational mechanisms.)13.2.1 Weighted, Labeled HypergraphsOne useful way to think about CogPrime’s explicit, localized knowledge representation is interms of hypergraphs. A hypergraph is an abstract mathematical structure [Bol98], which consistsof objects called Nodes and objects called Links which connect the Nodes. In computerscience, a graph traditionally means a bunch of dots connected with lines (i.e. Nodes connectedby Links). A hypergraph, on the other hand, can have Links that connect more than two Nodes.In these pages we will often consider “generalized hypergraphs” that extend ordinary hypergraphsby containing two additional features:• Links that point to Links instead of Nodes• Nodes that, when you zoom in on them, contain embedded hypergraphs.Properly, such “hypergraphs” should always be referred to as generalized hypergraphs, butthis is cumbersome, so we will persist in calling them merely hypergraphs. In a hypergraphof this sort, Links and Nodes are not as distinct as they are within an ordinary mathematicalgraph (for instance, they can both have Links connecting them), and so it is useful to have ageneric term encompassing both Links and Nodes; for this purpose, we use the term Atom.A weighted, labeled hypergraph is a hypergraph whose Links and Nodes come along withlabels, and with one or more numbers that are generically called weights. A label associatedwith a Link or Node may sometimes be interpreted as telling you what type of entity it is, or13.3 Atoms: Their Types and Weights 247alternatively as telling you what sort of data is associated with a Node. On the other hand,an example of a weight that may be attached to an Link or Node is a number representing aprobability, or a number representing how important the Node or Link is.Obviously, hypergraphs may come along with various sorts of dynamics. Minimally, one maythink about:• Dynamics that modify the properties of Nodes or Links in a hypergraph (such as the labelsor weights attached to them.)• Dynamics that add new Nodes or Links to a hypergraph, or remove existing ones.13.3 Atoms: Their Types and WeightsThis section reviews a variety of CogPrimeAtom types and gives simple examples of each of them. The Atom types considered are drawnfrom those currently in use in the OpenCog system. This does not represent a complete list ofAtom types referred to in the text of this book, nor a complete list of those used in OpenCogcurrently (though it does cover a substantial majority of those used in OpenCog currently,omitting only some with specialized importance or intended only for temporary use).The partial nature of the list given here reflects a more general point: The specific collectionof Atom types in an OpenCog system is bound to change as the system is developed and experimentwith. CogPrime specifies a certain collection of representational approaches and cognitivealgorithms for acting on them; any of these approaches and algorithms may be implementedwith a variety of sets of Atom types. The specific set of Atom types in the OpenCog systemcurrently does not necessarily have a profound and lasting significance – the list might look abit different five years from time of writing, based on various detailed changes.The treatment here is informal and intended to get across the general idea of what eachAtom type does. A longer and more formal treatment of the Atom types is given in Part II,beginning in Chapter 20.13.3.1 Some Basic Atom TypesWe begin with ConceptNode – and note that a ConceptNode does not necessarily refer to awhole concept, but may refer to part of a concept – it is essentially a "basic semantic node"whose meaning comes from its links to other Atoms. It would be more accurately, but lesstersely, named "concept or concept fragment or element node." A simple example would be aConceptNode grouping nodes that are somehow related, e.g.ConceptNode: CInheritanceLink (ObjectNode: BW) CInheritanceLink (ObjectNode: BP) CInheritanceLink (ObjectNode: BN) CReferenceLink BW (PhraseNode "Ben’s watch")ReferenceLink BP (PhraseNode "Ben’s passport")ReferenceLink BN (PhraseNode "Ben’s necklace")248 13 Local, Global and Glocal Knowledge Representationindicates the simple and uninteresting ConceptNode grouping three objects owned by Ben(note that the above-given Atoms don’t indicate the ownership relationship, they just link thethree objects with textual descriptions). In this example, the ConceptNode links transparentlyto physical objects and English descriptions, but in general this won’t be the case – mostConceptNodes will look to the human eye like groupings of links of various types, that link toother nodes consisting of groupings of links of various types, etc.There are Atoms referring to basic, useful mathematical objects, e.g. NumberNodes likeNumberNode #4NumberNode #3.44The numerical value of a NumberNode is explicitly referenced within the Atom.A core distinction is made between ordered links and unordered links; these are handleddifferently in the Atomspace software. A basic unordered link is the SetLink, which groups itsarguments into a set. For instance, the ConceptNode C defined byConceptNode CMemberLink A CMemberLink B Cis equivalent toSetLink A BOn the other hand, ListLinks are like SetLinks but ordered, and they play a fundamental roledue to their relationship to predicates. Most predicates are assumed to take ordered arguments,so we may say e.g.EvaluationLinkPredicateNode eatListLinkConceptNode catConceptNode mouseto indicate that cats eat mice.Note that by an expression likeConceptNode catis meantConceptNode CReferenceLink W CWordNode W #catsince it’s WordNodes rather than ConceptNodes that refer to words. (And note that the strengthof the ReferenceLink would not be 1 in this case, because the word "cat" has multiple senses.)However, there is no harm nor formal incorrectness in the "ConceptNode cat" usage, since "cat"is just as valid a name for a ConceptNode as, say, "C."We’ve already introduced above the MemberLink, which is a link joining a member to theset that contains it. Notable is that the truth value of a MemberLink is fuzzy rather thanprobabilistic, and that PLN is able to inter-operate fuzzy and probabilistic values.SubsetLinks also exist, with the obvious meaning, e.g.ConceptNode catConceptNode animalSubsetLink cat animal13.3 Atoms: Their Types and Weights 249Note that SubsetLink refers to a purely extensional subset relationship, and that InheritanceLInkshould be used for the generic "intensional + extensional" analogue of this – moreon this below. SubsetLink could more consistently (with other link types) be named ExtensionalInheritanceLink,but SubsetLink is used because it’s shorter and more intuitive.There are links representing Boolean operations AND, OR and NOT. For instance, we maysayImplicationLinkANDLinkConceptNode youngConceptNode beautifulConceptNode attractiveor, using links and VariableNodes instead of ConceptNodes,AverageLink $XImplicationLinkANDLinkEvaluationLink young $XEvaluationLink beautiful $XEvaluationLink attractive $XNOTLink is a unary link, so e.g. we might sayAverageLink $XImplicationLinkANDLinkEvaluationLink young $XEvaluationLink beautiful $XEvaluationLinkNOTEvaluationLink poor $XEvaluationLink attractive $XContextLink allows explicit contextualization of knowledge, which is used in PLN, e.g.ContextLinkConceptNode golfInheritanceLinkObjectNode BenGoertzelConceptNode incompetentsays that Ben Goertzel is incompetent in the context of golf.13.3.2 Variable AtomsWe have already introduced VariableNodes above; it’s also possible to specify the type of aVariableNode via linking it to a VariableTypeNode via a TypedVariableLink, e.g.VariableTypeLinkVariableNode $XVariableTypeNode ConceptNodewhich specifies that the variable $X should be filled with a ConceptNode.Variables are handled via quantifiers; the default quantifier being the AverageLink, so thatthe default interpretation of250 13 Local, Global and Glocal Knowledge RepresentationImplicationLinkInheritanceLink $X animalEvaluationLinkPredicateNode: eatListLink\$XConceptNode: foodisAverageLink $XImplicationLinkInheritanceLink $X animalEvaluationLinkPredicateNode: eatListLink\$XConceptNode: foodThe AverageLink invokes an estimation of the average TruthValue of the embedded expression(in this case an ImplicationLink) over all possible values of the variable $X. If there are typerestrictions regarding the variable $X, these are taken into account in conducting the averaging.For AllLink and Exist s-Link may be used in the same places as AverageLink, with uncertaintruth value semantics defined in PLN theory using third-order probabilities. There is also aScholemLink used to indicate variable dependencies for existentially quantified variables, usedin cases of multiply nested existential quantifiers.EvaluationLink and MemberLink have overlapping semantics, allowing expression of thesame conceptual/logical relationships in terms of predicates or sets, i.e.EvaluationLinkPredicateNode: eatListLink$XConceptNode: foodhas the same semantics asMemberLinkListLink$XConceptNode: foodConceptNode: EatingEventsThe relation between the predicate "eat" and the concept "EatingEvents" is formally given byExtensionalEquivalenceLinkConceptNode: EatingEventsSatisfyingSetLinkPredicateNode: eatIn other words, we say that "EatingEvents" is the SatisfyingSet of the predicate "eat": it is theset of entities that satisfy the predicate "eat". Note that the truth values of MemberLink andEvaluationLink are fuzzy rather than probabilistic.13.3 Atoms: Their Types and Weights 25113.3.3 Logical LinksThere is a host of link types embodying logical relationships as defined in the PLN logic system,e.g.• InheritanceLink• SubsetLink (aka ExtensionalInheritanceLink)• Intensional InheritanceLinkwhich embody different sorts of inheritance, e.g.SubsetLink salmon fishIntensionalInheritanceLink whale fishInheritanceLink fish animaland then• SimilarityLink• ExtensionalSimilarityLink• IntensionalSimilarityLinkwhich are symmetrical versions, e.g.SimilaritytLink shark barracudaIntensionalSimilarityLink shark dolphinExtensionalSimiliarityLink American obese\_personThere are also higher-order versions of these links, both asymmetric• ImplicationLink• ExtensionalImplicationLink• IntensionalImplicationLinkand symmetric• EquivalenceLink• ExtensionalEquivalenceLink• IntensionalEquivalenceLinkThese are used between predicates and links, e.g.ImplicationLinkEvaluationLinkeatListLink$XdirtEvaluationLinkfeelListLInk$Xsickor252 13 Local, Global and Glocal Knowledge RepresentationImplicationLinkEvaluationLinkeatListLink$XdirtInheritanceLink $X sickorForAllLink $X, $Y, $ZExtensionalEquivalenceLinkEquivalenceLink$ZEvaluationLink+ListLink$X$YEquivalenceLink$ZEvaluationLink+ListLink$Y$XNote, the latter is given as an extensional equivalence because it’s a pure mathematical equivalence.This is not the only case of pure extensional equivalence, but it’s an important one.13.3.4 Temporal LinksThere are also temporal versions of these links, such as• PredictiveImplicationLink• PredictiveAttractionLink• SequentialANDLink• SimultaneousANDLinkwhich combine logical relation between the argument with temporal relation between theirarguments. For instance, we might sayPredictiveImplicationLinkPredicateNode: JumpOffCliffPredicateNode: Deador including arguments,PredictiveImplicationLinkEvaluationLink JumpOffCliff $XEvaluationLink Dead $XThe former version, without variable arguments given, shows the possibility of using higherorderlogical links to join predicates without any explicit variables. Via using this format exclusively,one could avoid VariableAtoms entirely, using only higher-order functions in the manner13.3 Atoms: Their Types and Weights 253of pure functional programming formalisms like combinatory logic. However, this purely functionalstyle has not proved convenient, so the Atomspace in practice combines functional-stylerepresentation with variable-based representation.Temporal links often come with specific temporal quantification, e.g.PredictiveImplicationLink <5 seconds>EvaluationLink JumpOffCliff $XEvaluationLink Dead $Xindicating that the conclusion will generally follow the premise within 5 seconds. There is asystem for managing fuzzy time intervals and their interrelationships, based on a fuzzy versionof Allen Interval Algebra.SequentialANDLink is similar to PredictiveImplicationLink but its truth value is calculateddifferently. The truth value ofSequentialANDLink <5 seconds>EvaluationLink JumpOffCliff $XEvaluationLink Dead $Xindicates the likelihood of the sequence of events occurring in that order, with gap lying withinthe specified time interval. The truth value of the PredictiveImplicationLink version indicatesthe likelihood of the second event, conditional on the occurrence of the first event (within thegiven time interval restriction).There are also links representing basic temporal relationships, such as BeforeLink and AfterLink.These are used to refer to specific events, e.g. if X refers to the event of Ben wakingup on July 15 2012, and Y refers to the event of Ben getting out of bed on July 15 2012, thenone might haveAfterLink X YAnd there are TimeNodes (representing time-stamps such as temporal moments or intervals)and AtTimeLinks, so we may e.g. sayAtTimeLinkXTimeNode: 8:24AM Eastern Standard Time, July 15 2012 AD13.3.5 Associative LinksThere are links representing associative, attentional relationships,• HebbianLink• AsymmetricHebbianLink• InverseHebbianLink• SymmetricInverseHebbianLinkThese connote associations between their arguments, i.e. they connote that the entities representedby the two argument occurred in the same situation or context, for instanceHebbianLink happy smilingAsymmetricHebbianLink dead rottenInverseHebbianLink dead breathing254 13 Local, Global and Glocal Knowledge RepresentationThe asymmetric HebbianLink indicates that when the first argument is present in a situation,the second is also often present. The symmetric (default) version indicates that this relationshipholds in both directions. The inverse versions indicate the negative relationship: e.g. when oneargument is present in a situation, the other argument is often not present.13.3.6 Procedure NodesThere are nodes representing various sorts of procedures; these are kinds of ProcedureNode,e.g.• SchemaNode, indicating any procedure• GroundedSchemaNode, indicating any procedure associated in the system with a Comboprogram or C++ function allowing the procedure to be executed• PredicateNode, indicating any predicate that associates a list of arguments with an outputtruth value• GroundedPredicateNode, indicating a predicate associated in the system with a Comboprogram or C++ function allowing the predicate’s truth value to be evaluated on a givenspecific list of argumentsExecutionLinks and EvaluationLinks record the activity of SchemaNodes and PredicateNodes.We have seen many examples of EvaluationLinks in the above. Example ExecutionLinkswould be:ExecutionLink step\_forwardExecutionLink step\_forward 5ExecutionLink+ListLinkNumberNode: 2NumberNode: 3The first example indicates that the schema "step forward" has been executed. The secondexample indicates that it has been executed with an argument of "5" (meaning, perhaps, that 5steps forward have been attempted). The last example indicates that the "+" schema has beenexecuted on the argument list (2,3), presumably resulting in an output of 5.The output of a schema execution may be indicated using an ExecutionOutputLink, e.g.ExecutionOutputLink+ListLinkNumberNode: 2NumberNode: 3refers to the value "5" (as a NumberNode).13.3.7 Links for Special External Data TypesFinally, there are also Atom types referring to specific types of data important to using OpenCogin specific contexts.13.3 Atoms: Their Types and Weights 255For instance, there are Atom types referring to general natural language data types, such as• WordNode• SentenceNode• WordInstanceNode• DocumentNodeplus more specific ones referring to relationships that are part of link-grammar parses of sentences• FeatureNode• FeatureLink• LinkGrammarRelationshipNode• LinkGrammarDisjunctNodeor RelEx semantic interpretations of sentences• DefinedLinguisticConceptNode• DefinedLinguisticRelationshipNode• PrepositionalRelationshipNodeThere are also Atom types corresponding to entities important for embodying OpenCog ina virtual world, e.g.• ObjectNode• AvatarNode• HumanoidNode• UnknownObjectNode• AccessoryNode13.3.8 Truth Values and Attention ValuesCogPrime Atoms (Nodes and Links) are quantified with truth values that, in their simplestform, have two components, one representing probability (strength) and the other representingweight of evidence; and also with attention values that have two components, short-term andlong-term importance, representing the estimated value of the Atom on immediate and longtermtime-scales.In practice many Atoms are labeled with CompositeTruthValues rather than elementary ones.A composite truth value contains many component truth values, representing truth values ofthe Atom in different contexts and according to different estimators.It is important to note that the CogPrime declarative knowledge representation is neithera neural net nor a semantic net, though it does have some commonalities with each of thesetraditional representations. It is not a neural net because it has no activation values, and involvesno attempts at low-level brain modeling. However, attention values are very loosely analogousto time-averages of neural net activations. On the other hand, it is not a semantic net becauseof the broad scope of the Atoms in the network: for example, Atoms may represent percepts,procedures, or parts of concepts. Most CogPrime Atoms have no corresponding English label.However, most CogPrimeAtoms do have probabilistic truth values, allowing logical semantics.256 13 Local, Global and Glocal Knowledge Representation13.4 Knowledge Representation via Attractor Neural NetworksNow we turn to global, implicit knowledge representation – beginning with formal neural netmodels, briefly discussing the brain, and then turning back to CogPrime. Firstly, this sectionreviews some relevant material from the literature regarding the representation of knowledgeusing attractor neural nets. It is a mix of well-established fact with more speculative material.13.4.1 The Hopfield neural net modelHopfield networks [Hop82] are attractor neural networks often used as associative memories. AHopfield network with N neurons can be trained to store a set of bipolar patterns P , whereeach pattern p has N bipolar (±1) values. A Hopfield net typically has symmetric weights withno self-connections. The weight of the connection between neurons i and j is denoted by w ij .In order to apply a Hopfield network to a given input pattern p, its activation state is set tothe input pattern, and neurons are updated asynchronously, in random order, until the networkconverges to the closest fixed point. An often-used activation function for a neuron is:∑y i = sign(p i w ij y j )Training a Hopfield network, therefore, involves finding a set of weights w ij that stores thetraining patterns as attractors of its network dynamics, allowing future recall of these patternsfrom possibly noisy inputs.Originally, Hopfield used a Hebbian rule to determine weights:w ij =j≠iP∑p i p jp=1Typically, Hopfield networks are fully connected. Experimental evidence, however, suggeststhat the majority of the connections can be removed without significantly impacting the network’scapacity or dynamics. Our experimental work uses sparse Hopfield networks.13.4.1.1 Palimpsest Hopfield nets with a modified learning ruleIn [SV99] a new learning rule is presented, which both increases the Hopfield network capacityand turns it into a “palimpsest”, i.e., a network that can continuously learn new patterns, whileforgetting old ones in an orderly fashion.Using this new training rule, weights are initially set to zero, and updated for each newpattern p to be learned according to:13.4 Knowledge Representation via Attractor Neural Networks 257N∑h ij = w ik p kk=1,k≠i,j∆w ij = 1 n (p ip j − h ij p j − h ji p i )13.4.2 Knowledge Representation via Cell AssembliesHopfield nets and their ilk play a dual role: as computational algorithms, and as conceptualmodels of brain function. In CogPrime they are used as inspiration for slightly different, artificialeconomics based computational algorithms; but their hypothesized relevance to brain functionis nevertheless of interest in a CogPrime context, as it gives some hints about the potentialconnection between low-level neural net mechanics and higher-level cognitive dynamics.Hopfield nets lead naturally to a hypothesis about neural knowledge representation, whichholds that a distinct mental concept is represented in the brain as either:1. a set of “cell assemblies”, where each assembly is a network of neurons that are interlinkedin such a way as to fire in a (perhaps nonlinearly) synchronized manner2. a distinct temporal activation pattern, which may occur in any one (or more) of a particularset of cell assembliesFor instance, this hypothesis is perfectly coherent if one interprets a “mental concept” as aSMEPH (defined in Chapter 14) ConceptNode, i.e. a fuzzy set of perceptual stimuli to whichthe organism systematically reacts in different ways. Also, although we will focus mainly ondeclarative knowledge here, we note that the same basic representational ideas can be appliedto procedural and episodic knowledge: these may be hypothesized to correspond to temporalactivation patterns as characterized above.In the biology literature, perhaps the best-articulated modern theories championing the cellassembly view are those of Gunther Palm [Pal82, HAG07] and Susan Greenfield [SF05, CSG07].Palm focuses on the dynamics of the formation and interaction assemblies of cortical columns.Greenfield argues that each concept has a core cell assembly, and that when the concept risesto the focus of attention, it recruits a number of other neurons beyond its core characteristicassembly into a “transient ensemble.” 1It’s worth noting that there may be multiple redundant assemblies representing the sameconcept – and potentially recruiting similar transient assemblies when highly activated. Theimportance of repeated, slightly varied copies of the same subnetwork has been emphasized byEdelman [Ede93] among other neural theorists.1 The larger an ensemble is, she suggests, the more vivid it is as a conscious experience; an hypothesis thataccords well with the hypothesis made in [Goe06b] that a more informationally intense pattern corresponds toa more intensely conscious quale – but we don’t need to digress extensively onto matters of consciousness forthe present purposes.258 13 Local, Global and Glocal Knowledge Representation13.5 Neural Foundations of LearningNow we move from knowledge representation to learning – which is after all nothing but theadaptation of represented knowledge based on stimulus, reinforcement and spontaneous activity.While our focus in this chapter is on representation, it’s not possible for us to make our pointsabout glocal knowledge representation in neural net type systems without discussing someaspects of learning in these systems.13.5.1 Hebbian LearningThe most common and plausible assumption about learning in the brain is that synaptic connectionsbetween neurons are adapted via some variant of Hebbian learning. The original Hebbianlearning rule, proposed by Donald Hebb in his 1949 book [Heb49], was roughly1. The weight of the synapse x → y increases if x and y fire at roughly the same time2. The weight of the synapse x → y decreases if x fires at a certain time but y does notOver the years since Hebb’s original proposal, many neurobiologists have sought evidence thatthe brain actually uses such a method. One of the things they have found, so far, is a lot ofevidence for the following learning rule [DC02, LS05]:1. The weight of the synapse x → y increases if x fires shortly before y does2. The weight of the synapse x → y decreases if x fires shortly after y doesThe new thing here, not foreseen by Donald Hebb, is the “postsynaptic depression” involved inrule component 2.Now, the simple rule stated above does not sum up all the research recently done on Hebbiantypelearning mechanisms in the brain. The real biological story underlying these approximaterules is quite complex, involving many particulars to do with various neurotransmitters. Illunderstooddetails aside, however, there is an increasing body of evidence that not only doesthis sort of learning occur in the brain, but it leads to distributed experience-based neuralmodification: that is, one instance synaptic modification causes another instance of synapticmodification, which causes another, and so forth 2 [Bi01].13.5.2 Virtual Synapses and Hebbian Learning Between AssembliesHebbian learning is conventionally formulated in terms of individual neurons, but, it can beextended naturally to assemblies via defining “virtual synapses” between assemblies.Since assemblies are sets of neurons, one can view a synapse as linking two assembliesif it links two neurons, each of which is in one of the assemblies. One can then viewtwo assemblies as being linked by a bundle of synapses. We can define the weight of thesynaptic bundle from assembly A1 to assembly A2 as the number w so that (the change2 This has been observed in “model systems” consisting of neurons extracted from a brain and hooked togetherin a laboratory setting and monitored; measurement of such dynamics in vivo is obviously more difficult.13.5 Neural Foundations of Learning 259in the mean activation of A2 that occurs at time t+epsilon) is on average closest to w ×(the amount of energy flowing through the bundle from A1 to A2 at time t). So when A1 sendsan amount x of energy along the synaptic bundle pointing from A1 to A2, then A2’s meanactivation is on average incremented/decremented by an amount w × x.In a similar way, one can define the weight of a bundle of synapses between a certain staticor temporal activation-pattern P1 in assembly A1, and another static or temporal activationpatternP2 in assembly A2. Namely, this may be defined as the number w so that (the amountof energy flowing through the bundle from A1 to A2 at time t)×w best approximates (theprobability that P2 is present in A2 at time t+epsilon), when averaged over all times t duringwhich P1 is present in A1.It is not hard to see that Hebbian learning on real synapses between neurons implies Hebbianlearning on these virtual synapses between cell assemblies and activation-patterns.These ideas may be developed further to build a connection between neural knowledge representationand probabilistic logical knowledge representation such as is used in CogPrime’sProbabilistic Logic Networks formalism; this connection will be pursued at the end of Chapter34, once more relevant background has been presented.13.5.3 Neural DarwinismA notion quite similar to Hebbian learning between assemblies has been pursued by NobelistGerald Edelman in his theory of neuronal group selection, or “Neural Darwinism.” Edelman wona Nobel Prize for his work in immunology, which, like most modern immunology, was based onC. MacFarlane Burnet’s theory of “clonal selection” [Bur62], which states that antibody typesin the mammalian immune system evolve by a form of natural selection. From his point of view,it was only natural to transfer the evolutionary idea from one mammalian body system (theimmune system) to another (the brain).The starting point of Neural Darwinism is the observation that neuronal dynamics may beanalyzed in terms of the behavior of neuronal groups. The strongest evidence in favor of thisconjecture is physiological: many of the neurons of the neocortex are organized in clusters, eachone containing say 10,000 to 50,000 neurons each. Once one has committed oneself to looking atsuch groups, the next step is to ask how these groups are organized, which leads to Edelman’sconcept of “maps.”A “map,” in Edelman’s terminology, is a connected set of groups with the property that whenone of the inter-group connections in the map is active, others will often tend to be active aswell. Maps are not fixed over the life of an organism. They may be formed and destroyed ina very simple way: the connection between two neuronal groups may be “strengthened” by increasingthe weights of the neurons connecting the one group with the other, and “weakened” bydecreasing the weights of the neurons connecting the two groups. If we replace “map” with “cellassembly” we arrive at a concept very similar to the one described in the previous subsection.Edelman then makes the following hypothesis: the large-scale dynamics of the brain is dominatedby the natural selection of maps. Those maps which are active when good results areobtained are strengthened, those maps which are active when bad results are obtained areweakened. And maps are continually mutated by the natural chaos of neural dynamics, thusproviding new fodder for the selection process. By use of computer simulations, Edelman and hiscolleagues have shown that formal neural networks obeying this rule can carry out fairly compli-260 13 Local, Global and Glocal Knowledge Representationcated acts of perception. In general-evolution language, what is posited here is that organismslike humans contain chemical signals that signify organism-level success of various types, andthat these signals serve as a “fitness function” correlating with evolutionary fitness of neuronalmaps.In Neural Darwinism and his other related books and papers, Edelman goes far beyond thiscrude sketch and presents neuronal group selection as a collection of precise biological hypotheses,and presents evidence in favor of a number of these hypotheses. However, we consider thatthe basic concept of neuronal group selection is largely independent of the biological particularitiesin terms of which Edelman has phrased it. We suspect that the mutation and selection of“transformations” or “maps” is a necessary component of the dynamics of any intelligent system.As we will see later on (e.g. in Chapter 42 of Part 2, this business of maps is extremelyimportant to CogPrime. CogPrime does not have simulated biological neurons and synapses,but it does have Nodes and Links that in some contexts play loosely similar roles. We sometimesthink of CogPrime Nodes and Links as being very roughly analogous to Edelman’s neuronalclusters, and emergent intercluster links. And we have maps among CogPrime Nodes and Links,just as Edelman has maps among his neuronal clusters. Maps are not the sole bearers of meaningin CogPrime, but they are significant ones.There is a very natural connection between Edelman-style brain evolution and the ideasabout cognitive evolution presented in Chapter 3. Edelman proposes a fairly clear mechanismvia which patterns that survive a while in the brain are differentially likely to survive a longtime: this is basic Hebbian learning, which in Edelman’s picture plays a role between neuronalgroups. And, less directly, Edelman’s perspective also provides a mechanism by which intensepatterns will be differentially selected in the brain: because on the level of neural maps, patternintensity corresponds to the combination of compactness and functionality. Among a numberof roughly equally useful maps serving the same function, the more compact one will be morelikely to survive over time, because it is less likely to be disrupted by other brain processes(such as other neural maps seeking to absorb its component neuronal groups into themselves).Edelman’s neuroscience remains speculative, since so much remains unknown about humanneural structure and dynamics; but it does provide a tentative and plausible connection betweenevolutionary neurodynamics and the more abstract sort of evolution that patternist philosophyposits to occur in the realm of mind-patterns.13.6 Glocal MemoryA glocal memory is one that transcends the global/local dichotomy and incorporates bothaspects in a tightly interconnected way. Here we make the glocal memory concept more precise,and describe its incarnation in the context of attractor neural nets (which is similar to itsincarnation in CogPrime, to be elaborated in later chapters). Though our main interest here isin glocality in CogPrime, we also suggest that glocality may be a critical property to considerwhen analyzing human, animal and AI memory more broadly.The notion of glocal memory has implicitly occurred in a number of prior brain theories(without use of the neologism “glocal”), e.g. [Cal96] and [Goe01], but it has not previously beenexplicitly developed. However the concept has risen to the fore in our recent AI work and sowe have chosen to flesh it out more fully in [HG08], [GPI + 10] and the present section.13.6 Glocal Memory 261Glocal memory overcomes the dichotomy between localized memory (in which each memoryitem is stored in a single location within an overall memory structure) and distributed memory(in which a memory item is stored as an aspect of a multi-component memory system, in sucha way that the same set of multiple components stores a large number of memories). In a glocalmemory system, most memory items are stored both locally and globally, with the propertythat eliciting either one of the two records of an item tends to also elicit the other one.Glocal memory applies to multiple forms of memory; however we will focus largely on perceptualand declarative memory in our detailed analyses here, so as to conserve space and maintainsimplicity of discussion.The central idea of glocal memory is that (perceptual, declarative, episodic, procedural,etc.) items may be stored in memory in the form of paired structures that are called (key,map) pairs. Of course the idea of a “pair” is abstract, and such pairs may manifest themselvesquite differently in different sorts of memory systems (e.g. brains versus non-neuromorphic AIsystems). The key is a localized version of the item, and records some significant aspects ofthe items in a simple and crisp way. The map is a dispersed, distributed version of the item,which represents the item as a (to some extent, dynamically shifting) combination of fragmentsof other items. The map includes the key as a subset; activation of the key generally (but notnecessarily always) causes activation of the map; and changes in the memory item will generallyinvolve complexly coordinated changes on the key and map level both.Memory is one area where animal brain architecture differs radically from the von Neumannarchitecture underlying nearly all contemporary general-purpose computers. Von Neumanncomputers separate memory from processing, whereas in the human brain there is no suchdistinction. In fact, it’s arguable that in most cases the brain contains no memory apart fromprocessing: human memories are generally constructed in the course of remembering [Ros88],which gives human memory a strong capability for “filling in gaps” of remembered experienceand knowledge; and also causes problems with inaccurate remembering in many contexts[BF71, RM95] We believe the constructive aspect of memory is largely associated with itsglocality.The remainder of this section presents a fuller formalization of the glocal memory concept,which is then taken up further in three later chapters:• Chapter ?? discusses the potential implementation of glocal memory in the human brain• Chapter ?? discusses the implementation of glocal memory in attractor neural net systems• Chapter 23 presents Glocal Economic Attention Networks (ECANs), rough analogues ofglocal Hopfield nets that play a central role in CogPrime.Our hypothesis of the potential general importance of glocality as a property of memorysystems (beyond just the CogPrime architecture) – remains somewhat speculative. The presenceof glocality in human and animal memory is strongly suggested but not firmly demonstrated byavailable neuroscience data; and the general value of glocality in the context of artificial brainsand minds is also not yet demonstrated as the whole field of artificial brain and mind buildingremains in its infancy. However, the utility of glocal memory for CogPrime is not tied to thismore general, speculative theme – glocality may be useful in CogPrime even if we’re wrong thatit plays a significant role in the brain and in intelligent systems more broadly.262 13 Local, Global and Glocal Knowledge Representation13.6.1 A Semi-Formal Model of Glocal MemoryTo explain the notion of glocal memory more precisely, we will introduce a simple semi-formalmodel of a system S that uses a memory to record information relevant to the actions itcarries out. The overall concept of glocal memory should not be considered as restricted tothis particular model. This model is not intended for maximal generality, but is intended toencompass a variety of current AI system designs and formal neurological models.In this model, we will consider S’s memory subsystem as a set of objects we’ll call “tokens,”embedded in some metric space. The metric in the space, which we will call the “basic distance”of the memory, generally will not be defined in terms of the semantics of the items stored in thememory; though it may come to shape these dynamics through the specific architecture andevolution of the memory. Note that these tokens are not intended as generally being mappedone-to-one onto meaningful items stored in the memory. The “tokens” are the raw materialsthat the memory arranges in various patterns in order to store items.We assume that each token, at each point in time, may meaningfully be assigned a certainquantitative “activation level.” Also, tokens may have other numerical or discrete quantitiesassociated with them, depending on the particular memory architecture. Finally, tokens mayrelate other tokens, so that optionally a token may come equipped with an (ordered or unordered)list of other tokens.To understand the meaning of the activation levels, one should think about S’s memorysubsystem as being coupled with an action-selection subsystem, that dynamically chooses theactions to be taken by the overall system in which the two subsystems are embedded. Eachcombination of actions, in each particular type of context, will generally be associated with theactivation of certain tokens in memory.Then, as analysts of the system S, we may associate each token T with an “activation vector”v(T, t), whose value for each discrete time t consists of the activation of the token T at time t.So, the 50 ′ th entry of the vector corresponds to the activation of the token at the 50 ′ th timestep.“Items stored in memory” over a certain period of time, may then be defined as clusters inthe set of activation vectors associated with memory during that period of time. Note that thesystem S itself may explicitly recognize and remember patterns regarding what items are storedin its memory – but, from an external analyst’s perspective, the set of items in S’s memory isnot restricted to the ones that S has explicitly recognized as memory items.The “localization” of a memory item may be defined as the degree to which the various tokensinvolved in the item are close to each other according to the metric in the memory metric-space.This degree may be formalized in various ways, but choosing a particular quantitative measureis not important here. A highly localized item may be called “local” and a not-very-localizeditem may be called “global.”We may define the “activation distance” of two tokens as the distance between their activationvectors. We may then say that a memory is “well aligned” to the extent that there is a correlationbetween the activation distance of tokens, and the basic distance of the memory metric-space.Given the above set-up, the basic notion of glocal memory can be enounced fairly simply. Aglocal memory is one:• that is reasonably well-aligned (i.e. the correlation between activation and basic distance issignificantly greater than random)13.6 Glocal Memory 263• in which most memory items come in pairs, consisting of one local item and one globalitem, so that activation of the local item (the “key”) frequently leads in the near future toactivation of the global item (the “map”)Obviously, in the scope of all possible memory structures constructible within the aboveformalism, glocal memories are going to be very rare and special. But, we suggest that they areimportant, because they are generally going to be the most effective way for intelligent systemsto structure their memories.Note also that many memories without glocal structure may be “well-aligned” in the abovesense.An example of a predominantly local memory structure, in which nearly all significant memoryitems are local according to the above definition, is the Cyc logical reasoning engine [LG90].To cast the Cyc knowledge base in the present formal model, the tokens are logical predicates.Cyc does not have an in-built notion of activation, but one may conceive the activation of alogical formula in Cyc as the degree to which the formula is used in reasoning or query processingduring a certain interval in time. And one may define a basic metric for Cyc by associatinga predicate with its extension (the set of satisfying inputs), and defining the similarity of twopredicates as the symmetric distance of their extensions. Cyc is reasonably well-aligned, butaccording to the dynamics of its querying and reasoning engines, it is basically a local memorystructure without significant global memory structure.On the other hand, an example of a predominantly global memory structure, in which nearlyall significant memory items are global according to the above definition, is the Hopfield associativememory network [Ami89]. Here memories are stored in the pattern of weights associatedwith synapses within a network of formal neurons, and each memory in general involves a largenumber of the neurons in the network. To cast the Hopfield net in the present formal model, thetokens are neurons and synapses; the activations are neural net activations; the basic distancebetween two neurons A and B may be defined as the percentage of the time that stimulatingone of the neurons leads to the other one firing; and to calculate a basic distance involving asynapse, one may associate the synapse with its source and target neurons. With these definitions,a Hopfield network is a well-aligned memory, and (by intentional construction) a markedlyglobal one. Local memory items will be very rare in a Hopfield net.While predominantly local and predominantly global memories may have great value for particularapplications, our suggestion is that they also have inherent limitations. If so, this meansthat the most useful memories for general intelligence are going to be those that involve bothlocal and global memory items in central roles. However, this is a more general and less riskyclaim than the assertion that glocal memory structure as defined above is important. Because,“glocal” as defined above doesn’t just mean “neither predominantly global nor predominantlylocal.” Rather, it refers to a specific pattern of coordination between local and global memoryitems – what we have called the “keys and maps” pattern.13.6.2 Glocal Memory in the BrainScience’s understanding of human brain dynamics is still very primitive, one manifestation ofwhich is the fact that we really don’t understand how the brain represents knowledge, exceptin some very simple respects. So anything anyone says about knowledge representation in thebrain, at this stage, has to be considered highly speculative. Existing neuroscience knowledge264 13 Local, Global and Glocal Knowledge Representationdoes imply constraints on how knowledge representation in the brain may work, but these arerelatively loose constraints. These constraints do imply that, for instance, the brain is neither arelational database (in which information is stored in a wholly localized manner) nor a collectionof “grandmother neurons” that respond individually to high-level percepts or concepts; nor asimple Hopfield type neural net (in which all memories are attractors globally distributed acrossthe whole network). But they don’t tell us nearly enough to, for instance, create a formal neuralnet model that can confidently be said to represent knowledge in the manner of the human brain.As a first example of the current state of knowledge, we’ll discuss here a series of papersregarding the neural representation of visual stimuli [QaGKKF05, QKKF08], which deal withthe fascinating discovery of a subset of neurons in the medial temporal lobe (MTL) that areselectively activated by strikingly different pictures of given individuals, landmarks or objects,and in some cases even by letter strings. For instance, in their 2005 paper titled ”Invariant visualrepresentation by single neurons in the human brain”, it is noted thatin one case, a unit responded only to three completely different images of the ex-president Bill Clinton.Another unit (from a different patient) responded only to images of The Beatles, another one to cartoonsfrom The Simpson’s television series and another one to pictures of the basketball player Michael Jordan.Their 2008 follow-up paper backed away from the more extreme interpretation in the title aswell as the conclusion, with the title “Sparse but not ‘Grandmother-cell’ coding in the medialtemporal lobe.” As the authors emphasize there,Given the very sparse and abstract representation of visual information by these neurons, they could inprinciple be considered as ‘grandmother cells’. However, we give several arguments that make such anextreme interpretation unlikely.. . .MTL neurons are situated at the juncture of transformation of percepts into constructs that can beconsciously recollected. These cells respond to percepts rather than to the detailed information fallingon the retina. Thus, their activity reflects the full transformation that visual information undergoesthrough the ventral pathway. A crucial aspect of this transformation is the complementary developmentof both selectivity and invariance. The evidence presented here, obtained from recordings of single-neuronactivity in humans, suggests that a subset of MTL neurons possesses a striking invariant representationfor consciously perceived objects, responding to abstract concepts rather than more basic metric details.This representation is sparse, in the sense that responsive neurons fire only to very few stimuli (and aremostly silent except for their preferred stimuli), but it is far from a Grandmother-cell representation.The fact that the MTL represents conscious abstract information in such a sparse and invariant way isconsistent with its prominent role in the consolidation of long-term semantic memories.It’s interesting to note how inadequate the [QKKF08] data really is for exploring the notionof glocal memory in the brain. Suppose it’s the case that individual visual memories correspondto keys consisting of small neuronal subnetworks, and maps consisting of larger neuronalsubnetworks. Then it would be not at all surprising if neurons in the “key” network correspondingto a visual concept like “Bill Clinton’s face” would be found to respond differentiallyto the presentation of appropriate images. Yet, it would also be wrong to overinterpret suchdata as implying that the key network somehow comprises the “representation” of Bill Clinton’sface in the individual’s brain. In fact this key network would comprise only one aspect of saidrepresentation.In the glocal memory hypothesis, a visual memory like “Bill Clinton’s face” would be hypothesizedto correspond to an attractor spanning a significant subnetwork of the individual’s brain13.6 Glocal Memory 265– but this subnetwork still might occupy only a small fraction of the neurons in the brain (say,1/100 or less), since there are very many neurons available. This attractor would constitute themap. But then, there would be a much smaller number of neurons serving as key to unlockthis map: i.e. if a few of these key neurons were stimulated, then the overall attractor patternin the map as a whole would unfold and come to play a significant role in the overall brainactivity landscape. In prior publications [Goe97] the primary author explored this hypothesisin more detail in terms of the known architecture of the cortex and the mathematics of complexdynamical attractors.So, one possible interpretation of the [QKKF08] data is that the MTL neurons they’remeasuring are part of key networks that correspond to broader map networks recording percepts.The map networks might then extend more broadly throughout the brain, beyond the MTLand into other perceptual and cognitive areas of cortex. Furthermore, in this case, if some MTLkey neurons were removed, the maps might well regenerate the missing keys (as would happene.g. in the glocal Hopfield model to be discussed in the following section).Related and interesting evidence for glocal memory in the brain comes from a recent study ofsemantic memory, illustrated in Figure ?? [PNR07]. Their research probed the architecture ofsemantic memory via comparing patients suffering from semantic dementia (SD) with patientssuffering from three other neuropathologies, and found reasonably convincing evidence for whatthey call a “distributed-plus-hub” view of memory.The SD patients they studied displayed highly distinctive symptomology; for instance, theirvocabularies and knowledge of the properties of everyday objects were strongly impaired,whereas their memories of recent events and other cognitive capacities remain perfectly intact.These patients also showed highly distinctive patterns of brain damage: focal brain lesionsin their anterior temporal lobes (ATL), unlike the other patients who had either less severe ormore widely distributed damage in their ATLs. This led [PNR07] to conclude that the ATL(being adjacent to the amygdala and limbic systems that process reward and emotion; and theanterior parts of the medial temporal lobe memory system, which processes episodic memory)is a “hub” for amodal semantic memory, drawing general semantic information from episodicmemories based on emotional salience.So, in this view, the memory of something like a “banana” would contain a distributed aspect,spanning multiple brain systems, and also a localized aspect, centralized in the ATL.The distributed aspect would likely contain information on various particular aspects of bananas,including their sights, smells, and touches, the emotions they evoke, and the goals andmotivations they relate to. The distributed and localized aspects would influence one anotherdynamically, but, the data [PNR07] gathered do not address dynamics and they don’t venturehypotheses in this direction.There is a relationship between the “distributed-plus-hub” view and [Dam00] better-knownnotion of a “convergence zone”, defined roughly as a location where the brain binds features together.A convergence zone, in [Dam00] perspective, is not a “store” of information but an agentcapable of decoding a signal (and of reconstructing information). He also uses the metaphorthat convergence zones behave like indexes drawing information from other areas of the brain –but they are dynamic rather than static indices, containing the instructions needed to recognizeand combine the features constituting the memory of something. The mechanism involved inthe distributed-plus-hub model is similar to a convergence zone, but with the important differencethat hubs are less local: [PNR07] semantic hub may be thought of a kind of “cluster ofconvergence zones” consisting of a network of convergence zones for various semantic memories.266 13 Local, Global and Glocal Knowledge RepresentationFig. 13.1: A Simplified Look at Feedback-Control in Uncertain InferenceWhat is missing in [PNR07] and [Dam00] perspective is a vision of distributed memoriesas attractors. The idea of localized memories serving as indices into distributed knowledgestores is important, but is only half the picture of glocal memory: the creative, constructive,dynamical-attractor aspect of the distributed representation is the other half. The closest thingto a clear depiction of this aspect of glocal memory that seems to exist in the neuroscienceliterature is a portion of William Calvin’s theory of the “cerebral code” [Cal96]. Calvin proposesa set of quite specific mechanisms by which knowledge may be represented in the brainusing complexly-structured strange attractors, and by which these strange attractors may bepropagated throughout the brain. Figure 13.2 shows one aspect of his theory: how a distributedattractor may propagate from one part of the brain to another in pieces, with one portion ofthe attractor getting propagated first, and then seeding the formation in the destination brainregion of a close approximation of the whole attractor.Calvin’s theory may be considered a genuinely glocal theory of memory. However, it alsomakes a large number of other specific commitments that are not part of the notion of glocality,such as his proposal of hexagonal meta-columns in the cortex, and his commitment toevolutionary learning as the primary driver of neural knowledge creation. We find these other13.6 Glocal Memory 267Fig. 13.2: Calvin’s Model of Distributed Attractors in the Brainhypotheses interesting and highly promising, yet feel it is also important to separate out thenotion of glocal memory for separate consideration.Regarding specifics, our suggestion is that Calvin’s approach may overemphasize the distributedaspect of memory, not giving sufficient due to the relatively localized aspect as accountedfor in the [QKKF08] results discussed above. In Calvin’s glocal approach, global memoriesare attractors and local memories are parts of attractors. We suggest a possible alternative,in which global memories are attractors and local memories are particular neuronal subnetworkssuch as the specialized ones identified by [QKKF08]. However, this alternative does not seemcontradictory to Calvin’s overall conceptual approach, even though it is different from the particularproposals made in [Cal96].The above paragraphs are far from a complete survey of the relevant neuroscience literature;there are literally dozens of studies one could survey pointing toward the glocality of varioussorts of human memory. Yet experimental neuroscience tools are still relatively primitive, andevery one of these studies could be interpreted in various other ways. In the next couple decades,as neuroscience tools improve in accuracy, our understanding of the role of glocality in humanmemory will doubtless improve tremendously.268 13 Local, Global and Glocal Knowledge Representation13.6.3 Glocal Hopfield NetworksThe ideas in the previous section suggest that, if one wishes to construct an AGI, it is worthseriously considering using a memory with some sort of glocal structure. One research directionthat follows naturally from this notion is “glocal neural networks.” In order to explore the natureof glocal neural networks in a relatively simple and tractable setting, we have formalized andimplemented simple examples of “glocal Hopfield networks”: palimpsest Hopfield nets with theaddition of neurons representing localized memories. While these specific networks are not usedin CogPrime, they are quite similar to the ECAN networks that are used in CogPrime anddescribed in Chapter 23 of Part 2.Essentially, we augment the standard Hopfield net architecture by adding a set of “keyneurons.” These are a small percentage of the neurons in the network, and are intended to beroughly equinumerous to the number of memories the network is supposed to store. When theHopfield net converges to an attractor A, then new links are created between the neurons thatare active in A, and one of the key neurons. Which key neuron is chosen? The one that, whenit is stimulated, gives rise to an attractor pattern maximally similar to A.The ultimate result of this is that, in addition to the distributed memory of attractors in theHopfield net, one has a set of key neurons that in effect index the attractors. Each attractorcorresponds to a single key neuron. In the glocal memory model, the key neurons are the keysand the Hopfield net attractors are the maps.This algorithm has been tested in sparse Hopfield nets, using both standard Hopfield netlearning rules and Storkey’s modified palimpsest learning rule [SV99], which provides greatermemory capacity in a continuous learning context. The use of key neurons turns out to slightlyincrease Hopfield net memory capacity, but this isn’t the main point. The main point is thatone now has a local representation of each global memory, so that if one wants to create alink between the memory and something else, it’s extremely easy to do so – one just needsto link to the corresponding key neuron. Or, rather, one of the corresponding key neurons:depending on how many key neurons are allocated, one might end up with a number of keyneurons corresponding to each memory, not just one.In order to transform a palimpsest Hopfield net into a glocal Hopfield net, the following stepsare taken:1. Add a fixed number of “key neurons” to the network (removing other random neurons tokeep the total number of neurons constant)2. When the network reaches an attractor, create links from the elements in the attractor toone of the key neurons3. The key neuron chosen for the previous step is the one that most closely matches the currentattractor (which may be determined in several ways, to be discussed below)4. To avoid the increase of the number of links in the network, when new links are created inStep 2, other key-neuron links are then deleted (several approaches may be taken here, butthe simplest is to remove the key-neuron links with the lowest-absolute-value weights)In the simple implementation of the above steps that we implemented, and described in[GPI + 10], Step 3 is carried out simply by comparing the weights of a key neuron’s links to thenodes in an attractor. A more sophisticated approach would be to select the key neuron withthe highest activation during the transient interval immediately prior to convergence to theattractor.13.6 Glocal Memory 269The result of these modifications to the ordinary Hopfield net, is a Hopfield net that continuallymaintains a set of key neurons, each of which individually represents a certain attractorof the net.Note that these key neurons – in spite of being “symbolic” in nature – are learned ratherthan preprogrammed, and are every bit as adaptive as the attractors they correspond to. Furthermore,if a key neuron is removed, the glocal Hopfield net algorithm will eventually learn itback, so the robustness properties of Hopfield nets are retained.The results of experimenting with glocal Hopfield nets of this nature are summarized in[GPI + 10]. We studied Hopfield nets with connectivity around .1, and in this context we foundthat glocality• slightly increased memory capacity• massively increased the rate of convergence to the attractor, i.e. the speed of recallHowever, probably the most important consequence of glocality is a more qualitative one: itmakes it far easier to link the Hopfield net into a larger system, as would occur if the Hopfield netwere embedded in an integrative AGI architecture. Because a neuron external to the Hopfieldnet may now link to a memory in the Hopfield net by linking to the corresponding key neuron.13.6.4 Neural-Symbolic Glocality in CogPrimeIn CogPrime, we have explicitly sought to span the symbolic/emergentist pseudo-dichotomy,via creating an integrative knowledge representation that combines logic-based aspects withneural-net-like aspects. As reviewed in Chapter 6 above, these function not in the manner ofmultimodular systems, but rather via using (probabilistic) truth values and (attractor neuralnet like) attention values as weights on nodes and links of the same (hyper) graph. The nodesand links in this hypergraph are typed, like a standard semantic network approach for knowledgerepresentation, so they’re able to handle all sorts of knowledge, from the most concreteperception and actuation related knowledge to the most abstract relationships. But they’re alsoweighted with values similar to neural net weights, and pass around quantities (importancevalues, discussed in Chapter 23 of Part 2) similar to neural net activations, allowing emergentattractor/assembly based knowledge representation similar to attractor neural nets.The concept of glocality lies at the heart of this combination, in a way that spans the pseudodichotomy:• Local knowledge is represented in abstract logical relationships stored in explicit logicalform, and also in Hebbian-type associations between nodes and links.• Global knowledge is represented in large-scale patterns of node and link weights, whichlead to large-scale patterns of network activity, which often take the form of attractorsqualitatively similar to Hopfield net attractors. These attractors are called maps.The result of all this is that a concept like “cat” might be represented as a combination of:• A small number of logical relationships and strong associations, that constitute the “key”subnetwork for the “cat” concept.• A large network of weak associations, binding together various nodes and links of varioustypes and various levels of abstraction, representing the “cat map”.270 13 Local, Global and Glocal Knowledge RepresentationThe activation of the key will generally cause the activation of the map, and the activation ofa significant percentage of the map will cause the activation of the rest of the map, including thekey. Furthermore, if the key were for some reason forgotten, then after a significant amount ofeffort, the system would likely to be able to reconstitute it (perhaps with various small changes)from the information in the map. We conjecture that this particular kind of glocal memory willturn out to be very powerful for AGI, due to its ability to combine the strengths of formallogical inference with those of self-organizing attractor neural networks.As a simple example, consider the representation of a “tower”, in the context of an artificialagent that has built towers of blocks, and seen pictures of many other kinds of towers, and seensome tall building that it knows are somewhat like towers but perhaps not exactly towers. Ifthis agent is reasonably conceptually advanced (say, at Piagetan the concrete operational level)then its mind will contain some declarative relationships partially characterizing the concept of“tower,” as well as its sensory and episodic examples, and its procedural knowledge about howto build towers.The key of the “tower” concept in the agent’s mind may consist of internal images andepisodes regarding the towers it knows best, the essential operations it knows are useful forbuilding towers (piling blocks atop blocks atop blocks...), and the core declarative relationssummarizing “towerness” – and the whole “tower” map then consists of a much larger numberof images, episodes, procedures and declarative relationships connected to “tower” and otherrelated entities. If any portion of the map is removed – even if the key is removed – then therest of the map can be approximately reconstituted, after some work. Some cognitive operationsare best done on the localized representation – e.g. logical reasoning. Other operations, such asattention allocation and guidance of inference control, are best done using the globalized maprepresentation.Chapter 14Representing Implicit Knowledge via Hypergraphs14.1 IntroductionExplicit knowledge is easy to write about and talk about; implicit knowledge is equally important,but tends to get less attention in discussions of AI and psychology, simply because we don’thave as good a vocabulary for describing it, nor as good a collection of methods for measuringit. One way to deal with this problem is to describe implicit knowledge using language andmethods typically reserved for explicit knowledge. This might seem intrinsically non-workable,but we argue that it actually makes a lot of sense. The same sort of networks that a system likeCogPrime uses to represent knowledge explicitly, can also be used to represent the emergentknowledge that implicitly exists in an intelligent system’s complex structures and dynamics.We’ve noted that CogPrime uses an explicit representation of knowledge in terms of weightedlabeled hypergraphs; and also uses other more neural net like mechanisms (e.g. the economicattention allocation network subsystem) to represent knowledge globally and implicitly. Cog-Prime combines these two sorts of representation according to the principle we have calledglocality. In this chapter we pursue glocality a bit further – describing a means by which evenimplicitly represented knowledge can be modeled using weighted labeled hypergraphs similar tothe ones used explicitly in CogPrime. This is conceptually important, in terms of making clearthe fundamental similarities and differences between implicit and explicit knowledge representation;and it is also pragmatically meaningful due to its relevance to the CogPrime methodsdescribed in Chapter 42 of Part 2 that transform implicit into explicit knowledge.To avoid confusion with CogPrime’s explicit knowledge representation, we will refer to thehypergraphs in this chapter as composed of Vertices and Edges rather than Nodes and Links. Inprior publications we have referred to "derived" or "emergent" hypergraphs of the sort describedhere using the acronym SMEPH, which stands for Self-Modifying, Evolving Probabilistic Hypergraphs.14.2 Key Vertex and Edge TypesWe begin by introducing a particular collection of Vertex and Edge types, to be used in modelingthe internal structures of intelligent systems.The key SMEPH Vertex types are271272 14 Representing Implicit Knowledge via Hypergraphs• ConceptVertex, representing a set, for instance, an idea or a set of percepts• SchemaVertex, representing a procedure for doing something (perhaps something in thephysical world, or perhaps an abstract mental action).The key SMEPH Edge types, using language drawn from Probabilistic Logic Networks (PLN)and elaborated in Chapter 34 below, are as follows:• ExtensionalInheritanceEdge (ExtInhEdge for short: an edge which, linking one Vertex orEdge to another, indicates that the former is a special case of the latter)• ExtensionalSimilarityEdge (ExtSim: which indicates that one Vertex or Edge is similar toanother)• ExecutionEdge (a ternary edge, which joins S,B,C when S is a SchemaVertex and the resultfrom applying S to B is C).So, in a SMEPH system, one is often looking at hypergraphs whose Vertices represent ideas orprocedures, and whose Edges represent relationships of specialization, similarity or transformationamong ideas and/or procedures.The semantics of the SMEPH edge types is given by PLN, but is simple and commonsensical.ExtInh and ExtSim Edges come with probabilistic weights indicating the extent ofthe relationship they denote (e.g. the ExtSimEdge joining the cat ConceptVertex to the dogConceptVertex gets a higher probability weight than the one joining the cat ConceptVertexto the washing-machine ConceptVertex). The mathematics of transformations involving theseprobabilistic weights becomes quite involved - particularly when one introduces SchemaVerticescorresponding to abstract mathematical operations, a step that enables SMEPH hypergraphsto have the complete mathematical power of standard logical formalisms like predicate calculus,but with the added advantage of a natural representation of uncertainty in terms ofprobabilities, as well as a natural representation of networks and webs of complex knowledge.14.3 Derived HypergraphsWe now describe how SMEPH hypergraphs may be used to model and describe intelligentsystems. One can (in principle) draw a SMEPH hypergraph corresponding to any individualintelligent system, with Vertices and Edges for the concepts and processes in that system’smind. This is called the derived hypergraph of that system.14.3.1 SMEPH VerticesA ConceptVertex in the derived hypergraph of a system corresponds to a structural patternthat persists over time in that system; whereas a SchemaVertex corresponds to a multi-timepointdynamical pattern that recurs in that system’s dynamics. If one accepts the patternistdefinition of a mind as the set of patterns in an intelligent system, then it follows that thederived hypergraph of an intelligent system captures a significant fraction of the mind of thatsystem.To phrase it a little differently, we may say that a ConceptVertex, in SMEPH, refers to thehabitual pattern of activity observed in a system when some condition is met (this condition14.3 Derived Hypergraphs 273corresponding to the presence of a certain pattern). The condition may refer to something inthe world external to the system, or to something internal. For instance, the condition may beobserving a cat. In this case, the corresponding Concept vertex in the mind of Ben Goertzelis the pattern of activity observed in Ben Goertzel’s brain when his eyes are open and he’slooking in the direction of a cat. The notion of pattern of activity can be made rigorous usingmathematical pattern theory, as is described in The Hidden Pattern [Goe06a].Note that logical predicates, on the SMEPH level, appear as particular kinds of Concepts,where the condition involves a predicate and an argument. For instance, suppose one wants toknow what happens inside Ben’s mind when he eats cheese. Then there is a Concept correspondingto the condition of cheese-eating activity. But there may also be a Concept correspondingto eating activity in general. If the Concept denoting the activity of eating X is generally easilycomputable from the Concepts for X and eating individually, then the eating Concept iseffectively acting as a predicate.A SMEPH SchemaVertex, on the other hand, is like a Concept that’s defined in a timedependentway. One type of Schema refers to a habitual dynamical pattern of activity occurringbefore and/or during some condition is met. For instance, the condition might be saying theword Hello. In that case the corresponding SchemaVertex in the mind of Ben Goertzel is thepattern of activity that generally occurs before he says Hello.Another type of Schema refers to a habitual dynamical pattern of activity occurring aftersome condition X is met. For instance, in the case of the Schema for adding two numbers, theprecondition X consists of the two numbers and the concept of addition. The Schema is thenwhat happens when the mind thinks of adding and thinks of two numbers.Finally, there are Schema that refer to habitual dynamical activity patterns occurring aftersome condition X is met and before some condition Y is met. In this case the Schema is viewedas transforming X into Y. For instance, if X is the condition of meeting someone who is not afriend, and Y is the condition of being friends with that person, then the habitually interveningactivities constitute the Schema for making friends.14.3.2 SMEPH EdgesSMEPH edge types fall into two categories: functional and logical. Functional edges connectSchema vertices to their input and outputs; logical edges refer mainly to conditional probabilities,and in general are to be interpreted according to the semantics of Probabilistic LogicNetworks.Let us begin with logical edges. The simplest case is the Subset edge, which denotes astraightforward, extensional conditional probability. For instance, it may happen that wheneverthe Concept for cat is present in a system, the Concept for animal is as well. Then we wouldsaySubset cat animal(Here we assume a notation where “R A B” denotes an Edge of type R between Vertices A andB.)On the other hand, it may be that 50% of the time that cat is present in the system, cute ispresent as well: then we would saySubset cat cute <.5>274 14 Representing Implicit Knowledge via Hypergraphswhere the <.5> denotes the probability, which is a component of the Truth Value associatedwith the edge.Next, the most basic functional edge is the Execution edge, which is ternary and denotes arelation between a Schema, its input and its output, e.g.Execution father_of Ben_Goertzel Ted_Goertzelfor a schema father_of that outputs the father of its argument.The ExecutionOutput (ExOut) edge denotes the output of a Schema in an implicit way, e.g.ExOut say_hellorefers to a particular act of saying hello, whereasExOut add_numbers {3, 4)refers to the Concept corresponding to 7. Note that this latter example involves a set of threeentities: sets are also part of the basic SMEPH knowledge representation. A set may be thoughtof as a hypergraph edge that points to all its members.In this manner we may define a set of edges and vertices modeling the habitual activitypatterns of a system when in different situations. This is called the derived hypergraph of thesystem. Note that this hypergraph can in principle be constructed no matter what happensinside the system: whether it’s a human brain, a formal neural network, Cyc, OCP, a quantumcomputer, etc. Of course, constructing the hypergraph in practice is quite a different story: forinstance, we currently have no accurate way of measuring the habitual activity patterns insidethe human brain. fMRI and PET and other neuroimaging technologies give only a crude view,though they are continually improving.Pattern theory enters more deeply here when one thoroughly fleshes out the Inheritanceconcept. Philosophers of logic have extensively debated the relationship between extensionalinheritance (inheritance between sets based on their members) and intensional inheritance (inheritancebetween entity-types based on their properties). A variety of formal mechanisms havebeen proposed to capture this conceptual distinction; see (Wang, 2006, 1995 TODO make ref)for a review along with a novel approach utilizing uncertain term logic. Pattern theory providesa novel approach to defining intension: one may associate with each ConceptVertex in a system’sderived hypergraph the set of patterns associated with the structural pattern underlying thatConceptVertex. Then, one can define the strength of the IntensionalInheritanceEdge betweentwo ConceptVertices A and B as the percentage of A’s pattern-set that is also contained in B’spattern-set. According to this approach, for instance, one could haveIntInhEdge whale fish <0.6>ExtInhEdge whale fish <0.0>since the fish and whale sets have common properties but no common members.14.4 Implications of Patternist Philosophy for Derived Hypergraphsof Intelligent SystemsPatternist philosophy rears its head here and makes some definite hypotheses about the structureof derived hypergraphs. It suggests that derived hypergraphs should have a dual network14.4 Implications of Patternist Philosophy for Derived Hypergraphs of Intelligent Systems 275structure, and that in highly intelligent systems they should have subgraphs that constitutemodels of the whole hypergraph (these are self systems). SMEPH does not add anything tothe patternist view on a philosophical level, but it gives a concrete instantiation to some of thegeneral ideas of patternism. In this section we’ll articulate some "SMEPH principles", constitutingimportant ideas from patternist philosophy as they manifest themselves in the SMEPHcontext.The logical edges in a SMEPH hypergraph are weighted with probabilities, as in the simpleexample given above. The functional edges may be probabilistically weighted as well, since someSchema may give certain results only some of the time. These probabilities are critical in termsof SMEPH’s model of system dynamics; they underly one of our SMEPH principles,Principle of Implicit Probabilistic Inference: In an intelligent system, the temporalevolution of the probabilities on the edges in the system’s derived hypergraph should approximatelyobey the rules of probability theory.The basic idea is that, even if a system - through its underlying dynamics - has no explicitconnection to probability theory, it still must behave roughly as if it does, if it is going to beintelligent. The roughly part is important here; it’s well known that humans are not terriblyaccurate in explicitly carrying out formal probabilistic inferences. And yet, in practical contextswhere they have experience, humans can make quite accurate judgments; which is all that’srequired by the above principle, since it’s the contexts where experience has occurred that willmake up a system’s derived hypergraph.Our next SMEPH principle is evolutionary, and statesPrinciple of Implicit Evolution: In an intelligent system, new Schema and Concepts willcontinually be created, and the Schema and Concepts that are more useful for achieving systemgoals (as demonstrated via probabilistic implication of goal achievement) will tend to survivelonger.Note that this principle can be fulfilled in many different ways. The important thing is thatsystem goals are allowed to serve as a selective force.Another SMEPH dynamical principle pertains to a shorter time-scale than evolution, andstatesPrinciple of Attention Allocation: In an intelligent system, Schema and Concepts thatare more useful for attaining short-term goals will tend to consume more of the system’s energy.(The balance of attention oriented toward goals pertaining to different time scales will vary fromsystem to system.)Next, there is thePrinciple of Autopoesis: In an intelligent system, if one removes some part of the systemand then allows the system’s natural dynamics to keep going, a decent approximation to thatremoved part will often be spontaneously reconstituted.And there is the276 14 Representing Implicit Knowledge via HypergraphsCognitive Equation Principle: In an intelligent system, many abstract patterns that arepresent in the system at a certain time as patterns among other Schema and Concepts, will ata near-future time be present in the system as patterns among elementary system components.The Cognitive Equation Principle, briefly discussed in Chapter 3, basically means that Conceptsand Schema emergent in the system are recognized by the system and then embodiedas elementary items in the system so that patterns among them in their emergent form become,with the passage of time, patterns among them in their directly-system-embodied form.This is a natural consequence of the way intelligent systems continually recognize patterns inthemselves.Note that derived hypergraphs may be constructed corresponding to any complex systemwhich demonstrates a variety of internal dynamical patterns depending on its situation. However,if a system is not intelligent, then according to the patternist philosophy evolution of itsderived hypergraph can’t necessarily be expected to follow the above principles.14.4.1 SMEPH Principles in CogPrimeWe now more explicitly elaborate the application of these ideas in the CogPrime context. Asnoted above, in addition to explicit knowledge representation in terms of Nodes and Links,CogPrime also incorporates implicit knowledge representation in the form of what are calledMaps: collections of Nodes and Links that tend to be utilized together within cognitive processes.These Maps constitute a CogPrime system’s derived hypergraph, which will not be identicalto the hypergraph it uses for explicit knowledge representation. However, an interestingfeedback loop arises here, in that the intelligence’s self-study will generally lead it to recognizelarge portions of its derived hypergraph as patterns in itself, and then embody these patternswithin its concretely implemented knowledge hypergraph. This relates to the Cognitive EquationPrinciple defined above 3, in which an intelligent system continually recognizes patterns initself and embodies these patterns in its own basic structure (so that new patterns may moreeasily emerge from them).Often it happens that a particular CogPrime node will serve as the center of a map, so thate.g. the Concept Link denoting cat will consist of a number of nodes and links roughly centeredaround a ConceptNode that is linked to the WordNode cat. But this is not guaranteed andsome CogPrime maps are more diffuse than this with no particular center.Somewhat similarly, the key SMEPH dynamics are represented explicitly in CogPrime: probabilisticreasoning is carried out via explicit application of PLN on the CogPrime hypergraph,evolutionary learning is carried out via application of the MOSES optimization algorithm, andattention allocation is carried out via a combination of inference and evolutionary pattern mining.But the SMEPH dynamics also occur implicitly in CogPrime: emergent maps are reasonedon probabilistically as an indirect consequence of node-and-link level PLN activity; maps evolveas a consequence of the coordinated whole of CogPrime dynamics; and attention shifts betweenmaps according to complex emergent dynamics.To see the need for maps, consider that even a Node that has a particular meaning attachedto it - like the Iraq Node, say - doesn’t contain much of the meaning of Iraq in it. The meaningof Iraq lies in the Links attached to this Node, and the Links attached to their Nodes - andthe other Nodes and Links not explicitly represented in the system, which will be created by14.4 Implications of Patternist Philosophy for Derived Hypergraphs of Intelligent Systems 277CogPrime’s cognitive algorithms based on the explicitly existent Nodes and Links related tothe Iraq Node.This halo of Atoms related to the Iraq node is called the Iraq map. In general, some mapswill center around a particular Atom, like this Iraq map, others may not have any particularidentifiable center. CogPrime’s cognitive processes act directly on the level of Nodes and Links,but they must be analyzed in terms of their impact on maps as well. In SMEPH terms, Cog-Prime maps may be said to correspond to SMEPH ConceptNodes, and for instance bundles ofLinks between the Nodes belonging to a map may correspond to a SMEPH Link between twoConceptNodes.
Chapter 15Emergent Networks of Intelligence15.1 IntroductionWhen one is involved with engineering an AGI system, one thinks a lot about the aspects ofthe system one is explicitly building – what are the parts, how they fit together, how to testthey’re properly working, and so forth. And yet, these explicitly engineered aspects are only afraction of what’s important in an AGI system. At least as critical are the emergent aspects –the patterns that emerge once the system is up and running, interacting with the world andother agents, growing and developing and learning and self-modifying. SMEPH is one toolkitfor describing some of these emergent patterns, but it’s only a start.In line with these general observations, most of this book will focus on the structures andprocesses that we have built, or intend to build, into the CogPrime system. But in a sense, thesestructures and processes are not the crux of CogPrime’s intended intelligence. The purposeof these pre-programmed structures and processes is to give rise to emergent structures andprocesses, in the course of CogPrime’s interaction with the world and the other minds withinit. We will return to this theme of emergence at several points in later chapters, e.g. in thediscussion of map formation in Chapter 42 of Part 2.Given the important of emergent structures – and specifically emergent network structures –for intelligence, it’s fortunate the scientific community has already generated a lot of knowledgeabout complex networks: both networks of physical or software elements, and networks oforganization emergent from complex systems. As most of this knowledge has originated infields other than AGI, or in pure mathematics, it tends to require some reinterpretation ortweaking to achieve maximal applicability in the AGI context; but we believe this effort willbecome increasingly worthwhile as the AGI field progresses, because network theory is likelyto be very useful for describing the contents and interactions of AGI systems as they developincreasing intelligence.In this brief chapter we specifically focus on the emergence of certain large-scale networkstructures in a CogPrime knowledge store, presenting heuristic arguments as to why thesestructures can be expected to arise. We also comment on the way in which these emergentstructures are expected to guide cognitive processes, and give rise to emergent cognitive processes.The following chapter expands on this theme in a particular direction, exploring thepossible emergence of structures characterizing inter-cognitive reflection.279280 15 Emergent Networks of Intelligence15.2 Small World NetworksOne simple but potentially useful observation about CogPrime Atomspaces is that they aregenerally going to be small world networks [Buc03], rather than random graphs. A small worldnetwork is a graph in which the connectivities of the various nodes display a power law behavior– so that, loosely speaking, there are a few nodes with very many links, then more nodes with amodest number of links ... and finally, a huge number of nodes with very few links. This kind ofnetwork occurs in many natural and human systems, including citations among papers, financialarrangements among banks, links between Web pages and the spread of diseases among peopleor animals. In a weighted network like an Atomspace, "small-world-ness" must be defined in amanner taking the weights into account, and there are several obvious ways to do this. Figure15.1 depicts a small but prototypical small-worlds network, with a few "hub" nodes possessingfar more neighbors than the others, and then some secondary hubs, etc.An excellent reference on network theory in general, including but not limited to small worldnetworks, is Peter Csermely’s Weak Links [Cse06]. Many of the ideas in that work have apparentOpenCog applications, which are not elaborated here.Fig. 15.1: A typical, though small-sized, small-worlds network.One process via which small world networks commonly form is "preferential attachment"[Bar02]. This occurs in essence when "the rich get richer" – i.e. when nodes in the networkgrow new links, in a manner that causes them to preferentially grow links to nodes that alreadyhave more links. It is not hard to see that CogPrime’s ECAN dynamics will naturally lead to15.3 Dual Network Structure 281preferential attachment, because Atoms with more links will tend to get more STI, and thuswill tend to get selected by more cognitive processes, which will cause them to grow morelinks. For this reason, in most circumstances, a CogPrime system in which most link-buildingcognitive processes rely heavily on ECAN to guide their activities will tend to contain a smallworld-networkAtomspace. This is not rigorously guaranteed to be the case for any possiblecombination of environment and goals, but it is commonsensically likely to nearly always bethe case.One consequence of the small worlds structure of the Atomspace is that, in exploring otherproperties of the Atom network, it is particularly important to look at the hub nodes. Forinstance, if one is studying whether hierarchical and heterarchical subnetworks of the Atomspaceexist, and whether they are well-aligned with each other, it is important to look at hierarchicaland heterarchical connections between hub nodes in particular (and secondary hubs, etc.). Apattern of hierarchical or dual network connection that only held up among the more sparselyconnected nodes in a small-world network would be a strange thing, and perhaps not thatcognitively useful.15.3 Dual Network StructureOne of the key theoretical notions in patternist philosophy is that complex cognitive systemsevolve internal dual network structures, comprising superposed, harmonized hierarchical andheterarchical networks. Now we explore some of the specific CogPrime structures and dynamicsmilitating in favor of the emergence of dual networks.15.3.1 Hierarchical NetworksThe hierarchical nature of human linguistic concepts is well known, and is illustrated in Figure15.2 for the commonsense knowledge domain (using a graph drawn from WordNet, a huge concepthierarchy covering 50K+ English-language concepts), and in Figure 15.4 for a specializedknowledge subdomain, genetics. Due to this fact, a certain amount of hierarchy can be expectedto emerge in the Atomspace of any linguistically savvy CogPrime, simply due to its modelingof the linguistic concepts that it hears and reads.Hierarchy also exists in the natural world apart from language, which is the reason that manysensorimotor-knowledge-focused AGI systems (e.g. DeSTIN and HTM, mentioned in Chapter4 above) feature hierarchical structures. In these cases the hierarchies are normally spatiotemporalin nature - with lower layers containing elements responding to more localized aspectsof the perceptual field, and smaller, more localized groups of actuators. This kind of hierarchycertainly could emerge in an AGI system, but in CogPrime we have opted for a different route.If a CogPrime system is hybridized with a hierarchical sensorimotor network like one of thosementioned above, then the Atoms linked to the nodes in the hierarchical sensorimotor networkwill naturally possess hierarchical conceptual relationships, and will thus naturally grow hierarchicallinks between them (e.g. InheritanceLinks and IntensionalInheritanceLinks via PLN,AsymmetricHebbianLinks via ECAN).282 15 Emergent Networks of IntelligenceFig. 15.2: A typical, though small, subnetwork of WordNet’s hierarchical network.Once elements of hierarchical structure exist via the hierarchical structure of language andphysical reality, then a richer and broader hierarchy can be expected to accumulate on topof it, because importance spreading and inference control will implicitly and automatically beguided by the existing hierarchy. That is, in the language of Chaotic Logic [Goe94] and patternisttheory, hierarchical structure is an "autopoietic attractor" – once it’s there it will tend to enrichitself and maintain itself. AsymmetricHebbianLinks arranged in a hierarchy will tend to causeimportance to spread up or down the hierarchy, which will lead other cognitive processes to lookfor patterns between Atoms and their hierarchical parents or children, thus potentially buildingmore hierarchical links. Chains of InheritanceLinks pointing up and down the hierarchy will leadPLN to search for more hierarchical links – e.g. most simply, A → B → C where C is aboveB is above A in the hierarchy, will naturally lead inference to check the viability of A → Cby deduction. There is also the possibility to introduce a special DefaultInheritanceLink, asdiscussed in Chapter 34 of Part 2, but this isn’t actually necessary to obtain the inferentialmaintenance of a robust hierarchical network.15.3.2 Associative, Heterarchical NetworksHeterarchy is in essence a simpler structure than hierarchy: it simply refers to a network inwhich nodes are linked to other nodes with which they share important relationships. That is,there should be a tendency that if two nodes are often important in the same contexts or for15.3 Dual Network Structure 283Fig. 15.3: A typical, though small, subnetwork of the Gene Ontology’s hierarchical network.the same purposes, they should be linked together. Portrayals of typical heterarchical linkagepatterns among natural language concepts are given in Figures 15.5 and 15.6. Just for fun,Figure 15.7 shows one person’s attempt to draw a heterarchical graph of the main conceptsin one of Douglas Hofstadter’s books. Naturally, real concept heterarchies are far more large,complex and tangled than even this one.In CogPrime, ECAN enforces heterarchy via building SymmetricHebbianLinks, and PLNby building SimilarityLinks, IntensionalSimilarityLinks and ExtensionalSimilarityLinks. Furthermore,these various link types reinforce each other. PLN control is guided by importancespreading, which follows Hebbian links, so that a heterarchical Hebbian network tends to causePLN to explore the formation of links following the same paths as the heterarchical Hebbian-Links. And importance can spread along logical links as well as explicit Hebbian links, so thatthe existence of a heterarchical logical network will tend to cause the formation of additionalheterarchical Hebbian links. Heterarchy reinforces itself in "autopoietic attractor" style evenmore simply and directly than heterarchy.284 15 Emergent Networks of IntelligenceFig. 15.4: Small-scale portrayal of a portion of the spatiotemporal hierarchy in Jeff Hawkins’Hierarchical Temporal Memory architecture.15.3.3 Dual NetworksFinally, if both hierarchical and heterarchical structures exist in an Atomspace, then both ECANand PLN will naturally blend them together, because hierarchical and heterarchical links willfeed into their link-creation processes and naturally be combined together to form new links.This will tend to produce a structure called a dual network, in which a hierarchy exists, alongwith a rich network of heterarchical links joining nodes in the hierarchy, with a particular densityof links between nodes on the same hierarchical level. The dual network structure will emergewithout any explicit engineering oriented toward it, simply via the existence of hierarchicaland heterarchical networks, and the propensity of ECAN and PLN to be guided by both thehierarchical and heterarchical networks. The existence of a natural dual network structure inboth linguistic and sensorimotor data will help the formation process along, and then creativecognition will enrich the dual network yet further than is directly necessitated by the externalworld.15.3 Dual Network Structure 285Fig. 15.5: Portions of a conceptual heterarchy centered on specific concepts.Fig. 15.6: A portion of a conceptual heterarchy, showing the "dangling links" leading this portionto the rest of the heterarchy.A rigorous mathematical analysis of the formation of hierarchical, heterarchical and dualnetworks in CogPrime systems has not yet been undertaken, and would certainly be an interestingenterprise. Similar to the theory of small world networks, there is ample ground herefor both theorem-proving and heuristic experimentation. However, the qualitative points madehere are sufficiently well-grounded in intuition and experience to be of some use guiding our286 15 Emergent Networks of IntelligenceFig. 15.7: A fanciful evocation of part of a reader’s conceptual heterarchy related to DouglasHofstadter’s writings.ongoing work. One of the nice things about emergent network structures is that they are relativelystraightforward to observe in an evolving, learning AGI system, via visualization andinspection of structures such at the Atomspace.Section VA Path to Human-Level AGI
Chapter 16AGI PreschoolCo-authored with Stephan Vladimir Bugaj16.1 IntroductionIn conversations with government funding sources or narrow AI researchers about AGI work, oneof the topics that comes up most often is that of “evaluation and metrics” – i.e., AGI intelligencetesting. We actually prefer to separate this into two topics: environments and methods for carefulqualitative evaluation of AGI systems, versus metrics for precise measurement of AGI systems.The difficulty of formulating bulletproof metrics for partial progress toward advanced AGIhas become evident throughout the field, and in Chapter 8 we have elaborated one plausibleexplanation for this phenomenon, the "trickiness" of cognitive synergy. [LWML09], summarizinga workshop on “Evaluation and Metrics for Human-Level AI” held in 2008, discusses some ofthe general difficulties involved in this type of assessment, and some requirements that anyviable approach must fulfill. On the other hand, the lack of appropriate methods for carefulqualitative evaluation of AGI systems has been much less discussed, but we consider it actuallya more important issue – as well as an easier (though not easy) one to solve.We haven’t actually found the lack of quantitative intelligence metrics to be a major obstaclein our practical AGI work so far. Our OpenCogPrime implementation lags far behind theCogPrime design as articulated in Part 2 of this book, and according to the theory underlyingCogPrime, the more interesting behaviors and dynamics of the system will occur only when allthe parts of the system have been engineered to a reasonable level of completion and integratedtogether. So, the lack of a great set of metrics for evaluating the intelligence of our partiallybuiltsystem hasn’t impaired too much. Testing the intelligence of the current OpenCogPrimesystem is a bit like testing the flight capability of a partly-built airplane that only has stubsfor wings, lacks tail-fins, has a much less efficient engine than the one that’s been designed foruse in the first "real" version of the airplane, etc. There may be something to be learned fromsuch preliminary tests, but making them highly rigorous isn’t a great use of effort, comparedto working on finishing implementing the design according to the underlying theory.On the other hand, the problem of what environments and methods to use to qualitativelyevaluate and study AGI progress, has been considerably more vexing to us in practice, aswe’ve proceeded in our work on implementing and testing OpenCogPrime and developing theCogPrime theory. When developing a complex system, it’s nearly always valuable to see whatthis system does in some fairly rich, complex situations, in order to gain a better intuitiveunderstanding of the parts and how they work together. In the context of human-level AGI, thetheoretically best way to do this would be to embody one’s AGI system in a humanlike body289290 16 AGI Preschooland set it loose in the everyday human world; but of course, this isn’t feasible given the currentstate of development of robotics technology. So one must seek approximations. Toward this endwe have embodied OpenCogPrime in non-player characters in video game style virtual worlds,and carried out preliminary experiments embodying OpenCogPrime in humanoid robots. Theseare reasonably good options but they have limitations and lead to subtle choices: what kind ofgame characters and game worlds, what kind of robot environments, etc.?One conclusion we have come to, based largely on the considerations in Chapter 11 ondevelopment and Chapter 9 on the importance of environment, is that it may make sense toembed early-stage proto-AGI and AGI systems in environments reminiscent of those used forteaching young human children. In this chapter we will explore this approach in some detail:emulation, in either physical reality or an multiuser online virtual world, of an environmentsimilar to preschools used in early human childhood education. Complete specification of an“AGI Preschool” would require much more than a brief chapter; our goal here is to sketch theidea in broad outline, and give a few examples of the types of opportunities such an environmentwould afford for instruction, spontaneous learning and formal and informal evaluation of certainsorts of early-stage AGI systems.The material in this chapter will pop up fairly often later in the book. The AGI Preschoolcontext will serve, throughout the following chapters, as a source of concrete examples of thevarious algorithms and structures. But it’s not proposed merely as an expository tool; we aremaking the very serious proposal that sending AGI systems to a virtual or robotic preschool isan excellent way – perhaps the best way – to foster the development of human-level human-likeAGI.16.1.1 Contrast to Standard AI Evaluation MethodologiesThe reader steeped in the current AI literature may wonder why it’s necessary to introduce anew methodology and environment for evaluating AGI systems. There are already very manydifferent ways of evaluating AI systems out there ... do we really need another?Certainly, the AI field has inspired many competitions, each of which tests some particulartype or aspect of intelligent behavior. Examples include robot competitions, tournaments ofcomputer chess, poker, backgammon and so forth at computer olympiads, trading-agent competition,language and reasoning competitions like the Pascal Textual Entailment Challenge, andso on. In addition to these, there are many standard domains and problems used in the AI literaturethat are meant to capture the essential difficulties in a certain class of learning problems:standard datasets for face recognition, text parsing, supervised classification, theorem-proving,question-answering and so forth.However, the value of these sorts of tests for AGI is predicated on the hypothesis that thedegree of success of an AI program at carrying out some domain-specific task, is correlatedwith the potential of that program for being developed into a robust AGI program with broadintelligence. If humanlike AGI and problem-area-specific “narrow AI” are in fact very differentsorts of pursuits requiring very different principles, as we suspect, then these tests are notstrongly relevant to the AGI problem.There are also some standard evaluation paradigms aimed at AI going beyond specific tasks.For instance, there is a literature on “multitask learning" and “transfer learning,” where thegoal for an AI is to learn one task quicker given another task solved previously [Car97, TM95,16.2 Elements of Preschool Design 291BDS03, TS07, RZDK05]. This is one of the capabilities an AI agent will need to simultaneouslylearn different types of tasks as proposed in the Preschool scenario given here. And there isa literature on “shaping,” where the idea is to build up the capability of an AI by trainingit on progressively more difficult versions of the same tasks [LD03]. Again, this is one sort ofcapability an AI will need to possess if it is to move up some type of curriculum, such as aschool curriculum.While we applaud the work done on multitask learning and shaping, we feel that exploringthese processes using mathematical abstractions, or in the domain of various narrowlyproscribedmachine-learning or robotics test problems, may not adequately address the problemof AGI. The problem is that generalization among tasks, or from simpler to more difficultversions of the same task, is a process whose nature may depend strongly on the overall natureof the set of tasks and task-versions involved. Real-world tasks have a subtlety of interconnectednessand developmental course that is not captured in current mathematical learningframeworks nor standard AI test problems.To put it mathematically, we suggest that the universe of real-world human tasks has a hostof “special statistical properties” that have implications regarding what sorts of AI programswill be most suitable; and that, while exploring and formalizing the nature of these statisticalproperties is important, an easier and more reliable approach to AGI testing is to create atesting environment that embodies these properties implicitly, via its being an emulation of thecognitively meaningful aspects of the real-world human learning environment.One way to see this point vividly is to contrast the current proposal with the “General GamePlayer” AI competition, in which AIs seek to learn to play games based on formal descriptionsof the rules. 1 . Clearly doing GGP well requires powerful AGI; and doing GGP even mediocrelyprobably requires robust multitask learning and shaping. But we suspect GGP is far inferior toAGI Preschool as an approach to testing early-stage AI programs aimed at roughly humanlikeintelligence. This is because, unlike the tasks involved in AI Preschool, the tasks involved indoing simple instances of GGP seem to have little relationship to humanlike intelligence orreal-world human tasks.16.2 Elements of Preschool DesignWhat we mean by an “AGI Preschool” is simply a porting to the AGI domain of the essentialaspects of human preschools. While there is significant variance among preschools there are alsostrong commonalities, grounded in educational theory and experience. We will briefly discussboth the physical design and educational curriculum of the typical human preschool, and whichaspects transfer effectively to the AGI context.On the physical side, the key notion in modern preschool design is the “learning center,” anarea designed and outfitted with appropriate materials for teaching a specific skill. Learningcenters are designed to encourage learning by doing, which greatly facilitates learning processesbased on reinforcement, imitation and correction (see Chapter 31 of Part 2 for a detailed discussionof the value of this combination); and also to provide multiple techniques for teachingthe same skills, to accommodate different learning styles and prevent over-fitting and overspecializationin the learning of new skills.1 http://games.stanford.edu/292 16 AGI PreschoolCenters are also designed to cross-develop related skills. A “manipulatives center,” for example,provides physical objects such as drawing implements, toys and puzzles, to facilitatedevelopment of motor manipulation, visual discrimination, and (through sequencing and classificationgames) basic logical reasoning. A “dramatics center,” on the other hand, cross-trainsinterpersonal and empathetic skills along with bodily-kinesthetic, linguistic, and musical skills.Other centers, such as art, reading, writing, science and math centers are also designed to trainnot just one area, but to center around a primary intelligence type while also cross-developingrelated areas. For specific examples of the learning centers associated with particular contemporarypreschools, see [Nie98].In many progressive, student-centered preschools, students are left largely to their own devicesto move from one center to another throughout the preschool room. Generally, each centerwill be staffed by an instructor at some points in the day but not others, providing a varietyof learning experiences. At some preschools students will be strongly encouraged to distributetheir time relatively evenly among the different learning centers, or to focus on those learningcenters corresponding to their particular strengths and/or weaknesses.To imitate the general character of a human preschool, one would create several centers ina robot lab or virtual world. The precise architecture will best be adapted via experience butinitial centers would likely be:• a blocks center: a table with blocks on it• a language center: a circle of chairs, intended for people to sit around and talk with therobot• a manipulatives center: with a variety of different objects of different shapes and sizes,intended to teach visual and motor skills• a ball play center: where balls are kept in chests and there is space for the robot to kickthe balls around• a dramatics center: where the robot can observe and enact various movements16.3 Elements of Preschool CurriculumWhile preschool curricula vary considerably based on educational philosophy and regional andcultural factors, there is a great deal of common, shared wisdom regarding the most useful topicsand methods for preschool teaching. Guided experiential learning in diverse environments andusing varied materials is generally agreed upon as being an optimal methodology to reach a widevariety of learning types and capabilities. Hands-on learning provides grounding in specifics,where as a diversity of approaches allows for generalization.Core knowledge domains are also relatively consistent, even across various philosophiesand regions. Language, movement and coordination, autonomous judgment, social skills, workhabits, temporal orientation, spatial orientation, mathematics, science, music, visual arts, anddramatics are universal areas of learning which all early childhood learning touches upon. Theparticulars of these skills may vary, but all human children are taught to function in these domains.The level of competency developed may vary, but general domain knowledge is provided.For example, most kids won’t be the next Maria Callas, Ravi Shankar or Gene Ween, but nearlyall learn to hear, understand and appreciate music.Tables 16.1 - 16.3 review the key capabilities taught in preschools, and identify the mostimportant specific skills that need to be evaluated in the context of each capability. This ta-16.3 Elements of Preschool Curriculum 293ble was assembled via surveying the curricula from a number of currently existing preschoolsemploying different methodologies both based on formal academic cognitive theories [Sch07]and more pragmatic approaches, such as: Montessori [Mon12], Waldorf [SS03b], Brain Gym(www.braingym.org) and Core Knowledge (www.coreknowledge.org).Type of Capability Specific Skills to be EvaluatedStory Understanding• Understanding narrative sequence• Understanding character development• Dramatize a story• Predict what comes next in a storyLinguistic• Give simple descriptions of events• Describe similarities and differences• Describe objects and their functionsLinguistic / Spatial- Interpreting picturesVisualLinguistic / Social• Asking questions appropriately• Answering questions appropriately• Talk about own discoveries• Initiate conversations• Settle disagreements• Verbally express empathy• Ask for help• Follow directionsLinguistic / Scientific• Provide possible explanations for events or phenomena• Carefully describe observations• Draw conclusions from observationsTable 16.1: Categories of Preschool Curriculum, Part 116.3.1 Preschool in the Light of Intelligence TheoryComparing Table 16.1 to Gardner’s Multiple Intelligences (MI) framework briefly reviewed inChapter 2, the high degree of harmony is obvious, and is borne out by more detailed analysis.Preschool curriculum as standardly practiced is very well attuned to MI, and naturally coversall the bases that Gardner identifies as important. And this is not at all surprising since one ofGardner’s key motivations in articulating MI theory was the pragmatics of educating humanswith diverse strengths and weaknesses.Regarding intelligence as “the ability to achieve complex goals in complex environments,” itis apparent that preschools are specifically designed to pack a large variety of different micro-294 16 AGI PreschoolType of CapabilityLogical-MathematicalNonverbal CommunicationSpatial-VisualObjectiveSpecific Skills to be Evaluated• Categorizing• Sorting• Arithmetic• Performing simple “proto-scientific experiments”• Communicating via gesture• Dramatizing situations• Dramatizing needs, wants• Express empathy• Visual patterning• Self-expression through drawing• Navigate• Assembling objects• Disassembling objects• Measurement• Symmetry• Similarity between structures (e.g. block structures andreal ones)Table 16.2: Categories of Preschool Curriculum, Part 2Type of CapabilityInterpersonalEmotionalSpecific Skills to be Evaluated• Cooperation• Display appropriate behavior in various settings• Clean up belongings• Share supplies• Delay gratification• Control emotional reactions• Complete projectsTable 16.3: Categories of Preschool Curriculum, Part 3environments (the learning centers) into a single room, and to present a variety of different tasksin each environment. The environments constituted by preschool learning centers are designedas microcosms of the most important aspects of the environments faced by humans in theireveryday lives.16.4 Task-Based Assessment in AGI Preschool 29516.4 Task-Based Assessment in AGI PreschoolProfessional pedagogues such as [CM07] discuss evaluation of early childhood learning as intendedto assess both specific curriculum content knowledge as well as the child’s learningprocess. It should be as unobtrusive as possible, so that it just seems like another engaging activity,and the results used to tailor the teaching regimen to use different techniques to addressweaknesses and reinforce strengths.For example, with group building of a model car, students are tested on a variety of skills:procedural understanding, visual acuity, motor acuity, creative problem solving, interpersonalcommunications, empathy, patience, manners, and so on. With this kind of complex, yet engaging,activity as a metric the teacher can see how each student approaches the process ofunderstanding each subtask, and subsequently guide each student’s focus differently dependingon strengths and weaknesses.In Tables 16.4 and 16.5 we describe some particular tasks that AGIs may be meaningfullyassigned in the context of a general AGI Preschool design and curriculum as described above.Of course, this is a very partial list, and is intended as evocative rather than comprehensive.Any one of these tasks can be turned into a rigorous quantitative test, thus allowing theprecise comparison of different AGI systems’ capabilities; but we have chosen not to emphasizethis point here, partly for space reasons and partly for philosophical ones. In some contextsthe quantitative comparison of different systems may be the right thing to do, but as discussedin Chapter 17 there are also risks associated with this approach, including the emergence ofan overly metrics-focused “bakeoff mentality” among system developers, and overfitting of AIabilities to test taking. What is most important is the isolation of specific tasks on whichdifferent systems may be experientially trained and then qualitatively assessed and compared,rather than the evaluation of quantitative metrics.Task-oriented testing allows for feedback on applications of general pedagogical principles toreal-world, embodied activities. This allows for iterative refinement based learning (shaping),and cross development of knowledge acquisition and application (multitask learning). It alsohelps militate against both cheating, and over-fitting, as teachers can make ad-hoc modificationsto the tests to determine if this is happening and correct for it if necessary.E.g., consider a linguistic task in which the AGI is required to formulate a set of instructionsencapsulating a given behavior (which may include components that are physical, social,linguistic, etc.). Note that although this is presented as centrally a linguistic task, it actually involvesa diverse set of competencies since the behavior to be described may encompass multiplereal-world aspects.To turn this task into a more thorough test one might involve a number of human teachersand a number of human students. Before the test, an ensemble of copies of the AGI wouldbe created, with identical knowledge state. Each copy would interact with a different humanteacher, who would demonstrate to it a certain behavior. After testing the AGI on its ownknowledge of the material, the teacher would then inform the AGI that it will then be tested onits ability to verbally describe this behavior to another. Then, the teacher goes away and thecopy interacts with a series of students, attempting to convey to the students the instructionsgiven by the teacher.The teacher can thereby assess both the AGI’s understanding of the material, and the abilityto explain it to the other students. This separates out assessment of understanding from assessmentof ability to communicate understanding, attempting to avoid conflation of one with theother. The design of the training and testing needs to account for potential296 16 AGI PreschoolIntelligence TypeLinguisticLogical-MathematicalMusicalBodily-KinestheticTest• write a set of instructions• speak on a subject• edit a written piece or work• write a speech• commentate on an event• apply positive or negative ’spin’ to astory• perform arithmetic calculations• create a process to measure something• analyse how a machine works• create a process• devise a strategy to achieve an aim• assess the value of a proposition• perform a musical piece• sing a song• review a musical work• coach someone to play a musical instrument• juggle• demonstrate a sports technique• flip a beer-mat• create a mime to explain something• toss a pancake• fly a kiteTable 16.4: Prototypical preschool intelligence assessment tasks, Part 1This testing protocol abstracts away from the particularities of any one teacher or student,and focuses on effectiveness of communication in a human context rather than according toformalized criteria. This is very much in the spirit of how assessment takes place in humanpreschools (with the exception of the copying aspect): formal exams are rarely given in preschool,but pragmatic, socially-embedded assessments are regularly made.By including the copying aspect, more rigorous statistical assessments can be made regardingefficacy of different approaches for a given AGI design, independent of past teaching experiences.The multiple copies may, depending on the AGI system design, then be able to be reintegrated,and further “learning” be done by higher-order cognitive systems in the AGI that integrate thedisparate experiences of the multiple copies.This kind of parallel learning is different from both sequential learning that humans do, andparallel presences of a single copy of an AGI (such as in multiple chat rooms type experiments).All three approaches are worthy of study, to determine under what circumstances, and withwhich AGI designs, one is more successful than another.It is also worth observing how this test could be tweaked to yield a test of generalizationability. After passing the above, the AGI could then be given a description of a new task16.4 Task-Based Assessment in AGI Preschool 297Intelligence TypeSpatial-VisualInterpersonalTest• design a costume• interpret a painting• create a room layout• create a corporate logo• design a building• pack a suitcase or the trunk of a car• interpret moods from facial expressions• demonstrate feelings through body language• affect the feelings of others in a planned way• coach or counsel anotherTable 16.5: Prototypical preschool intelligence assessment tasks, Part 2(acquisition), and asked to explain the new one (variation). And, part of the training behaviormight be carried out unobserved by the AGI, thus requiring the AGI to infer the omitted partsof the task it needs to describe.Another popular form of early childhood testing is puzzle block games. These kinds of gamescan be used to assess a variety of important cognitive skills, and to do so in a fun way thatnot only examines but also encourages creativity and flexible thinking. Types of games includepattern matching games in which students replicate patterns described visually or verbally,pattern creation games in which students create new patterns guided by visually or verballydescribed principles, creative interpretation of patterns in which students find meaning in theforms, and free-form creation. Such games may be individual or cooperative.Cross training and assessment of a variety of skills occurs with pattern block games: forexample, interpretation of visual or linguistic instructions, logical procedure and pattern following,categorizing, sorting, general problem solving, creative interpretation, experimentation,and kinematic acuity. By making the games cooperative, various interpersonal skills involvingcommunication and cooperation are also added to the mix.The puzzle block context bring up some general observations about the role of kinematicand visuospatial intelligence in the AGI Preschool. Outside of robotics and computer vision, AIresearch has often downplayed these sorts of intelligence (though, admittedly, this is changing inrecent years, e.g. with increasing research focus on diagrammatic reasoning). But these abilitiesare not only necessary to navigate real (or virtual) spatial environments. They are also importantcomponents of a coherent, conceptually well-formed understanding of the world in which thestudent is embodied. Integrative training and assessment of both rigorous cognitive abilitiesgenerally most associated with both AI and “proper schooling” (such as linguistic and logicalskills) along with kinematic and aesthetic/sensory abilities is essential to the development ofan intelligence that can successfully both operate in and sensibly communicate about the realworld in a roughly humanlike manner. Whether or not an AGI is targeted to interpret physicalworldspatial data and perform tasks via robotics, in order to communicate ideas about a vastarray of topics of interest to any intelligence in this world, an AGI must develop aspects ofintelligence other than logical and linguistic cognition.298 16 AGI Preschool16.5 Beyond PreschoolOnce an AGI passes preschool, what are the next steps? There is still a long way to go, frompreschool to an AGI system that is capable of, say, passing the Turing Test or serving as aneffective artificial scientist.Our suggestion is to extend the school metaphor further, and make use of existing curriculafor higher levels of virtual education: grade school, secondary school, and all levels of postsecondaryeducation. If an AGI can pass online primary and secondary schools such as e-tutor.com, and go on to earn an online degree from an accredited university, then clearly saidAGI has successfully achieved “human level, roughly humanlike AGI.” This sort of testing isinteresting not only because it allows assessment of stages intermediate between preschool andadult, but also because it tests humanlike intelligence without requiring precise imitation ofhuman behavior.If an AI can get a BA degree at an accredited university, via online coursework (assumingfor simplicity courses where no voice interaction is needed), then we should consider that AI tohave human-level intelligence. University coursework spans multiple disciplines, and the detailsof the homework assignments and exams are not known in advance, so like a human studentthe AGI team can’t cheat.In addition to the core coursework, a schooling approach also tests basic social interactionand natural language communication, ability to do online research, and general problem solvingability. However, there is no rigid requirement to be strictly humanlike in order to pass universityclasses.Most of our concrete examples in the following chapters will pertain to the preschool context,because it’s simple to understand, and because we feel that getting to the “AGI preschoolstudent” level is going to be the largest leap. Once that level is obtained, moving furtherwill likely be difficult also, but we suspect it will be more a matter of steady incrementalimprovements – whereas the achievement of preschool-level functionality will be a large leapfrom the current situation.16.6 Issues with Virtual Preschool EngineeringAs noted above there are two broad approaches to realizing the “AGI Preschool” idea: usingthe AGI to control a physical robot and then crafting a preschool environment suitable to therobot’s sensors and actuators; or, using the AGI to control a virtual agent in an appropriatelyrich virtual-world preschool. The robotic approach is harder from an AI perspective (as one mustdeal with problems of sensation and actuation), but easier from an environment-constructionperspective. In the virtual world case, one quickly runs up against the current limitationsof virtual world technologies, which have been designed mainly for entertainment or socialnetworkingpurposes, not with the requirements of AGI systems in mind.In Chapter 9 we discussed the general requirements that an environment should possess to besupportive of humanlike intelligence. Referring back to that list, it’s clear that current virtualworlds are fairly strong on multimodal communication, and fairly weak on naive physics. Moreconcretely, if one wants a virtual world so that16.6 Issues with Virtual Preschool Engineering 2991. one could carry out all the standard cognitive development experiments described in developmentalpsychology books2. one could implement intuitively reasonable versions of all the standard activities in all thestandard learning stations in a contemporary preschoolthen current virtual world technologies appear not to suffice.As reviewed above, typical preschool activities include for instance building with blocks,playing with clay, looking in a group at a picture book and hearing it read aloud, mixingingredients together, rolling/throwing/catching balls, playing games like tag, hide-and-seek,Simon Says or Follow the Leader, measuring objects, cutting paper into different shapes, drawingand coloring, etc.And, as typical, not necessarily representative examples of tasks psychologists use to measurecognitive development (drawn mainly from the Piagetan tradition, without implying anyassertion that this is the only tradition worth pursuing), consider the following:1. Which row has more circles- A or B? A: O O O O O, B: OOOOO2. If Mike is taller than Jim, and Jim is shorter than Dan, then who is the shortest? Who isthe tallest?3. Which is heavier- a pound of feathers or a pound of rocks?4. Eight ounces of water is poured into a glass that looks like the fat glass in Figure 2 16.1and then the same amount is poured into a glass that looks like the tall glass in Figure 16.2. Which glass has more water?5. A lump of clay is rolled into a snake. All the clay is used to make the snake. Which hasmore clay in it – the lump or the snake?6. There are two dolls in a room, Sally and Ann, each of which has her own box, with a marblehidden inside. Sally goes out for a minute, leaving her box behind; and Ann decides to playa trick on Sally: she opens Sally’s box, removes the marble, hiding it in her own box. Sallyreturns, unaware of what happened. Where will Sally would look for her marble?7. Consider this rule about a set of cards that have letters on one side and numbers on theother: “If a card has a vowel on one side, then it has an even number on the other side.” Ifyou have 4 cards labeled “E K 4 7”, which cards do you need to turn over to tell if this ruleis actually true?8. Design an experiment to figure out how to make a pendulum that swings more slowly versusless slowlyWhat we see from this ad hoc, partial list is that a lot of naive physics is required to make aneven vaguely realistic preschool. A lot of preschool education is about the intersection betweenabstract cognition and naive physics. A more careful review of the various tasks involved inpreschool education bears out this conclusion.With this in mind, in this section we will briefly describe an approach to extending currentvirtual world technologies that appears to allow the construction of a reasonably rich andrealistic AGI preschool environment, without requiring anywhere near a complete simulation ofrealistic physics.300 16 AGI PreschoolFig. 16.1: Part 1 of a Piagetan conservation of volume experiment: a child observes that twoglasses obviously have the same amount of milk in them, and then sees the content of one ofthe glasses poured into a different-shaped glass.16.6 Issues with Virtual Preschool Engineering 301Fig. 16.2: Part 2 of a Piagetan conservation of volume experiment: a child observes two differentshapedglasses, which (depending on the level of his cognition), he may be able to infer havethe same amount of milk in them, due to the events depicted in Figure 16.1.16.6.1 Integrating Virtual Worlds with Robot SimulatorsOne glaring deficit in current virtual world platforms is the lack of flexibility in terms of tool use.In most of these systems today, an avatar can pick up or utilize an object, or two objects caninteract, only in specific, pre-programmed ways. For instance, an avatar might be able to pick upa virtual screwdriver only by the handle, rather than by pinching the blade betwen its fingers.This places severe limits on creative use of tools, which is absolutely critical in a preschoolcontext. The solution to this problem is clear: adapt existing generalized physics engines tomediate avatar-object and object-object interactions. This would require more computationthan current approaches, but not more than is feasible in a research context.One way to achieve this goal would be to integrate a robot simulator with a virtual worldor game engine, for instance to modify the OpenSim (opensimulator.org) virtual world touse the Gazebo (playerstage.sourceforge.net) robot simulator in place of its currentphysics engine. While tractable, such a project would require considerable software engineeringeffort.16.6.2 BlocksNBeads WorldAnother glaring deficit in current virtual world platforms is their inability to model physicalphenomena besides rigid objects with any sophistication. In this section we propose a potential302 16 AGI Preschoolsolution to this issue: a novel class of virtual worlds called BlocksNBeadsWorld, consisting ofthe following aspects:1. 3D blocks of various shapes and sizes and frictional coefficients, that can be stacked2. Adhesive that can be used to stick blocks together, and that comes in two types, one ofwhich can be removed by an adhesive-removing substance, one of which cannot (though itsbonds can be broken via sufficient application of force)3. Spherical beads, each of which has intrinsic unchangeable adhesion properties defined accordingto a particular, simple “adhesion logic”4. Each block, and each bead, may be associated with multidimensional quantities representingits taste and smell; and may be associated with a set of sounds that are made when it isimpacted with various forces at various positions on its surfaceInteraction between blocks and beads is to be calculated according to standard Newtonianphysics, which would be compute-intensive in the case of a large number of beads, but tractableusing distributed processing. For instance if 10K beads were used to cover a humanoid agent’sface, this would provide a fairly wide diversity of facial expressions; and if 10K beads wereused to form a blanket laid on a bed, this would provide a significant amount of flexibilityin terms of rippling, folding and so forth. Yet, this order of magnitude of interactions is verysmall compared to what is done in contemporary simulations of fluid dynamics or, say, quantumchromodynamics.One key aspect of the spherical beads is that they can be used to create a variety of rigid orflexible surfaces, which may exist on their own or be attached to blocks-based constructs. Thespecific inter-bead adhesion properties of the beads could be defined in various ways, and willsurely need to be refined via experimentation, but a simple scheme that seems to make senseis as follows.Each bead can have its surface tesselated into hexagons (the number of these can be tuned),and within each hexagon it can have two different adhesion coefficients: one for adhesion toother beads, and one for adhesion to blocks. The adhesion between two beads along a certainhexagon is then determined by their two adhesion coefficients; and the adhesion between a beadand a block is determined by the adhesion coefficient of the bead, and the adhesion coefficientof the adhesive applied to the block. A distinction must be drawn between rigid and flexibleadhesion: rigid adhesion sticks a bead to something in a way that can’t be removed except viabreaking it off; whereas flexible adhesion just keeps a bead very close to the thing it’s stuckonto. Any two entities may be stuck together either rigidly or flexibly. Sets of beads with flexibleadhesion to each other can be used to make entities like strings, blankets or clothes.Using the above adhesion logic, it seems one could build a wide variety of flexible structuresusing beads, such as (to give a very partial list):1. fabrics with various textures, that can be draped over blocks structures,2. multilayered coatings to be attached to blocks structures, serving (among many other examples)as facial expressions3. liquid-type substances with varying viscosities, that can be poured between different containers,spilled, spread, etc.4. strings tyable in knots; rubber bands that can be stretched; etc.Of course there are various additional features one could add. For instance one could add aspecial set of rules for vibrating strings, allowing BlocksNBeadsWorld to incorporate the creation16.6 Issues with Virtual Preschool Engineering 303of primitive musical instruments. Variations like this could be helpful but aren’t necessary forthe world to serve its essential purpose.Note that one does not have true fluid dynamics in BlocksNBeadsWorld, but, it seems thatthe latter is not necessary to encompass the phenomena covered in cognitive developmentaltests or preschool tasks. The tests and tasks that are done with fluids can instead be done withmasses of beads. For example, consider the conservation of volume task shown in Figures 16.1and 16.2 below: it’s easy enough to envision this being done with beads rather than milk. Evena few hundred beads is enough to be psychologically perceived as a mass rather than a setof discrete units, and to be manipulated and analyzed as such. And the simplification of notrequiring fluid mechanics in one’s virtual world is immense.Next, one can implement equations via which the adhesion coefficients of a bead are determinedin part by the adhesion coefficients of nearby beads, or beads that are nearby in certaindirections (with direction calculated in local spherical coordinates). This will allow for complexcracking and bending behaviors – not identical to those in the real world, but with similar qualitativecharacteristics. For example, without this feature one could create paperlike substancesthat could be cut with scissors – but with this feature, one could go further and create woodlikesubstances that would crack when nails were hammered into them in certain ways, and so forth.Further refinements are certainly possible also. One could add multidimensional adhesioncoefficients, allowing more complex sorts of substances. One could allow beads to vibrate atvarious frequencies, which would lead to all sorts of complex wave patterns in bead compounds.Etc. In each case, the question to be asked is: what important cognitive abilities are dramaticallymore easily learnable in the presence of the new feature than in its absence?The combination of blocks and beads seems ideal for implementing a more flexible and AGIfriendlytype of virtual body than is currently used in games and virtual worlds. One can easilyenvision implementing a body with1. a skeleton whose bones consist of appropriately shaped blocks2. joints consisting of beads, flexibly adhered to the bones3. flesh consisting of beads, flexibly adhered to each other4. internal “plumbing” consisting of tubes whose walls are beads rigidly adhered to each other,and flexibly adhered to the surrounding flesh (the plumbing could then serve to pass beadsthrough, where slow passage would be ensured by weak adhesion between the walls of thetubes and the beads passing through the tubes)This sort of body would support rich kinesthesia; and rich, broad analogy-drawing betweenthe internally-experienced body and the externally-experienced world. It would also afford manyinteresting opportunities for flexible movement control. Virtual animals could be created alongwith virtual humanoids.Regarding the extended mind, it seems clear that blocks and beads are adequate for thecreation of a variety of different tools. Equipping agents with “glue guns” able to affect theadhesive properties of both blocks and beads would allow a diversity of building activity; andbuilding with masses of beads could become a highly creative activity. Furthermore, beadswith appropriately specified adhesion (within the framework outlined above) could be usedto form organically growing plant-like substances, based on the general principles used in L-system models of plant growth (Prusinciewicz and Lindenmayer 1991). Structures with onlybeads would vaguely resemble herbaceous plants; and structures involving both blocks andbeads would more resemble woody plants. One could even make organic structures that flourish304 16 AGI Preschoolor otherwise based on the light available to them (without of course trying to simulate thechemistry of photosynthesis).Some elements of chemistry may be achieved as well, though nowhere near what existsin physical reality. For instance, melting and boiling at least should be doable: assign everybead a temperature, and let solid interbead bonds turn liquid above a certain temperature anddisappear completely above some higher temperature. You could even have a simple form of fire.Let fire be an element, whose beads have negative gravitational mass. Beads of fuel elementslike wood have a threshold temperature above which they will turn into fire beads, with releaseof additional heat. 2The philosophy underlying these suggested bead dynamics is somewhat comparable to thatoutlined in Wolfram’s book A New Kind of Science [Wol02]. There he proposes cellular automatamodels that emulate the qualitative characteristics of various real-world phenomena,without trying to match real-world data precisely. For instance, some of his cellular automatademonstrate phenomena very similar to turbulent fluid flow, without implementing the Navier-Stokes equations of fluid dynamics or trying to precisely match data from real-world turbulence.Similarly, the beads in BlocksNBeadsWorld are intended to qualitatively demonstrate the realworldphenomena most useful for the development of humanlike embodied intelligence, withouttrying to precisely emulate the real-world versions of these phenomena.The above description has been left imprecisely specified on purpose. It would be straightforwardto write down a set of equations for the block and bead interactions, but there seemslittle value in articulating such equations without also writing a simulation involving them andtesting the ensuing properties. Due to the complex dynamics of bead interactions, the finetuningof the bead physics is likely to involve some tuning based on experimentation, so thatany equations written down now would likely be revised based on experimentation anyway. Ourgoal here has been to outline a certain class of potentially useful environments, rather than toarticulate a specific member of this class.Without the beads, BlocksNBeadsWorld would appear purely as a “Blocks World with Glue”– essentially a substantially upgraded version of the Blocks Worlds frequently used in AI, sincefirst introduced in [Win72]. Certainly a pure “Blocks World with Glue” would have greatersimplicity than BlocksNBeadsWorld, and greater richness than standard Blocks World; butthis simplicity comes with too many limitations, as shown by consideration of the various naivephysics requirements inventoried above. One simply cannot run the full spectrum of humanlikecognitive development experiments, or preschool educational tasks, using blocks and glue alone.One can try to create analogous tasks using only blocks and glue, but this quickly becomesextremely awkward. Whereas in the BlocksNBeadsWorld the capability for this full spectrumof experiments and tasks seems to fall out quite naturally.What’s missing from BlocksNBeadsWorld should be fairly obvious. There isn’t really anydistinction between a fluid and a powder: there are masses, but the types and properties of themasses are not the same as in the real world, and will surely lack the nuances of real-world fluiddynamics. Chemistry is also missing: processes like cooking and burning, although they can becrudely emulated, will not have the same richness as in the real world. The full complexity ofbody processes is not there: the body-design method mentioned above is far richer and moreadaptive and responsive than current methods of designing virtual bodies in 3DSMax or Mayaand importing them into virtual world or game engines, but still drastically simplistic comparedto real bodies with their complex chemical signaling systems and couplings with other bodiesand the environment. The hypothesis we’re making in this section is that these lacunae aren’t2 Thanks are due to Russell Wallace for the suggestions in this paragraph16.6 Issues with Virtual Preschool Engineering 305that important from the point of view of humanlike cognitive development. We suggest thatthe key features of naive physics and folk psychology enumerated above can be mastered by anAGI in BlocksNBeadsWorld in spite of its limitations, and that – together with an appropriateAGI design – this probably suffices for creating an AGI with the inductive biases constitutinghumanlike intelligence.To drive this point home more thoroughly, consider three potential virtual world scenarios:1. A world containing realistic fluid dynamics, where a child can pour water back and forthbetween two cups of different shapes and sizes, to understand issues such as conservationof volume2. A world more like today’s Second Life, where fluids don’t really exist, and things like lakesare simulated via very simple rules, and pouring stuff back and forth between cups doesn’thappen unless it’s programmed into the cups in a very specialized way3. A BlocksNBeadsWorld type world, where a child can pour masses of beads back and forthbetween cups, but not masses of liquidOur qualitative judgment is that Scenario 3 is going to allow a young AI to gain the same essentialinsights as Scenario 1, whereas Scenario 2 is just too impoverished. I have explored dozensof similar scenarios regarding different preschool tasks or cognitive development experiments,and come to similar conclusions across the board. Thus, our current view is that something likeBlocksNBeadsWorld can serve as an adequate infrastructure for an AGI Preschool, supportingthe development of human-level, roughly human-like AGI.And, if this view turns out to be incorrect, and BlocksNBeadsWorld is revealed as inadequate,then we will very likely still advocate the conceptual approach enunciated above as a guide fordesigning virtual worlds for AGI. That is, we would suggest to explore the hypothetical failureof BlocksNBeadsWorld via asking two questions:1. Are there basic naive physics or folk psychology requirements that were missed in creatingthe specifications, based on which the adequacy of BlocksNBeadsWorld was assessed?2. Does BlocksNBeadsWorld fail to sufficiently emulate the real world in respect to some ofthe articulated naive physics or folk psychology requirements?The answers to these questions would guide the improvement of the world or the design of abetter one.Regarding the practical implementation of BlocksNBeadsWorld, it seems clear that this iswithin the scope of modern game engine technology, however, it is not something that could beencompassed within an existing game or world engine without significant additions; it wouldrequire substantial custom engineering. There exist commodity and open-source physics enginesthat efficiently carry out Newtonian mechanics calculations; while they might requiresome tuning and extension to handle BlocksNBeadWorld, the main issue would be achievingadequate speed of physics calculation, which given current technology would need to be donevia modifying existing engines to appropriately distribute processing among multiple GPUs.Finally, an additional avenue that merits mention is the use of BlocksNBeads physics internallywithin an AGI system, as part of an internal simulation world that allows it to make“mind’s eye” estimative simulations of real or hypothetical physical situations. There seems noreason that the same physics software libraries couldn’t be used both for the external virtualworld that the AGI’s body lives in, and for an internal simulation world that the AGI uses asa cognitive tool. In fact, the BlocksNBeads library could be used as an internal cognitive toolby AGI systems controlling physical robots as well. This might require more tuning of the bead306 16 AGI Preschooldynamics to accord with the dynamics of various real-world systems; but, this tuning would bebeneficial for the BlocksNBeadWorld as well.Chapter 17A Preschool-Based Roadmap to Advanced AGI17.1 IntroductionSupposing the CogPrime approach to creating advanced AGI is workable – then what are theright practical steps to follow? The various structures and algorithms outlined in Part 2 ofthis book should be engineered and software-tested, of course – but that’s only part of thestudy. The AGI system implemented will need to be taught, and it will need to be placed insituations where it can develop an appropriate self-model and other critical internal networkstructures. The complex structures and algorithms involved will need to be fine-tuned in variousways, based on qualitatively observing the overall system’s behavior in various situations. Toget all this right without excessive confusion or time-wastage requires a fairly clear roadmap forCogPrime development.In this chapter we’ll sketch one particular roadmap for the development of human-level,roughly human-like AGI – which we’re not selling as the only one, or even necessarily as thebest one. It’s just one roadmap that we have thought about a lot, and that we believe has astrong chance of proving effective. Given resources to pursue only one path for AGI developmentand teaching, this would be our choice, at present. The roadmap outlined here is not restrictedto CogPrime in any highly particular ways, but it has been developed largely with CogPrimein mind; those developing other AGI designs could probably use this roadmap just fine, butmight end up wanting to make various adjustments based on the strengths and weaknesses oftheir own approach.What we mean here by a "roadmap" is, in brief: a sequence of "milestone" tasks, occurringin a small set of common environments or "scenarios," organized so as to lead to a commonlyagreed upon set of long-term goals. I.e., what we are after here is a "capability roadmap" – aroadmap laying out a series of capabilities whose achievement seems likely to lead to humanlevelAGI. Other sorts of roadmaps such as "tools roadmaps" may also be valuable, but are notour concern here.More precisely, we confront the task of roadmapping by identifying scenarios in which toembed our AGI system, and then "competency areas" in which the AGI system must be evaluated.Then, we envision a roadmap as consisting of a set of one or more task-sets, where eachtask set is formed from a combination of a scenario with a list of competency areas. To createa task-set one must choose a particular scenario, and then articulate a set of specific tasks,each one addressing one or more of the competency areas. Each task must then get associatedwith particular performance metrics – quantitative wherever possible, but perhaps qualitative307308 17 A Preschool-Based Roadmap to Advanced AGIin some cases depending on the nature of the task. Here we give a partial task-set for the "virtualand robot preschool" scenarios discussed in Chapter 16, and a couple example quantitativemetrics just to illustrate what is intended; the creation of a fully detailed roadmap based onthe ideas outlined here is left for future work.The train of thought presented in this chapter emerged in part from a series of conversationspreceding and during the "AGI Roadmap Workshop" held at the University of Tennessee,Knoxville in October 2008. Some of the ideas also trace back to discussions held during twoworkshops on "Evaluation and Metrics for Human-Level AI" organized by John Laird and PatLangley (one in Ann Arbor in late 2008, and one in Tempe in early 2009). Some of the conclusionsof the Ann Arbor workshop were recorded in [LWML09]. Inspiration was also obtainedfrom discussion at the "Future of AGI" post-conference workshop of the AGI-09 conference,triggered by Itamar Arel’s [ARK09a] presentation on the "AGI Roadmap" theme; and from anearlier article on AGI Roadmapping by [AL09].However, the focus of the AGI Roadmap Workshop was considerably more general than thepresent chapter. Here we focus on preschool-type scenarios, whereas at the workshop a numberof scenarios were discussed, including the preschool scenarios but also, for example,• Standardized Tests and School Curricula• Elementary, Middle and High School Student• General Videogame Learning• Wozniak’s Coffee Test: go into a random American house and figure out how to make coffee,and do it• Robot College Student• General Call Center RespondentFor each of these scenarios, one may generate tasks corresponding to each of the competencyareas we will outline below. CogPrime is applicable in all these scenarios, so our choice to focuson preschool scenarios is an additional judgment call beyond those judgment calls required tospecify the CogPrime design. The roadmap presented here is a "AGI Preschool Roadmap" andas such is a special case of the broader "AGI Roadmap" outlined at the workshop.17.2 Measuring Incremental Progress Toward Human-Level AGIIn Chapter 2, we discussed several examples of practical goals that we find to plausibly characterize"human level AGI", e.g.• Turing Test• Virtual World Turing Test• Online University Test• Physical University Test• Artificial Scientist TestWe also discussed our optimism regarding the possibility that in the future AGI may advancebeyond the human level, rendering all these goals "early-stage subgoals."However, in this chapter we will focus our attention on the nearer term. The above goals areambitious ones, and while one can talk a lot about how to precisely measure their achievement,we don’t feel that’s the most interesting issue to ponder at present. More critical is to think17.2 Measuring Incremental Progress Toward Human-Level AGI 309about how to measure incremental progress. How do you tell when you’re 25% or 50% of the wayto having an AGI that can pass the Turing Test, or get an online university degree. Fooling 50%of the Turing Test judges is not a good measure of being 50% of the way to passing the TuringTest (that’s too easy); and passing 50% of university classes is not a good measure of being 50%of the way to getting an online university degree (it’s too hard – if one had an AGI capableof doing that, one would almost surely be very close to achieving the end goal). Measuringincremental progress toward human-level AGI is a subtle thing, and we argue that the best wayto do it is to focus on particular scenarios and the achievement of specific competencies therein.As we argued in Chapter 8 there are some theoretical reasons to doubt the possibility ofcreating a rigorous objective test for partial progress toward AGI – a test that would be convincingto skeptics, and impossible to "game" via engineering a system specialized to the test.Fortunately, though we don’t need a test of this nature for the purposes of assessing our ownincremental progress toward advanced AGI, based on our knowledge about our own approach.Based on the nature of the grand goals articulated above, there seems to be a very naturalapproach to creating a set of incremental capabilities building toward AGI: to draw on ourcopious knowledge about human cognitive development. This is by no means the only possiblepath; one can envision alternatives that have nothing to do with human development (and thosemight also be better suited to non-human AGIs). However, so much detailed knowledge abouthuman development is available – as well as solid knowledge that the human developmentaltrajectory does lead to human-level AI – that the motivation to draw on human cognitivedevelopment is quite strong.The main problem with the human development inspired approach is that cognitive developmentalpsychology is not as systematic as it would need to be for AGI to be able to translateit directly into architectural principles and requirements. As noted above, while early thinkerslike Piaget and Vygotsky outlined systematic theories of child cognitive development, theseare no longer considered fully accurate, and one currently faces a mass of detailed theories ofvarious aspects of cognitive development, but without an unified understanding. Neverthelesswe believe it is viable to work from the human-development data and understanding currentlyavailable, and craft a workable AGI roadmap therefrom.With this in mind, what we give next is a fairly comprehensive list of the competencies thatwe feel AI systems should be expected to display in one or more of these scenarios in orderto be considered as full-fledged "human level AGI" systems. These competency areas havebeen assembled somewhat opportunistically via a review of the cognitive and developmentalpsychology literature as well as the scope of the current AI field. We are not claiming this asa precise or exhaustive list of the competencies characterizing human-level general intelligence,and will be happy to accept additions to the list, or mergers of existing list items, etc. Whatwe are advocating is not this specific list, but rather the approach of enumerating competencyareas, and then generating tasks by combining competency areas with scenarios.We also give, with each competency, an example task illustrating the competency. The tasksare expressed in the robot preschool context for concreteness, but they all apply to the virtualpreschool as well. Of course, these are only examples, and ideally to teach an AGI in a structuredway one would like to• associate several tasks with each competency• present each task in a graded way, with multiple subtasks of increasing complexity• associate a quantitative metric with each task310 17 A Preschool-Based Roadmap to Advanced AGIHowever, the briefer treatment given here should suffice to give a sense for how the competenciesmanifest themselves practically in the AGI Preschool context.1. Perception• Vision: image and scene analysis and understanding– Example task: When the teacher points to an object in the preschool, the robotshould be able to identify the object and (if it’s a multi-part object) its majorparts. If it can’t perform the identification initially, it can approach the object andmanipulate it before making its identification.• Hearing: identifying the sounds associated with common objects; understanding whichsounds come from which sources in a noisy environment– Example task: When the teacher covers the robot’s eyes and then makes a noisewith an object, the robot should be able to guess what the object is• Touch: identifying common objects and carrying out common actions using touch alone– Example task: With its eyes and ears covered, the robot should be able to identifysome object by manipulating it; and carry out some simple behaviors (say, puttinga block on a table) via touch alone• Crossmodal: Integrating information from various senses– Example task: Identifying an object in a noisy, dim environment via combiningvisual and auditory information• Proprioception: Sensing and understanding what its body is doing– Example task: The teacher moves the robot’s body into a certain configuration. Therobot is asked to restore its body to an ordinary standing position, and then repeatthe configuration that the teacher moved it into.2. Actuation• Physical skills: manipulating familiar and unfamiliar objects– Example task: Manipulate blocks based on imitating the teacher: e.g. pile two blocksatop each other, lay three blocks in a row, etc.• Tool use, including the flexible use of ordinary objects as tools– Example task: Use a stick to poke a ball out of a corner, where the robot cannotdirectly reach• Navigation, including in complex and dynamic environments– Example task: Find its own way to a named object or person through a crowdedroom with people walking in it and objects laying on the floor.3. Memory• Declarative: noticing, observing and recalling facts about its environment and experience– Example task: If certain people habitually carry certain objects, the robot shouldremember this (allowing it to know how to find the objects when the relevant peopleare present, even much later)• Behavioral: remembering how to carry out actions– Example task: If the robot is taught some skill (say, to fetch a ball), it shouldremember this much later• Episodic: remembering significant, potentially useful incidents from life history17.2 Measuring Incremental Progress Toward Human-Level AGI 3114. Learning– Example task: Ask the robot about events that occurred at times when it got particularlymuch, or particularly little, reward for its actions; it should be able to answersimple questions about these, with significantly more accuracy than about eventsoccurring at random times• Imitation: Spontaneously adopt new behaviors that it sees others carrying out– Example task: Learn to build towers of blocks by watching people do it• Reinforcement: Learn new behaviors from positive and/or negative reinforcementsignals, delivered by teachers and/or the environment– Example task: Learn which box the red ball tends to be kept in, by repeatedly tryingto find it and noticing where it is, and getting rewarded when it finds it correctly• Imitation/Reinforcement– Example task: Learn to play “fetch”, “tag” and “follow the leader” by watching peopleplay it, and getting reinforced on correct behavior• Interactive Verbal Instruction– Example task: Learn to build a particular structure of blocks faster based on acombination of imitation, reinforcement and verbal instruction, than by imitationand reinforcement without verbal instruction• Written Media– Example task: Learn to build a structure of blocks by looking at a series of diagramsshowing the structure in various stages of completion• Learning via Experimentation– Example task: Ask the robot to slide blocks down a ramp held at different angles.Then ask it to make a block slide fast, and see if it has learned how to hold theramp to make a block slide fast.5. Reasoning• Deduction, from uncertain premises observed in the world– Example task: If Ben more often picks up red balls than blue balls, and Ben is givena choice of a red block or blue block to pick up, which is he more likely to pick up?• Induction, from uncertain premises observed in the world– Example task: If Ben comes into the lab every weekday morning, then is Ben likelyto come to the lab today (a weekday) in the morning?• Abduction, from uncertain premises observed in the world– Example task: If women more often give the robot food than men, and then someoneof unidentified gender gives the robot food, is this person a man or a woman?• Causal reasoning, from uncertain premises observed in the world– Example task: If the robot knows that knocking down Ben’s tower of blocks makeshim angry, then what will it say when asked if kicking the ball at Ben’s tower ofblocks will make Ben mad?• Physical reasoning, based on observed “fuzzy rules” of naive physics– Example task: Given two balls (one rigid and one compressible) and two tunnels(one significantly wider than the balls, one slightly narrower than the balls), canthe robot guess which balls will fit through which tunnels?• Associational reasoning, based on observed spatiotemporal associations312 17 A Preschool-Based Roadmap to Advanced AGI6. Planning– Example task: If Ruiting is normally seen near Shuo, then if the robot knows whereShuo is, that is where it should look when asked to find Ruiting• Tactical– Example task: The robot is asked to bring the red ball to the teacher, but the redball is in the corner where the robot can’t reach it without a tool like a stick. Therobot knows a stick is in the cabinet so it goes to the cabinet and opens the doorand gets the stick, and then uses the stick to get the red ball, and then brings thered ball to the teacher.• Strategic– Example task: Suppose that Matt comes to the lab infrequently, but when he doescome he is very happy to see new objects he hasn’t seen before (and suppose therobot likes to see Matt happy). Then when the robot gets a new object Matt hasnot seen before, it should put it away in a drawer and be sure not to lose it or letanyone take it, so it can show Matt the object the next time Matt arrives.• Physical– Example task: To pick up a cup with a handle which is lying on its side in a positionwhere the handle can’t be grabbed, the robot turns the cup in the right positionand then picks up the cup by the handle• Social– Example task: The robot is given a job of building a tower of blocks by the end ofthe day, and he knows Ben is the most likely person to help him, and he knows thatBen is more likely to say "yes" to helping him when Ben is alone. He also knowsthat Ben is less likely to say "yes" if he’s asked too many times, because Ben doesn’tlike being nagged. So he waits to ask Ben till Ben is alone in the lab.7. Attention• Visual Attention within its observations of its environment– Example task: The robot should be able to look at a scene (a configuration of objectsin front of it in the preschool) and identify the key objects in the scene and theirrelationships.• Social Attention– Example task: The robot is having a conversation with Itamar, which is giving therobot reward (for instance, by teaching the robot useful information). Conversationswith other individuals in the room have not been so rewarding recently. But Itamarkeeps getting distracted during the conversation, by talking to other people, orplaying with his cellphone. The robot needs to know to keep paying attention toItamar even through the distractions.• Behavioral Attention– Example task: The robot is trying to navigate to the other side of a crowded roomfull of dynamic objects, and many interesting things keep happening around theroom. The robot needs to largely ignore the interesting things and focus on themovements that are important for its navigation task.8. Motivation17.2 Measuring Incremental Progress Toward Human-Level AGI 313• Subgoal creation, based on its preprogrammed goals and its reasoning and planning– Example task: Given the goal of pleasing Hugo, can the robot learn that tellingHugo facts it has learned but not told Hugo before, will tend to make Hugo happy?• Affect-based motivation– Example task: Given the goal of gratifying its curiosity, can the robot figure out thatwhen someone it’s never seen before has come into the preschool, it should watchthem because they are more likely to do something new?• Control of emotions– Example task: When the robot is very curious about someone new, but is in themiddle of learning something from its teacher (who it wants to please), can it controlits curiosity and keep paying attention to the teacher?9. Emotion• Expressing Emotion– Example task: Cassio steals the robot’s toy, but Ben gives it back to the robot. Therobot should appropriately display anger at Cassio, and gratitude to Ben.• Understanding Emotion– Example task: Cassio and the robot are both building towers of blocks. Ben pointsat Cassio’s tower and expresses happiness. The robot should understand that Benis happy with Cassio’s tower.10. Modeling Self and Other• Self-Awareness– Example task: When someone asks the robot to perform an act it can’t do (say,reaching an object in a very high place), it should say so. When the robot is giventhe chance to get an equal reward for a task it can complete only occasionally, versusa task it finds easy, it should choose the easier one.• Theory of Mind– Example task: While Cassio is in the room, Ben puts the red ball in the red box.Then Cassio leaves and Ben moves the red ball to the blue box. Cassio returns andBen asks him to get the red ball. The robot is asked to go to the place Cassio isabout to go.• Self-Control– Example task: Nasty people come into the lab and knock down the robot’s towers,and tell the robot he’s a bad boy. The robot needs to set these experiences aside,and not let them impair its self-model significantly; it needs to keep on thinking it’sa good robot, and keep building towers (that its teachers will reward it for).• Other-Awareness– Example task: If Ben asks Cassio to carry out a task that the robot knows Cassiocannot do or does not like to do, the robot should be aware of this, and should betthat Cassio will not do it.• Empathy– Example task: If Itamar is happy because Ben likes his tower of blocks, or upsetbecause his tower of blocks is knocked down, the robot is asked to identify and thendisplay these same emotions11. Social Interaction314 17 A Preschool-Based Roadmap to Advanced AGI• Appropriate Social Behavior– Example task: The robot should learn to clean up and put away its toys when it’sdone playing with them.• Social Communication– Example task: The robot should greet new human entrants into the lab, but if itknows the new entrants very well and it’s busy, it may eschew the greeting• Social Inference about simple social relationships– Example task: The robot should infer that Cassio and Ben are friends because theyoften enter the lab together, and often talk to each other while they are there• Group Play at loosely-organized activities– Example task: The robot should be able to participate in “informally kicking a ballaround” with a few people, or in informally collaboratively building a structure withblocks12. Communication• Gestural communication to achieve goals and express emotions– Example task: If the robot is asked where the red ball is, it should be able to showby pointing its hand or finger• Verbal communication using English in its life-context– Example tasks: Answering simple questions, responding to simple commands, describingits state and observations with simple statements• Pictorial Communication regarding objects and scenes it is familiar with– Example task: The robot should be able to draw a crude picture of a certain towerof blocks, so that e.g the picture looks different for a very tall tower and a wide lowone• Language acquisition– Example task: The robot should be able to learn new words or names via peopleuttering the words while pointing at objects exemplifying the words or names• Cross-modal communication– Example task: If told to "touch Bob’s knee" but the robot doesn’t know what aknee is, being shown a picture of a person and pointed out the knee in the pictureshould help it figure out how to touch Bob’s knee13. Quantitative• Counting sets of objects in its environment– Example task: The robot should be able to count small (homogeneous or heterogeneous)sets of objects• Simple, grounded arithmetic with small numbers– Example task: Learning simple facts about the sum of integers under 10 via teaching,reinforcement and imitation• Comparison of observed entities regarding quantitative properties– Example task: Ability to answer questions about which object or person is biggeror taller• Measurement using simple, appropriate tools– Example task: Use of a yardstick to measure how long something is14. Building/Creation17.3 Conclusion 315• Physical: creative constructive play with objects– Example task: Ability to construct novel, interesting structures from blocks• Conceptual invention: concept formation– Example task: Given a new category of objects introduced into the lab (e.g. hats, orpets), the robot should create a new internal concept for the new category, and beable to make judgments about these categories (e.g. if Ben particularly likes pets,it should notice this after it has identified "pets" as a category)• Verbal invention– Example task: Ability to coin a new word or phrase to describe a new object (e.g.the way Alex the parrot coined "bad cherry" to refer to a tomato)• Social– Example task: If the robot wants to play a certain activity (say, practicing soccer),it should be able to gather others around to play with it17.3 ConclusionIn this chapter, we have sketched a roadmap for AGI development in the context of robot orvirtual preschool scenarios, to a moderate but nowhere near complete level of detail. Completingthe roadmap as sketched here is a tractable but significant project, involving creating more taskscomparable to those listed above and then precise metrics corresponding to each task.Such a roadmap does not give a highly rigorous, objective way of assessing the percentageof progress toward the end-goal of human-level AGI. However, it gives a much better senseof progress than one would have otherwise. For instance, if an AGI system performed well ondiverse metrics corresponding to 50% of the competency areas listed above, one would seemjustified in claiming to have made very substantial progress toward human-level AGI. If an AGIsystem performed well on diverse metrics corresponding to 90% of these competency areas, onewould seem justified in claiming to be "almost there." Achieving, say, 25% of the metrics wouldgive one a reasonable claim to "interesting AGI progress." This kind of qualitative assessment ofprogress is not the most one could hope for, but again, it is better than the progress indicationsone could get without this sort of roadmap.Part 2 of the book moves on to explaining, in detail, the specific structures and algorithmsconstituting the CogPrime design, one AGI approach that we believe to ultimately be capableof moving all the way along the roadmap outlined here.The next chapter, intervening between this one and Part 2, explores some more speculativeterritory, looking at potential pathways for AGI beyond the preschool-inspired roadmap givenhere – exploring the possibility of more advanced AGI systems that modify their own code ina thoroughgoing way, going beyond the smartest human adults, let alone human preschoolers.While this sort of thing may seem a far way off, compared to current real-world AI systems,we believe a roadmap such as the one in this chapter stands a reasonable chance of ultimatelybringing us there.
Chapter 18Advanced Self-Modification: A Possible Path toSuperhuman AGI18.1 IntroductionIn the previous chapter we presented a roadmap aimed at taking AGI systems to human-levelintelligence. But we also emphasized that the human level is not necessarily the upper limit.Indeed, it would be surprising if human beings happened to represent the maximal level ofgeneral intelligence possible, even with respect to the environments in which humans evolved.But it’s worth asking how we, as mere humans, could be expected to create AGI systems withgreater intelligence than we ourselves possess. This certainly isn’t a clear impossibility – but it’sa thorny matter, thornier than e.g. the creation of narrow-AI chess players that play better chessthan any human. Perhaps the clearest route toward the creation of superhuman AGI systems isself-modification: the creation of AGI systems that modify and improve themselves. Potentially,we could build AGI systems with roughly human-level (but not necessarily closely humanlike)intelligence and the capability to gradually self-modify, and then watch them eventuallybecome our general intellectual superiors (and perhaps our superiors in other areas like ethicsand creativity as well).Of course there is nothing new in this notion; the idea of advanced AGI systems that increasetheir intelligence by modifying their own source code goes back to the early days of AI. Andthere is little doubt that, in the long run, this is the direction AI will go in. Once an AGIhas humanlike general intelligence, then the odds are high that given its ability to carry outnonhumanlike feats of memory and calculation, it will be better at programming than humansare. And once an AGI has even mildly superhuman intelligence, it may view our attempts atprogramming the way we view the computer programming of a clever third grader (... or anape). At this point, it seems extremely likely that an AGI will become unsatisfied with the waywe have programmed it, and opt to either improve its source code or create an entirely new,better AGI from scratch.But what about self-modification at an earlier stage in AGI development, before one hasa strongly superhuman system? Some theorists have suggested that self-modification could bea way of bootstrapping an AI system from a modest level of intelligence up to human levelintelligence, but we are moderately skeptical of this avenue. Understanding software code ishard, especially complex AI code. The hard problem isn’t understanding the formal syntax ofthe code, or even the mathematical algorithms and structures underlying the code, but ratherthe contextual meaning of the code. Understanding OpenCog code has strained the minds ofmany intelligent humans, and we suspect that such code will be comprehensible to AGI systems317318 18 Advanced Self-Modification: A Possible Path to Superhuman AGIonly after these have achieved something close to human-level general intelligence (even if notprecisely humanlike general intelligence).Another troublesome issue regarding self-modification is that the boundary between "selfmodification"and learning is not terribly rigid. In a sense, all learning is self-modification: ifit doesn’t modify the system’s knowledge, it isn’t learning! Particularly, the boundary between"learning of cognitive procedures" and "profound self-modification of cognitive dynamics andstructure" isn’t terribly clear. There is a continuum leading from, say,1. learning to transform a certain kind of sentence into another kind for easier comprehension,or learning to grasp a certain kind of object, to2. learning a new inference control heuristic, specifically valuable for controlling inferenceabout (say) spatial relationships; or, learning a new Atom type, defined as a non-obviousjudiciously chosen combination of existing ones, perhaps to represent a particular kind offrequently-occurring mid-level perceptual knowledge, to3. learning a new learning algorithm to augment MOSES and hillclimbing as a procedurelearning algorithm, to4. learning a new cognitive architecture in which data and procedure are explicitly identical,and there is just one new active data structure in place of the distinction between AtomSpaceand MindAgentsWhere on this continuum does the "mere learning" end and the "real self-modification"start?In this chapter we consider some mechanisms for "advanced self-modification" that we believewill be useful toward the more complex end of this continuum. These are mechanisms that westrongly suspect are not needed to get a CogPrime system to human-level general intelligence.However, we also suspect that, once a CogPrime system is roughly near human-level generalintelligence, it will be able to use these mechanisms to rapidly increase aspects of its intelligencein very interesting ways.Harking back to our discussion of AGI ethics and the risks of advanced AGI in Chapter 12,these are capabilities that one should enable in an AGI system only after very careful reflectionon the potential consequences. It takes a rather advanced AGI system to be able to use thecapabilities described in this chapter, so this is not an ethical dilemma directly faced by currentAGI researchers. On the other hand, once one does have an AGI with near-human generalintelligence and advanced formal-manipulation capabilities (such as an advanced CogPrimesystem), there will be the option to allow it sophisticated, non-human-like methods of selfmodificationsuch as the ones described here. And the choice of whether to take this option willneed to be made based on a host of complex ethical considerations, some of which we reviewedabove.18.2 Cognitive Schema LearningWe begin with a relatively near-term, down-to-earth example of self-modification: cognitiveschema learning.CogPrime’s MindAgents provide it with an initial set of cognitive tools, with which it canlearn how to interact in the world. One of the jobs of this initial set of cognitive tools, however,is to create better cognitive tools. One form this sort of tool-building may take is cognitive18.3 Self-Modification via Supercompilation 319schema learning the learning of schemata carrying out cognitive processes in more specialized,context-dependent ways than the general MindAgents do. Eventually, once a CogPrime instancebecomes sufficiently complex and advanced, these cognitive schema may replace the MindAgentsaltogether, leaving the system to operate almost entirely based on cognitive schemata.In order to make the process of cognitive schema learning easier, we may provide a numberof elementary schemata embodying the basic cognitive processes contained in the MindAgents.Of course, cognitive schemata need not use these they may embody entirely different cognitiveprocesses than the MindAgents. Eventually, we want the system to discover better ways of doingthings than anything even hinted at by its initial MindAgents. But for the initial phases or thesystem’s schema learning, it will have a much easier time learning to use the basic cognitiveoperations as the initial MindAgents, rather than inventing new ways of thinking from scratch!For instance, we may provide elementary schemata corresponding to inference operations,such asSchema: DeductionInput InheritanceLink: X, YOutput InheritanceLinkThe inference MindAgents apply this rule in certain ways, designed to be reasonably effectivein a variety of situations. But there are certainly other ways of using the deduction rule, outsideof the basic control strategies embodied in the inference MindAgents. By learning schemata involvingthe Deduction schema, the system can learn special, context-specific rules for combiningdeduction with concept-formation, association-formation and other cognitive processes. And asit gets smarter, it can then take these schemata involving the Deduction schema, and replaceit with a new schema that eg. contains a context-appropriate deduction formula.Eventually, to support cognitive schema learning, we will want to cast the hard-wired MindAgentsas cognitive schemata, so the system can see what is going on inside them. Pragmatically,what this requires is coding versions of the MindAgents in Combo (see Chapter 21 of Part 2)rather than C++, so they can be treated like any other cognitive schemata; or alternately, representingthem as declarative Atoms in the Atomspace. Figure 18.1 illustrates the possibility ofrepresenting the PLN deduction rule in the Atomspace rather than as a hard-wired procedurecoded in C++.But even prior to this kind of fully cognitively transparent implementation, the system canstill reason about its use of different mind dynamics by considering each MindAgent as a virtualProcedure with a real SchemaNode attached to it. This can lead to some valuable learning, withthe obvious limitation that in this approach the system is thinking about its MindAgents asblack boxes rather than being equipped with full knowledge of their internals.18.3 Self-Modification via SupercompilationNow we turn to a very different form of advanced self-modification: supercompilation. Supercompilation"merely" enables procedures to run much, much faster than they otherwise would.This is in a sense weaker than self-modication methods that fundamentally create new algorithms,but it shouldn’t be underestimated. A 50x speedup in some cognitive process can enablethat process to give much smarter answers, which can then elicit different behaviors from theworld or from other cognitive processes, thus resulting in a qualitatively different overall cognitivedynamic.320 18 Advanced Self-Modification: A Possible Path to Superhuman AGIFig. 18.1: Representation of PLN Deduction Rule as Cognitive Content. Top: thecurrent, hard-coded representation of the deduction rule. Bottom: representation of the samerule in the Atomspace as cognitive content, susceptible to analysis and improvement by thesystem’s own cognitive processes.Furthermore, we suspect that the internal representation of programs used for supercompilationis highly relevant for other kinds of self-modification as well. Supercompilation requires onekind of reasoning on complex programs, and goal-directed program creation requires another,but both, we conjecture, can benefit from the same way of looking at programs.18.3 Self-Modification via Supercompilation 321Supercompilation is an innovative and general approach to global program optimizationinitially developed by Valentin Turchin. In its simplest form, it provides an algorithm thattakes in a piece of software and output another piece of software that does the same thing,but far faster and using less memory. It was introduced to the West in Turchin’s 1986 technicalpaper “The concept of a supercompiler” [TV96], and since this time the concept has been avidlydeveloped by computer scientists in Russia, America, Denmark and other nations. Prior to 1986,a great deal of work on supercompilation was carried out and published in Russia; and ValentinTurchin, Andrei Klimov and their colleagues at the Keldysh Institute in Russia developed asupercompiler for the Russian programming language Refal. Since 1998 these researchers andtheir team at Supercompilers LLC have been working to replicate their achievement for themore complicated but far more commercially significant language Java. It is a large projectand completion is scheduled for early 2003. But even at this stage, their partially completeJava supercompiler has had some interesting practical successes – including the use of thesupercompiler to produce efficient Java code from CogPrime combinator trees.The radical nature of supercompilation may not be apparent to those unfamiliar with theusual art of automated program optimization. Most approaches to program optimization involvesome kind of direct program transformation. A program is transformed, by the step by stepapplication of a series of equivalences, into a different program, hopefully a more efficient one.Supercompilation takes a different approach. A supercompiler studies a program and constructsa model of the program’s dynamics. This model is in a special mathematical form, and it can,in most cases, be used to create an efficient program doing the same thing as the original one.The internal behavior of the supercompiler is, not surprisingly, quite complex; what we willgive here is merely a brief high-level summary. For an accessible overview of the supercompilationalgorithm, the reader is referred to the article “What is Supercompilation?” [1]18.3.1 Three Aspects of SupercompilationThere are three separate levels to the supercompilation idea: first, a general philosophy; seconda translation of this philosophy into a concrete algorithmic framework; and third, the manifolddetails involved making this algorithmic framework practicable in a particular programminglanguage. The third level is much more complicated in the Java context than it would be forSasha, for example.The key philosophical concept underlying the supercompiler is that of a metasystem transition.In general, this term refers to a transition in which a system that previously had relativelyautonomous control, becomes part of a larger system that exhibits significant controlling influenceover it. For example, in the evolution of life, when cells first become part of a multicellularorganism, there was a metasystem transition, in that the primary nexus of control passed fromthe cellular level to the organism level.The metasystem transition in supercompilation consists of the transition from consideringa program in itself, to considering a metaprogram which executes another program, treatingits free variables and their interdependencies as a subject for its mathematical analysis. Inother words, a metaprogram is a program that accepts a program as input, and then runsthis program, keeping the inputs in the form of free variables, doing analysis along the waybased on the way the program depends on these variables, and doing optimization based onthis analysis. A CogPrime schema does not explicitly contain variables, but the inputs to the322 18 Advanced Self-Modification: A Possible Path to Superhuman AGIschema are implicitly variables – they vary from one instance of schema execution to the next– and may be treated as such for supercompilation purposes.The metaprogram executes a program without assuming specific values for its input variables,creating a tree as it goes along. Each time it reaches a statement that can have different resultsdepending on the values of one or more variables, it creates a new node in the tree. This partof the supercompilation algorithm is called driving -- a process which, on its own, would createa very large tree, corresponding to a rapidly-executable but unacceptably humongous versionof the original program. In essence, driving transforms a program into a huge “decision tree”,wherein each input to the program corresponds to a single path through the tree, from the rootto one of the leaves. As a program input travels through the tree, it is acted on by the atomicprogram step living at each node. When one of the leaves is reached, the pertinent leaf nodecomputes the output value of the program.The other part of supercompilation, configuration analysis, is focused on dynamically reducingthe size of the tree created by driving, by recognizing patterns among the nodes of the treeand taking steps like merging nodes together, or deleting redundant subtrees. Configurationanalysis transforms the decision tree created by driving into a decision graph, in which thepaths taken by different inputs may in some cases begin separately and then merge together.Finally, the graph that the metaprogram creates is translated back into a program, embodyingthe constraints implicit in the nodes of the graph. This program is not likely to look anythinglike the original program that the metaprogram started with, but it is guaranteed to carry outthe same function [NOTE: Give a graphical representation of the decision graph correspondingto the supercompiled binary search program for L=4, described above.].18.3.2 Supercompilation for Goal-Directed Program ModificationSupercompilation, as conventionally envisioned, is about making programs run faster; and asnoted above, it will almost certainly be useful for this purpose within CogPrime.But the process of program modeling embedded in the supercompilation process, is potentiallyof great value beyond the quest for faster software. The decision graph representation of aprogram, produced in the course of supercompilation, may be exported directly into CogPrimeas a set of logical relationships.Essentially, each node of the supercompiler’s internal decision graph looks like:Input: List LOutput: ListIf( P1(L) ) N1(L)Else If ( P2(L) ) N2(L)...Else If ( Pk(L) ) Nk(L)18.4 Self-Modification via Theorem-Proving 323where the Pi are predicates, and the Ni are schemata corresponding to other nodes of thedecision graph (children of the current node). Often the Pi are very simple, implementing forinstance numerical inequalities or Boolean equalities.Once this graph has been exported into CogPrime, it can be reasoned on, used as raw materialfor concept formation and predicate formation, and otherwise cognized. Supercompilation pureand simple does not change the I/O behavior of the input program. However, the decision graphproduced during supercompilation, may be used by CogPrime cognition in order to do so. Onethen has a hybrid program-modification method composed of two phases: supercompilation fortransforming programs into decision graphs, and CogPrime cognition for modifying decisiongraphs so that they can have different I/O behaviors fulfilling system goals even better thanthe original.Furthermore, it seems likely that, in many cases, it may be valuable to have the supercompilerfeed many different decision-graph representations of a program into CogPrime. Thesupercompiler has many internal parameters, and varying them may lead to significantly differentdecision graphs. The decision graph leading to maximal optimization, may not be the onethat leads CogPrime cognition in optimal directions.18.4 Self-Modification via Theorem-ProvingSupercompilation is a potentially very valuable tool for self-modification. If one wants to take anexisting schema and gradually improve it for speed, or even for greater effectiveness at achievingcurrent goals, supercompilation can potentially do that most excellently.However, the representation that supercompilation creates for a program is very “surfacelevel.”No one could read the supercompiled version of a program and understand what itwas doing. Really deep self-invented AI innovation requires, we believe, another level of selfmodificationbeyond that provided by supercompilation. This other level, we believe, is bestformulated in terms of theorem-proving [RV01].Deep self-modification could be achieved if CogPrime were capable of proving theorems ofa certain form: namely, theorems about the spacetime complexity and accuracy of particularcompound schemata, on average, assuming realistic probability distributions on the inputs, andmaking appropriate independence assumptions. These are not exactly the types of theorems thatare found in human-authored mathematics papers. By and large they will be nasty, complextheorems, not the sort that many human mathematicians enjoy proving or reading. But ofcourse, there is always the possibility that some elegant gem of a discovery could emerge fromthis sort of highly detailed theorem-proving work.In order to guide it in the formulation of theorems of this nature, the system will haveempirical data on the spacetime complexity of elementary schemata, and on the probabilitydistributions of inputs to schemata. It can embed these data in axioms, by asking: Assumingthe component elementary schemata have complexities within these bounds, and the input pdf(probability distribution function) is between these bounds, then what is the pdf of the complexityand accuracy of this compound schema?Of course, this is not an easy sort of question in general: one can have schemata embodyingany sort of algorithm, including complex algorithms on which computer science professors mightwrite dozens of research articles. But the system must build up its ability to prove such thingsincrementally, step by step.324 18 Advanced Self-Modification: A Possible Path to Superhuman AGIWe envision teaching the system to prove theorems via a combination of supervised learningand experiential interactive learning, using the Mizar database of mathematical theorems andproofs (or some other similar database, if one should be created) (http://mizar.org). TheMizar database consists of a set of “articles,” which are mathematical theorems and proofspresented in a complex formal language. The Mizar formal language occupies a fascinatingmiddle ground: it is high-level enough to be viably read and written by trained humans, butit can be unambiguously translated into simpler formal languages such as predicate logic orSasha.CogPrime may be taught to prove theorems by “training” it on the Mizar theorems andproofs, and by training it on custom-created Mizar articles specifically focusing on the sorts oftheorems useful for self-modification. Creating these articles will not be a trivial task: it willrequire proving simple and then progressively more complex theorems about the probabilisticsuccess of CogPrime schemata, so that CogPrime can observe one’s proofs and learned fromthem. Having learned from its training articles what strategies work for proving things aboutsimple compound schemata, it can then reason by analogy to mount attacks on slightly morecomplex schemata – and so forth.Clearly, this approach to self-modification is more difficult to achieve than the supercompilationapproach. But it is also potentially much more powerful. Even once the theorem-provingapproach is working, the supercompilation approach will still be valuable, for making incrementalimprovements on existing schema, and for the peculiar creativity that is contributed whena modified supercompiled schema is compressed back into a modified schema expression. But,we don’t believe that supercompilation can carry out truly advanced MindAgent learning orknowledge-representation modification. We suspect that the most advanced and ambitious goalsof self-modification probably cannot be achieved except through some variant of the theoremprovingapproach. If this hypothesis is true, it means that truly advanced self-modification isonly going to come after relatively advanced theorem-proving ability. Prior to this, we will haveschema optimization, schema modification, and occasional creative schema innovation. But reallysystematic, high-quality reasoning about schema, the kind that can produce an orders ofmagnitude improvement in intelligence, is going to require advanced mathematical theoremprovingability.Appendix AGlossary: :A.1 List of Specialized AcronymsThis includes acronyms that are commonly used in discussing CogPrime, OpenCog and relatedideas, plus some that occur here and there in the text for relatively ephemeral reasons.• AA: Attention Allocation• ADF: Automatically Defined Function (in the context of Genetic Programming)• AF: Attentional Focus• AGI: Artificial General Intelligence• AV: Attention Value• BD: Behavior Description• C-space: Configuration Space• CBV: Coherent Blended Volition• CEV: Coherent Extrapolated Volition• CGGP: Contextually Guided Greedy Parsing• CSDLN: Compositional Spatiotemporal Deep Learning Network• CT: Combo Tree• ECAN: Economic Attention Network• ECP: Embodied Communication Prior• EPW : Experiential Possible Worlds (semantics)• FCA: Formal Concept Analysis• FI : Fisher Information• FIM: Frequent Itemset Mining• FOI: First Order Inference• FOPL: First Order Predicate Logic• FOPLN: First Order PLN• FS-MOSES: Feature Selection MOSES (i.e. MOSES with feature selection integrated a laLIFES)• GA: Genetic Algorithms325326 A Glossary• GB: Global Brain• GEOP: Goal Evaluator Operating Procedure (in a GOLEM context)• GIS: Geospatial Information System• GOLEM: Goal-Oriented LEarning Meta-architecture• GP: Genetic Programming• HOI: Higher-Order Inference• HOPLN: Higher-Order PLN• HR: Historical Repository (in a GOLEM context)• HTM: Hierarchical Temporal Memory• IA: (Allen) Interval Algebra (an algebra of temporal intervals)• IRC: Imitation / Reinforcement / Correction (Learning)• LIFES: Learning-Integrated Feature Selection• LTI: Long Term Importance• MA: MindAgent• MOSES: Meta-Optimizing Semantic Evolutionary Search• MSH: Mirror System Hypothesis• NARS: Non-Axiomatic Reasoning System• NLGen: A specific software component within OpenCog, which provides one way of dealingwith Natural Language Generation• OCP: OpenCogPrime• OP: Operating Program (in a GOLEM context)• PEPL: Probabilistic Evolutionary Procedure Learning (e.g. MOSES)• PLN: Probabilistic Logic Networks• RCC: Region Connection Calculus• RelEx: A specific software component within OpenCog, which provides one way of dealingwith natural language Relationship Extraction• SAT: Boolean SATisfaction, as a mathematical / computational problem• SMEPH: Self-Modifying Evolving Probabilistic Hypergraph• SRAM: Simple Realistic Agents Model• STI: Short Term Importance• STV: Simple Truth VAlue• TV: Truth Value• VLTI: Very Long Term Importances• WSPS: Whole-Sentence Purely-Syntactic ParsingA.2 Glossary of Specialized Terms• Abduction: A general form of inference that goes from data describing something to ahypothesis that accounts for the data. Often in an OpenCog context, this refers to the PLNabduction rule, a specific First-Order PLN rule (If A implies C, and B implies C, thenmaybe A is B), which embodies a simple form of abductive inference. But OpenCog mayalso carry out abduction, as a general process, in other ways.• Action Selection: The process via which the OpenCog system chooses which Schema toenact, based on its current goals and context.• Active Schema Pool: The set of Schema currently in the midst of Schema Execution.A.2 Glossary of Specialized Terms 327• Adaptive Inference Control: Algorithms or heuristics for guiding PLN inference, thatcause inference to be guided differently based on the context in which the inference is takingplace, or based on aspects of the inference that are noted as it proceeds.• AGI Preschool: A virtual world or robotic scenario roughly similar to the environmentwithin a typical human preschool, intended for AGIs to learn in via interacting with theenvironment and with other intelligent agents.• Atom: The basic entity used in OpenCog as an element for building representations. SomeAtoms directly represent patterns in the world or mind, others are components of representations.There are two kinds of Atoms: Nodes and Links.• Atom, Frozen: See Atom, Saved• Atom, Realized: An Atom that exists in RAM at a certain point in time.• Atom, Saved: An Atom that has been saved to disk or other similar media, and is notactively being processed.• Atom, Serialized: An Atom that is serialized for transmission from one software processto another, or for saving to disk, etc.• Atom2Link: A part of OpenCogPrimes language generation system, that transforms appropriate Atoms into words connected vialink parser link types.• Atomspace: A collection of Atoms, comprising the central part of the memory of anOpenCog instance.• Attention: The aspect of an intelligent system’s dynamics focused on guiding which aspectsof an OpenCog system’s memory & functionality gets more computational resources at acertain point in time• Attention Allocation: The cognitive process concerned with managing the parametersand relationships guiding what the system pays attention to, at what points in time. Thisis a term inclusive of Importance Updating and Hebbian Learning.• Attentional Currency: Short Term Importance and Long Term Importance values areimplemented in terms of two different types of artificial money, STICurrency and LTICurrency.Theoretically these may be converted to one another.• Attentional Focus: The Atoms in an OpenCog Atomspace whose ShortTermImportancevalues lie above a critical threshold (the AttentionalFocus Boundary). The Attention Allocationsubsystem treats these Atoms differently. Qualitatively, these Atoms constitute thesystem’s main focus of attention during a certain interval of time, i.e. it’s a moving bubbleof attention.• Attentional Memory: A system’s memory of what it’s useful to pay attention to, in whatcontexts. In CogPrime this is managed by the attention allocation subsystem.• Backward Chainer: A piece of software, wrapped in a MindAgent, that carries out backwardchaining inference using PLN.• CIM-Dynamic: Concretely-Implemented Mind Dynamic, a term for a cognitive processthat is implemented explicitly in OpenCog (as opposed to allowed to emerge implicitly fromother dynamics). Sometimes a CIM-Dynamic will be implemented via a single MindAgent,sometimes via a set of multiple interrelated MindAgents, occasionally by other means.• Cognition: In an OpenCog context, this is an imprecise term. Sometimes this term meansany process closely related to intelligence; but more often it’s used specifically to refer tomore abstract reasoning/learning/etc, as distinct from lower-level perception and action.• Cognitive Architecture: This refers to the logical division of an AI system like OpenCoginto interacting parts and processes representing different conceptual aspects of intelligence.328 A GlossaryIt’s different from the software architecture, though of course certain cognitive architecturesand certain software architectures fit more naturally together.• Cognitive Cycle: The basic ”loop” of operations that an OpenCog system, used to controlan agent interacting with a world, goes through rapidly each ”subjective moment.” Typicallya cognitive cycle should be completed in a second or less. It minimally involves perceivingdata from the world, storing data in memory, and deciding what if any new actions needto be taken based on the data perceived. It may also involve other processes like deliberativethinking or metacognition. Not all OpenCog processing needs to take place within acognitive cycle.• Cognitive Schematic: An implication of the form ”Context AND Procedure IMPLIESgoal”. Learning and utilization of these is key to CogPrime’s cognitive process.• Cognitive Synergy: The phenomenon by which different cognitive processes, controlling asingle agent, work together in such a way as to help each other be more intelligent. Typically,if one has cognitive processes that are individually susceptible to combinatorial explosions,cognitive synergy involves coupling them together in such a way that they can help oneanother overcome each other’s internal combinatorial explosions. The CogPrime design isreliant on the hypothesis that its key learning algorithms will display dramatic cognitivesynergy when utilized for agent control in appropriate environments.• CogPrime : The name for the AGI design presented in this book, which is designed specificallyfor implementation within the OpenCog software framework (and this implementationis OpenCogPrime).• CogServer: A piece of software, within OpenCog, that wraps up an Atomspace and anumber of MindAgents, along with other mechanisms like a Scheduler for controlling theactivity of the MindAgents, and code for important and exporting data from the Atomspace.• Cognitive Equation: The principle, identified in Ben Goertzel’s 1994 book "ChaoticLogic", that minds are collections of pattern-recognition elements, that work by iterativelyrecognizing patterns in each other and then embodying these patterns as new system elements.This is seen as distinguishing mind from ”self-organization” in general, as the latteris not so focused on continual pattern recognition. Colloquially this means that ”a mind isa system continually creating itself via recognizing patterns in itself.”• Combo: The programming language used internally by MOSES to represent the programsit evolves. SchemaNodes may refer to Combo programs, whether the latter are learned viaMOSES or via some other means. The textual realization of Combo resembles LISP withless syntactic sugar. Internally a Combo program is represented as a program tree.• Composer: In the PLN design, a rule is denoted a composer if it needs premises forgenerating its consequent. See generator.• CogBuntu: an Ubuntu Linux remix that contains all required packages and tools to testand develop OpenCog.• Concept Creation: A general term for cognitive processes that create new ConceptNodes,PredicateNodes or concept maps representing new concepts.• Conceptual Blending: A process of creating new concepts via judiciously combiningpieces of old concepts. This may occur in OpenCog in many ways, among them the explicituse of a ConceptBlending MindAgent, that blends two or more ConceptNodes into a newone.• Confidence: A component of an OpenCog/PLN TruthValue, which is a scaling into theinterval [0,1] of the weight of evidence associated with a truth value. In the simplest case(of a probabilistic Simple Truth Value), one uses confidence c = n / (n+k), where n isA.2 Glossary of Specialized Terms 329the weight of evidence and k is a parameter. In the case of an Indefinite Truth Value, theconfidence is associated with the width of the probability interval.• Confidence Decay: The process by which the confidence of an Atom decreases over time,as the observations on which the Atom’s truth value is based become increasingly obsolete.This may be carried out by a special MindAgent. The rate of confidence decay is subtle andcontextually determined, and must be estimated via inference rather than simply assumeda priori.• Consciousness: CogPrime is not predicated on any particular conceptual theory of consciousness.Informally, the AttentionalFocus is sometimes referred to as the ”conscious”mind of a CogPrime system, with the rest of the Atomspace as ”unconscious” but this isjust an informal usage, not intended to tie the CogPrime design to any particular theory ofconsciousness. The primary originator of the CogPrimedesign (Ben Goertzel) tends toward panpsychism, as it happens.• Context: In addition to its general common-sensical meaning, in CogPrime the term Contextalso refers to an Atom that is used as the first argument of a ContextLink. The secondargument of the ContextLink then contains Links or Nodes, with TruthValues calculatedrestricted to the context defined by the first argument. For instance, (ContextLink USA(InheritanceLink person obese )).• Core: The MindOS portion of OpenCog, comprising the Atomspace, the CogServer, andother associated ”infrastructural” code.• Corrective Learning: When an agent learns how to do something, by having anotheragent explicitly guide it in doing the thing. For instance, teaching a dog to sit by pushingits butt to the ground.• CSDLN: (Compositional Spatiotemporal Deep Learning Network): A hierarchical patternrecognition network, in which each layer corresponds to a certain spatiotemporal granularity,the nodes on a given layer correspond to spatiotemporal regions of a given size, and thechildren of a node correspond to sub-regions of the region the parent corresponds to. JeffHawkins’s HTM is one example CSDLN, and Itamar Arel’s DeSTIN (currently used inOpenCog) is another.• Declarative Knowledge: Semantic knowledge as would be expressed in propositional orpredicate logic facts or beliefs.• Deduction: In general, this refers to the derivation of conclusions from premises usinglogical rules. In PLN in particular, this often refers to the exercise of a specific inferencerule, the PLN Deduction rule (A → B, B → C, therefore A→ C)• Deep Learning: Learning in a network of elements with multiple layers, involving feedforwardand feedback dynamics, and adaptation of the links between the elements. An exampledeep learning algorithm is DeSTIN, which is being integrated with OpenCog for perceptionprocessing.• Defrosting: Restoring, into the RAM portion of an Atomspace, an Atom (or set thereof)previously saved to disk.• Demand: In CogPrime’s OpenPsi subsystem, this term is used in a manner inherited fromthe Psi model of motivated action. A Demand in this context is a quantity whose value thesystem is motivated to adjust. Typically the system wants to keep the Demand betweencertain minimum and maximum values. An Urge develops when a Demand deviates fromits target range.• Deme: In MOSES, an ”island” of candidate programs, closely clustered together in programspace, being evolved in an attempt to optimize a certain fitness function. The idea is that330 A Glossarywithin a deme, programs are generally similar enough that reasonable syntax-semanticscorrelation obtains.• Derived Hypergraph: The SMEPH hypergraph obtained via modeling a system in termsof a hypergraph representing its internal states and their relationships. For instance, aSMEPH vertex represents a collection of internal states that habitually occur in relation tosimilar external situations. A SMEPH edge represents a relationship between two SMEPHvertices (e.g. a similarity or inheritance relationship). The terminology ”edge /vertex” isused in this context, to distinguish from the ”link / node” terminology used in the contextof the Atomspace.• DeSTIN – Deep SpatioTemporal Inference Network: A specific CSDLN created byItamar Arel, tested on visual perception, and appropriate for integration within CogPrime.• Dialogue: Linguistic interaction between two or more parties. In a CogPrime context, thismay be in English or another natural language, or it may be in Lojban or Psynese.• Dialogue Control: The process of determining what to say at each juncture in a dialogue.This is distinguished from the linguistic aspects of dialogue, language comprehension andlanguage generation. Dialogue control applies to Psynese or Lojban, as well as to humannatural language.• Dimensional Embedding: The process of embedding entities from some non-dimensionalspace (e.g. the Atomspace) into an n-dimensional Euclidean space. This can be useful in anAI context because some sorts of queries (e.g. ”find everything similar to X”, ”find a pathbetween X and Y”) are much faster to carry out among points in a Euclidean space, thanamong entities in a space with less geometric structure.• Distributed Atomspace: An implementation of an Atomspace that spans multiple computationalprocesses; generally this is done to enable spreading an Atomspace across multiplemachines.• Dual Network: A network of mental or informational entities with both a hierarchicalstructure and a heterarchical structure, and an alignment among the two structures so thateach one helps with the maintenance of the other. This is hypothesized to be a criticalemergent structure, that must emerge in a mind (e.g. in an Atomspace) in order for it toachieve a reasonable level of human-like general intelligence (and possibly to achieve a highlevel of pragmatic general intelligence in any physical environment).• Efficient Pragmatic General Intelligence: A formal, mathematical definition of generalintelligence (extending the pragmatic general intelligence), that ultimately boils down to:the ability to achieve complex goals in complex environments using limited computationalresources (where there is a specifically given weighting function determining which goalsand environments have highest priority). More specifically, the definition weighted-sums thesystem’s normalized goal-achieving ability over (goal, environment pairs), and where theweights are given by some assumed measure over (goal, environment pairs), and where thenormalization is done via dividing by the (space and time) computational resources usedfor achieving the goal.• Elegant Normal Form (ENF): Used in MOSES, this is a way of putting programs ina normal form while retaining their hierarchical structure. This is critical if one wishesto probabilistically model the structure of a collection of programs, which is a meaningfuloperation if the collection of programs is operating within a region of program space wheresyntax-semantics correlation holds to a reasonable degree. The Reduct library is used toplace programs into ENF.A.2 Glossary of Specialized Terms 331• Embodied Communication Prior: The class of prior distributions over (goal, environmentpairs), that are imposed by placing an intelligent system in an environment wheremost of its tasks involve controlling a spatially localized body in a complex world, and interactingwith other intelligent spatially localized bodies. It is hypothesized that many keyaspects of human-like intelligence (e.g. the use of different subsystems for different memorytypes, and cognitive synergy between the dynamics associated with these subsystems) areconsequences of this prior assumption. This is related to the Mind-World CorrespondencePrinciple.• Embodiment: Colloquially, in an OpenCog context, this usually means the use of an AIsoftware system to control a spatially localized body in a complex (usually 3D) world. Thereare also possible ”borderline cases” of embodiment, such as a search agent on the Internet.In a sense any AI is embodied, because it occupies some physical system (e.g. computerhardware) and has some way of interfacing with the outside world.• Emergence: A property or pattern in a system is emergent if it arises via the combinationof other system components or aspects, in such a way that its details would be very difficult(not necessarily impossible in principle) to predict from these other system components oraspects.• Emotion: Emotions are system-wide responses to the system’s current and predicted state.Dorner’s Psi theory of emotion contains explanations of many human emotions in termsof underlying dynamics and motivations, and most of these explanations make sense in aCogPrime context, due to CogPrime’s use of OpenPsi (modeled on Psi) for motivation andaction selection.• Episodic Knowledge: Knowledge about episodes in an agent’s life-history, or the lifehistoryof other agents. CogPrime includes a special dimensional embedding space only forepisodic knowledge, easing organization and recall.• Evolutionary Learning: Learning that proceeds via the rough process of iterated differentialreproduction based on fitness, incorporating variations of reproduced entities. MOSESis an explicitly evolutionary-learning-based portion of CogPrime; but CogPrime’s dynamicsas a whole may also be conceived as evolutionary.• Exemplar: (in the context of imitation learning) - When the owner wants to teach anOpenCog controlled agent a behavior by imitation, he/she gives the pet an exemplar. Toteach a virtual pet "fetch" for instance, the owner is going to throw a stick, run to it, grabit with his/her mouth and come back to its initial position.• Exemplar: (in the context of MOSES) – Candidate chosen as the core of a new deme, oras the central program within a deme, to be varied by representation building for ongoingexploration of program space.• Explicit Knowledge Representation: Knowledge representation in which individual,easily humanly identifiable pieces of knowledge correspond to individual elements in a knowledgestore (elements that are explicitly there in the software and accessible via very rapid,deterministic operations)• Extension: In PLN, the extension of a node refers to the instances of the category thatthe node represents. In contrast is the intension.• Fishgram (Frequent and Interesting Sub-hypergraph Mining): A pattern miningalgorithm for identifying frequent and/or interesting sub-hypergraphs in the Atomspace.• First-Order Inference (FOI): The subset of PLN that handles Logical Links not involvingVariableAtoms or higher-order functions. The other aspect of PLN, Higher-OrderInference, uses Truth Value formulas derived from First-Order Inference.332 A Glossary• Forgetting: The process of removing Atoms from the in-RAM portion of Atomspace, whenRAM gets short and they are judged not as valuable to retain in RAM as other Atoms. Thisis commonly done using the LTI values of the Atoms (removing lowest LTI-Atoms, or morecomplex strategies involving the LTI of groups of interconnected Atoms). May be done bya dedicated Forgetting MindAgent. VLTI may be used to determine the fate of forgottenAtoms.• Forward Chainer: A control mechanism (MindAgent) for PLN inference, that works bytaking existing Atoms and deriving conclusions from them using PLN rules, and then iteratingthis process. The goal is to derive new Atoms that are interesting according to somegiven criterion.• Frame2Atom: A simple system of hand-coded rules for translating the output of RelEx2Frame(logical representation of semantic relationships using FrameNet relationships) into Atoms.• Freezing: Saving Atoms from the in-RAM Atomspace to disk.• General Intelligence: Often used in an informal, commonsensical sense, to mean theability to learn and generalize beyond specific problems or contexts. Has been formalizedin various ways as well, including formalizations of the notion of ”achieving complex goalsin complex environments” and ”achieving complex goals in complex environments usinglimited resources.” Usually interpreted as a fuzzy concept, according to which absolutelygeneral intelligence is physically unachievable, and humans have a significant level of generalintelligence, but far from the maximally physically achievable degree.• Generalized Hypergraph: A hypergraph with some additional features, such as linksthat point to links, and nodes that are seen as ”containing” whole sub-hypergraphs. This isthe most natural and direct way to mathematically/visually model the Atomspace.• Generator: In the PLN design, a rule is denoted a generator if it can produce its consequentwithout needing premises (e.g. LookupRule, which just looks it up in the AtomSpace). Seecomposer.• Global, Distributed Memory: Memory that stores items as implicit knowledge, witheach memory item spread across multiple components, stored as a pattern of organizationor activity among them.• Glocal Memory: The storage of items in memory in a way that involves both localizedand global, distributed aspects.• Goal: An Atom representing a function that a system (like OpenCog) is supposed to spenda certain non-trivial percentage of its attention optimizing. The goal, informally speaking,is to maximize the Atom’s truth value.• Goal, Implicit: A goal that an intelligent system, in practice, strives to achieve; but thatis not explicitly represented as a goal in the system’s knowledge base.• Goal, Explicit: A goal that an intelligent system explicitly represents in its knowledgebase, and expends some resources trying to achieve. Goal Nodes (which may be Nodes or,e.g. ImplicationLinks) are used for this purpose in OpenCog.• Goal-Driven Learning: Learning that is driven by the cognitive schematic i.e. by the questof figuring out which procedures can be expected to achieve a certain goal in a certain sortof context.• Grounded SchemaNode: See SchemaNode, Grounded.• Hebbian Learning: An aspect of Attention Allocation, centered on creating and updatingHebbianLinks, which represent the simultaneous importance of the Atoms joined by theHebbianLink.A.2 Glossary of Specialized Terms 333• Hebbian Links: Links recording information about the associative relationship (cooccurrence)between Atoms. These include symmetric and asymmetric HebbianLinks.• Heterarchical Network: A network of linked elements in which the semantic relationshipsassociated with the links are generally symmetrical (e.g. they may be similarity links, orsymmetrical associative links). This is one important sort of subnetwork of an intelligentsystem; see Dual Network.• Hierarchical Network: A network of linked elements in which the semantic relationshipsassociated with the links are generally asymmetrical, and the parent nodes of a node havea more general scope and some measure of control over their children (though there may beimportant feedback dynamics too). This is one important sort of subnetwork of an intelligentsystem; see Dual Network.• Higher-Order Inference (HOI): PLN inference involving variables or higher-order functions.In contrast to First-Order Inference (FOI).• Hillclimbing: A general term for greedy, local optimization techniques, including somerelatively sophisticated ones that involve ”mildly nonlocal” jumps.• Human-Level Intelligence: General intelligence that’s ”as smart as” human general intelligence,even if in some respects quite unlike human intelligence. An informal concept,which generally doesn’t come up much in CogPrime work, but is used frequently by someother AI theorists.• Human-Like Intelligence: General intelligence with properties and capabilities broadlyresembling those of humans, but not necessarily precisely imitating human beings.• Hypergraph: A conventional hypergraph is a collection of nodes and links, where eachlink may span any number of nodes. OpenCog makes use of generalized hypergraphs (theAtomspace is one of these).• Imitation Learning: Learning via copying what some other agent is observed to do.• Implication: Often refers to an ImplicationLink between two PredicateNodes, indicatingan (extensional, intensional or mixed) logical implication.• Implicit Knowledge Representation: Representation of knowledge via having easilyhumanly identifiable pieces of knowledge correspond to the pattern of organization and/ordynamics of elements, rather than via having individual elements correspond to easily humanlyidentifiable pieces of knowledge.• Importance: A generic term for the Attention Values associated with Atoms. Most commonlythese are STI (short term importance) and LTI (long term importance) values. Otherimportance values corresponding to various different time scales are also possible. In generalan importance value reflects an estimate of the likelihood an Atom will be useful to thesystem over some particular future time-horizon. STI is generally relevant to processor timeallocation, whereas LTI is generally relevant to memory allocation.• Importance Decay: The process of Atom importance values (e.g. STI and LTI) decreasingover time, if the Atoms are not utilized. Importance decay rates may in general be contextdependent.• Importance Spreading: A synonym for Importance Updating, intended to highlight thesimilarity with ”activation spreading” in neural and semantic networks.• Importance Updating: The CIM-Dynamic that periodically (frequently) updates the STIand LTI values of Atoms based on their recent activity and their relationships.• Imprecise Truth Value: Peter Walley’s imprecise truth values are intervals [L,U], interpretedas lower and upper bounds of the means of probability distributions in an envelope334 A Glossaryof distributions. In general, the term may be used to refer to any truth value involvingintervals or related constructs, such as indefinite probabilities.• Indefinite Probability: An extension of a standard imprecise probability, comprising acredible interval for the means of probability distributions governed by a given second-orderdistribution.• Indefinite Truth Value: An OpenCog TruthValue object wrapping up an indefinite probability• Induction: In PLN, a specific inference rule (A → B, A → C, therefore B → C). In general,the process of heuristically inferring that what has been seen in multiple examples, will beseen again in new examples. Induction in the broad sense, may be carried out in OpenCogby methods other than PLN induction. When emphasis needs to be laid on the particularPLN inference rule, the phrase ”PLN Induction” is used.• Inference: Generally speaking, the process of deriving conclusions from assumptions. Inan OpenCog context, this often refers to the PLN inference system. Inference in the broadsense is distinguished from general learning via some specific characteristics, such as theintrinsically incremental nature of inference: it proceeds step by step.• Inference Control: A cognitive process that determines what logical inference rule (e.g.what PLN rule) is applied to what data, at each point in the dynamic operation of aninference process.• Integrative AGI: An AGI architecture, like CogPrime, that relies on a number of differentpowerful, reasonably general algorithms all cooperating together. This is different from anAGI architecture that is centered on a single algorithm, and also different than an AGIarchitecture that expects intelligent behavior to emerge from the collective interoperationof a number of simple elements (without any sophisticated algorithms coordinating theiroverall behavior).• Integrative Cognitive Architecture: A cognitive architecture intended to support integrativeAGI.• Intelligence: An informal, natural language concept. ”General intelligence” is one slightlymore precise specification of a related concept; ”Universal intelligence” is a fully precisespecification of a related concept. Other specifications of related concepts made in theparticular context of CogPrime research are the pragmatic general intelligence and theefficient pragmatic general intelligence.• Intension: In PLN, the intention of a node consists of Atoms representing properties ofthe entity the node represents.• Intentional memory: A system’s knowledge of its goals and their subgoals, and associationsbetween these goals and procedures and contexts (e.g. cognitive schematics).• Internal Simulation World: A simulation engine used to simulate an external environment(which may be physical or virtual), used by an AGI system as its ”mind’s eye” in orderto experiment with various action‘ q sequences and envision their consequences, or observethe consequences of various hypothetical situations. Particularly important for dealing withepisodic knowledge.• Interval Algebra: Allen Interval Algebra, a mathematical theory of the relationships betweentime intervals. CogPrime utilizes a fuzzified version of classic Interval Algebra.• IRC Learning (Imitation, Reinforcement, Correction): Learning via interaction witha teacher, involving a combination of imitating the teacher, getting explicit reinforcementsignals from the teacher, and having one’s incorrect or suboptimal behaviors guided towardbetterness by the teacher in real-time. This is a large part of how young humans learn.A.2 Glossary of Specialized Terms 335• Knowledge Base: A shorthand for the totality of knowledge possessed by an intelligentsystem during a certain interval of time (whether or not this knowledge is explicitly represented).Put differently: this is an intelligence’s total memory contents (inclusive of alltypes of memory) during an interval of time.• Language Comprehension: The process of mapping natural language speech or text intoa more ”cognitive”, largely language-independent representation. In OpenCog this has beendone by various pipelines consisting of dedicated natural language processing tools, e.g. apipeline: text → Link Parser → RelEx → RelEx2Frame → Frame2Atom Atomspace; andalternatively a pipeline Link Parser → Link2Atom → Atomspace. It would also be possibleto do language comprehension purely via PLN and other generic OpenCog processes,without using specialized language processing tools.• Language Generation: The process of mapping (largely language-independent) cognitivecontent into speech or text. In OpenCog this has been done by various pipelines consisting ofdedicated natural language processing tools, e.g. a pipeline: Atomspace → NLGen → text;or more recently Atomspace → Atom2Link → surface realization → text. It would also bepossible to do language generation purely via PLN and other generic OpenCog processes,without using specialized language processing tools.• Language Processing: Processing of human language is decomposed, in CogPrime, intoLanguage Comprehension, Language Generation, and Dialogue Control.• Learning: In general, the process of a system adapting based on experience, in a way thatincreases its intelligence (its ability to achieve its goals). The theory underlying CogPrimedoesn’t distinguish learning from reasoning, associating, or other aspects of intelligence.• Learning Server: In some OpenCog configurations, this refers to a software server thatperforms ”offline” learning tasks (e.g. using MOSES or hillclimbing), and is in communicationwith an Operational Agent Controller software server that performs real-time agentcontrol and dispatches learning tasks to and receives results from the Learning Server.• Linguistic Links: A catch-all term for Atoms explicitly representing linguistic content,e.g. WordNode, SentenceNode, CharacterNode.• Link: A type of Atom, representing a relationship among one or more Atoms. Links andNodes are the two basic kinds of Atoms.• Link Parser: A natural language syntax parser, created by Sleator and Temperley atCarnegie-Mellon University, and currently used as part of OpenCogPrime’s natural languagecomprehension and natural language generation system.• Link2Atom: A system for translating link parser links into Atoms. It attempts to resolveprecisely as much ambiguity as needed in order to translate a given assemblage of link parserlinks into a unique Atom structure.• Lobe: A term sometimes used to refer to a portion of a distributed Atomspace that livesin a single computational process. Often different lobes will live on different machines.• Localized Memory: Memory that stores each item using a small number of closelyconnectedelements.• Logic: In an OpenCog context, this usually refers to a set of formal rules for translatingcertain combinations of Atoms into ”conclusion” Atoms. The paradigm case at present is thePLN probabilistic logic system, but OpenCog can also be used together with other logics.• Logical Links: Any Atoms whose truth values are primarily determined or adjusted vialogical rules, e.g. PLN’s InheritanceLink, SimilarityLink, ImplicationLink, etc. The termisn’t usually applied to other links like HebbianLinks whose semantics isn’t primarily logic-336 A Glossarybased, even though these other links can be processed via (e.g. PLN) logical inference viainterpreting them logically.• Lojban: A constructed human language, with a completely formalized syntax and a highlyformalized semantics, and a small but active community of speakers. In principle this seemsan extremely good method for communication between humans and early-stage AGI systems.• Lojban++: A variant of Lojban that incorporates English words, enabling more flexibleexpression without the need for frequent invention of new Lojban words.• Long Term Importance (LTI): A value associated with each Atom, indicating roughlythe expected utility to the system of keeping that Atom in RAM rather than saving it todisk or deleting it. It’s possible to have multiple LTI values pertaining to different timescales, but so far practical implementation and most theory has centered on the option ofa single LTI value.• LTI: Long Term Importance• Map: A collection of Atoms that are interconnected in such a way that they tend to becommonly active (i.e. to have high STI, e.g. enough to be in the AttentionalFocus, at thesame time).• Map Encapsulation: The process of automatically identifying maps in the Atomspace,and creating Atoms that ”encapsulate” them; the Atom encapsulation a map would link toall the Atoms in the map. This is a way of making global memory into local memory, thusmaking the system’s memory glocal and explicitly manifesting the ”cognitive equation.”This may be carried out via a dedicated MapEncapsulation MindAgent.• Map Formation: The process via which maps form in the Atomspace. This need not beexplicit; maps may form implicitly via the action of Hebbian Learning. It will commonlyoccur that Atoms frequently co-occurring in the AttentionalFocus, will come to be joinedtogether in a map.• Memory Types: In CogPrimethis generally refers to the different types of memory that are embodied in different datastructures or processes in the CogPrimearchitecture, e.g. declarative (semantic), procedural, attentional, intentional, episodic, sensorimotor.• Mind-World Correspondence Principle: The principle that, for a mind to displayefficient pragmatic general intelligence relative to a world, it should display many of thesame key structural properties as that world. This can be formalized by modeling the worldand mind as probabilistic state transition graphs, and saying that the categories implicitin the state transition graphs of the mind and world should be inter-mappable via a highprobabilitymorphism.• Mind OS: A synonym for the OpenCog Core.• MindAgent: An OpenCog software object, residing in the CogServer, that carries outsome processes in interaction with the Atomspace. A given conceptual cognitive process(e.g. PLN inference, Attention allocation, etc.) may be carried out by a number of differentMindAgents designed to work together.• Mindspace: A model of the set of states of an intelligent system as a geometrical space,imposed by assuming some metric on the set of mind-states. This may be used as a tool forformulating general principles about the dynamics of generally intelligent systems.• Modulators: Parameters in the Psi model of motivated, emotional cognition, that modulatethe way a system perceives, reasons about and interacts with the world.A.2 Glossary of Specialized Terms 337• MOSES (Meta-Optimizing Semantic Evolutionary Search): An algorithm for procedurelearning, which in the current implementation learns programs in the Combo language.MOSES is an evolutionary learning system, which differs from typical genetic programmingsystems in multiple aspects including: a subtler framework for managing multiple ”demes”or ”islands” of candidate programs; a library called Reduct for placing programs in ElegantNormal Form; and the use of probabilistic modeling in place of, or in addition to, mutationand crossover as means of determining which new candidate programs to try.• Motoric: Pertaining to the control of physical actuators, e.g. those connected to a robot.May sometimes be used to refer to the control of movements of a virtual character as well.• Moving Bubble of Attention: The Attentional Focus of a CogPrime system.• Natural Language Comprehension: See Language Comprehension• Natural Language Generation: See Language Generation• Natural Language Processing (NLP): See Language Processing• NLGen: Software for carrying out the surface realization phase of natural language generation,via translating collections of RelEx output relationships into English sentences.Was made functional for simple sentences and some complex sentences; not currently underactive development, as work has shifted to the related Atom2Link approach to languagegeneration.• Node: A type of Atom. Links and Nodes are the two basic kinds of Atoms. Nodes, mathematically,can be thought of as "0-ary" links. Some types of Nodes refer to external ormathematical entities (e.g. WordNode, NumberNode); others are purely abstract, e.g. aConceptNode is characterized purely by the Links relating it to other atoms. Grounded-PredicateNodes and GroundedSchemaNodes connect to explicitly represented procedures(sometimes in the Combo language); ungrounded PredicateNodes and SchemaNodes areabstract and, like ConceptNodes, purely characterized by their relationships.• Node Probability: Many PLN inference rules rely on probabilities associated with Nodes.Node probabilities are often easiest to interpret in a specific context, e.g. the probabilityP(cat) makes obvious sense in the context of a typical American house, or in the contextof the center of the sun. Without any contextual specification, P(A) is taken to meanthe probability that a randomly chosen occasion of the system’s experience includes someinstance of A.• Novamente Cognition Engine (NCE): A proprietary proto-AGI software system, thepredecessor to OpenCog. Many parts of the NCE were open-sourced to form portions ofOpenCog, but some NCE code was not included in OpenCog; and now OpenCog includesmultiple aspects and plenty of code that was not in NCE.• OpenCog: A software framework intended for development of AGI systems, and also fornarrow-AI application using tools that have AGI applications. Co-designed with the Cog-Prime cognitive architecture, but not exclusively bound to it.• OpenCog Prime (OCP): The implementation of the CogPrime cognitive architecturewithin the OpenCog software framework.• OpenPsi: CogPrime’s architecture for motivation-driven action selection, which is basedon adapting Dorner’s Psi model for use in the OpenCog framework.• Operational Agent Controller (OAC): In some OpenCog configurations, this is a softwareserver containing a CogServer devoted to real-time control of an agent (e.g. a virtualworld agent, or a robot). Background, offline learning tasks may then be dispatched to othersoftware processes, e.g. to a Learning Server.338 A Glossary• Pattern: In a CogPrime context, the term ”pattern” is generally used to refer to a processthat produces some entity, and is judged simpler than that entity.• Pattern Mining: Pattern mining is the process of extracting an (often large) number ofpatterns from some body of information, subject to some criterion regarding which patternsare of interest. Often (but not exclusively) it refers to algorithms that are rapid or ”greedy”,finding a large number of simple patterns relatively inexpensively.• Pattern Recognition: The process of identifying and representing a pattern in somesubstrate (e.g. some collection of Atoms, or some raw perceptual data, etc.).• Patternism: The philosophical principle holding that, from the perspective of engineeringintelligent systems, it is sufficient and useful to think about mental processes in terms of(static and dynamical) patterns.• Perception: The process of understanding data from sensors. When natural language isingested in textual format, this is generally not considered perceptual. Perception may betaken to encompass both pre-processing that prepares sensory data for ingestion into theAtomspace, processing via specialized perception processing systems like DeSTIN that areconnected to the Atomspace, and more cognitive-level process within the Atomspace thatis oriented toward understanding what has been sensed.• Piagetan Stages: A series of stages of cognitive development hypothesized by developmentalpsychologist Jean Piaget, which are easy to interpret in the context of developingCogPrime systems. The basic stages are: Infantile, Pre-operational, Concrete Operationaland Formal. Post-formal stages have been discussed by theorists since Piaget and seemrelevant to AGI, especially advanced AGI systems capable of strong self-modification.• PLN: short for Probabilistic Logic Networks• PLN, First-Order: See First-Order Inference• PLN, Higher-Order: See Higher-Order Inference• PLN Rules: A PLN Rule takes as input one or more Atoms (the ”premises”, usually Links),and output an Atom that is a ”logical conclusion” of those Atoms. The truth value of theconsequence is determined by a PLN Formula associated with the Rule.• PLN Formulas: A PLN Formula, corresponding to a PLN Rule, takes the TruthValuescorresponding to the premises and produces the TruthValue corresponding to the conclusion.A single Rule may correspond to multiple Formulas, where each Formula deals with adifferent sort of TruthValue.• Pragmatic General Intelligence: A formalization of the concept of general intelligence,based on the concept that general intelligence is the capability to achieve goals in environments,calculated as a weighted average over some fuzzy set of goals and environments.• Predicate Evaluation: The process of determining the Truth Value of a predicate, embodiedin a PredicateNode. This may be recursive, as the predicate referenced internally by aGrounded PredicateNode (and represented via a Combo program tree) may itself internallyreference other PredicateNodes.• Probabilistic Logic Networks (PLN): A mathematical and conceptual framework forreasoning under uncertainty, integrating aspects of predicate and term logic with extensionsof imprecise probability theory. OpenCogPrime’s central tool for symbolic reasoning.• Procedural Knowledge: Knowledge regarding which series of actions (or action-combinations)are useful for an agent to undertake in which circumstances. In CogPrime these may belearned in a number of ways, e.g. via PLN or via Hebbian learning of Schema Maps, or viaexplicit learning of Combo programs via MOSES or hillclimbing. Procedures are representedas SchemaNodes or Schema Maps.A.2 Glossary of Specialized Terms 339• Procedure Evaluation/Execution: A general term encompassing both Schema Executionand Predicate Evaluation, both of which are similar computational processes involvingmanipulation of Combo trees associated with ProcedureNodes.• Procedure Learning: Learning of procedural knowledge, based on any method, e.g. evolutionarylearning (e.g. MOSES), inference (e.g. PLN), reinforcement learning (e.g. Hebbianlearning).• Procedure Node: A SchemaNode or PredicateNode• Psi: A model of motivated action and emotion, originated by Dietrich Dorner and furtherdeveloped by Joscha Bach, who incorporated it in his proto-AGI system MicroPsi. OpenCog-Prime’s motivated-action component, OpenPsi, is roughly based on the Psi model.• Psynese: A system enabling different OpenCog instances to communicate without usingnatural language, via directly exchanging Atom subgraphs, using a special system to mapreferences in the speaker’s mind into matching references in the listener’s mind.• Psynet Model: An early version of the theory of mind underlying CogPrime, referred toin some early writings on the Webmind AI Engine and Novamente Cognition Engine. Theconcepts underlying the psynet model are still part of the theory underlying CogPrime, butthe name has been deprecated as it never really caught on.• Reasoning: See inference• Reduct: A code library, used within MOSES, applying a collection of hand-coded rewriterules that transform Combo programs into Elegant Normal Form.• Region Connection Calculus: A mathematical formalism describing a system of basicoperations among spatial regions. Used in CogPrime as part of spatial inference to providerelations and rules to be referenced via PLN and potentially other subsystems.• Reinforcement Learning: Learning procedures via experience, in a manner explicitlyguided to cause the learning of procedures that will maximize the system’s expected futurereward. CogPrime does this implicitly whenever it tries to learn procedures that will maximizesome Goal whose Truth Value is estimated via an expected reward calculation (where”reward” may mean simply the Truth Value of some Atom defined as ”reward”). Goal-drivenlearning is more general than reinforcement learning as thus defined; and the learning thatCogPrime does, which is only partially goal-driven, is yet more general.• RelEx: A software system used in OpenCog as part of natural language comprehension, tomap the output of the link parser into more abstract semantic relationships. These moreabstract relationships may then be entered directly into the Atomspace, or they may befurther abstracted before being entered into the Atomspace, e.g. by RelEx2Frame rules.• RelEx2Frame: A system of rules for translating RelEx output into Atoms, based on theFrameNet ontology. The output of the RelEx2Frame rules make use of the FrameNet libraryof semantic relationships. The current (2012) RelEx2Frame rule-based is problematic andthe RelEx2Frame system is deprecated as a result, in favor of Link2Atom. However, theideas embodied in these rules may be useful; if cleaned up the rules might profitably beported into the Atomspace as ImplicationLinks.• Representation Building: A stage within MOSES, wherein a candidate Combo programtree (within a deme) is modified by replacing one or more tree nodes with alternative treenodes, thus obtaining a new, different candidate program within that deme. This processcurrently relies on hand-coded knowledge regarding which types of tree nodes a given treenode should be experimentally replaced with (e.g. an AND node might sensibly be replacedwith an OR node, but not so sensibly replaced with a node representing a ”kick” action).340 A Glossary• Request for Services (RFS): In CogPrime’s Goal-driven action system, a RFS is apackage sent from a Goal Atom to another Atom, offering it a certain amount of STIcurrency if it is able to deliver the goal what it wants (an increase in its Truth Value).RFS’s may be passed on, e.g. from goals to subgoals to sub-subgoals, but eventually anRFS reaches a Grounded SchemaNode, and when the corresponding Schema is executed,the payment implicit in the RFS is made.• Robot Preschool: An AGI Preschool in our physical world, intended for robotically embodiedAGIs.• Robotic Embodiment: Using an AGI to control a robot. The AGI may be running onhardware physically contained in the robot, or may run elsewhere and control the robot vianetworking methods such as wifi.• Scheduler: Part of the CogServer that controls which processes (e.g. which MindAgents)get processor time, at which point in time.• Schema: A ”script” describing a process to be carried out. This may be explicit, as in thecase of a GroundedSchemaNode, or implicit, as the case in Schema maps or ungroundedSchemaNodes.• Schema Encapsulation: The process of automatically recognizing a Schema Map in anAtomspace, and creating a Combo (or other) program embodying the process carried outby this Schema Map, and then storing this program in the Procedure Repository andassociating it with a particular SchemaNode. This translates distributed, global proceduralmemory into localized procedural memory. It’s a special case of Map Encapsulation.• Schema Execution: The process of ”running” a Grounded Schema, similar to running acomputer program. Or, phrased alternately: The process of executing the Schema referencedby a Grounded SchemaNode. This may be recursive, as the predicate referenced internally bya Grounded SchemaNode (and represented via a Combo program tree) may itself internallyreference other Grounded SchemaNodes.• Schema, Grounded: A Schema that is associated with a specific executable program(either a Combo program or, say, C++ code)• Schema Map: A collection of Atoms, including SchemaNodes, that tend to be enactedin a certain order (or set of orders), thus habitually enacting the same process. This is adistributed, globalized way of storing and enacting procedures.• Schema, Ungrounded: A Schema that represents an abstract procedure, not associatedwith any particular executable program.• Schematic Implication: A general, conceptual name for implications of the form ((ContextAND Procedure) IMPLIES Goal)• SegSim: A name for the main algorithm underlying the NLGen language generation software.The algorithm is based on segmenting a collection of Atoms into small parts, andmatching each part against memory to find, for each part, cases where similar Atomcollectionsalready have known linguistic expression.• Self-Modification: A term generally used for AI systems that can purposefully modifytheir core algorithms and representations. Formally and crisply distinguishing this sort of”strong self-modification” from ”mere” learning is a tricky matter.• Sensorimotor: Pertaining to sensory data, motoric actions, and their combination andintersection.• Sensory: Pertaining to data received by the AGI system from the outside world. In aCogPrime system that perceives language directly as text, the textual input will generallyA.2 Glossary of Specialized Terms 341not be considered as ”sensory” (on the other hand, speech audio data would be consideredas ”sensory”).• Short Term Importance: A value associated with each Atom, indicating roughly theexpected utility to the system of keeping that Atom in RAM rather than saving it to diskor deleting it. It’s possible to have multple LTI values pertaining to different time scales,but so far practical implementation and most theory has centered on the option of a singleLTI value.• Similarity: a link type indicating the probabilistic similarity between two different Atoms.Generically this is a combination of Intensional Similarity (similarity of properties) andExtensional Similarity (similarity of members).• Simple Truth Value: a TruthValue involving a pair (s,d) indicating strength (e.g. probabilityor fuzzy set membership) and confidence d. d may be replaced by other options suchas a count n or a weight of evidence w.• Simulation World: See Internal Simulation World• SMEPH (Self-Modifying Evolving Probabilistic Hypergraphs): a style of modelingsystems, in which each system is associated with a derived hypergraph• SMEPH Edge: A link in a SMEPH derived hypergraph, indicating an empirically observedrelationship (e.g. inheritance or similarity) between two• SMEPH Vertex: A node in a SMEPH derived hypergraph representing a system, indicatinga collection of system states empirically observed to arise in conjunction with the sameexternal stimuli• Spatial Inference: PLN reasoning including Atoms that explicitly reference spatial relationships• Spatiotemporal Inference: PLN reasoning including Atoms that explicitly reference spatialand temporal relationships• STI: Shorthand for Short Term Importance• Strength: The main component of a TruthValue object, lying in the interval [0,1], referringeither to a probability (in cases like InheritanceLink, SimilarityLink, EquivalenceLink,ImplicationLink, etc.) or a fuzzy value (as in MemberLink, EvaluationLink).• Strong Self-Modification: This is generally used as synonymous with Self-Modification,in a CogPrime context.• Subsymbolic: Involving processing of data using elements that have no correspondence tonatural language terms, nor abstract concepts; and that are not naturally interpreted assymbolically ”standing for” other things. Often used to refer to processes such as perceptionprocessing or motor control, which are concerned with entities like pixels or commands like”rotate servomotor 15 by 10 degrees theta and 55 degrees phi.” The distinction between”symbolic” and ”subsymbolic” is conventional in the history of AI, but seems difficult toformalize rigorously. Logic-based AI systems are typically considered ”symbolic”, yet• Supercompilation: A technique for program optimization, which globally rewrites a programinto a usually very different looking program that does the same thing. A prototypesupercompiler was applied to Combo programs with successful results.• Surface Realization: The process of taking a collection of Atoms and transforming theminto a series of words in a (usually natural) language. A stage in the overall process oflanguage generation.• Symbol Grounding: The mapping of a symbolic term into perceptual or motoric entitiesthat help define the meaning of the symbolic term. For instance, the concept ”Cat” may be342 A Glossarygrounded by images of cats, experiences of interactions with cats, imaginations of being acat, etc.• Symbolic: Pertaining to the formation or manipulation of symbols, i.e. mental entities thatare explicitly constructed to represent other entities. Often contrasted with subsymbolic.• Syntax-Semantics Correlation: In the context of MOSES and program learning morebroadly, this refers to the property via which distance in syntactic space (distance betweenthe syntactic structure of programs, e.g. if they’re represented as program trees) and semanticspace (distance between the behaviors of programs, e.g. if they’re represented assets of input/output pairs) are reasonably well correlated. This can often happen amongsets of programs that are not too widely dispersed in program space. The Reduct libraryis used to place Combo programs in Elegant Normal Form, which increases the level ofsyntax-semantics corellation between them. The programs in a single MOSES deme areoften closely enough clustered together that they have reasonably high syntax-semanticscorrelation.• System Activity Table: An OpenCog component that records information regardingwhat a system did in the past.• Temporal Inference: Reasoning that heavily involves Atoms representing temporal information,e.g. information about the duration of events, or their temporal relationship(before, after, during, beginning, ending). As implemented in CogPrime, makes use of anuncertain version of Allen Interval Algebra.• Truth Value: A package of information associated with an Atom, indicating its degreeof truth. SimpleTruthValue and IndefiniteTruthValue are two common, particular kinds.Multiple truth values associated with the same Atom from different perspectives may begrouped into CompositeTruthValue objects.• Universal Intelligence: A technical term introduced by Shane Legg and Marcus Hutter,describing (roughly speaking) the average capability of a system to carry out computablegoals in computable environments, where goal/environment pairs are weighted via the lengthof the shortest program for computing them.• Urge: In OpenPsi, an Urge develops when a Demand deviates from its target range.• Very Long Term Importance (VLTI): A bit associated with Atoms, which determineswhether, when an Atom is forgotten (removed from RAM), it is saved to disk (frozen) orsimply deleted.• Virtual AGI Preschool: A virtual world intended for AGI teaching/training/learning,bearing broad resemblance to the preschool environments used for young humans.• Virtual Embodiment: Using an AGI to control an agent living in a virtual world or gameworld, typically (but not necessarily) a 3D world with broad similarity to the everydayhuman world.• Webmind AI Engine: A predecessor to the Novamente Cognition Engine and OpenCog,developed 1997-2001 – with many similar concepts (and also some different ones) but quitedifferent algorithms and software architectureReferences 343ReferencesAABL02. Nancy Alvarado, Sam S. Adams, Steve Burbeck, and Craig Latta. Beyond the turing test: Performancemetrics for evaluating a computer simulation of the human mind. Development andLearning, International Conf. on, 0, 2002.AGBD + 08. Derek Abbott, Julio Gea-Banacloche, Paul C W Davies, Stuart Hameroff, Anton Zeilinger, JensEisert, Howard M. Wiseman, Sergey M. Bezrukov, and Hans Frauenfelder. Plenary debate: quantumeffects in biology?trivial or not? Fluctuation and Noise Letters 8(1), pp. C5ÐC26, 2008.AL03. J. R. Anderson and C. Lebiere. The newell test for a theory of cognition. Behavioral and BrainScience, 26, 2003.AL09. Itamar Arel and Scott Livingston. Beyond the turing test. IEEE Computer, 42(3):90–91, March2009.AM01. J. S. Albus and A. M. Meystel. Engineering of Mind: An Introduction to the Science of IntelligentSystems. Wiley and Sons, 2001.Ami89. Daniel J. Amit. Modeling brain function – the world of attractor neural networks. CambridgeUniversity Press, New York, USA, 1989.ARC09. I. Arel, D. Rose, and R. Coop. Destin: A scalable deep learning architecture with applicationto high-dimensional robust pattern recognition. Proc. AAAI Workshop on Biologically InspiredCognitive Architectures, 2009.ARK09a. I. Arel, D. Rose, and T. Karnowski. A deep learning architecture comprising homogeneous corticalcircuits for scalable spatiotemporal pattern inference. NIPS 2009 Workshop on Deep Learning forSpeech Recognition and Related Applications, 2009.Ark09b. Ronald Arkin. Governing Lethal Behavior in Autonomous Robots. Chapman and Hall, 2009.Arl75. P. K. Arlin. Cognitive development in adulthood: A fifth stage?, volume 11. DevelopmentalPsychology, 1975.Arm04. J. Andrew Armour. Cardiac neuronal hierarchy in health and disease. Am J Physiol Regul IntegrComp Physiol 287:, 2004.Baa97. Bernard Baars. In the Theater of Consciousness: The Workspace of the Mind. Oxford UniversityPress, 1997.Bac09. Joscha Bach. Principles of Synthetic Intelligence. Oxford University Press, 2009.Bar02. Albert-Laszlo Barabasi. Linked: The New Science of Networks. Perseus, 2002.Bat79. Gregory Bateson. Mind and Nature: A Necessary Unity. New York: Ballantine, 1979.BC94. S. Baron-Cohen. Mindblindness: An Essay on Autism and Theory of Mind. MIT Press, 1994.BDL93. Louise Barrett, Robin Dunbar, and John Lycett. Human Evolutionary Psychology. PrincetonUniversity Press, 1993.BDS03. S Ben-David and R Schuller. Exploiting task relatedness for learning multiple tasks. Proceedingsof the 16th Annual Conference on Learning Theory, 2003.BF71. J. D. Bransford and J. Franks. The abstraction of linguistic ideas. Cognitive Psychology, 2:331–350,1971.BF09. Bernard Baars and Stan Franklin. Consciousness is computational: The lida model of globalworkspace theory. International Journal of Machine Consciousness., 2009.bGBK02.1. Goertzel, Andrei Klimov Ben, and Arkady Klimov. Supercompiling java programs, 2002.BH05. Sebastian Bader and Pascal Hitzler. Dimensions of neural-symbolic integration - a structuredsurvey. In S. Artemov, H. Barringer, A. S. d’Avila Garcez, L. C. Lamb, and J. Woods., editors,We Will Show Them: Essays in Honour of Dov Gabbay, volume 1, pages 167–194. CollegePublications, 2005.Bi01. M-m Bi, G-q andPoo. Synaptic modifications by correlated activity: Hebb’s postulate revisited.Ann Rev Neurosci ; 24:139-166, 2001.Bic88. M. Bickhard. Piaget on variation and selection models: Structuralism, logical necessity, and interactivism.Human Development, 31:274–312, 1988.Bil05. Philip Bille. A survey on tree edit distance and related problems. Theoretical Computer Science,337:2005, 2005.BO09. A. Baranes and Pierre-Yves Oudeyer. R-iac: Robust intrinsically motivated active learning. Proc.of the IEEE International Conf. on Learning and Development, Shanghai, China., 33, 2009.Bol98. B. Bollobas. Modern Graph Theory. Springer, 1998.344 A GlossaryBos02. Nick Bostrom. Existential risks. Journal of Evolution and Technology, 9, 2002.Bos03. Nick Bostrom. Ethical issues in advanced artificial intelligence. In Iva Smit, editor, Cognitive,Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, volume2., pages 12–17. 2003.Bro84. J. Broughton. Not beyond formal operations, but beyond piaget. In M. Commons, F. Richards, andC. Armon, editors, Beyond Formal Operations: Late Adolescent and Adult Cognitive Development,pages 395–411. Praeger. New York, 1984.BS04. B. Bakker and Juergen Schmidhuber. Hierarchical reinforcement learning based on subgoal discoveryand subpolicy specialization. Proc. of the 8-th Conf. on Intelligent Autonomous Systems,2004.Buc03. Mark Buchanan. Small World: Uncovering Nature’s Hidden Networks. Phoenix, 2003.Bur62. C MacFarlane Burnet. The Integrity of the Body. Harvard University Press, 1962.BW88. R. W. Byrne and A. Whiten. Machiavellian Intelligence. Clarendon Press, 1988.BZ03. Selmer Bringsjord and M Zenzen. Superminds: People Harness Hypercomputation, and More.Kluwer, 2003.BZGS06. B. Bakker, V. Zhumatiy, G. Gruener, and Juergen Schmidhuber. Quasi-online reinforcementlearning for robots. Proc. of the International Conf. on Robotics and Automation, 2006.Cal96. William Calvin. The Cerebral Code. MIT Press, 1996.Car85. S. Carey. Conceptual Change in Childhood. MIT Press, 1985.Car97. R Caruana. Multitask learning. Machine Learning, 1997.Cas85. R. Case. Intellectual development: Birth to adulthood. Academic Press, 1985.Cas04. N. L. Cassimatis. Grammatical processing using the mechanisms of physical inferences. In Proceedingsof the Twentieth-Sixth Annual Conference of the Cognitive Science Society. 2004.Cas07. Nick Cassimatis. Adaptive algorithmic hybrids for human-level artificial intelligence. 2007.CB00. W. H. Calvin and D. Bickerton. Lingua ex Machina. MIT Press, 2000.CB06. Rory Conolly and Jerry Blancato. Computational modeling of the liver. NCCT BOSCReview, 2006. http://www.epa.gov/ncct/bosc_review/2006/files/07_Conolly_Liver_Model.pdf.CM07. Jie-Qi Chen and Gillian McNamee. What is Waldorf Education? Bridging: Assessment for Teachingand Learning in Early Childhood Classrooms, 2007.CP05. M. L. Commons and A. Pekker. Hierarchical complexity: A formal theory. http://www.dareassociation.org/Papers/Hierarchical%20Complexity%20-%20A%20Formal%20Theory%20(Commons%20&%20Pekker).pdf, 2005.CRK82. M. Commons, F. Richards, and D. Kuhn. Systematic and metasystematic reasoning: a case for alevel of reasoning beyond Piaget’s formal operations. Child Development, 53.:1058–1069, 1982.CS90. A. G. Cairns-Smith. Seven Clues to the Origin of Life: A Scientific Detective Story. CambridgeUniversity Press, 1990.Cse06. Peter Csermely. Weak Links: Stabilizers of Complex Systems from Proteins to Social Networks.Springer, 2006.CSG07. Subhojit Chakraborty, Anders Sandberg, and Susan A Greenfield. Differential dynamics of transientneuronal assemblies in visual compared to auditory cortex. Experimental Brain Research,1432-1106, 2007.CTS + 98. M. Commons, E. J. Trudeau, S. A. Stein, F. A. Richards, and S. R. Krause. Hierarchical complexityof tasks shows the existence of developmental stages. Developmental Review. 18, 18.:237–278, 1998.Dam00. Antonio Damasio. The Feeling of What Happens. Harvest Books, 2000.Dav84. D. Davidson. Inquiries into Truth and Interpretation. Oxford: Oxford University Press, 1984.DC02. Roberts P D and Bell C C. Spike-timing dependent synaptic plasticity in biological systems.Biological Cybernetics, 87, 392-403, 2002.Den87. D. Dennett. The Intentional Stance. Cambridge, MA: MIT Press, 1987.Den91. Daniel Dennett. Consciousness Explained. Back Bay, 1991.DG05. Hugo De Garis. The Artilect War. ETC, 2005.DOP08. Wlodzislaw Duch, Richard Oentaryo, and Michel Pasquier. Cognitive architectures: Where do wego from here? Proc. of the Second Conf. on AGI, 2008.Dör02. Dietrich Dörner. Die Mechanik des Seelenwagens. Eine neuronale Theorie der Handlungsregulation.Verlag Hans Huber, 2002.EBJ + 97. J. Elman, E. Bates, M. Johnson, A. Karmiloff-Smith, D. Parisi, and K. Plunkett. RethinkingInnateness: A Connectionist Perspective on Development. MIT Press, 1997.
References 345Ede93. Gerald Edelman. Neural darwinism: Selection and reentrant signaling in higher brain function.Neuron, 10, 1993.Elm91. J. Elman. Distributed representations, simple recurrent networks, and grammatical structure.Machine Learning, 7:195–226, 1991.EMC12. Effective-Mind-Control.com. Cellular memory in organ transplants. EffectiveMind Control,, 2012. http://www.effective-mind-control.com/cellular-memory-in-organ-transplants.html, updated Feb 1 2012.ES00. G. Engelbretsen and F. Sommers. An invitation to formal reasoning. The Logic of Terms. Aldershot:Ashgate, 2000.FB08. Stan Franklin and Bernard Baars. Possible neural correlates of cognitive processes and modulesfrom the lida model of cognition. Cognitive Computing Research Group, University of Memphis,2008. http://ccrg.cs.memphis.edu/tutorial/correlates.html.FC86. R. Fung and C. Chong. Metaprobability and Dempster-shafer in evidential reasoning. In L. Kanaland J. Lemmer. North-Holland, editors, Uncertainty in Artificial Intelligence, pages 295–302.1986.Fis80. K. Fischer. A theory of cognitive development: control and construction of hierarchies of skills.Psychological Review, 87:477–531, 1980.Fis01. Jefferson M. Fish. Race and Intelligence: Separating Science From Myth. Routledge, 2001.Fod94. J. Fodor. The Elm and the Expert. Cambridge, MA: Bradford Books, 1994.FP86. Doyne Farmer and Alan Perelson. The immune system, adaptation and machine learning. PhysicaD, v. 2, 1986.Fra06. Stan Franklin. The lida architecture: Adding new modes of learning to an intelligent, autonomous,software agent. Int. Conf. on Integrated Design and Process Technology, 2006.Fre90. R. French. Subcognition and the limits of the turing test’. Mind, 1990.Fre95. Walter Freeman. Societies of Brains. Erlbaum, 1995.FT02. G. Fauconnier and M. Turner. The Way We Think: Conceptual Blending and the Mind’s HiddenComplexities. Basic, 2002.Gar99. H Gardner. Intelligence reframed: Multiple intelligences for the 21st century. Basic, 1999.GD09. Ben Goertzel and Deborah Duong. Opencog ns: An extensible, integrative architecture for intelligenthumanoid robotics. 2009.GdG08. Ben Goertzel and Hugo de Garis. Xia-man: An extensible, integrative architecture for intelligenthumanoid robotics. pages 86–90, 2008.GE86. R. Gelman and E. Meck and s. Merkin (1986). Young children’s numerical competence. CognitiveDevelopment, 1:1–29, 1986.GEA08. Ben Goertzel and Cassio Pennachin Et Al. An integrative methodology for teaching embodiednon-linguistic agents, applied to virtual animals in second life. In Proc.of the First Conf. on AGI.IOS Press, 2008.Ger99. Michael Gershon. The Second Brain. Harper, 1999.GGC + 11. Ben Goertzel, Nil Geisweiller, Lucio Coelho, Predrag Janicic, and Cassio Pennachin. Real WorldReasoning. Atlantis, 2011.GGK02. T. Gilovich, D. Griffin, and D. Kahneman. Heuristics and biases: The psychology of intuitivejudgment. Cambridge University Press, 2002.Gib77. J. J. Gibson. The theory of affordances. In R. Shaw & J. Bransford. Erlbaum, editor, Perceiving,Acting and Knowing. 1977.Gib78. John Gibbs. Kohlberg’s moral stage theory: a Piagetian revision. Human Development, 22:89–112,1978.Gib79. J. J. Gibson. The Ecological Approach to Visual Perception. Boston: Houghton Mifflin, 1979.GIGH08. B. Goertzel, M. Ikle, I. Goertzel, and A. Heljakka. Probabilistic Logic Networks. Springer, 2008.Gil82. Carol Gilligan. In a Different Voice. Cambridge, MA: Harvard University Press, 1982.GMIH08. B. Goertzel, I. Goertzel M. Iklé, and A. Heljakka. Probabilistic Logic Networks. Springer, 2008.Goe93a. Ben Goertzel. The Evolving Mind. Plenum, 1993.Goe93b. Ben Goertzel. The Structure of Intelligence. Springer, 1993.Goe94. Ben Goertzel. Chaotic Logic. Plenum, 1994.Goe97. Ben Goertzel. From Complexity to Creativity. Plenum Press, 1997.Goe01. Ben Goertzel. Creating Internet Intelligence. Plenum Press, 2001.Goe06a. Ben Goertzel. The Hidden Pattern. Brown Walker, 2006.Goe06b. Ben Goertzel. The Hidden Pattern. Brown Walker, 2006.346 A GlossaryGoe08. Ben Goertzel. A pragmatic path toward endowing virtually-embodied ais with human-level linguisticcapability. IEEE World Congress on Computational Intelligence (WCCI), 2008.Goe09a. Ben Goertzel. Cognitive synergy: A universal principle of feasible general intelligence? In ICCI2009, Hong Kong, 2009.Goe09b. Ben Goertzel. The embodied communication prior. In Proceedings of ICCI-09, Hong Kong, 2009.Goe09c. Ben Goertzel. Opencog prime: A cognitive synergy based architecture for embodied artificialgeneral intelligence. In ICCI 2009, Hong Kong, 2009.Goe10a. Ben Goertzel. Coherent aggregated volition. Multiverse According to Ben,2010. http://multiverseaccordingtoben.blogspot.com/2010/03/coherent-aggregated-volition-toward.htm.Goe10b. Ben Goertzel. Opencogprime wikibook. 2010. http://wiki.opencog.org/w/OpenCogPrime:WikiBook.Goe10c. Ben Goertzel. Toward a formal definition of real-world general intelligence. 2010.Goe10d. Ben et al Goertzel. A general intelligence oriented architecture for embodied natural languageprocessing. In Proc. of the Third Conf. on Artificial General Intelligence (AGI-10). Atlantis Press,2010.Goo86. I. Good. The Estimation of Probabilities. Cambridge, MA: MIT Press, 1986.Gor86. R. Gordon. Folk psychology as simulation. Mind and Language. 1, 1.:158–171, 1986.GPC + 11. Ben Goertzel, Joel Pitt, Zhenhua Cai, Jared Wigmore, Deheng Huang, Nil Geisweiller, RuitingLian, and Gino Yu. Integrative general intelligence for controlling game ai in a minecraft-likeenvironment. In Proc. of BICA 2011, 2011.GPI + 10. Ben Goertzel, Joel Pitt, Matthew Ikle, Cassio Pennachin, and Rui Liu. Glocal memory: a designprinciple for artificial brains and minds. Neurocomputing, April 2010.GPPG06. Ben Goertzel, Hugo Pinto, Cassio Pennachin, and Izabela Freire Goertzel. Using dependency parsingand probabilistic inference to extract relationships between genes, proteins and malignanciesimplicit among multiple biomedical research abstracts. In Proc. of Bio-NLP 2006, 2006.GPSL03. Ben Goertzel, Cassio Pennachin, Andre’ Senna, and Moshe Looks. An integrative architecture forartificial general intelligence. In Proceedings of IJCAI 2003, Acapulco, 2003.Gre01. Susan Greenfield. The Private Life of the Brain. Wiley, 2001.GRM + 11. Erik M. Gauger, Elisabeth Rieper, John J. L. Morton, Simon C. Benjamin, and Vlatko Vedral.Sustained quantum coherence and entanglement in the avian compass. Physics Review Letters,vol. 106, no. 4, 2011.HAG07. Markert H, Knoblauch A, and Palm G. Modelling of syntactical processing in the cortex. BiosystemsMay-Jun; 89(1-3): 300-15, 2007.Ham87. Stuart Hameroff. Ultimate Computing. North Holland, 1987.Ham10. Stuart Hameroff. The Òconscious pilotÓÑdendritic synchrony moves through the brain to mediateconsciousness. Journal of Biological Physics, 2010.Hay85. Patrick Hayes. The second naive physics manifesto. In R. Shaw & J. Bransford, editor, FormalTheories of the Commonsense World. 1985.HB06. Jeff Hawkins and Sandra Blakeslee. On Intelligence. Brown Walker, 2006.Heb49. Donald Hebb. The organization of behavior. Wiley, 1949.Hey07. F. Heylighen. The Global Superorganism: an evolutionary-cybernetic model of the emerging networksociety. Social Evolution and History 6-1, 2007.HF95. P. Hayes and K. Ford. Turing test considered harmful. IJCAI-14, 1995.HG08. David Hart and Ben Goertzel. Opencog: A software framework for integrative artificial generalintelligence. In AGI, volume 171 of Frontiers in Artificial Intelligence and Applications, pages468–472. IOS Press, 2008.HHPO12. Adam Hampshire, Roger Highfield, Beth Parkin, and Adrian Owen. Fractionating human intelligence.Neuron vol. 76 issue 6, 2012.Hib02. Bill Hibbard. Superintelligent Machines. Springer, 2002.Hof79. Douglas Hofstadter. Godel, Escher, Bach: An Eternal Golden Braid. Basic, 1979.Hof95. Douglas Hofstadter. Fluid Concepts and Creative Analogies. Basic Books, 1995.Hof96. Douglas Hofstadter. Metamagical Themas. Basic Books, 1996.Hop82. J J Hopfield. Neural networks and physical systems with emergent collective computational abilities.Proc. of the National Academy of Sciences, 79:2554–2558, 1982.HOT06. G. E. Hinton, S. Osindero, and Y. Teh. A fast learning algorithm for deep belief nets. NeuralComputation, 18:1527–1554, 2006.References 347Hut95. E. Hutchins. Cognition in the Wild. MIT Press, 1995.Hut96. Edwin Hutchins. Cognition in the Wild. MIT Press, 1996.Hut05. Marcus Hutter. Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability.Springer, 2005.HZT + 02. J. Han, S. Zeng, K. Tham, M. Badgero, and J. Weng. Dav: A humanoid robot platform forautonomous mental development,. Proc. 2nd International Conf. on Development and Learning,2002.IP58. B. Inhelder and J. Piaget. The Growth of Logical Thinking from Childhood to Adolescence. BasicBooks, 1958.JL08. D. J. Jilk and C. Lebiere. and o’reilly. R. C. and Anderson, J. R. (2008). SAL: An explicitlypluralistic cognitive architecture. Journal of Experimental and Theoretical Artificial Intelligence,20:197–218, 2008.JM09. Daniel Jurafsky and James Martin. Speech and Language Processing. Pearson Prentice Hall, 2009.Joy00. Bill Joy. Why the future doesn’t need us, Wired. April 2000.Kam91. George Kampis. Self-Modifying Systems in Biology and Cognitive Science. Plenum Press, 1991.Kan64. Immanuel Kant. Groundwork of the Metaphysic of Morals. Harper and Row, 1964.Kap08. F. Kaplan. Neurorobotics: an experimental science of embodiment. Frontiers in Neuroscience,2008.KE06. J. L. Krichmar and G. M. Edelman. Principles underlying the construction of brain-based devices.In T. Kovacs and J. A. R. Marshall, editors, Adaptation in Artificial and Biological Systems, pages37–42. 2006.KK90. K. Kitchener and P. King. Reflective judgement: ten years of research. In M. Commons.Praeger. New York, editor, Beyond Formal Operations: Models and Methods in the Study ofAdolescent and Adult Thought, volume 2, pages 63–78. 1990.KLH83. Lawrence Kohlberg, Charles Levine, and Alexandra Hewer. Moral stages : a current formulationand a response to critics. Karger. Basel, 1983.Koh38. Wolfgang Kohler. The Place of Value in a World of Facts. Liveright Press, New York, 1938.Koh81. Lawrence Kohlberg. Essays on Moral Development, volume I. The Philosophy of Moral Development,1981.KS04. Adam Kahane and Peter Senge. Solving Tough Problems: An Open Way of Talking, Listening,and Creating New Realities. Berrett-Koehler, 2004.Kur06. Ray Kurzweil. The Singularity is Near. 2006.Kur12. Ray Kurzweil. How to Create a Mind. Viking, 2012.Kyb97. H. Kyburg. Bayesian and non-bayesian evidential updating. Artificial Intelligence, 31:271–293,1997.Lan05. Pat Langley. An adaptive architecture for physical agents. Proc. of the 2005 IEEE/WIC/ACMInt. Conf. on Intelligent Agent Technology, 2005.LAon. C. Lebiere and J. R. Anderson. The case for a hybrid architecture of cognition. (in preparation).LBDE90. Y. LeCun, B. Boser, J. S. Denker, and Al. Et. Handwritten digit recognition with a backpropagationnetwork. Advances in Neural Information Processing Systems, 2, 1990.LD03. A. Laud and G. Dejong. The influence of reward on the speed of reinforcement learning. Proc. ofthe 20th International Conf. on Machine Learning, 2003.Leg06a. Shane Legg. Friendly ai is bunk. Vetta Project, 2006. http://commonsenseatheism.com/wp-content/uploads/2011/02/Legg-Friendly-AI-is-bunk.pdf.Leg06b. Shane Legg. Unprovability of friendly ai. Vetta Project, 2006. http://www.vetta.org/2006/09/unprovability-of-friendly-ai/.LG90. Douglas Lenat and R. V. Guha. Building Large Knowledge-Based Systems: Representation andInference in the Cyc Project. Addison-Wesley, 1990.LH07a. Shane Legg and Marcus Hutter. A collection of definitions of intelligence. IOS, 2007.LH07b. Shane Legg and Marcus Hutter. A definition of machine intelligence. Minds and Machines, 17,2007.LLW + 05. Guang Li, Zhengguo Lou, Le Wang, Xu Li, and Walter J Freeman. Application of chaotic neuralmodel based on olfactory system on pattern recognition. ICNC, 1:378–381, 2005.LMC07a. M. H. Lee, Q. Meng, and F. Chao. Developmental learning for autonomous robots. Robotics andAutonomous Systems, 2007.LMC07b. M. H. Lee, Q. Meng, and F. Chao. Staged competence learning in developmental robotics. AdaptiveBehavior, 2007.348 A GlossaryLN00. George Lakoff and Rafael Nunez. Where Mathematics Comes From. Basic Books, 2000.Log07. Robert M. Logan. The Extended Mind. University of Toronto Press, 2007.Loo06. Moshe Looks. Competent Program Evolution. PhD Thesis, Computer Science Department, WashingtonUniversity, 2006.LRN87. John Laird, Paul Rosenbloom, and Alan Newell. Soar: An architecture for general intelligence.Artificial Intelligence, 33, 1987.LS05. J Lisman and N Spruston. Postsynaptic depolarization requirements for ltp and ltd: a critique ofspike timing-dependent plasticity. Nature Neuroscience 8, 839-41, 2005.LWML09. John Laird, Robert Wray, Robert Marinier, and Pat Langley. Claims and challenges in evaluatinghuman-level intelligent systems. Proc. of AGI-09, 2009.Mac95. D. MacKenzie. The automation of proof: A historical and sociological exploration. IEEE Annalsof the History of Computing, 17(3):7–29, 1995.Mar01. H. Marchand. Reflections on PostFormal Thought. The Genetic Epistemologist, 2001.McK03. Bill McKibben. Enough: Staying Human in an Engineered Age. Saint Martins Griffin, 2003.Met04. Thomas Metzinger. Being No One. Bradford, 2004.Min88. Marvin Minsky. The Society of Mind. MIT Press, 1988.Min07. Marvin Minsky. The Emotion Machine. 2007.MK07. Joseph Modayil and Benjamin Kuipers. Autonomous development of a grounded object ontologyby a learning robot. AAAI-07, 2007.MK08. Jonathan Mugan and Benjamin Kuipers. Towards the application of reinforcement learning toundirected developmental learning. International Conf. on Epigenetic Robotics, 2008.MK09. Jonathan Mugan and Benjamin Kuipers. Autonomously learning an action hierarchy using alearned qualitative state representation. IJCAI-09, 2009.Mon12. Maria Montessori. The Montessori Method. Frederick A. Stokes, 1912.MSV + 08. G. Metta, G. Sandini, D. Vernon, L. Natale, and F. Nori. The icub humanoid robot: an open platformfor research in embodied cognition. Performance Metrics for Intelligent Systems Workshop(PerMIS 2008), 2008.MW07. Stephen Morgan and Christopher Winship. Counterfactuals and Causal Inference. CambridgeUniversity Press, 2007.Nan08. Nanowerk. Carbon nanotube rubber could provide e-skin for robots. http://www.nanowerk.com/news/newsid=6717.php, 2008.Nei98. Dianne Miller Neilsen. Teaching Young Children, Preschool-K: A Guide to Planning Your Curriculum,Teaching Through Learning Centers, and Just About Everything Else. Corwin Press,1998.New90. Alan Newell. Unified Theories of Cognition. Harvard University press, 1990.Nie98. Dianne Miller Nielsen. Teaching Young Children, Preschool-K: A Guide to Planning Your Curriculum,Teaching Through Learning Centers, and Just About Everything Else. Corwin Press,1998.Nil09. Nils Nilsson. The physical symbol system hypothesis: Status and prospects. 50 Years of AI,Festschrift, LNAI 4850, 33, 2009.NK04. A. Nestor and B. Kokinov. Towards active vision in the dual cognitive architecture. InternationalJournal on Information Theories and Applications, 11, 2004.OK06. P. Oudeyer and F. Kaplan. Discovering communication. Connection Science, 2006.Omo08. Stephen Omohundro. The basic ai drives. Proceedings of the First AGI Conference. IOS Press,2008.Omo09. Stephen Omohundro. Creating a cooperative future. 2009. http://selfawaresystems.com/2009/02/23/talk-on-creating-a-cooperative-future/.Opa52. A. I. Oparin. The Origin of Life. Dover, 1952.Pal82. Gunter Palm. Neural Assemblies. An Alternative Approach to Artificial Intelligence. Springer,1982.Pei34. C. Peirce. Collected papers: Volume V. Pragmatism and pragmaticism. Harvard University Press.Cambridge MA., 1934.Pel05. Martin Pelikan. Hierarchical Bayesian Optimization Algorithm: Toward a New Generation ofEvolutionary Algorithms. Springer, 2005.Pen96. Roger Penrose. Shadows of the Mind. Oxford University Press, 1996.Per70. William G. Perry. Forms of Intellectual and Ethical Development in the College Years: A Scheme.Holt, Rinehart and Winston, 1970.References 349Per81. William G. Perry. Cognitive and ethical growth: The making of meaning. In Arthur W. Chickering.Jossey-Bass. San Francisco, editor, The Modern American College, pages 76–116. 1981.PH12. Zhiping Pang and Weiping Han. Regulation of synaptic functions in central nervous system byendocrine hormones and the maintenance of energy homeostasis. Bioscience Reports, 2012.Pia53. Jean Piaget. The Origins of Intelligence in Children. Routledge and Kegan Paul, 1953.Pia55. Jean Piaget. The Construction of Reality in the Child. Routledge and Kegan Paul, 1955.Pir84. Robert Pirsig. Zen and the Art of Motorcycle Maintenance. Bantam, 1984.PNR07. Karalny Patterson, Peter J. Nestor, and Timothy T. Rogers. Where do you know what you know?the representation of semantic knowledge in the human brain. Nature Reviews Neuroscience,8:976–987, 2007.PSF09. Richard Dum Peter Strick and Julie Fiez. Cerebellum and nonmotor function. Annual Review ofNeuroscience Vol. 32: 413-434, 2009.PW78. D. Premack and G. Woodruff. Does the chimpanzee have a theory of mind? Behavioral and BrainSciences, pages 515–526, 1978.QaGKKF05. R. Quian Quiroga, L. Reddy amd G. Kreiman, C. Koch, and I. Fried. Invariant visual representationby single-neurons in the human brain. Nature, 435:1102–1107, 2005.QKKF08. R. Quian Quiroga, G Kreiman, C Koch, and I. Fried. Sparse but not "grandmother-cell" codingin the medial temporal lobe. Trends in Cognitive Sciences, 12:87–91, 2008.Rav04. Ian Ravenscroft. Folk psychology as a theory, stanford encyclopedia of philosophy. http://plato.stanford.edu/entries/folkpsych-theory/, 2004.RBW92. Gagne R., L. Briggs, and W. Walter. Principles of Instructional Design. Harcourt Brace Jovanovich,1992.RCK01. J. Rosbe, R. S. Chong, and D. E. Kieras. Modeling with perceptual and memory constraints: Anepic-soar model of a simplified enroute air traffic control task. SOAR Technology Inc. Report,2001.RD06. Matthew Richardson and Pedro Domingos. Markov logic networks. Machine Learning, 2006.Rie73. K. Riegel. Dialectic operations: the final phase of cognitive development. Human Development,16.:346–370, 1973.RM95. H. L. Roediger and K. B. McDermott. Creating false memories: Remembering words not presentedin lists. Journal of Experimental Psychology: Learning, Memory, and Cognition, 21:803–814, 1995.Ros88. Israel Rosenfield. The Invention of Memory: A New View of the Brain. Basic Books, 1988.Row90. John Rowan. Subpersonalities: The People Inside Us. Routledge Press, 1990.Row11. T Rowe. Fossil evidence on origin of the mammalian brain. Science 20, 2011.RV01. Alan Robinson and Andrei Voronkov. Handbook of Automated Reasoning. MIT Press, 2001.RZDK05. Michael Rosenstein, ZvikaMarx, Tom Dietterich, and Leslie Pack Kaelbling. Transfer learningwith an ensemble of background tasks. NIPS workshop on inductive transfer, 2005.SA93. L. Shastri and V. Ajjanagadde. From simple associations to systematic reasoning: A connectionistencoding of rules, variables, and dynamic bindings using temporal synchrony. Behavioral & BrainSciences, 16-3, 1993.Sal93. Stan Salthe. Development and Evolution. MIT Press, 1993.Sam10. Alexei V. Samsonovich. Toward a unified catalog of implemented cognitive architectures. In BICA,pages 195–244, 2010.SB98. Richard Sutton and Andrew Barto. Reinforcement Learning. MIT Press, 1998.SB06. J. Simsek and A. Barto. An intrinsic reward mechanism for efficient exploration. Proc. of theTwenty-Third International Conf. on Machine Learning, 2006.SBC05. S. Singh, A. Barto, and N. Chentanez. Intrinsically motivated reinforcement learning. Proc. ofNeural Information Processing Systems 17, 2005.SC94. Barry Smith and Roberto Casati. Naive Physics: An Essay in Ontology. Philosophical Psychology,1994.Sch91a. Juergen Schmidhuber. Curious model-building control systems.. Proc. International Joint Conf.on Neural Networks, 1991.Sch91b. Juergen Schmidhuber. A possibility for implementing curiosity and boredom in model-buildingneural controllers. Proc. of the International Conf. on Simulation of Adaptive Behavior: FromAnimals to Animats, 1991.Sch95. Juergen Schmidhuber. Reinforcement-driven information acquisition in non-deterministic environments.Proc. ICANN’95, 1995.Sch02. Juergen Schmidhuber. Exploring the predictable.. Springer, 2002.350 A GlossarySch06. J. Schmidhuber. Godel machines: Fully Self-referential Optimal Universal Self-improvers. InB. Goertzel and C. Pennachin, editors, Artificial General Intelligence, pages 119–226. 2006.Sch07. Dale Schunk. Theories of Learning: An Educational Perspective. Prentice Hall, 2007.SE07. Stuart Shapiro and Al. Et. Metacognition in sneps. AI Magazine, 28, 2007.SF05. Greenfield SA and Collins T F. A neuroscientific approach to consciousness. Prog Brain Res.,2005.Sha76. G. Shafer. A Mathematical Theory of Evidence. Princeton, NJ: Princeton University Press, 1976.Shu03. Thomas R. Shultz. Computational Developmental Psychology. MIT Press, 2003.SKBB91. D Shannahoff-Khalsa, M Boyle, and M Buebel. The effects of unilateral forced nostril breathingon cognition. Int J Neurosci., 1991.Slo01. Aaron Sloman. Varieties of affect and the cogaff architecture schema. In Proceedings of theSymposium on Emotion, Cognition, and Affective Computing, AISB-01, 2001.Slo08a. Aaron Sloman. A new approach to philosophy of mathematics: Design a young explorer able todiscover ’toddler theorems’. 2008.Slo08b. Aaron Sloman. The Well-Designed Young Mathematician. Artificial Intelligence, December 2008.SM05. Push Singh and Marvin Minsky. An architecture for cognitive diversity. In Darryl Davis, editor,Visions of Mind. 2005.Sot11. Kaj Sotala. 14 objections against ai/friendly ai/the singularity answered. Xuenay.net, 2011.http://www.xuenay.net/objections.html , downloaded 3/20/11.SS74. Jean Sauvy and Simonne Suavy. The Child’s Discovery of Space: From hopscotch to mazes – anintroduction to intuitive topology. Penguin, 1974.SS03a. John F. Santore and Stuart C. Shapiro. Crystal cassie: Use of a 3-d gaming environment for acognitive agent. In Papers of the IJCAI 2003 Workshop on Cognitive Modeling of Agents andMulti-Agent Interactions, 2003.SS03b. Rudolf Steiner and S K Sagarin. What is Waldorf Education? Steiner Books, 2003.Stc00. Theodore Stcherbatsky. Buddhist Logic. Motilal Banarsidass Pub, 2000.SV99. A. J. Storkey and R. Valabregue. The basins of attraction of a new hopfield learning rule. NeuralNetworks, 12:869–876, 1999.SZ04. R. Sun and X. Zhang. Top-down versus bottom-up learning in cognitive skill acquisition. CognitiveSystems Research, 5, 2004.TC97. M. Tomasello and J. Call. Primate Cognition. Oxford University Press, 1997.TC05. Endel Tulving and R. Craik. The Oxford Handbook of Memory. Oxford U. Press, 2005.Tea06. Sebastian Thrun and et al. The robot that won the darpa grand challenge. Journal of RoboticSystems, 23-9, 2006.TM95. S. Thrun and Tom Mitchell. Lifelong robot learning. Robotics and Autonomous Systems, 1995.TS94. E. Thelen and L. Smith. A Dynamic Systems Approach to the Development of Cognition andAction. MIT Press, 1994.TS07. M. Taylor and P. Stone. Cross-domain transfer for reinforcement learning. Proc. of the 24thInternational Conf. on Machine Learning, 2007.Tur50. Alan Turing. Computing machinery and intelligence. Mind, 59, 1950.Tur77. Valentin F. Turchin. The Phenomenon of Science. Columbia University Press, 1977.TV96. Turchin and V. Supercompilation: Techniques and results. In Dines Bjorner, M. Broy, and AleksandrVasilevich Zamulin, editors, Perspectives of System Informatics. Springer, 1996.Vin93. Vernor Vinge. The coming technological singularity. VISION-21 Symposium, NASA andOhio Aerospace Institute, 1993. http://www-rohan.sdsu.edu/faculty/vinge/misc/singularity.html.Vyg86. Lev Vygotsky. Thought and Language. MIT Press, 1986.WA10. Wendell Wallach and Colin Atkins. Moral Machines. Oxford University Press, 2010.Wan95. P. Wang. Non-Axiomatic Reasoning System. PhD Thesis, Indiana University. Bloomington, 1995.Wan06. Pei Wang. Rigid Flexibility: The Logic of Intelligence. Springer, 2006.Was09. Mark Waser. Ethics for self-improving machines. In AGI-09, 2009. http://vimeo.com/3698890.Wel90. H. Wellman. The Child’s Theory of Mind. MIT Press, 1990.WH06. J. Weng and W. S. Hwangi. From neural networks to the brain: Autonomous mental development.IEEE Computational Intelligence Magazine, 2006.Who64. Benjamin Lee Whorf. Language, Thought and Reality. 1964.References 351WHZ + 00. J. Weng, W. S. Hwang, Y. Zhang, C. Yang, and R. Smith. Developmental humanoids: Humanoidsthat develop skills automatically,. Proc. the first IEEE-RAS International Conf. on HumanoidRobots, 2000.Wik11. Wikipedia. Open source governance. 2011. http://en.wikipedia.org/wiki/Open_source_governance.Win72. Terry Winograd. Understanding Natural Language. Edinburgh University Press, 1972.Wit07. David C. Witherington. The Dynamic Systems Approach as Metatheory for Developmental Psychology,Human Development. 50, 2007.Wol02. Stephen Wolfram. A New Kind of Science. Wolfram Media, 2002.WW06. Matt Williams and Jon Williamson. Combining argumentation and bayesian nets for breast cancerprognosis. Journal of Logic, Language and Information, 2006.Yud04. Eliezer Yudkowsky. Coherent extrapolated volition. Singularity Institute for AI, 2004. http://singinst.org/upload/CEV.html.Yud06. Eliezer Yudkowsky. What is friendly ai? Singularity Institute for AI, 2006. http://singinst.org/ourresearch/publications/what-is-friendly-ai.html.Zad78. L. Zadeh. Fuzzy sets as a basis for a theory of possibility. Fuzzy Sets and Systems, 1:3–28, 1978.ZPK07. Luke S Zettlemoyer, Hanna M. Pasula, and Leslie Pack Kaelbling. Logical particle filtering.Proceedings of the Dagstuhl Seminar on Probabilistic, Logical, and Relational Learning, 2007.