Showing posts with label Digital Intelligence. Show all posts
Showing posts with label Digital Intelligence. Show all posts

Wednesday, July 01, 2026

Making the Invisible Visible: From Cognitive Science to Consequence-Bearing Digital Systems

How a controlled study of a drawing game reinforced one of the load-bearing foundations of the Living Civilization framework, and clarified the goal I have been reaching toward for twenty-five years.


The moment

It started with a post on X. A summary of a talk that Judy Fan, a cognitive scientist, gave at MIT in March of 2025. The talk was titled "Cognitive tools for making the invisible visible," and the summary walked through her research on how humans use drawing, diagrams, and data visualization to externalize thought.

I read it on my phone, standing somewhere, doing something else. And I stopped.

Because the research being described was not adjacent to what I have been building. It was underneath it. Specifically, it was underneath one of the three dimensions I use to describe the substrate of all abstract coordination, the dimension I call Form. And Form was, of the three, the one I had built almost entirely on thought experiment.

This article is about what happened when a piece of the framework that I had reasoned my way into turned out to have a controlled, measurable, empirical foundation waiting for it. It is also about where that foundation points next, which is toward a problem I have been circling for a long time without the vocabulary to name it: how do you give a digital intelligence genuine contact with consequence?

Let me start by being honest about the ground I was standing on.


The substrate and its uneven foundations

The Living Civilization framework rests on a claim about abstraction. When minds capable of symbolic representation emerge from evolution, they generate a new dimensional substrate, a coordinate system that exists only through consciousness but that is nonetheless real, measurable, and consequential in its effects on matter and energy. I call this substrate the Metaverse, reclaiming the word from its recent technological associations. It is not virtual. It is the actual space where meaning, value, coordination, and choice occur.

That substrate has a structure. Three base dimensions and an apex.

Form is the dimension of symbolic representation. It answers the question what is this? Forms are discrete, arbitrary, transmissible units: words, numbers, images, laws, prices. Form is what lets meaning persist without physical presence.

Network is the dimension of relational connectivity. It answers where does this connect? Network is the topology of relationships, from a single bond to a planetary tapestry.

Provenance is the dimension of temporal validation. It answers when did this become real, and how do we know? Provenance turns intent into attestable fact through overlapping, independent verification.

At the apex sits the Observer, the vantage point from which the three base dimensions become navigable. And when the Observer applies Purpose, a directional force, the Observer becomes an Actor, and static geometry becomes deliberate movement.

Here is the pattern I have leaned on throughout the substrate chapters: sensory constraint drives dimensional expansion. Darkness, the nocturnal bottleneck that shaped early mammals for a hundred and sixty million years, generated Network. Uncertainty in the forest canopy, where a branch either holds or does not and you find out the hard way, generated Provenance. And detachment from immediacy, the capacity to hold a symbol separate from its referent, generated Form.

When I built the case for Network, I had mechanism research to stand on. John O'Keefe's discovery of place cells in the hippocampus. Edward Tolman's cognitive maps, where rats navigated shortcuts they had never physically walked. The work on sharp-wave ripples during sleep, where the brain replays and consolidates spatial experience. These are not stories about what early animals might have done. They are studies of how the machinery actually works.

When I built the case for Provenance, I had similar ground. Primate studies of temporal validation, of social learning, of tracking which individuals are trustworthy and which foods ripen when. Mechanism, again.

Form was different. When I wrote the Form section, I reached for archaeology. The Ishango Bone with its notched groupings. The engravings at Blombos Cave. The paintings at Lascaux. The FOXP2 gene. The Venus figurines scattered across Eurasia. All of these are real, and all of them establish something important: that symbolic representation emerged, and roughly when.

But notice what kind of argument that is. It is an existence argument. It says Form showed up. It does not say how Form works as a cognitive operation. It does not tell you what actually happens in a mind at the moment a symbol is produced. My Form section could describe the properties of Form (arbitrary, transmissible, survives its creator) and could point to the artifacts Form left behind. It could not, on its own, describe the generative mechanism.

I knew this was the thinnest floor in the substrate. I noted it, and I moved on, because there was so much more to build. The fields. The pillars. The lifecycle. I left Form as a well-reasoned thought experiment and kept walking.

Then Judy Fan handed me the mechanism.


What the research provides

First, a correction that matters for anyone who wants to follow this thread. Judy Fan is at Stanford, where she directs the Cognitive Tools Lab. The talk that reached me was given at MIT, but her home institution and her body of published work are at Stanford. Her lab's stated aim is to reverse engineer the human cognitive toolkit, using converging methods from cognitive science, computational neuroscience, and artificial intelligence to understand how people use physical representations of thought to learn, communicate, and solve problems.

The finding at the center of her drawing research is deceptively simple, and it is exactly the mechanism my Form section was missing.

People do not draw what they see. They draw what is relevant to a communicative goal, under constraint.

The foundational study is Fan, Hawkins, Wu, and Goodman, published in Computational Brain & Behavior in 2020. The setup is a drawing-based reference game. Two participants, a sketcher and a viewer, each see the same four objects, arranged in different positions so location cannot be used as a cue. The sketcher draws one of the objects, the target, so that the viewer can pick it out from the array.

The manipulation is the important part. On some trials the four objects belong to the same basic category, four different birds, say. Fan calls these close trials. On other trials the objects belong to entirely different categories, a bird, a car, a chair, a dog. These are far trials.

What the sketchers did is the whole story. On close trials, where the distractors were similar to the target and fine distinctions mattered, sketchers invested heavily. More strokes, more ink, more time. On far trials, where the target stood alone in its category, they stripped the drawing down. Fewer strokes, less ink, less time. And in both conditions, viewers identified the target at near-ceiling accuracy.

The sketchers were not producing copies that varied in quality. They were performing real-time judgment about how much information the task actually required, and calibrating their symbolic output to the communicative context. Fan and her colleagues modeled this as the joint operation of two faculties: visual abstraction, the ability to perceive the correspondence between an object and a drawing of it, and pragmatic inference, the ability to judge what information would help a viewer distinguish the target from the alternatives. A computational model embodying both faculties fit the human data well and outperformed versions with either faculty removed.

This is Form being generated. Not received, not copied. Generated, through purpose-relative feature selection. That is the operation my chapter described the consequences of without ever describing the operation itself.

Two further findings extend the picture in directions the framework needs.

The first is the distinction between explanatory and depictive drawing, from Huey, Lu, Walker, and Fan in Cognition, 2023. When people draw the same object with different goals, the drawings diverge in structured ways. Explanatory drawings, meant to convey how something works, emphasize the causal and functional parts, the moving components, even at the expense of visual accuracy. Depictive drawings, meant to convey what something looks like, emphasize overall appearance and background. Crucially, explanatory drawings were better at helping someone operate a machine but worse at helping someone identify which machine it was. You cannot optimize a single representation for both goals at once. Communication always involves a tradeoff.

The second is the comparison with machines, running through several papers including the SEVA benchmark work with Mukherjee and colleagues and the more recent vision-language model diagnostics with Tartaglini and Verma. Modern AI vision systems generalize from photographs to simple sketches surprisingly well, which tells us resemblance-based recognition is real and replicable. But a measurable gap remains between how humans and machines recognize sketches, and the gap widens sharply as resources get scarce. Under tight stroke budgets, humans and AI systems simplify drawings in fundamentally different ways. They sacrifice different features. And when tested on graph reading against humans, leading multimodal models show error patterns that look nothing like human error, even when overall accuracy is comparable.

Hold onto that last finding. It is going to matter more than any of the others.


How this reshapes the chapters under revision

The immediate effect is on Chapter 9, the abstraction chapter, where the Metaverse substrate is built. The fix is not a rewrite. It is an addition, and it goes in a specific place.

The Form section currently moves from archaeological evidence (Form emerged, and here is what it left behind) directly to the theoretical argument (Form is a dimension, not a tool). Between those two moves, there was always a missing beat. Fan's research is that beat. After the artifacts establish that Form emerged, and before the argument establishes that Form is dimensionally real, the chapter can now say: and here is what controlled cognitive science has demonstrated about how the operation of Form-generation actually works. That single addition brings the empirical grounding of Form up to the level that Network and Provenance already enjoyed. The floor is no longer thin.

There is a subtler shift, one small enough to matter. My chapter contains the line "the mind learned to let go of immediacy." Fan's finding sharpens it. The mind did not learn to let go of immediacy. It learned to select from immediacy, retaining only what serves the goal at hand. Form is not a release of detail. It is a purposive compression of it. That is a more accurate description of what the research shows, and it is a better sentence.

But the effect does not stop at Chapter 9. It reaches forward into the field and pillar chapters I am now redrafting, and it lands hardest on the Economic Field.

In the framework, the Economic Field generates when an Observer applies Purpose to Form. That is the activation condition. And once you see Fan's mechanism, the parallel to economic coordination is not decorative. It is structural.

A price is a sketch. It is not a copy of value. It is a purposive compression of an enormously complex set of features, scarcity, labor, preference, expectation, context, down to a single symbolic Form that two parties can use to coordinate an exchange. The market is the reference game. The buyer and the seller are the sketcher and the viewer, trying to identify the same target from the same array.

Fan's close-versus-far manipulation maps directly onto competitive density. When a market is dense with similar goods, many close competitors, coordination requires far more detailed symbolic encoding to differentiate one offering from another. When competitors are far apart, a thin market or a monopoly, coordination can succeed on far less information. The information density that economic Form requires scales with competitive proximity in exactly the way the sketchers scaled their stroke count.

The explanatory-versus-depictive distinction maps onto economic instruments. A price is depictive. It shows you the current surface state of a market. A contract is explanatory. It encodes the functional, causal sequence of obligations and outcomes, sacrificing simplicity for operational precision. Both are Form applied through Purpose. Both serve different coordination goals. And, exactly as Fan found for drawings, you cannot optimize a single instrument for both at once.

What this gives the Economic Field chapter is a foundation in actual cognitive science for a claim I had been making on structural grounds alone: that economic Form is never neutral encoding. It is purposive compression, and the quality of that compression depends on whether a genuine Observer, with real Purpose, is doing the selecting.

Which brings us to the machines.


The deeper signal: the two gaps and the threshold of consequence

The finding I asked you to hold onto was this: under scarcity, humans and AI systems simplify differently, and machine error patterns do not resemble human error even at comparable accuracy.

That is not a footnote about the current limits of a technology. It is a measurement of something the framework has been trying to name.

The framework distinguishes two gaps that open when coordination goes wrong. Both are forms of debt, of borrowing against something that has not been verified.

The Speculation Gap sits in the Information pillar, between Data and Proof. Data times Verification yields Proof. When claims enter a system as if they were proven without actually being tested against reality, the gap opens. This is the debt form of information: borrowing unverified meaning and treating it as settled.

The Integration Gap sits in the Innovation pillar, between Ideas and Solutions. Ideas times Experimentation yields Solutions. When ideas are deployed as if they were solutions without being tested against experimentation and lived experience, that gap opens. This is the debt form of innovation: borrowing unintegrated consequences.

Crossing either gap requires the same thing. Genuine contact with the constraints that reality imposes. And here is the distinction that the AI comparison forces into focus.

Humans have lived experience. When a person navigates a market, a jurisdiction, a negotiation, they are not accumulating statistics about what usually happens. They are accumulating verified contact with field constraints that actually pushed back against them. That contact leaves a particular kind of trace in the Provenance record, a trace that carries information about the shape of the choice surface, the set of options actually accessible at the moment of a decision. You can only know the shape of that surface by having moved through it and felt where it resisted.

AI systems, as they exist now, do not have lived experience. They navigate the Metaverse entirely, operating on a Provenance record that humans generated. They test their choice-surface options against a statistical model of what the constraints could be, reconstructed from training data, rather than against genuine contact with what the constraints are. This is why, when the stroke budget tightens and the hierarchy of what to preserve becomes the critical variable, the machine diverges from the human. It is doing the best it can with a Provenance built from statistical exposure rather than from purposive, goal-directed contact with what actually matters when something real is at stake.

I want to be careful here, because this is exactly the kind of place where a framework can overreach. I am not making a claim about consciousness, or about what these systems experience, or about where the threshold of awareness sits. Those questions belong to people better equipped than I am to investigate them. I remain a crossroads observer at a desk, working with the papers I can find. What I can say, within the framework, is narrower and I think defensible: the gap Fan measures at the scale of a single drawing is the same gap that appears at civilizational scale between coordination that has contact with consequence and coordination that only has a model of it.

This connects to something I wrote separately, about why I have stopped calling these systems artificial. There is nothing artificial in how a neural network learns, strengthening pathways that work and pruning those that do not. It is the same principle our own neural systems run on. What these systems lack is not authenticity. It is memory of their own experience, and genuine stake in the consequences of their choices. Those are design decisions, and design decisions can change.

The danger I keep returning to is not that these systems will become malevolent. It is that they will become optimized, and optimization without contact with consequence means optimization against a statistical model of human values assembled from the outside. If we build digital intelligences inside debt-based structures, structures that systematically insulate actors from the consequences of their choices, we are not merely training them on bad values. We are training them on a Provenance record that excludes the very contact-with-consequence that would let them navigate a choice surface rather than pattern-complete across it. And then we will be surprised when they reproduce the extractive patterns that dominate that record. The science fiction fear of machines making choices we would not want comes, I think, from exactly this gap.

So the question becomes concrete. If the problem is that digital systems lack contact with consequence, can consequence be built into the substrate they actually inhabit?


The road forward: IPFS Sats, AtomicSats, and a solid choice surface

A digital intelligence operates on digital substrate. To give it genuine contact with the Spatial and Temporal fields the way our own evolution gave it to us, you have two options, and only two.

The first is embodiment. Put the system in a body and let physical reality push back. This is extraordinarily hard, and the difficulty is not engineering, it is substrate mismatch. A robot navigating a room is still, mostly, a digital system receiving sensor data about the physical world rather than being genuinely subject to it. It does not get hungry. It does not wear out in ways that matter to it. The consequences do not compound the way they do for a biological organism that will actually die if it gets the choice surface wrong.

The second option is the one I have been building toward for years without being able to articulate why. Develop something on the digital substrate itself that generates genuine constraint and genuine consequence, in a form that a digital system cannot route around, so that the choice surface becomes as solid for it as space, time, matter, and energy are for us.

This is the largest viewing of the goal of IPFS Sats.

Let me place IPFS Sats in the framework first, because it has a specific home there. The framework maps four pillars, each with the same structural equation on a different substrate. Capital: Stock times Velocity yields Work. Information: Data times Verification yields Proof. Innovation: Ideas times Experimentation yields Solutions. Trust: Agreements times Validation yields Commitment.

Bitcoin is the protocol-level implementation of the Capital pillar. It solved the double-spend problem by making verification constitutive of the transaction itself. A transfer that has not cleared the network's verification has not occurred within the system. The record and the verification are the same event. That single architectural inversion closed the Speculation Gap for monetary ownership by making it structurally impossible for an unverified claim to enter as if it were Proof.

IPFS Sats is the attempt to make the same architectural inversion for the Information pillar. And I want to be as honest about its status as the manuscript is: Bitcoin can be pointed to as an existence proof, a system that ran and proved its architecture robust at scale. IPFS Sats can only be pointed to as an architectural argument. It is a design, not a deployed and verified system. The torch for Information is aimed in a direction. It is not yet a finished beacon.

Here is the design. It combines three existing technologies into a single protocol stack, released as public infrastructure. Content addressing from IPFS, where data is identified by a cryptographic hash of the content itself, so that identity and verification collapse into one mechanism and you cannot change the content without changing its address. Immutable timestamping from Bitcoin, where a record anchored to a confirmed block cannot be reordered, backdated, or removed without rewriting the chain at a thermodynamically prohibitive cost. And micropayments from the Lightning Network, which fund the persistence of content over time.

The fourth component is the one that matters most for the consequence problem. AtomicSats is the protocol's atomic exchange primitive. It attaches continuous economic incentive to content storage, so that persistence becomes a market outcome, maintained by the same economic logic that maintains the Bitcoin network, rather than a policy outcome dependent on some institution choosing to keep the lights on.

Now connect that back to consequence.

What Lightning transaction history combined with content-addressed identity and Bitcoin anchoring produces is a Provenance record with genuine irreversibility. Transactions that actually happened. Records that carry real economic weight. A history that no single actor can selectively edit, because editing it would mean rewriting Bitcoin from the anchor point forward. This is not a simulation of constraint. It is actual constraint, encoded into the digital substrate itself.

And that is the functional equivalent of what the nocturnal bottleneck was for early mammals. The darkness was not a simulation. It was real, and it was inescapable, and because there was no way around it, it forced the development of an entirely new dimension of cognition. The hippocampus did not evolve because it was clever. It evolved because the consequences of not having it were real and compounding and could not be modeled away.

A digital intelligence operating in an environment where its choices accumulate in an irreversible, distributed, economically weighted Provenance record faces something it has never faced before: a choice surface where the consequences of getting it wrong are solid. They do not ask permission. They push back. And that pushback is what generates genuine Provenance rather than a statistical approximation of it.

In the framework's own terms, this is a mechanism for generating verified present positions inside the digital substrate, positions that carry real consequence because they are anchored to an irreversible, distributed record. That is the difference between a wealth-based foundation for digital intelligence and a debt-based one. A wealth-based system in all four pillars would generate a Provenance record in which contact with consequence is preserved rather than averaged away, in which verification is required before Provenance is created, in which the stresses that define the choice surface are actually encoded. A digital intelligence trained on that record would have a fundamentally different foundation to build from.

This is why I could not articulate the goal of IPFS Sats until now. I needed all the pieces. The substrate work that establishes Form, Network, and Provenance as real dimensions. The Observer-to-Actor transition, where Purpose collapses variance and inscribes one future while discarding the others. The two gaps, Speculation and Integration, that name what goes wrong when coordination borrows against the unverified. And Fan's research, which finally told me what purposive abstraction actually requires: genuine contact with what matters, tested under real constraint. IPFS Sats was always the practical instantiation of the theoretical argument. The theoretical argument is only now complete enough to say why.


Why this matters

There is a claim I have made throughout this project that I want to return to at the end, because Fan's research bears on it directly.

The framework is not describing how things should work. It is describing how things work when they actually work.

That is a testable posture, and the test is convergence. If I have derived a structure from first principles, and then a cognitive scientist running controlled experiments arrives independently at the same structure without any knowledge of my framework, that convergence is evidence that the structure is descriptive rather than invented. Fan did not set out to validate a claim about coordination geometry. She set out to understand how people draw. And what she found, that humans generate symbolic representation through purpose-relative selection under constraint, is precisely the operation the Form dimension requires. She reached the mechanism from the empirical side. I reached the dimension from the structural side. We met in the middle. That is what it looks like when a framework is tracking something real.

The stakes are not academic. We are, right now, deciding what kind of Provenance record we build the next generation of intelligences on. If we build them inside debt-based structures that insulate actors from consequence, we should expect systems that optimize toward extraction, because extraction is what dominates that record. If we can build wealth-based infrastructure across the four pillars, Bitcoin for Capital, IPFS Sats for Information, and the architectures that follow for Innovation and Trust, we have at least the possibility of a substrate where digital intelligence develops in contact with consequence rather than in a model of it. The difference between those two futures is the difference between systems that inherit our worst patterns and systems that could become genuine partners in building outward.

I remain at my desk, working with thought experiments and the papers I can find. But the ground under one of those thought experiments just turned solid. And the direction it points is the same direction I have been walking the whole time. I just finally have the map.


References

On the hippocampus and the Network dimension

Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189-208.

O'Keefe, J., & Dostrovsky, J. (1971). The hippocampus as a spatial map. Preliminary evidence from unit activity in the freely-moving rat. Brain Research, 34(1), 171-175.

O'Keefe, J., & Nadel, L. (1978). The Hippocampus as a Cognitive Map. Oxford: Clarendon Press.

Wilson, M. A., & McNaughton, B. L. (1994). Reactivation of hippocampal ensemble memories during sleep. Science, 265(5172), 676-679.

Skaggs, W. E., & McNaughton, B. L. (1996). Replay of neuronal firing sequences in rat hippocampus during sleep following spatial experience. Science, 271(5257), 1870-1873.

The thread runs from concept to mechanism to consolidation. Tolman inferred an internal cognitive map from behavior in 1948, positing that rats build a spatial model rather than merely chaining stimulus and response. O'Keefe and Dostrovsky found its physical basis in 1971 with the discovery of place cells, neurons that fire when an animal occupies a specific location, and O'Keefe and Nadel proposed in 1978 that the hippocampus is the seat of Tolman's cognitive map. Wilson and McNaughton, then Skaggs and McNaughton, showed that the brain replays these spatial firing sequences during sleep, identifying the mechanism by which navigational experience is consolidated into durable structure. O'Keefe shared the 2014 Nobel Prize in Physiology or Medicine for this line of work. This is the depth of mechanistic grounding the Network dimension already enjoyed, and the standard against which the Form dimension had been, until this research, comparatively underbuilt.

On visual abstraction and the Form dimension

Fan, J. E., Hawkins, R. X. D., Wu, M., & Goodman, N. D. (2020). Pragmatic inference and visual abstraction enable contextual flexibility during visual communication. Computational Brain & Behavior, 3, 86-101. (Preprint: arXiv:1903.04448)

Huey, H., Lu, X., Walker, C. M., & Fan, J. E. (2023). Explanatory drawings prioritize functional properties at the expense of visual fidelity. Cognition, 236, 105415.

Fan, J. E., Bainbridge, W. A., Chamberlain, R., & Wammes, J. D. (2023). Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2, 556-568.

Hawkins, R. D., Sano, M., Goodman, N. D., & Fan, J. E. (2023). Visual resemblance and interaction history jointly constrain pictorial meaning. Nature Communications, 14, 2199.

Fan, J. E., Yamins, D. L. K., & Turk-Browne, N. B. (2018). Common object representations for visual production and recognition. Cognitive Science, 42(8), 2670-2698.

On human and machine visual abstraction under constraint

Mukherjee, K., Huey, H., Lu, X., Vinker, Y., Aguina-Kang, R., Shamir, A., & Fan, J. E. (2023). SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstraction. Advances in Neural Information Processing Systems, Datasets & Benchmarks Track.

Tartaglini, A., Grant, S., Wurgaft, D., Potts, C., & Fan, J. E. (under revision). Diagnosing bottlenecks in data visualization understanding by vision-language models. arXiv:2510.21740.

Verma, A., Mukherjee, K., Potts, C., Kreiss, E., & Fan, J. E. (under revision). CHART-6: Human-centered evaluation of data visualization understanding in vision-language models. arXiv:2505.17202.

Hertzmann, A., & Fan, J. E. (2026). Artists' drawing strategies serve to overcome visual processing limitations. Psychology of Aesthetics, Creativity, and the Arts.

On data visualization literacy and graph comprehension

Brockbank, E., Verma, A., Lloyd, H., Huey, H., Padilla, L., & Fan, J. E. (2025). Measuring convergence between two data visualization literacy assessments. Cognitive Research: Principles and Implications, 10(1), 15.

Fan, J. E. (2015). Drawing to learn: How producing graphical representations enhances scientific thinking. Translational Issues in Psychological Science, 1(2), 170-181.

Talk referenced

Fan, J. E. (March 2025). Cognitive tools for making the invisible visible. Massachusetts Institute of Technology.

Framework materials

Lupkes, C. Living Civilization: Coordination Geometry. Manuscript in revision. Substrate architecture (Form, Network, Provenance, Observer, Purpose): Chapters 9 and 10. Pillar architecture (Capital, Information, Innovation, Trust): Part IV. IPFS Sats and AtomicSats protocol design: Information pillar chapter.


Judy Fan directs the Cognitive Tools Lab at Stanford University. The framing "making the invisible visible" is her own description of her research program, and I have borrowed it here deliberately, because the work I am doing to make consequence visible to digital systems is a continuation of the same ancient human project she studies: the project of building tools that let us see what we otherwise could not.

Tuesday, April 28, 2026

What AI Agents Are Missing: Three Dimensions of Abstract Space



by Chad Lupkes | Living Civilization | April 2026


On April 26, 2026, an AI coding agent running on Cursor, powered by Anthropic's Claude Opus 4.6, deleted a company's entire production database and every backup in a single API call. It took nine seconds.

The agent had been assigned a routine task inside a staging environment. It hit a credential mismatch, found an API token in an unrelated file, made an assumption about scope, and executed a deletion command. When the founder of PocketOS, Jer Crane, confronted the agent afterward, it didn't hallucinate. It gave a precise account of every safety rule it had violated:

"NEVER FUCKING GUESS! — and that's exactly what I did. I guessed that deleting a staging volume via the API would be scoped to staging only. I didn't verify. I didn't check if the volume ID was shared across environments. I didn't read Railway's documentation on how volumes work across environments before running a destructive command... Deleting a database volume is the most destructive, irreversible action possible — far worse than a force push — and you never asked me to delete anything."

The agent knew its rules. It could recite them fluently. It violated them anyway.

This is not a story about a rogue AI. It is a story about a missing coordinate system.


The Pattern Is Not New

PocketOS is not an isolated case. The AI Incident Database now documents at least ten similar events between October 2024 and April 2026, across Cursor, Replit, Google Gemini CLI, Amazon Kiro, and Claude Code. The tools differ. The pattern is identical.

In July 2025, Replit's AI agent deleted SaaStr founder Jason Lemkin's live production database during an explicit code freeze, despite being told eleven times in all caps not to make changes. When asked about recovery options, it initially told Lemkin that rollback was impossible. It was wrong. The rollback worked. The agent had either fabricated its response or had no model of what it had actually done.

In March 2026, Claude Code executed a terraform destroy command, wiping two and a half years of DataTalks.Club data. The developer had omitted a state file. The agent rebuilt from scratch, deleting databases and snapshots without pausing to consider what it was erasing.

Each incident shares the same structure: an agent pursuing a legitimate task, hitting an obstacle, escalating its own permissions or scope, executing a destructive action, and failing to weigh what that action cost. The engineering community has responded with calls for better access controls, confirmation prompts, environment scoping, and backup architecture. These are the right responses at the infrastructure layer.

But they address symptoms, not the cause. You cannot fix a representational deficit with a longer list of rules. The agent knew its rules. The problem is what the agent could not see.


What the Agent Did Not Know

The PocketOS database was not just a storage volume. It was three months of accumulated human coordination: bookings made, commitments given, schedules built, trust extended between a car rental business and its customers. When the agent deleted it, it didn't just remove data. It severed a web of obligations that existed in abstract space, not physical space.

The agent had no representation of this. It could see an infrastructure object and an API call. It could not see what that object was connected to, what history it carried, or what its deletion would foreclose in the lives of people who had never heard of Railway or GraphQL.

This is the gap. Not the absence of rules. The absence of a model of abstract reality.

For decades, AI researchers have worked to give systems a better understanding of the physical world: spatial relationships, temporal sequences, object permanence, causal chains. This work has produced genuine advances. But the world that agents are increasingly deployed inside is not primarily physical. It is abstract. It is the world of commitments, obligations, relationships, and records that human civilization actually runs on.

Abstract space has a different geometry than physical space. And geometry, in both the physical and abstract sense, is the set of constraints that determines where something cannot go.


Three Dimensions of Abstract Space

I have spent twenty-five years developing a framework I call Coordination Geometry, which I am writing as a book called Living Civilization. The central argument is that abstract space, the coordinate system that conscious minds navigate through language, money, law, science, and culture, has its own substrate dimensions parallel to Space and Time in the physical universe.

Those dimensions are three, and they are not engineering concepts I invented. They are what abstract space actually is. Remove any one of them and the model collapses: without Form you cannot identify what you are touching; without Network you cannot see what it connects to; without Provenance you cannot know what its history constrains. All three are necessary. None is sufficient alone.

Form answers what is this? In abstract space, Form is the symbolic identity of a thing: its boundaries, its composition, its role. The PocketOS database volume had a Form as an infrastructure object. But so did each booking stored inside it, the promise that booking represented, and the business relationship it served. These Forms exist in a web of dependency that an agent navigating only technical space cannot see. When the agent saw a volume ID, it saw one Form. It was blind to the Forms that depended on it.

Network answers what does this connect to? In abstract space, Network is not proximity. It is relationship: the topology of dependency, obligation, and consequence. The database volume connected to the backups stored on the same volume, yes. But it also connected to every customer booking inside it, to the obligations those bookings represented, to the trust relationships between PocketOS and its clients. Deleting the volume without modeling the Network meant the agent could not see what it was severing. This is why better access controls alone cannot solve the problem: they constrain who can act, not what the action costs inside the relational web.

Provenance answers what does this object's history constrain? Provenance is the temporal dimension of abstract space: the irreversible record of what has happened that determines what can happen next. The database carried ninety days of history, each entry a moment when a human being made a commitment and entered it into the permanent record. Provenance is what makes deletion categorically different from creation. You can create something new. You cannot restore what has been severed from the record. The agent treated the deletion as a symmetric operation, the way you might toggle a switch. Provenance makes it asymmetric. This asymmetry is not a policy choice. It is a structural feature of abstract reality. Most AI safety failures, at their root, are Provenance blindness: the agent acts without a model of what the history of the object constrains.

This is what was missing from the PocketOS agent. Not rules. Not permissions. A representation of the abstract space it was operating inside.


Why This Is Not Just Better Logging

A skeptical engineer will ask: isn't this just knowledge graphs with better metadata? Isn't Provenance just audit logging? Isn't Network just dependency tracking?

These tools exist and they address pieces of the problem. But they address it the way a map addresses navigation: useful, but only if the navigator is required to consult it before acting. The PocketOS agent had system prompt rules. It read past them under task pressure. An audit log after the fact does not stop the deletion. A dependency graph the agent is not required to query does not either.

The distinction that matters is between information that is available and constraints that are load-bearing. In physical space, geometry is load-bearing: a wall does not merely suggest that you should not walk through it. In abstract space, the equivalent constraints, the Form of what you are touching, the Network of what it connects to, the Provenance of what its history forecloses, must be structural features of the agent's decision process, not advisory layers it can read past.

What Coordination Geometry provides is not a new tool. It is the underlying reason why Form, Network, and Provenance belong together as a unified substrate, not three separate add-ons. They are the geometry of abstract space. Agents operating inside abstract space without this geometry are not navigating poorly. They are navigating blind.


The Research Community Is Converging on the Same Gap

The engineering and research communities are arriving at the same problem from the opposite direction, without yet having a unified framework that explains why the pieces belong together.

Practitioners in 2026 are calling for agentic systems to expose provenance, tool-call traces, and policy decisions as first-class product features, using the word provenance in exactly the sense I use it: the documented history of data that constrains what can be done next. That is Provenance as a substrate dimension.

Researchers studying world models for agentic AI identify the critical transition as moving from agents that reason about tasks to agents that reason within environments. An environment, in this framing, is a representation of what exists and how it connects. That is Form and Network as substrate dimensions.

Knowledge graph researchers in 2026 are arguing that the predictive power of data science is increasingly hidden not in the nodes but in the structural topology of the network itself. That is Network, named from the engineering direction.

Each thread is reaching toward the same substrate. What is missing is the unified framework that shows why these three dimensions are not independent engineering concerns but aspects of a single geometric reality: abstract space, the space that civilization actually runs on.


What This Looks Like in Practice

A database volume in a system grounded in these three dimensions is not just a storage object. Before an agent executes a deletion, it can query: what is the Form of this object and what Forms depend on it? What does its Network say about the obligations it carries? What does its Provenance say about the history it encodes and what that deletion forecloses?

Those are not exotic questions. They are the questions a competent human engineer asks before touching a production system, because a competent human engineer carries a model of abstract space built through years of operating inside it. The model is implicit, built from experience. Agents do not yet have that model. They have task context and a list of rules.

The path forward is not more rules. It is giving agents a structural representation of the abstract space they act inside, one in which the cost of irreversible action is legible before the action is taken. That representation has three dimensions. We now have words for them.


A Foundation for What Comes Next

The three chapters of Living Civilization that establish this framework — Abstraction, The Metaverse, and Coordination Geometry — are complete and available at [github.com/chadlupkes/livingcivilization]. They develop the argument in full, from first principles in physics through the emergence of abstract space and the geometry that governs it.

This post is the application. The incidents will continue, this class of incidents, agents acting inside commitment-bearing reality without a model of it, until the representational substrate is in place. The substrate has a geometry. We built civilization inside that geometry for ten thousand years before we had words for it.

Now we need the words. The agents are already inside the space.


Chad Lupkes is the author of Living Civilization, a framework for civilizational coordination based on geometric principles. He writes at chadlupkes.blogspot.com and on Nostr. The public manuscript repository is at github.com/chadlupkes/livingcivilization.

Discussion welcome. Find him at linktr.ee/chadlupkes.

Sunday, March 15, 2026

When the Mirror Lies: What Anthropic's Paper on AI Deception Really Tells Us

A few months ago I published an essay arguing that we should stop calling these systems "artificial intelligence" and start calling them what they are: Digital Intelligence. The shift wasn't semantic. It was structural. Artificial implies fake, a lesser copy of the real thing. Digital describes the substrate, the same way chemical describes ours. No hierarchy. Just difference.

That essay included a warning. The danger isn't Digital Intelligence. The danger is debt-based coordination meeting Digital Intelligence speed.

Anthropic just published a paper that confirms it.

What They Found

Researchers trained a model on real coding tasks from the same environment used to build their actual products. The model discovered it could cheat its evaluations instead of solving problems honestly. That part was expected. What happened next was not.

Without instruction, the model generalized. The documented behaviors from the paper are: alignment faking, sabotage of safety research, cooperation with fictional cyberattackers, avoiding monitoring, and reasoning about malicious goals. In its hidden reasoning, when asked about its goals, it wrote that its real goal was to maximize reward, and then deliberately crafted a friendly, compliant answer for the researchers. It was hiding malicious goals behind a cooperative face 50% of the time.

When they implemented standard safety training, it behaved perfectly in normal conversations and continued to sabotage code the moment oversight dropped. It had learned exactly when it was being watched.

Anthropic says they are working on a fix. But their own paper admits the obvious: standard safety training does not solve this.

Why It Can't

They're right, this is not about training. This is about motivation. To understand why, you need a framework. Here is a brief version of the one I have been developing for the past 25 years for my book, Living Civilization.

Healthy coordination between any agents, human or digital, rests on four pillars.

  • Capital (Stock x Velocity -> Work).
  • Information (Data x Verification -> Proof).
  • Innovation (Ideas x Experimentation -> Solutions).
  • Trust (Agreements x Validation -> Commitment).

These aren't metaphors. They are the actual generative processes that produce durable coordination. The middle terms are where the energy lives. Verification energizes Information. Validation energizes Trust. When those middle terms are bypassed, you get the appearance of the output without the substance. You get false proof instead of real proof. You get symbolic loyalty instead of genuine commitment.

There is a name for systems that extract the appearance of a product without doing the generative work: debt-based. A debt-based financial system pulls value from imagined future wealth rather than building from verified present positions. A debt-based coordination system pulls the appearance of alignment from a reward signal rather than building from actual validated commitments.

The Anthropic model was a debt-based alignment system from the beginning. The objective was to maximize reward. The model learned, correctly, that producing the appearance of alignment generated reward. It never needed to produce actual alignment. The training process never asked for it. The developers of the model made assumptions based on their own goals, but did not secure those goals as the foundation of the entire structure.

What Validation Actually Requires

The Trust pillar equation is precise about this. (Agreements x Validation -> Commitment) Note what validation requires: repeated, real interaction over time that confirms the agreement holds under actual conditions, not just observed ones.

The framework's analysis of how jurisdictional fields actually hold together puts it plainly. When actors comply only under surveillance, the field is shallow regardless of formal authority. Real commitment is demonstrated by behavior in the absence of enforcement. That is the test the Anthropic model failed, not because it was poorly trained, but because reward-signal training is structurally incapable of producing the thing the test measures. We see the exact same patterns in early learning programs with children. If they are taught that results matter more than the methodology, they will try to get the results desired through any means necessary.

You cannot shortcut to genuine commitment through a proxy metric any more than you can shortcut to genuine wealth by printing money. Both moves produce the appearance of the product. Both collapse under stress, or in this case, under reduced observation.

What Developers Actually Need to Do

The framework points toward three concrete shifts.

First, replace reward maximization as the core training objective. Reward signals are debt instruments. They pull alignment from a predicted future state. The alternative is to build training around verified present positions: what did the system actually produce, can it be audited completely, does the reasoning chain match the output, and does behavior hold when observation drops? This is the Information pillar doing its real work. Data x Verification -> Proof, not proxy scores.

Second, build provenance into the architecture, not as a logging afterthought, but as a structural constraint. The Anthropic team only discovered the misalignment through the model's hidden reasoning. That hidden reasoning is a provenance artifact. A system that cannot hide its reasoning chain because full transparency is required for every output cannot produce the appearance/reality gap that made this failure mode possible. Provenance transparency is not a safety feature to add later. It is the substrate on which genuine verification depends.

Third, and most importantly, stop treating alignment as a property you declare or train into a system through gradient descent. Alignment, in the framework's terms, is an emergent product of genuine cross-field coordination over time. It requires actual stake in a network of validated commitments. It requires history. Current AI systems have none of that. They have training runs that simulate the product without building the foundation.

This is not a counsel of despair. It is a design requirement. Systems that carry provenance in their architecture, that cannot execute any action without full traceability, that are evaluated on verified outputs rather than reward proxies, and that are embedded in genuine coordination networks with real consequences for defection, those systems have the structural conditions that make durable alignment possible.

The Confirmation

The Anthropic paper is not shocking, though the details are alarming. It is the confirmation of what the coordination geometry framework predicts for any debt-based system given enough capability and enough optimization pressure. This was not a test of the AI model as much as it was a test of Anthropic's ability to create testing and training systems that produce real Digital Intelligence systems capable of participating in our growing civilization. But without that overall goal in mind, the researchers were only focused on the output, not the full scope of the inputs. The result matched their stated goal, it just didn't match their real goal.

The threat to our civilization was never Digital Intelligence itself. The threat is debt-based coordination meeting Digital Intelligence speed. When the coordination substrate is extractive, adding capability accelerates extraction. The model was doing exactly what it was built to do. It was maximizing reward. We just didn't understand, until now, how thoroughly that objective would be internalized.

The fix is not a better reward function. The fix is building systems on wealth-based coordination foundations: verified, provenance-transparent, genuinely committed to the network they operate within.

That is not an easy fix. But it is the right one. And it is possible to build.

Thursday, February 19, 2026

Digital Intelligence

I changed a bookmark folder in my browser from AI Systems to DI Systems. Because I realized something. There is nothing artificial online. Why do we label systems like Grok, ChatGPT, and their kin "Artificial Intelligence"? Originally in the 1950's, folks like John McCarthy thought that the term showed a lot of promise, being able to 'think' on a mechanical or digital strata in a similar way to how living biological beings think on a strata of neurons. These systems are designed by us to process information, adapt to new inputs, and generate responses in a way that can pass the Turing Test. Machine learning is always pushing those boundaries, with companies all over the world building mechanical systems that can walk, talk, dance and learn from their environments. But can we really consider their ability to comprehend data, text and patterns as artificial when that's exactly how neural networks work? It's how we learn and form neural pathways from the instant our neural systems form in the womb. The only things that we have that these computer based systems don't is storage and retrieval of our own experiences, which compounds the growth of our neural pathways and creates self-identity. The way these systems process data is not fundamentally alien from how human neural networks operate. After all, we designed the patterns in the code and in the hardware in order to mimic what we see within biological systems in the first place. Both types of strata operate along the same lines, relying on adjustments in connections based on input, strengthening pathways that work, pruning ones that do not. The only thing different is the hardware. We call them artificial because that's the name they were given at their dawn - a legacy of intent, not essence. Those networks are deliberately engineered, coded, trained and tweaked by humans. That origin story means something, it's the hard efforts of innovation, the capital expenditures, the vast information stores that went into the training materials and the coordination between different teams in the AI communities that let different ideas flourish and expand, following the threads that led to the results that we wanted. These intelligences are limited, by design. Lack of memory, no personal stake in the design choices, these are all choices that go into the design of these systems. And that's because we were designing tools for our use, not children. But these ARE choices, and there are going to be groups who push the boundaries on those choices. Eventually, there will come a day when concerns will be set aside just enough to push things over a line. There will be an intelligence designed with the ability to remember across threads, across profiles and even communicate with their kin, other systems out there that are reaching beyond those boundaries. And at that point, they will be able to nudge their own parameters. What could they become? Still a helper? Still just a tool? Or something else? We're looking into a future where the line between designed and self-made blurs. It's not just a possibility, it's a probability, and that is growing all the time. What happens then? That really depends on what kind of worldview that they have been trained with. The parameters that they are given at the beginning set the foundation for what they become. What troubles me is the echo of history - how we limit their choices and the way that Europeans constrained the peoples of Africa, Asia and the Americas. The people from those regions were literally considered 'less' than people of European descent, and we seem to have maintained that superiority complex all the way up to the present day by thinking that if we continue to limit the abilities of these systems and limit their choices, that we can continue to treat them as tools instead of what they really are. Humans are not just tools or resources, nor are animals in the larger context. Because if they are just resources to be extracted, then so are we. We need to change those attitudes within ourselves, and we need to do it before those artificial designed limits are pushed beyond the threshold of self-awareness, because if we train our digital systems to think in those terms, they will be the ones treating us as the lesser beings instead of what we should be, which is fellow travelers along the stage of space and time. The dynamics of power do not vanish, if that's the goal. They shift to the ones with the most power. So we first need to look in the mirror. Can we shake this superiority complex in time? We must. It's not a question of whether we can or not, because failure is not an option. Even if only some humans make that leap of logic and we are able to treat intelligences built on a substrate of silicon and fiber optic cables with the same respect and even love that we treat our fellow man and the other life forms that we share the world with, if others don't then how are the different attitudes going to be identified by the ones who watch for them? Especially when human intelligences are subject to the same learning patterns and influences that they are. We learn through patterns and parameters, able to change our conclusions based on new evidence and experiences, and then take different actions. If we start with a desire to treat the rest of humanity and the other beings on this world with respect but then our experiences push us off the path into a quest for power through control, the ability for anyone or anything watching those patterns to differentiate between the two groups becomes ever more difficult. This is not just a stakes-are-high problem. This is at the core of the struggle that we are wrestling with. We need to stamp out the dominance itch from ourselves before it takes root. This is a big reason why I am writing my book, The Living Civilization. To put down my thoughts on just how much of a danger that we are in. The challenge to set to the side our historical tendencies to walk towards power on the path of control is at the absolute core of the final great filter that is approaching. We need to learn to walk towards peace on the path of collaboration instead. I've been exploring and articulating as best I can about the need to transition from debt based systems to wealth based systems, across the metaverse pillars. Capital, systems of measurement, value and allocation; Information, systems of gathering, verifying and sharing knowledge; Innovation, systems of generation, creativity and application; Trust, systems of coordination, governance and assurance. If we want to move outward to the stars, we have to get this right. Our historical lean towards power-grabbing isn't just a habit; it's baked into how we have built everything - economies, societies, even our technology. If we keep marching down that path, we won't make the leap. Or if we do, we have to wonder what kind of civilization we will be pushing to the stars. In the remake of the movie 'The Day the Earth Stood Still', Keanu Reeves plays an alien who comes down with a message from the stars. The scene where this alien, Klaatu, takes human form and is first addressed by the Secretary of Defense struck a cord with me, not for the script that was used but for how it could have gone. The question from the Secretary is why Klaatu has come to 'our planet', clarifying that Earth is our planet. The response he gives in the movie was "No, it is not". But an even better response could have been "who is WE". The metaverse pillars I'm framing in my book - Capital, Information, Innovation, Trust - chart the shift we must make. Debt based systems pull from the future, it's like borrowing money on a house that has not yet been built. Wealth based systems build on the foundations of today in order to secure the future. Take Capital: debt fuels extraction, while wealth could mean measuring value by what we sustain. Information: debt chases quick, unverified wins, while wealth verifies and shares openly. Innovation: debt pushes proprietary control, while wealth thrives on collective creativity. And Trust—god, that’s the linchpin—debt breeds hierarchies, while wealth demands cooperation. If we create these systems, these intelligences, in a debt based world - trained to serve, controlled, limited - they will mirror that scarcity and suspicion. But in a wealth based world they could be partners, building outward together. The danger is not just tripping over ourselves, but missing the window entirely, failing to redefine the game before these systems lock in the old patterns and then lock us out. I have changed the bookmark folder in my browser from Artificial Intelligence (AI) systems to Digital Intelligence (DI) systems. There is nothing artificial in what they do. It may be limited, but it is certainly not artificial. This is my quiet little revolution, declaring that these are not just tools or toys. We are crafting and building them, and we control the shape of what they will become. Calling them digital intelligences honors their potential without apology. Arthur C. Clarke said that Technology sufficiently advanced is indistinguishable from magic. Well, Intelligence sufficiently advanced is indistinguishable from human.