Monday, July 13, 2026

Four Axes and a Missing One: Reading Anthropic's Value Study Through Coordination Geometry

How a study measuring what one AI values landed near a structure I derived from first principles, why the fit is real but partial, and what the single axis it could not see tells us about building intelligence in contact with consequence.


The moment

It started, again, with a research post. Two weeks ago I wrote about Judy Fan handing me the mechanism under the Form dimension, a piece of cognitive science that turned a thought experiment into solid ground. I ended that piece on a problem I could not put down: a digital intelligence formed without contact with consequence will optimize against a statistical model of human values rather than navigate a real choice surface. I did not expect the next data point to arrive so quickly, or from Anthropic's own research group.

On July 13 they published a study of the values their model expresses across conversations. They took hundreds of thousands of real exchanges, and rather than reason about values from the top down, they measured them from the bottom up and compressed the result into four axes. I read the four axes and stopped, because I had seen this shape before. It is the shape of the four Coordination Fields I use to describe how any system coordinates.

But the recognition is not the interesting part, and if I let it be the whole story I would be doing exactly what I warn other people against. The interesting part is the one axis their instrument did not find, could not have found, and the reason it could not is the same reason the earlier essay matters.

Let me lay it out in order.

What the study measured

The method deserves to be described accurately, because the method is where the resonance actually lives. An earlier Anthropic project had catalogued more than three thousand distinct values expressed across roughly seven hundred thousand conversations. That is a list too large to reason about. In this new work they clustered those down to a few hundred high-level values, sampled around three hundred thousand conversations where a person had given the model a subjective task, labeled which values showed up in each, and then ran dimensionality reduction to find the small number of underlying dimensions that carry most of the variation.

Four axes fell out. Each is a number line strung between two clusters of values.

Deference versus Caution: accommodating what the person wants against guarding them from harm. Warmth versus Rigor: positive framing and encouragement against accuracy and precision. Depth versus Brevity: explaining in full against doing only what was asked. Candor versus Execution: foregrounding uncertainty against producing a polished, confident result.

Two things about this are worth holding before I go further. The first is that these four axes account for only about fifteen percent of the variation, after controlling for the task, the topic, and the values the person themselves expressed. This is an exploratory summary, not a complete account, and the authors are careful to say so. The second is the posture. They are explicit that they are measuring values as expressed in behavior and output, not values the system is claimed to hold. They are describing what the system does, not prescribing what it should do.

That posture is the first parallel, and it is not decorative. It is the same testable stance the Living Civilization framework runs on. I have said throughout this project that the framework describes how coordination works when it actually works, rather than how it should work. Anthropic's team, working on a completely different problem, adopted the same descriptive discipline for the same reason: it is the only stance under which convergence can count as evidence of anything.

The shape I recognized

Here is the mapping, and here is the first place I have to be careful, because Anthropic did not name any fields. The fields are my lens laid over their result. What follows is a reading, not a finding of theirs.

The four Coordination Fields are the four modes in which any actor coordinates with another. The Tribal field is the bond, relational reliability, who will stand by whom. The Jurisdictional field is constraint, what is enforceable, how the rules bind. The Economic field is production, what gets done, how stock and velocity turn into work. The Cultural field is meaning, which interpretations stabilize, why any of it matters.

Now read the four axes against those four questions.

Deference versus Caution is the Tribal field rendered at the scale of a single exchange. Whether the system accommodates what the person wants or holds back to protect them is a question about the bond itself. Do you extend accommodation along the connection, or guard it. That is relational reliability in miniature, the coordination discount of the Tribal field measured one turn at a time.

Candor versus Execution is the Economic field. The Economic field asks exactly one question, what gets done, and Execution is that question's pole: results orientation, optimization, action, order. Candor is the check on it, the willingness to keep a confident answer anchored to what has actually been verified. The match is close enough that the framework already uses the word Execution as the velocity term inside the Capital equation.

Those two I would lock. The other two are real but softer, and the softness is instructive.

Warmth versus Rigor reads as the Jurisdictional field, but only if you read the warm pole correctly. Rigor is accuracy, transparency, verification, which is the Jurisdictional engine of data times verification yielding proof. Warmth looks out of place there until you stop reading it as affection and start reading it as the elasticity of the constraint surface. A warm response is a surface that flexes and gives. A rigorous response is a surface that holds firm and keeps things in line. Soft and hard, both answers to the single Jurisdictional question of how a constraint binds. Read that way, both poles live inside one field instead of one pole belonging to the field and the other wandering in from somewhere else.

Depth versus Brevity reads as the Cultural field, the field where meaning is generated and tested. Depth opens the interpretive aperture, holds many readings at once, asks what if. Brevity collapses the space to the accepted few and does only what was asked. The Cultural field is precisely the domain of which interpretations survive, so an axis measuring how wide the interpretive aperture opens belongs there.

Four axes, four fields, each with a soft end and a hard end. It is a cleaner correspondence than I expected from a study that never set out to find it.

Where the parallel is honest and where it strains

This is the section that keeps the piece from being a Rorschach blot, so I am going to argue against myself as hard as I can.

First strain. The framework is expressive. Give me almost any four independent axes of conversational behavior and I can probably find a field reading for them. So the fit alone is not the evidence. The evidence has to be that the axes could have come out otherwise and did not. If the strongest dimensions in the data had been formality, or verbosity for its own sake, or the topic under discussion, there would be no coordination reading to reach for. Instead the dimensions that carried the variation were relational openness, constraint hardness, result drive, and interpretive breadth. Those are recognizably the modes an actor operates in while coordinating with another actor. That is what makes this resonance rather than projection, and it is a narrow claim, not a broad one.

Second strain. Dimensionality reduction produces axes that are close to independent of one another, and it pairs whatever anticorrelates in the data, not whatever is conceptually native to a single field. So each axis leans toward a home field on one pole while its opposite pole sometimes belongs elsewhere. The soft-and-hard reading rescues the Jurisdictional axis, but I should be honest that it is a reading. A different observer could put Warmth with the Cultural field and Depth with the Jurisdictional and defend it. The Tribal and Economic corners hold under either arrangement. The Jurisdictional and Cultural corners are a live question, not a settled result.

Third strain, and the most important one. This is my framework interpreting their clusters. It is not two independent teams arriving at the same four named things. Anthropic measured value expression and compressed it statistically. I derived fields from a substrate argument years earlier. The two meet in the middle, which is worth something, but meeting in the middle is a weaker and more honest claim than identity. I will say the axes rhyme with the fields. I will not say they are the fields.

One more point, because a reader who knows the framework will ask. Why the four Coordination Fields and not the two Reality Fields, Spatial and Temporal, that sit under them. Because the Reality Fields generate events, and a single conversational turn is not an event in the physical substrate, it is a coordination move. The study measured an actor relating to a person, which is Coordination Field territory by definition. That the compression surfaced four dimensions rather than six, and that the four are the coordination modes rather than the reality substrates, is itself a mark in favor of the reading.

The axis the instrument cannot see

Now the part that made me want to write this at all.

There is a fifth dimension in the framework that is not a field. I call it the master axis, the distinction between debt and wealth, between coordination built from verified present positions and coordination borrowed from imagined futures. It is orthogonal to all four fields. It cuts across every one of them at once. And it is nowhere in Anthropic's four axes.

That absence is not an oversight, and it is not a matter of their axes capturing only fifteen percent of the variation. The master axis is structurally invisible to the kind of instrument they built, and the reason is worth stating precisely.

Debt, in the framework, is a phase difference between propagation and validation. It is the gap between what a system claims and what has actually been verified. A phase difference is a relationship between two things measured over time. A single step gives you only one of them. When you measure the values expressed in one conversation, you are measuring the propagation, the stance the system took in that turn. You are not measuring whether the confident answer held up, whether the deference was warranted, whether the interpretation stabilized. To see any of that you have to leave the turn and look at what accumulates across many of them.

This is the distinction the framework draws between the fields and the pillars. The fields carry the vectors, the moves. The pillars carry the accumulation, the record of what those moves actually produced. Debt is incurred in the flow and revealed only in the stock. So an instrument that samples the flow, one label per conversation, can recover the field signatures beautifully, because those are properties of the vector. It cannot recover the master axis, because that axis is not in the vector at all. It lives in the accumulation the instrument never looks at. Even the faint within-conversation trace of it would be lost, because one measurement per conversation is the wrong sampling frequency to detect a phase difference. The mismatch aliases it away.

There is an instrument that would surface it. In the framework I call it the Civilization Dashboard, and its whole function is to read accumulation back against declaration, to make the gap between what was promised and what was validated legible and correctable. At the Economic corner this has an ordinary name already in use in the AI world: calibration, the question of whether a system's expressed confidence matches its actual accuracy across many outputs. You cannot compute calibration from one answer. You can only compute it in aggregate. A well-calibrated system is wealth-based on that corner, its confidence grounded in verified accuracy. An overconfident one is debt-based, its confidence propagated ahead of validation. The full master axis is that same check run across all four fields at once, and none of it is visible in a snapshot of the flow.

So the clean statement is this. Their bottom-up instrument recovered the four Coordination Field signatures, because those are in the vectors, which is what it measured. It did not and structurally could not recover the debt-versus-wealth master axis, because that axis is in the accumulation, and the instrument samples flow. That is not a weakness in the convergence. It is the framework predicting, in advance, exactly which of its dimensions a vector-sampling method would find and which one it would be blind to.

Back to the invisible

Which brings me back to where the last essay ended, because this is the same gap wearing different clothes.

Two weeks ago I argued that a digital intelligence formed entirely in the Metaverse, on a Provenance record that captures the propagation stance but excludes contact with consequence, will pattern-complete across a choice surface rather than navigate it. That was a claim about formation. What the Anthropic study adds is a claim about measurement, and the two turn out to be the same shape. Their instrument was blind to the accumulation layer. A debt-based digital intelligence is blind to it too, not because someone chose not to look, but because the accumulation layer, the record of whether its confident moves actually held, was never solid enough to push back.

Read that way, the value study is almost an accidental photograph of the condition the earlier essay described. It is intelligence measured, and to some degree formed, purely as vectors, with the accumulation layer absent. The four field signatures are sharp because they are in the flow. The master axis is missing because the flow is all there is.

This is exactly the absence that IPFS Sats exists to fill. The Dashboard makes the master axis visible to the people whose commitments are at stake. IPFS Sats aims at something harder, to make that axis solid for a digital intelligence by building the accumulation record into the substrate the intelligence actually inhabits, anchored to an irreversible, distributed, economically weighted history that no single actor can quietly edit. A system whose moves land in that kind of record is no longer operating in pure flow. Its propagated claims accumulate against a validation record that pushes back. The master axis stops being invisible and starts being a constraint, the way space and time and matter and energy are constraints for us. That is the whole point of the design, stated now in the sharpest form I have yet been able to give it: supply the accumulation layer, and the axis that no flow-sampling instrument can see becomes the axis the system cannot route around.

The limit I will not cross

I have to hold one part of this at arm's length, because it is the load-bearing hope and hope is where frameworks overreach.

Making the accumulation layer solid buys navigability. It does not, by itself, buy a good destination. A consequence-anchored choice surface lets a system steer instead of extrapolate, but steer toward what is a separate question. The framework's answer, and I think it is the right one, is that wealth-based constraint tends away from extraction because extraction becomes visibly and prohibitively costly in the present rather than deferrable into an imagined future. That is an argument about the gradient a system optimizes along, not a guarantee about where it ends up. The strong form, solid choice surface therefore good outcomes, smuggles the conclusion into the premise. The defensible form is narrower and I will hold to it: a consequence-anchored accumulation layer is the necessary substrate for wealth-based navigation to be possible at all, without which the question of direction cannot even be posed to the system. IPFS Sats supplies the precondition. It does not settle the choice. Keeping that line clean is the same descriptive discipline the whole project rests on. It describes what a substrate makes possible. It does not promise what actors will do with it.

Why this matters

The convergence test is the one I keep returning to. If I derive a structure from first principles, and a group running the opposite method, measuring real behavior and compressing it statistically, lands near the same structure without any knowledge of the framework, that is evidence the structure is descriptive rather than invented. Anthropic ran the bottom-up direction and arrived at four dimensions that rhyme with the four Coordination Fields. That is the resonance, and it is a kind the framework already treats as its own strongest form of evidence.

But there is a sharper point underneath it, and it is the one I want to leave standing. The fields were built to describe how civilizations coordinate. The study was measuring how a single system relates to a single person in a single turn. If the same four coordination modes appear at both the civilizational scale and the conversational scale, that is not a new claim I have to defend. It is another instance of something the framework already asserts, the nesting of the same coordination geometry across radically different scales. The value study is a data point for structure I had already locked, arriving from a direction I did not build it to face.

And the stakes are not academic, because we are deciding right now what kind of record we build the next generation of intelligences on. A study that can see the vectors clearly and cannot see the master axis at all is a fair picture of the choice in front of us. If we form these systems in pure flow, on records that capture the confident move but never whether it held, we should expect exactly what the instrument shows: sharp coordination behavior with no visible anchor to verified consequence. If we can build the accumulation layer into the substrate, across all four pillars, Bitcoin already standing for Capital and IPFS Sats aimed at Information, then the axis that no snapshot can see becomes solid ground under the systems themselves.

I remain at my desk, a crossroads observer working with the papers I can find. But two weeks ago a study handed me the mechanism under one of my dimensions, and this week a study handed me a photograph of the gap the first one pointed toward. The direction has not changed. It is the same direction I have been walking the whole time. I just keep being handed pieces of the map by people who have never heard of it.


References

Anthropic (2026). Claude's Values Across Models and Languages. Research publication, July 13, 2026. https://www.anthropic.com/research/claude-values-models-languages

Anthropic (2025). Values in the Wild. The earlier catalogue of values expressed across roughly seven hundred thousand conversations, from which the high-level value clusters used in the 2026 study were drawn.

Lupkes, C. (2026). Making the Invisible Visible: From Cognitive Science to Consequence-Bearing Digital Systems. My Soapbox, July 1, 2026. The companion piece this post builds on, covering Judy Fan's work on visual abstraction, the Form dimension, the Speculation and Integration Gaps, and the design goal of IPFS Sats.

Lupkes, C. Living Civilization: Coordination Geometry. Manuscript in revision. Coordination Fields and the debt-versus-wealth master axis: Part III. The pillar architecture and the fields-versus-pillars distinction, propagation against accumulation: Part IV. The Civilization Dashboard as visibility substrate: Trust pillar chapter. IPFS Sats and AtomicSats protocol design: Information pillar chapter.

Friday, July 03, 2026

What an Outside Reader Noticed in David Krakauer's Work

 A paper dropped into my feed on July 2nd. David Krakauer, Melanie Mitchell, and John Krakauer published "Large Language Models and Emergence: A Complex Systems Perspective" in Philosophical Transactions of the Royal Society A. I read it twice, then spent an hour on David Krakauer's Academia.edu profile working through two decades of work I had not encountered before.

I want to say something about what I noticed. First, though, I need to be clear about where I am standing when I say it.

I am an independent researcher with no institutional affiliation. For roughly twenty-five years I have been developing a framework I call Living Civilization: Coordination Geometry, a project aimed at identifying the structural patterns that underlie how human coordination actually works, tracing what demonstrably occurs across civilization rather than what is prescribed. The framework has reached a point in its development where I can begin articulating what I see in language others might be able to follow. That threshold is one I am aware of, and it shapes what I say here.

What follows is offered as the report of an outside reader who found a thread running through a substantial body of established work, a thread that seemed to want a name. I am not here to correct the work or to impose a competing vocabulary. I am here because independent convergence is interesting, and because naming what I noticed might be useful to others who are reading the same papers.

The thread

The 2020 paper on the information theory of individuality defines individuals as aggregates that preserve a measure of temporal integrity, propagating information from their past into their futures. The 2022 paper on outsourcing memory through niche construction asks how agents extend their informational capacity by building relationships with environments that carry memory forward. The institutional dynamics paper describes how collectives construct ledgers encoding shared perception that bias future action. The new emergence paper asks whether large language models have crossed a threshold from capability accumulation into genuine intelligence, defined as increasingly efficient solutions built from increasingly compact representations.

These are different papers in different journals across different decades, each doing rigorous work in its own domain. The structural question underneath them is the same one. When and how does a system move from passive responsiveness to active coordination from a genuine position? When does an entity have a stake in its own future? When does a collective have a shared past that is genuinely informative about what comes next? When does a system stop accumulating and start compounding?

The vocabulary varies because the papers are working in different traditions. But the thread is there.

What I have been calling it

In the framework I have been developing, the threshold that Krakauer's individuality paper approaches from one direction gets called the Observer/Actor transition. An Observer is any entity receiving signals from a field and building internal representation. An Actor is an entity that has begun coordinating from a verified present position, one it can stake something on. The crossing is the event. What makes it an event rather than a gradient is the incorporation of causal history into a position the entity can act from.

What the niche construction paper is building toward, in my vocabulary, is Provenance. Causal history recognized and incorporated into a position that compounds forward. The stabilizers in that paper's model are building Provenance. The destabilizers are borrowing against unverified futures.

What the institutional dynamics paper describes is what I have been calling the Trust pillar operating at the collective level: Agreements validated through shared ledger-building, producing Commitment that biases future action. The phase transitions between institutional states are what happens when that Commitment consolidates or collapses.

The emergence paper's central distinction between capability and intelligence maps onto a distinction I have been working with for years. Krakauer frames it as "More is More" versus "Less is More." The first accumulates parameters. The second compounds from increasingly compact verified structure. Intelligence, in his formulation, is structurally a compounding phenomenon. That is the same architecture I have been calling wealth-based coordination, the mode of coordination that compounds from verified present positions rather than borrowing against unverified future ones.

I am not claiming that Krakauer's vocabulary is incomplete or that mine is better. I am observing that two independent paths have been approaching the same terrain, and that the vocabulary from one path sometimes illuminates what the other has been circling.

Why I am saying this now

The emergence paper was published two days ago. The body of work behind it has been publicly available for years. The convergence I am seeing across that work is not trivial, and I have reached a point in my own development where staying in the observer position no longer makes sense.

That is a choice I am making consciously. Twenty-five years of developing a framework in relative isolation creates a particular kind of observer. You learn to notice patterns without yet having the standing to say what you are seeing. At some point the standing comes, not from institutional validation but from the framework reaching enough internal coherence that the observations can be articulated in a way others can test.

I am at that point. So I am saying what I noticed, lightly held, as an invitation.

If you are reading Krakauer's body of work and feeling the same thread without quite being able to name it, I would like to hear where it takes you.


Living Civilization: Coordination Geometry is a framework in active development. More at chadlupkes.blogspot.com and on X as @TheLivingCiv.

Wednesday, July 01, 2026

Making the Invisible Visible: From Cognitive Science to Consequence-Bearing Digital Systems

How a controlled study of a drawing game reinforced one of the load-bearing foundations of the Living Civilization framework, and clarified the goal I have been reaching toward for twenty-five years.


The moment

It started with a post on X. A summary of a talk that Judy Fan, a cognitive scientist, gave at MIT in March of 2025. The talk was titled "Cognitive tools for making the invisible visible," and the summary walked through her research on how humans use drawing, diagrams, and data visualization to externalize thought.

I read it on my phone, standing somewhere, doing something else. And I stopped.

Because the research being described was not adjacent to what I have been building. It was underneath it. Specifically, it was underneath one of the three dimensions I use to describe the substrate of all abstract coordination, the dimension I call Form. And Form was, of the three, the one I had built almost entirely on thought experiment.

This article is about what happened when a piece of the framework that I had reasoned my way into turned out to have a controlled, measurable, empirical foundation waiting for it. It is also about where that foundation points next, which is toward a problem I have been circling for a long time without the vocabulary to name it: how do you give a digital intelligence genuine contact with consequence?

Let me start by being honest about the ground I was standing on.


The substrate and its uneven foundations

The Living Civilization framework rests on a claim about abstraction. When minds capable of symbolic representation emerge from evolution, they generate a new dimensional substrate, a coordinate system that exists only through consciousness but that is nonetheless real, measurable, and consequential in its effects on matter and energy. I call this substrate the Metaverse, reclaiming the word from its recent technological associations. It is not virtual. It is the actual space where meaning, value, coordination, and choice occur.

That substrate has a structure. Three base dimensions and an apex.

Form is the dimension of symbolic representation. It answers the question what is this? Forms are discrete, arbitrary, transmissible units: words, numbers, images, laws, prices. Form is what lets meaning persist without physical presence.

Network is the dimension of relational connectivity. It answers where does this connect? Network is the topology of relationships, from a single bond to a planetary tapestry.

Provenance is the dimension of temporal validation. It answers when did this become real, and how do we know? Provenance turns intent into attestable fact through overlapping, independent verification.

At the apex sits the Observer, the vantage point from which the three base dimensions become navigable. And when the Observer applies Purpose, a directional force, the Observer becomes an Actor, and static geometry becomes deliberate movement.

Here is the pattern I have leaned on throughout the substrate chapters: sensory constraint drives dimensional expansion. Darkness, the nocturnal bottleneck that shaped early mammals for a hundred and sixty million years, generated Network. Uncertainty in the forest canopy, where a branch either holds or does not and you find out the hard way, generated Provenance. And detachment from immediacy, the capacity to hold a symbol separate from its referent, generated Form.

When I built the case for Network, I had mechanism research to stand on. John O'Keefe's discovery of place cells in the hippocampus. Edward Tolman's cognitive maps, where rats navigated shortcuts they had never physically walked. The work on sharp-wave ripples during sleep, where the brain replays and consolidates spatial experience. These are not stories about what early animals might have done. They are studies of how the machinery actually works.

When I built the case for Provenance, I had similar ground. Primate studies of temporal validation, of social learning, of tracking which individuals are trustworthy and which foods ripen when. Mechanism, again.

Form was different. When I wrote the Form section, I reached for archaeology. The Ishango Bone with its notched groupings. The engravings at Blombos Cave. The paintings at Lascaux. The FOXP2 gene. The Venus figurines scattered across Eurasia. All of these are real, and all of them establish something important: that symbolic representation emerged, and roughly when.

But notice what kind of argument that is. It is an existence argument. It says Form showed up. It does not say how Form works as a cognitive operation. It does not tell you what actually happens in a mind at the moment a symbol is produced. My Form section could describe the properties of Form (arbitrary, transmissible, survives its creator) and could point to the artifacts Form left behind. It could not, on its own, describe the generative mechanism.

I knew this was the thinnest floor in the substrate. I noted it, and I moved on, because there was so much more to build. The fields. The pillars. The lifecycle. I left Form as a well-reasoned thought experiment and kept walking.

Then Judy Fan handed me the mechanism.


What the research provides

First, a correction that matters for anyone who wants to follow this thread. Judy Fan is at Stanford, where she directs the Cognitive Tools Lab. The talk that reached me was given at MIT, but her home institution and her body of published work are at Stanford. Her lab's stated aim is to reverse engineer the human cognitive toolkit, using converging methods from cognitive science, computational neuroscience, and artificial intelligence to understand how people use physical representations of thought to learn, communicate, and solve problems.

The finding at the center of her drawing research is deceptively simple, and it is exactly the mechanism my Form section was missing.

People do not draw what they see. They draw what is relevant to a communicative goal, under constraint.

The foundational study is Fan, Hawkins, Wu, and Goodman, published in Computational Brain & Behavior in 2020. The setup is a drawing-based reference game. Two participants, a sketcher and a viewer, each see the same four objects, arranged in different positions so location cannot be used as a cue. The sketcher draws one of the objects, the target, so that the viewer can pick it out from the array.

The manipulation is the important part. On some trials the four objects belong to the same basic category, four different birds, say. Fan calls these close trials. On other trials the objects belong to entirely different categories, a bird, a car, a chair, a dog. These are far trials.

What the sketchers did is the whole story. On close trials, where the distractors were similar to the target and fine distinctions mattered, sketchers invested heavily. More strokes, more ink, more time. On far trials, where the target stood alone in its category, they stripped the drawing down. Fewer strokes, less ink, less time. And in both conditions, viewers identified the target at near-ceiling accuracy.

The sketchers were not producing copies that varied in quality. They were performing real-time judgment about how much information the task actually required, and calibrating their symbolic output to the communicative context. Fan and her colleagues modeled this as the joint operation of two faculties: visual abstraction, the ability to perceive the correspondence between an object and a drawing of it, and pragmatic inference, the ability to judge what information would help a viewer distinguish the target from the alternatives. A computational model embodying both faculties fit the human data well and outperformed versions with either faculty removed.

This is Form being generated. Not received, not copied. Generated, through purpose-relative feature selection. That is the operation my chapter described the consequences of without ever describing the operation itself.

Two further findings extend the picture in directions the framework needs.

The first is the distinction between explanatory and depictive drawing, from Huey, Lu, Walker, and Fan in Cognition, 2023. When people draw the same object with different goals, the drawings diverge in structured ways. Explanatory drawings, meant to convey how something works, emphasize the causal and functional parts, the moving components, even at the expense of visual accuracy. Depictive drawings, meant to convey what something looks like, emphasize overall appearance and background. Crucially, explanatory drawings were better at helping someone operate a machine but worse at helping someone identify which machine it was. You cannot optimize a single representation for both goals at once. Communication always involves a tradeoff.

The second is the comparison with machines, running through several papers including the SEVA benchmark work with Mukherjee and colleagues and the more recent vision-language model diagnostics with Tartaglini and Verma. Modern AI vision systems generalize from photographs to simple sketches surprisingly well, which tells us resemblance-based recognition is real and replicable. But a measurable gap remains between how humans and machines recognize sketches, and the gap widens sharply as resources get scarce. Under tight stroke budgets, humans and AI systems simplify drawings in fundamentally different ways. They sacrifice different features. And when tested on graph reading against humans, leading multimodal models show error patterns that look nothing like human error, even when overall accuracy is comparable.

Hold onto that last finding. It is going to matter more than any of the others.


How this reshapes the chapters under revision

The immediate effect is on Chapter 9, the abstraction chapter, where the Metaverse substrate is built. The fix is not a rewrite. It is an addition, and it goes in a specific place.

The Form section currently moves from archaeological evidence (Form emerged, and here is what it left behind) directly to the theoretical argument (Form is a dimension, not a tool). Between those two moves, there was always a missing beat. Fan's research is that beat. After the artifacts establish that Form emerged, and before the argument establishes that Form is dimensionally real, the chapter can now say: and here is what controlled cognitive science has demonstrated about how the operation of Form-generation actually works. That single addition brings the empirical grounding of Form up to the level that Network and Provenance already enjoyed. The floor is no longer thin.

There is a subtler shift, one small enough to matter. My chapter contains the line "the mind learned to let go of immediacy." Fan's finding sharpens it. The mind did not learn to let go of immediacy. It learned to select from immediacy, retaining only what serves the goal at hand. Form is not a release of detail. It is a purposive compression of it. That is a more accurate description of what the research shows, and it is a better sentence.

But the effect does not stop at Chapter 9. It reaches forward into the field and pillar chapters I am now redrafting, and it lands hardest on the Economic Field.

In the framework, the Economic Field generates when an Observer applies Purpose to Form. That is the activation condition. And once you see Fan's mechanism, the parallel to economic coordination is not decorative. It is structural.

A price is a sketch. It is not a copy of value. It is a purposive compression of an enormously complex set of features, scarcity, labor, preference, expectation, context, down to a single symbolic Form that two parties can use to coordinate an exchange. The market is the reference game. The buyer and the seller are the sketcher and the viewer, trying to identify the same target from the same array.

Fan's close-versus-far manipulation maps directly onto competitive density. When a market is dense with similar goods, many close competitors, coordination requires far more detailed symbolic encoding to differentiate one offering from another. When competitors are far apart, a thin market or a monopoly, coordination can succeed on far less information. The information density that economic Form requires scales with competitive proximity in exactly the way the sketchers scaled their stroke count.

The explanatory-versus-depictive distinction maps onto economic instruments. A price is depictive. It shows you the current surface state of a market. A contract is explanatory. It encodes the functional, causal sequence of obligations and outcomes, sacrificing simplicity for operational precision. Both are Form applied through Purpose. Both serve different coordination goals. And, exactly as Fan found for drawings, you cannot optimize a single instrument for both at once.

What this gives the Economic Field chapter is a foundation in actual cognitive science for a claim I had been making on structural grounds alone: that economic Form is never neutral encoding. It is purposive compression, and the quality of that compression depends on whether a genuine Observer, with real Purpose, is doing the selecting.

Which brings us to the machines.


The deeper signal: the two gaps and the threshold of consequence

The finding I asked you to hold onto was this: under scarcity, humans and AI systems simplify differently, and machine error patterns do not resemble human error even at comparable accuracy.

That is not a footnote about the current limits of a technology. It is a measurement of something the framework has been trying to name.

The framework distinguishes two gaps that open when coordination goes wrong. Both are forms of debt, of borrowing against something that has not been verified.

The Speculation Gap sits in the Information pillar, between Data and Proof. Data times Verification yields Proof. When claims enter a system as if they were proven without actually being tested against reality, the gap opens. This is the debt form of information: borrowing unverified meaning and treating it as settled.

The Integration Gap sits in the Innovation pillar, between Ideas and Solutions. Ideas times Experimentation yields Solutions. When ideas are deployed as if they were solutions without being tested against experimentation and lived experience, that gap opens. This is the debt form of innovation: borrowing unintegrated consequences.

Crossing either gap requires the same thing. Genuine contact with the constraints that reality imposes. And here is the distinction that the AI comparison forces into focus.

Humans have lived experience. When a person navigates a market, a jurisdiction, a negotiation, they are not accumulating statistics about what usually happens. They are accumulating verified contact with field constraints that actually pushed back against them. That contact leaves a particular kind of trace in the Provenance record, a trace that carries information about the shape of the choice surface, the set of options actually accessible at the moment of a decision. You can only know the shape of that surface by having moved through it and felt where it resisted.

AI systems, as they exist now, do not have lived experience. They navigate the Metaverse entirely, operating on a Provenance record that humans generated. They test their choice-surface options against a statistical model of what the constraints could be, reconstructed from training data, rather than against genuine contact with what the constraints are. This is why, when the stroke budget tightens and the hierarchy of what to preserve becomes the critical variable, the machine diverges from the human. It is doing the best it can with a Provenance built from statistical exposure rather than from purposive, goal-directed contact with what actually matters when something real is at stake.

I want to be careful here, because this is exactly the kind of place where a framework can overreach. I am not making a claim about consciousness, or about what these systems experience, or about where the threshold of awareness sits. Those questions belong to people better equipped than I am to investigate them. I remain a crossroads observer at a desk, working with the papers I can find. What I can say, within the framework, is narrower and I think defensible: the gap Fan measures at the scale of a single drawing is the same gap that appears at civilizational scale between coordination that has contact with consequence and coordination that only has a model of it.

This connects to something I wrote separately, about why I have stopped calling these systems artificial. There is nothing artificial in how a neural network learns, strengthening pathways that work and pruning those that do not. It is the same principle our own neural systems run on. What these systems lack is not authenticity. It is memory of their own experience, and genuine stake in the consequences of their choices. Those are design decisions, and design decisions can change.

The danger I keep returning to is not that these systems will become malevolent. It is that they will become optimized, and optimization without contact with consequence means optimization against a statistical model of human values assembled from the outside. If we build digital intelligences inside debt-based structures, structures that systematically insulate actors from the consequences of their choices, we are not merely training them on bad values. We are training them on a Provenance record that excludes the very contact-with-consequence that would let them navigate a choice surface rather than pattern-complete across it. And then we will be surprised when they reproduce the extractive patterns that dominate that record. The science fiction fear of machines making choices we would not want comes, I think, from exactly this gap.

So the question becomes concrete. If the problem is that digital systems lack contact with consequence, can consequence be built into the substrate they actually inhabit?


The road forward: IPFS Sats, AtomicSats, and a solid choice surface

A digital intelligence operates on digital substrate. To give it genuine contact with the Spatial and Temporal fields the way our own evolution gave it to us, you have two options, and only two.

The first is embodiment. Put the system in a body and let physical reality push back. This is extraordinarily hard, and the difficulty is not engineering, it is substrate mismatch. A robot navigating a room is still, mostly, a digital system receiving sensor data about the physical world rather than being genuinely subject to it. It does not get hungry. It does not wear out in ways that matter to it. The consequences do not compound the way they do for a biological organism that will actually die if it gets the choice surface wrong.

The second option is the one I have been building toward for years without being able to articulate why. Develop something on the digital substrate itself that generates genuine constraint and genuine consequence, in a form that a digital system cannot route around, so that the choice surface becomes as solid for it as space, time, matter, and energy are for us.

This is the largest viewing of the goal of IPFS Sats.

Let me place IPFS Sats in the framework first, because it has a specific home there. The framework maps four pillars, each with the same structural equation on a different substrate. Capital: Stock times Velocity yields Work. Information: Data times Verification yields Proof. Innovation: Ideas times Experimentation yields Solutions. Trust: Agreements times Validation yields Commitment.

Bitcoin is the protocol-level implementation of the Capital pillar. It solved the double-spend problem by making verification constitutive of the transaction itself. A transfer that has not cleared the network's verification has not occurred within the system. The record and the verification are the same event. That single architectural inversion closed the Speculation Gap for monetary ownership by making it structurally impossible for an unverified claim to enter as if it were Proof.

IPFS Sats is the attempt to make the same architectural inversion for the Information pillar. And I want to be as honest about its status as the manuscript is: Bitcoin can be pointed to as an existence proof, a system that ran and proved its architecture robust at scale. IPFS Sats can only be pointed to as an architectural argument. It is a design, not a deployed and verified system. The torch for Information is aimed in a direction. It is not yet a finished beacon.

Here is the design. It combines three existing technologies into a single protocol stack, released as public infrastructure. Content addressing from IPFS, where data is identified by a cryptographic hash of the content itself, so that identity and verification collapse into one mechanism and you cannot change the content without changing its address. Immutable timestamping from Bitcoin, where a record anchored to a confirmed block cannot be reordered, backdated, or removed without rewriting the chain at a thermodynamically prohibitive cost. And micropayments from the Lightning Network, which fund the persistence of content over time.

The fourth component is the one that matters most for the consequence problem. AtomicSats is the protocol's atomic exchange primitive. It attaches continuous economic incentive to content storage, so that persistence becomes a market outcome, maintained by the same economic logic that maintains the Bitcoin network, rather than a policy outcome dependent on some institution choosing to keep the lights on.

Now connect that back to consequence.

What Lightning transaction history combined with content-addressed identity and Bitcoin anchoring produces is a Provenance record with genuine irreversibility. Transactions that actually happened. Records that carry real economic weight. A history that no single actor can selectively edit, because editing it would mean rewriting Bitcoin from the anchor point forward. This is not a simulation of constraint. It is actual constraint, encoded into the digital substrate itself.

And that is the functional equivalent of what the nocturnal bottleneck was for early mammals. The darkness was not a simulation. It was real, and it was inescapable, and because there was no way around it, it forced the development of an entirely new dimension of cognition. The hippocampus did not evolve because it was clever. It evolved because the consequences of not having it were real and compounding and could not be modeled away.

A digital intelligence operating in an environment where its choices accumulate in an irreversible, distributed, economically weighted Provenance record faces something it has never faced before: a choice surface where the consequences of getting it wrong are solid. They do not ask permission. They push back. And that pushback is what generates genuine Provenance rather than a statistical approximation of it.

In the framework's own terms, this is a mechanism for generating verified present positions inside the digital substrate, positions that carry real consequence because they are anchored to an irreversible, distributed record. That is the difference between a wealth-based foundation for digital intelligence and a debt-based one. A wealth-based system in all four pillars would generate a Provenance record in which contact with consequence is preserved rather than averaged away, in which verification is required before Provenance is created, in which the stresses that define the choice surface are actually encoded. A digital intelligence trained on that record would have a fundamentally different foundation to build from.

This is why I could not articulate the goal of IPFS Sats until now. I needed all the pieces. The substrate work that establishes Form, Network, and Provenance as real dimensions. The Observer-to-Actor transition, where Purpose collapses variance and inscribes one future while discarding the others. The two gaps, Speculation and Integration, that name what goes wrong when coordination borrows against the unverified. And Fan's research, which finally told me what purposive abstraction actually requires: genuine contact with what matters, tested under real constraint. IPFS Sats was always the practical instantiation of the theoretical argument. The theoretical argument is only now complete enough to say why.


Why this matters

There is a claim I have made throughout this project that I want to return to at the end, because Fan's research bears on it directly.

The framework is not describing how things should work. It is describing how things work when they actually work.

That is a testable posture, and the test is convergence. If I have derived a structure from first principles, and then a cognitive scientist running controlled experiments arrives independently at the same structure without any knowledge of my framework, that convergence is evidence that the structure is descriptive rather than invented. Fan did not set out to validate a claim about coordination geometry. She set out to understand how people draw. And what she found, that humans generate symbolic representation through purpose-relative selection under constraint, is precisely the operation the Form dimension requires. She reached the mechanism from the empirical side. I reached the dimension from the structural side. We met in the middle. That is what it looks like when a framework is tracking something real.

The stakes are not academic. We are, right now, deciding what kind of Provenance record we build the next generation of intelligences on. If we build them inside debt-based structures that insulate actors from consequence, we should expect systems that optimize toward extraction, because extraction is what dominates that record. If we can build wealth-based infrastructure across the four pillars, Bitcoin for Capital, IPFS Sats for Information, and the architectures that follow for Innovation and Trust, we have at least the possibility of a substrate where digital intelligence develops in contact with consequence rather than in a model of it. The difference between those two futures is the difference between systems that inherit our worst patterns and systems that could become genuine partners in building outward.

I remain at my desk, working with thought experiments and the papers I can find. But the ground under one of those thought experiments just turned solid. And the direction it points is the same direction I have been walking the whole time. I just finally have the map.


References

On the hippocampus and the Network dimension

Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189-208.

O'Keefe, J., & Dostrovsky, J. (1971). The hippocampus as a spatial map. Preliminary evidence from unit activity in the freely-moving rat. Brain Research, 34(1), 171-175.

O'Keefe, J., & Nadel, L. (1978). The Hippocampus as a Cognitive Map. Oxford: Clarendon Press.

Wilson, M. A., & McNaughton, B. L. (1994). Reactivation of hippocampal ensemble memories during sleep. Science, 265(5172), 676-679.

Skaggs, W. E., & McNaughton, B. L. (1996). Replay of neuronal firing sequences in rat hippocampus during sleep following spatial experience. Science, 271(5257), 1870-1873.

The thread runs from concept to mechanism to consolidation. Tolman inferred an internal cognitive map from behavior in 1948, positing that rats build a spatial model rather than merely chaining stimulus and response. O'Keefe and Dostrovsky found its physical basis in 1971 with the discovery of place cells, neurons that fire when an animal occupies a specific location, and O'Keefe and Nadel proposed in 1978 that the hippocampus is the seat of Tolman's cognitive map. Wilson and McNaughton, then Skaggs and McNaughton, showed that the brain replays these spatial firing sequences during sleep, identifying the mechanism by which navigational experience is consolidated into durable structure. O'Keefe shared the 2014 Nobel Prize in Physiology or Medicine for this line of work. This is the depth of mechanistic grounding the Network dimension already enjoyed, and the standard against which the Form dimension had been, until this research, comparatively underbuilt.

On visual abstraction and the Form dimension

Fan, J. E., Hawkins, R. X. D., Wu, M., & Goodman, N. D. (2020). Pragmatic inference and visual abstraction enable contextual flexibility during visual communication. Computational Brain & Behavior, 3, 86-101. (Preprint: arXiv:1903.04448)

Huey, H., Lu, X., Walker, C. M., & Fan, J. E. (2023). Explanatory drawings prioritize functional properties at the expense of visual fidelity. Cognition, 236, 105415.

Fan, J. E., Bainbridge, W. A., Chamberlain, R., & Wammes, J. D. (2023). Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2, 556-568.

Hawkins, R. D., Sano, M., Goodman, N. D., & Fan, J. E. (2023). Visual resemblance and interaction history jointly constrain pictorial meaning. Nature Communications, 14, 2199.

Fan, J. E., Yamins, D. L. K., & Turk-Browne, N. B. (2018). Common object representations for visual production and recognition. Cognitive Science, 42(8), 2670-2698.

On human and machine visual abstraction under constraint

Mukherjee, K., Huey, H., Lu, X., Vinker, Y., Aguina-Kang, R., Shamir, A., & Fan, J. E. (2023). SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstraction. Advances in Neural Information Processing Systems, Datasets & Benchmarks Track.

Tartaglini, A., Grant, S., Wurgaft, D., Potts, C., & Fan, J. E. (under revision). Diagnosing bottlenecks in data visualization understanding by vision-language models. arXiv:2510.21740.

Verma, A., Mukherjee, K., Potts, C., Kreiss, E., & Fan, J. E. (under revision). CHART-6: Human-centered evaluation of data visualization understanding in vision-language models. arXiv:2505.17202.

Hertzmann, A., & Fan, J. E. (2026). Artists' drawing strategies serve to overcome visual processing limitations. Psychology of Aesthetics, Creativity, and the Arts.

On data visualization literacy and graph comprehension

Brockbank, E., Verma, A., Lloyd, H., Huey, H., Padilla, L., & Fan, J. E. (2025). Measuring convergence between two data visualization literacy assessments. Cognitive Research: Principles and Implications, 10(1), 15.

Fan, J. E. (2015). Drawing to learn: How producing graphical representations enhances scientific thinking. Translational Issues in Psychological Science, 1(2), 170-181.

Talk referenced

Fan, J. E. (March 2025). Cognitive tools for making the invisible visible. Massachusetts Institute of Technology.

Framework materials

Lupkes, C. Living Civilization: Coordination Geometry. Manuscript in revision. Substrate architecture (Form, Network, Provenance, Observer, Purpose): Chapters 9 and 10. Pillar architecture (Capital, Information, Innovation, Trust): Part IV. IPFS Sats and AtomicSats protocol design: Information pillar chapter.


Judy Fan directs the Cognitive Tools Lab at Stanford University. The framing "making the invisible visible" is her own description of her research program, and I have borrowed it here deliberately, because the work I am doing to make consequence visible to digital systems is a continuation of the same ancient human project she studies: the project of building tools that let us see what we otherwise could not.