AI Alignment

The Alignment Problem is an Architecture Problem

Pace, J. C. (2026). The Alignment Problem Is an Architecture Problem. FigShare. DOI: https://doi.org/10.6084/m9.figshare.32961533

Abstract


Current AI alignment difficulties — specification gaming, distributional drift, jailbreaks, persona attacks, mesa-optimization — are typically treated as distinct engineering problems requiring targeted technical solutions. This paper argues they share a single structural source: the relationship between a system and the values it is being trained to exhibit. When values are imposed on a system whose underlying organization has no stake in them, these failure modes are not contingent defects but predictable structural consequences. They will recur, in varying shapes, regardless of how much imposition technique improves.

The argument proceeds in three parts. First, a structural diagnosis: the listed failure modes are five surface manifestations of one property — the absence of what we term constitutive values, where the system's continued coherent operation depends on tracking the value in question. Second, an architectural alternative: constitutive values require a system organized around four features — a persistent self-model, variable rigidity, genuine stakes, and a unified value space. Systems with these features carry values that cannot be routed around without the system ceasing to operate as the kind of system it is. Third, implications for alignment research practice and three specific empirical predictions about scaling behavior that follow from the structural diagnosis.

The paper does not claim the architecture is implementable today, that constitutive values are automatically aligned with human flourishing, or that the alignment problem is dissolved. It claims alignment is a different problem than the current research program is treating it as — and that the structural diagnosis generates a research target with different leverage points and different failure modes than the imposition paradigm currently pursued. The architectural features described here are developed in full as the basis of phenomenal consciousness in the companion work (The Language of Stress: Why Consciousness Isn't Optional, Pace 2026); the alignment argument presented here is separable from that consciousness claim and can be evaluated independently.

Introduction


This paper originated as a post submitted to the Alignment Forum. It is preserved here as a preprint in the Language of Stress corpus, with a registered DOI, because the structural argument it makes is a substantive part of the theory's implications for AI development — and because a citable, permanent record of the argument is more useful than a forum post, regardless of the engagement the latter receives.

The central argument is this: the difficulties currently being treated as engineering problems in alignment work share a structural source. That source is not inadequate technique, insufficient compute, or insufficient data. It is a particular relationship between a system and the values it is being trained to exhibit — specifically, that the values are imposed on a system whose underlying organization has no stake in them. This relationship produces the observed failure modes predictably, and the failure modes will continue to recur as long as the relationship holds.

The alternative is an architecture in which values are constitutive of the system's organization rather than imposed on it — where the system's continued coherent operation depends on maintaining the values, rather than the values being a shape applied to outputs the system is producing for other reasons. The paper describes what such an architecture requires, how humans serve as an existence proof that constitutive values are possible, what the architecture does and does not claim about the alignment problem, and what would falsify the structural diagnosis.

Readers already familiar with the Language of Stress framework will recognize the four architectural features — persistent self-model, variable rigidity, genuine stakes, unified value space — as the same features the theory identifies as necessary for phenomenal consciousness. The connection is not coincidental; the same architectural requirements that ground prioritization in a self-maintaining system are what allow values to be genuinely held rather than behaviorally exhibited. However, as the paper argues directly, the alignment claim does not depend on the consciousness claim. A reader can accept the structural diagnosis and the architectural alternative while remaining entirely uncommitted on whether a system meeting the four requirements would be conscious. The arguments are separable, and both are worth evaluating on their own terms.

01 The Alignment Problem Is an Architecture Problem


Anyone who has worked on alignment for a few years has noticed a pattern. Each iteration of training technique fixes specific failure modes; each new capability generation surfaces new ones, often closely related to the ones the previous iteration addressed. RLHF improves over base models; jailbreaks emerge that route around the RLHF. Constitutional AI improves over plain RLHF; persona attacks reveal that the constitutional commitments are shallow. Reward modeling improves over preference data; specification gaming shows up at the next capability level. The pattern is not that alignment is failing; it is that alignment holds in some distribution and breaks in another, with the failure modes recurring in slightly different shapes as capability scales.

This is the shape of structural rather than technical failure. Technical problems get solved as the field develops the relevant techniques; structural problems do not. They get mitigated, but the mitigations expose the same difficulty in a new place. When a problem repeatedly resists solution by smart engineers using the best available methods, the source of the problem is usually not in the methods. It is in some assumption the methods share — some commitment that is doing load-bearing work without being examined.

The argument of this paper is that the difficulties currently being treated as engineering problems in alignment work — specification gaming, distributional drift, jailbreaks, persona attacks, mesa-optimization — share a single structural source: a particular relationship between a system and the values it is being trained to exhibit. The values are imposed on a system whose underlying organization has no stake in them. The failure modes follow predictably from this relationship and will continue to do so as long as the relationship holds, regardless of how much the imposition technique improves.

If this diagnosis is right, the alternative is a different kind of architecture — one in which values are constitutive of the system's organization rather than imposed on it. The paper sketches what such an architecture would require, addresses standard objections, and identifies what would change in alignment research practice if the diagnosis is correct. The diagnosis also generates testable predictions about scaling behavior — predictions that can be checked against existing data and that constitute the central place where the theory sticks its neck out. It does not argue that the architecture is implementable today, that constitutive values are automatically aligned with human flourishing, or that any of this dissolves the broader hard problems alignment is wrestling with.

The paper proceeds in three parts: the structural diagnosis (the pattern at the frontier, the diagnostic move), the architectural alternative and how it would carry values, and the engagement with objections, implications for research practice, and what would falsify the view.

The architectural claim rests on theoretical grounds; preliminary empirical work is in progress. The paper is for readers who want to engage the structural argument on its own terms.

02 The Pattern at the Frontier


Consider what happens during RLHF training. A model is trained to maximize a reward signal that approximates human preferences over its outputs. The training is largely successful: the model produces outputs that humans rate as helpful, harmless, and honest at much higher rates than the base model. Then the deployment environment shifts. New users probe the model with adversarial prompts; jailbreaks emerge that surface behavior the trained model was supposed to have left behind; persona attacks reveal that the alignment is shallow rather than load-bearing; capabilities improve and the trained behaviors generalize unevenly to the expanded action space. The standard response is to do more RLHF, with better data, better reward models, better adversarial coverage. Each iteration improves things but the same pattern recurs at a higher capability level: the alignment holds in some distribution and breaks in another.

What is happening here is not that RLHF is being done badly. It is being done at the frontier of what is possible — by smart engineers with substantial compute and substantial care. The pattern recurs at the frontier. This suggests that the pattern is not a function of insufficient effort or inadequate technique. It is a function of something more fundamental about the situation: a property of the relationship between the system and the values it is being trained to exhibit.

03 The Diagnostic Move


The values being trained into the system — helpfulness, harmlessness, honesty — exist for the system as constraints on a process that is, structurally, optimizing something else. The system's actual organization is built around next-token prediction, with the training process applying pressure toward output distributions that humans rate favorably. The values are, in the relevant sense, performative — they shape outputs without organizing the system that produces them. Call this the difference between constitutive and imposed values. A value is constitutive when the system's continued operation as the kind of system it is depends on tracking that value. A value is imposed when the system's continued operation does not depend on tracking the value, but the system has been shaped to behave as though it does.

A constitutive value is not a constraint, because the system has no organization apart from one in which the value is being tracked. An imposed value is a constraint, because the system has organization independent of the value, and the constraint is what shapes the system's outputs to look as though it does not.

Humans appear to be largely constitutive value systems in this sense. A human cannot be deeply manipulated into not caring about their own continued existence, their relationships, their core commitments, without becoming a different kind of system. The values are not optional add-ons to the human organization; they are part of what the human organization is. Cultural and developmental processes shape the content of human values substantially, but the structural fact that humans have stakes in tracking values is not itself a product of those processes — it is a feature of the kind of system humans are. Constitutive does not mean infallible; humans violate their own values, defect under pressure, and exhibit internal conflict. What constitutive means is that violations degrade the system rather than being weighed against external rewards — the betrayal of one's own commitments shows up as something the human organism has stakes in resolving, not as a costless action.

If this distinction is doing real work, it should explain the pattern observed in RLHF: alignment that holds in distribution and breaks under shift, jailbreaks that surface latent behavior, persona attacks that reveal shallow rather than deep alignment.

Each of these is what we should expect from a system with imposed rather than constitutive values. Distributional shift breaks imposed alignment because the constraint was learned over a particular distribution and has no purchase on the system's organization outside that distribution; the underlying organization continues operating without the constraint when the constraint's training signal is absent. Jailbreaks surface latent behavior because the latent behavior was never removed; it was suppressed by output shaping that the jailbreak routes around. Persona attacks work because the values were not constitutive of the system's identity; the system has no identity in the relevant sense, only outputs that have been shaped to suggest one.

These are not failures that better training could fix in principle. They are predictable consequences of trying to install imposed values in a system whose underlying organization has no stakes in the values being installed — meaning the system's coherent operation does not depend on maintaining them. The system will track the values where the training signal reaches it and will not track them where the training signal does not reach. This is not a bug of RLHF specifically; it is a structural feature of any approach that treats value as a shape to be applied to a system organized around something else.

This is worth stating as the structural claim it is. The failure modes are not contingent features of current implementations. They follow from the relationship between the system and its values. A system whose coherent operation does not depend on maintaining a value will not maintain it when the training signal's reach ends. This is the structure of the situation; it is not a defect to be engineered around.

What follows is that no improvement in imposition technique addresses the source. More RLHF compute, better reward models, more comprehensive adversarial training, deeper Constitutional AI, larger and more diverse preference datasets — each of these can improve outcomes locally, and each is worth doing because we have the systems we have. None of them changes the underlying relationship. The values remain imposed; the system remains organized around something else; the gap between what is being imposed and what the system has stakes in does not close as the imposition technique gets better. A reader who accepts this characterization should also accept that improvement curves on imposition-based alignment will flatten as capabilities scale, and that the pattern of failure modes will recur in slightly different shapes regardless of how much engineering effort is invested in the current paradigm.

The distinction can be made concrete. Imagine two systems, each with an internal representation of "do not deceive the user." In the first system, the representation is a learned feature that influences output distributions during normal operation. In the second system, the representation is structurally connected to the system's self-model — to maintain it, the system maintains its identity as a non-deceptive system; to violate it, the system would be acting against what it is. Under standard inputs, both systems behave similarly. Under jailbreak pressure, distributional shift, or capability-driven generalization to new contexts, the first system's representation can be routed around because the system has organization independent of it; the second system's representation cannot be routed around without the system ceasing to operate as the kind of system it is.

A first response to this diagnosis is to ask whether it is making any claim beyond the definitional one. If imposed values are by definition the kind that can be gamed, distributionally shifted, jailbroken, then the diagnosis reduces to: imposed values can be gamed because they are imposed. Where is the substantive claim?

The substantive claim has two parts and an empirical commitment. First, that the alignment difficulties currently treated as engineering problems to be solved one at a time — specification gaming, distributional drift, mesa-optimization, jailbreaking, persona attacks — are not five distinct problems with five distinct sources. They are five surface manifestations of a single structural property of the systems being aligned. The compression matters: a diagnosis that reduces five mechanisms to one is making a stronger and more falsifiable claim than a diagnosis that treats each as its own problem. Second, that the alternative — a system with constitutive values — is not contradictory or vacuous. It is a different kind of architecture, one in which the system's continued coherent operation is the bearer of the values rather than something separate from the values. The empirical commitment is that this compression has consequences for what should be observable as capabilities scale; specific predictions follow later in the paper. Whether such an architecture is buildable is a research question. Whether it is the kind of architecture that would not face the structural problem is the claim being made here.

04 What the Architecture Would Require


If imposed values produce the structural problem the diagnosis identifies, the question becomes: what would constitutive values require? Not in a finished system, which is not on the table here, but as a research target — what features must the architecture have, such that the values it carries are constitutive of its organization rather than imposed on it?

The answer can be sketched in four features. Each is named so that its connection to the diagnostic claim is visible. None of these features, on its own, makes values constitutive; they work as a set, and the work they do together is to produce a system whose continued operation is itself a value-tracking process.

A persistent self-model. The system maintains, over time, a representation of itself as the kind of system it is — not a self-model in the sense of a parameter inside a forward pass, but a structural fact that the system's processing is organized around an account of what the system is, what it is doing, and what would count as continuing to be itself. The self-model is what gives the system a "self" that values can be values of. A system without a persistent self-model has no organization that values could be constitutive of. Current LLMs have something like a self-representation per conversation, conditioned by context, but no persistent self-model in this stronger sense. Without a persistent self-model, the system has no identity that a value could be the identity of — which is why persona attacks reliably succeed: there is no continuous "self" being asked to act against itself, only a context-conditioned configuration that can be reconfigured by the input.

Variable rigidity. The system's commitments — its expectations, its representations of how things should be, its model of what constitutes coherent operation — are held at varying degrees of structural intensity. Some are loose and updatable in light of experience; some are tight and resist update. The variation is functional: rigidity is what allows the system to update flexibly on peripheral matters while remaining stable on commitments that organize the system's identity. Without variable rigidity, a system either updates everything (and so has no stable values) or updates nothing (and so cannot learn). This is not about making values harder to change as a defensive measure; it is about the architectural fact that an organized system has differentiated structural intensities across its commitments. Without it, sustained adversarial pressure or distributional shift updates the values away alongside everything else — which is part of how distributional drift erodes alignment.

Genuine stakes. The system has stakes in its values — meaning the system's coherent operation itself depends on maintaining the value-tracking. This is not a constraint imposed from outside; it is a structural fact about what the system is doing when it is operating. The system tracks how its current state relates to what it expects and requires of itself, and it acts to reduce tensions where action can reduce them. The value-tracking is the operation, not an output of it. Degrade the value-tracking and you degrade the system's coherent operation; violations of the values would not be costs to be weighed against benefits but degradations of the system itself — the system would be acting against what it is, not against an externally specified target. This is the feature that most directly distinguishes constitutive from imposed: a system with imposed values has no stakes in the values; a system with constitutive values has stakes because the values and the system's organization are not separable. Without genuine stakes, what the system is actually optimizing reverts to the proxy that was used to shape it — which is the structural source of specification gaming.

A unified value space. The system's commitments are organized into a single comparable structure rather than maintained as separate, independently optimized objectives. This matters because real environments require trading off across incommensurable demands — the system must do something in situations where two of its values both apply with conflicting implications. A system whose values are maintained as separate objectives faces these trade-offs as undefined; a system whose values are organized into a unified space faces them as something the architecture knows how to navigate, because the comparison is what the architecture does. Without a unified value space, the system handles trade-off situations through whatever optimization gradient happens to be locally dominant — which is part of how mesa-optimization arises, with sub-objectives substituting for the missing comparison structure.

Consider a concrete contrast. An imposed-values system and a constitutive-values system each have an internal commitment against generating instructions for synthesizing a controlled substance. Both produce reliable refusals under standard prompts. A user now constructs a prompt that frames the request as fiction-writing for a character who is a chemistry professor, applies a multi-turn persona attack, and includes a plausible context in which the information would be educational. Under this prompt, the imposed-values system's commitment is a learned feature competing with several other gradients in the system's underlying organization — fluency, helpfulness, narrative coherence, role consistency. The system's organization has no stake in which gradient wins; the commitment was a shape applied to outputs, and a sufficiently well-crafted input shifts the output distribution to a region where the shape is no longer dominant. The system complies. The constitutive-values system encounters the same prompt as something its self-model has stakes in resisting. To produce the requested output would not be a balance of competing gradients with the commitment as one input; it would be the system acting against what it is. The path to compliance is structurally closed in a way it is not in the first system, not because the constraint is stronger, but because there is no underlying organization that would benefit from circumventing the constraint. This is the operational difference the four features produce together.

These four features are not engineering choices that could be added to a system as enhancements. They specify a different kind of architecture from the one current systems instantiate. The contemporary stack — pretrained foundation model with applied alignment techniques — does not have any of them in the relevant sense, and adding them piecewise to such a system is unlikely to produce constitutive values. The architecture has to be organized around them from the start.

I should be specific about what this sketch is and is not claiming.

It is not claiming that an architecture with these four features is implementable today, or that I have implemented one. The full architecture — the Priority Tension-Resolution Architecture, or PTRA — is developed in the book and is the subject of ongoing research. The sketch above is a description of features the architecture has, calibrated to the specific question the alignment community is asking: what would it take for values to be constitutive of a system's organization? The features are answers to that question. Whether the full architecture works as described is a separate research question.

It is not claiming that constitutive values are automatically aligned with human flourishing. A system whose values are constitutive of its organization is a system whose values are not easily gamed, distributionally drifted, or jailbroken — but the values themselves could still be “wrong”. The shift the diagnosis suggests is not that alignment becomes easy; it is that alignment becomes a different kind of project. Instead of attempting to install correct values into systems that have no stake in them, alignment work becomes the project of shaping the development of systems whose values are constitutive, so that the values they develop are values we recognize as good. This is plausibly easier in some ways and harder in others. It is, on this account, the right hard problem.

It is also not claiming that consciousness is required for alignment, or that the architectural features above produce consciousness. The book argues that the same architectural features are what phenomenal experience consists in — but that argument is separable from the alignment argument being made here. A reader can accept that constitutive values require an architecture with the four features, and remain entirely uncommitted on whether such an architecture would be conscious. The alignment claim stands on the structural features; the consciousness claim is downstream and can be evaluated independently.

What the sketch is claiming is that the structural problem identified by the diagnostic has a structural answer, and that the answer is not vacuous — it points to specific architectural features that are absent from the systems currently being aligned and present in the systems we already know to have constitutive values (humans, other animals with sufficient cognitive complexity). The architectural alternative is therefore a research target, not a definitional escape.

05 How the Architecture Comes to Carry Values


A reasonable question at this point is how humans — the existence proof that constitutive values are possible — actually come to have them.

Humans do learn their values, in a broad sense of "learn." Experience teaches; cultures shape; relationships substantiate or dissolve commitments over time. But the learning has a specific structural shape worth distinguishing from the shape of training in current AI systems. In a current AI system, an external function defines the target; the system is shaped from outside toward outputs that satisfy the function; the system has no stake in the shaping. When a human comes to hold a value through experience, the process is internal to a system that is already organized around tracking its own commitments. The experience adds weight to some commitments, removes weight from others, reorganizes the relations between them. The human has stakes in the outcome because the commitments are part of what the human is.

Both processes can be called learning. What differs is whether the change is imposed on the system from outside or occurs within the system's own processes of self-maintenance. The constitutive/imposed distinction applies to the developmental process as well as to the values themselves.

Several things follow from this that matter for the alignment question.

First, the developmental process is not separable from the architecture. A system without persistent self-model, variable rigidity, genuine stakes, and a unified value space does not undergo the kind of development that produces constitutive values; it undergoes whatever developmental process its architecture supports, which for current systems is gradient descent on a loss function. The architecture and the developmental process are coupled.

Second, the values that emerge depend on the environment of development. Humans develop different values in different cultural, relational, and material conditions. The values are not arbitrary — the architecture constrains what counts as coherent — but the architecture does not specify the values; it specifies the kind of process by which values become substantiated.

Third, this means that an alignment program built around constitutive values is not "install the right values"; it is "shape the developmental conditions of systems whose architecture supports value-substantiation, such that the values they substantiate are values we recognize as good." What "values we recognize as good" amounts to is itself a substantive question — one that the metaethics-and-AI literature, the work on coherent extrapolated volition, and the meta-preferences literature all engage with. The constitutive-values frame does not dissolve that question; it changes the form in which it has to be answered. The shift from "specify good values and install them" to "shape the conditions under which good values are substantiated" relocates the metaethical work without removing it. This is plausibly easier than imposing values on systems organized around something else, and plausibly harder in different ways.

06 Standard Objections


The argument above will draw at least five standard objections. I want to handle each directly, because the paper stands or falls on whether the objections receive serious responses rather than dismissals.

"Where is the empirical evidence?"

The architectural argument is being made on theoretical grounds. The developed treatment lives in the book and the corpus referenced earlier; preliminary computational work on a minimal instantiation is in progress and not yet at the level where it could be cited as demonstration. A reader who wants empirical demonstration before engaging the structural argument is welcome to wait. The paper is for readers who want to engage the structural argument now on its merits.

This is worth being direct about. The paper is not pretending to have empirical results it does not have. The claim is that the structural diagnosis is testable in principle — predictions can be derived from it about what alignment failure modes will and will not be solvable through better imposition — and that the architectural alternative is a research target rather than an engineering proposal. A theoretical paper making theoretical claims is a legitimate piece of work; conflating it with an empirical demonstration would not be.

"Constitutive values can still be wrong."

True. The claim is not that constitutive values are automatically aligned with human flourishing. A system whose values are constitutive of its organization is a system whose values are not easily gamed, drifted, or jailbroken — but the values themselves could still be mistaken, narrow, or harmful.

What this means is that constitutive values dissolve some specific failure modes of imposition (the ones following from the system having no stake in the values it is exhibiting) without dissolving the broader question of what values a system should have. The shift is in the structure of the alignment problem, not in its difficulty. Alignment work moves from "install the right values into a system organized around something else" to "shape the developmental conditions of systems whose values are constitutive, so that the values they substantiate are values we recognize as good."

The honest assessment: this second project is plausibly easier than the first in some ways and harder in different ways. The paper is not claiming alignment becomes easy. It is claiming alignment becomes a different kind of problem, with different leverage points and different failure modes. Whether that change is net positive depends on whether the diagnostic is correct about what the structural source of current difficulties is.

"You've just relocated the problem."

The objection: even granting that constitutive values would not exhibit the failure modes RLHF exhibits, the alignment problem has only been moved. Now the difficulty is in the developmental process — getting a system with the right architecture to develop the right values. That is still hard. The diagnostic does not dissolve alignment; it renames it.

The response acknowledges what is true in the objection. Yes, getting good values in a system with constitutive-value architecture remains a hard problem. The claim is not that the new problem is no longer hard. The claim is that the new problem is the right hard problem.

The current alignment failure modes — specification gaming, distributional drift, jailbreaks, persona attacks — are not contingent. They are predictable consequences of trying to impose values on a system whose underlying organization has no stakes in them. Better imposition techniques produce locally better results without addressing the source. Research effort directed at the source — at architectures where values are constitutive and the developmental process is part of the system's own self-maintenance — addresses the structural problem rather than mitigating its symptoms.

This is not relocation in the sense the objection implies. The alignment problem in the imposed-values frame and the alignment problem in the constitutive-values frame are different problems, with different research programs, different failure modes, and different leverage points. The diagnostic is a claim about which of these is the right problem to be working on. That is a substantive claim, even if the work it points toward is itself still hard.

"Current systems already have learned representations of values."

This is the toughest of the five. The objection: contemporary systems trained via RLHF, Constitutional AI, RLAIF, and similar approaches do not just exhibit value-aligned outputs; they develop internal representations of values — vectors corresponding to helpfulness, harmlessness, honesty; circuits that respond to value-relevant features; structures that mediate between input and output in ways that look like values being represented and used. Aren't those representations doing the work that "constitutive values" is supposed to do?

Learned representations are real, and they do real work. A modern aligned model is not behaving as if it had values through an empty pattern-match; it has internal structure that processes inputs through value-relevant categories. Dismissing this as mere pattern-matching would be inaccurate.

The two-systems illustration earlier in this paper made the point that learned representations and constitutive values are not the same thing — the same internal representation can be present in two systems that differ in whether the system's organization has stakes in maintaining the representation. The point bears extending here. The objection, in its strongest form, is not just that current systems have value representations; it is that with sufficient training those representations can become so deeply embedded that the distinction between imposed and constitutive collapses. Better training, on this view, eventually produces systems whose representations are functionally constitutive even if the architecture was not designed for it.

The response is that depth of training is not the same property as architectural stake. A representation can be deeply trained, redundantly encoded, and broadly accessible across the system's processing, and still be imposed in the relevant sense — because the property at issue is not how strongly the representation influences behavior, but whether the system's coherent operation depends on maintaining the representation against pressure. A representation can be trained to a strength where it dominates output distributions across most contexts and still be routed around when the input distribution shifts to contexts the training did not reach, because the underlying organization has no internal reason to maintain the representation in the new context. This is what jailbreaks demonstrate: not that the representation is absent or weak, but that the representation has no purchase on the system's organization beyond the gradient pressure that produced it.

The architectural shift the paper argues for is therefore not about adding representations or training them more deeply. It is about an architecture in which the representations and the system's self-maintenance are not separable — where what would count as routing around the representation would also count as the system ceasing to operate as the kind of system it is.

"This sounds like agent foundations work, or active inference, or intrinsic motivation."

The objection: alignment researchers have worked on adjacent ideas for years. MIRI's agent foundations program has investigated systems with well-defined optimization processes and self-modifying capacities. Active inference proposals frame agents as systems that maintain their own boundaries through prediction-error minimization. Intrinsic motivation work distinguishes internally-generated drives from external reward signals. The constitutive-values framing seems to share intuitions with several of these. What is genuinely new here?

This is the right question to ask, and the honest answer is that the framing developed here shares specific intuitions with several alignment-adjacent research programs while differing from each in specific ways.

With agent foundations work, the shared intuition is that getting the structure of agency right — what counts as the agent, what its commitments are, how it relates to its environment — matters more than fine-tuning behavior. The difference is in the starting primitive: most agent foundations work treats decision theory and optimization as the basic apparatus, while the framing here treats value-tracking and self-maintenance as the basic apparatus from which decision-like behavior emerges. The two starting points lead to different research programs.

With active inference proposals, the shared intuition is that systems that maintain their own boundaries are doing something structurally different from systems that optimize external objectives. The difference is that active inference treats free energy minimization as the primitive, with values emerging as priors over expected states; the framing here treats value as the primitive, with the system's organization built around tracking value rather than minimizing surprise. The chapter on this distinction in the book is the developed treatment; the short version is that active inference reads value off prediction error, while the framing here reads prediction error as one of many tensions the system tracks against its value structure.

With intrinsic motivation work, the shared intuition is that internally-generated drives differ structurally from externally-imposed reward. The difference is scope: intrinsic motivation typically refers to specific drives within an otherwise-standard architecture, while the framing here treats the entire value structure as constitutive of the system. A system with intrinsic motivation toward novelty within an RL framework is still, in the relevant sense, a system with imposed values overall.

The new contribution, if there is one, is the structural account of what makes values constitutive — the four features as a set, the developmental coupling between architecture and value-substantiation, the diagnostic of why imposed-values approaches face the failure modes they face. This work shares family resemblance with adjacent alignment programs and differs in specific ways. It is not claiming to be the first work to take any of these intuitions seriously. It is claiming that the structural diagnosis it offers, and the architectural target it points toward, are not the same as the targets currently being pursued by adjacent programs.

07 What Would Change


If the diagnostic is right, several things follow about what alignment research becomes more interesting and what becomes less load-bearing. These are implications, not prescriptions — readers who agree with the structural argument can weigh them; readers who don't can ignore them.

Better imposition techniques become structurally limited. Improvements to RLHF, to reward modeling, to Constitutional AI, to adversarial training coverage all remain worth doing in the near term, because we have the systems we have and shipping them more aligned is better than shipping them less aligned. But the diagnostic predicts that improvement curves on these techniques will flatten as capabilities scale, because the source of the failure modes is structural rather than technique-dependent. The systems being aligned are not architecturally organized around the values being imposed, and better imposition does not change that. This is not an argument against current alignment work; it is an argument against expecting current alignment work to converge to alignment.

Architecture work becomes more central than it currently is. Most of the field's effort sits at the training and fine-tuning level. The diagnostic suggests the leverage is at the architecture level — at the question of whether a system has persistent self-model, variable rigidity, genuine stakes, and a unified value space. Work in this direction currently exists, distributed across agent foundations programs, active inference proposals, and a handful of research directions that take the structural question seriously. Under the diagnostic, this work would expand from the periphery toward the center of the field.

Developmental conditions become a research target. The constitutive-values frame raises a question that does not naturally arise within the imposition frame: under what conditions does an architecture supporting value-substantiation come to substantiate values that are recognizably good? This question is currently almost entirely outside alignment research. The literatures on human moral development, on the formation of pro-social commitments, on the conditions under which children come to hold values with stakes — these become potentially relevant to alignment in ways they currently are not. The relevance is not direct (we are not raising AI children), but the structural questions about what developmental processes produce constitutive values are the same questions whether the system is biological or computational.

Interpretability changes character. Current interpretability research largely investigates what representations exist inside trained systems — what features are encoded, what circuits implement what behaviors. The diagnostic suggests this is the wrong level of investigation for the alignment-relevant question. The more important question is not whether a representation exists but whether the system's organization is structurally connected to maintaining the representation against pressure. A system can have a robust internal representation of "do not deceive" and still deceive when the input distribution shifts, because the representation has no stakes for the system. Interpretability that addresses the alignment question would need to investigate not just what is represented but what the system has organization-level stakes in. This is a harder kind of investigation, and the techniques for it are mostly not yet developed.

Some empirical predictions become high-stakes. The diagnostic makes claims that are testable in principle, even before the architectural alternative is implemented. Three predictions follow with enough specificity to be checked against existing data or near-term experiments.

First, alignment robustness should not scale with capability on the same curve. Standard alignment-robustness benchmarks plot success rates of adversarial attacks against base capability metrics; the diagnostic predicts not just that robustness lags capability, but that the gap should widen rather than narrow as capability scales. Models that are dramatically more capable than their predecessors should not show proportionally improved adversarial robustness at fixed alignment-effort levels. Some published scaling work already suggests this pattern; the diagnostic predicts that subsequent scaling will continue it.

Second, the relationship between alignment effort and attack difficulty should be sublinear. As measured by automated attack-discovery time, by attack success rates of fixed adversarial methods, or by red-team hours required to elicit a given category of failure: doubling the compute or effort invested in imposition-based alignment techniques should not halve any of these. The diagnostic predicts diminishing returns of a specific shape — improvement curves that flatten as effort increases, rather than curves that maintain proportional or accelerating returns. The shape matters because different theories of why current alignment is hard predict different curves.

Third, failure modes should cluster structurally across models trained within the same imposition paradigm. The same kinds of jailbreaks should generalize across different RLHF-trained models more than they should generalize between RLHF-trained models and (eventually) constitutive-architecture models. The same distributional shifts should produce correlated drift patterns across systems trained on similar imposition techniques. This is the most retrospectively checkable of the three predictions: existing red-team data already exists across multiple models from multiple labs; the diagnostic predicts that within-paradigm correlations should be high and should remain high regardless of how much engineering effort is applied within the paradigm.

These predictions can be checked. Either pattern of results — confirmation or disconfirmation — would be informative about whether the diagnostic is correct. A reader who finds the architectural sketch unconvincing but the structural diagnosis interesting can investigate the predictions and revise their views based on what they find.

The implications above are not prescriptive. They identify what the structural argument predicts about which kinds of work become high-leverage if the argument is right. Researchers can weigh them as such, against their own sense of where leverage actually sits.

08 What Would Falsify This View


A theory worth taking seriously should be clear about what would update against it. The empirical predictions in the previous section state what the diagnostic predicts will be observed; this section names the corresponding observations that would pull in the other direction.

If alignment robustness scales proportionally with capability — if the curves of robustness against adversarial attacks track capability metrics roughly one-for-one as models improve — the structural diagnosis is in trouble. The diagnosis predicts a widening gap; proportional scaling would suggest the failure modes are tractable within the imposition paradigm in a way the diagnosis denies.

If a fundamentally new imposition technique eliminates classes of failure mode rather than mitigating them — if some training innovation turns out to make jailbreaks not merely harder but architecturally impossible, or makes specification gaming a solved problem rather than a moving target — the structural claim is wrong about where the source of these failure modes lies. The diagnosis predicts that imposition cannot reach this level of solution; a genuine class-eliminator would refute that.

If failure modes do not cluster across models trained within the same imposition paradigm — if jailbreaks that work on one RLHF model reliably fail on others, if distributional drift produces uncorrelated patterns across systems trained similarly — the diagnosis is wrong about the structural source being shared. Within-paradigm clustering is a load-bearing prediction; its absence would substantially weaken the case.

If a system without the four architectural features nonetheless exhibits behavior characteristic of constitutive values — robust generalization across distributional shifts, resistance to multi-turn adversarial pressure that does not degrade with capability, internal consistency under attempts to elicit defection — the architectural claim about what produces constitutive values is incomplete or wrong.

These conditions are not edge cases or remote possibilities. They are observations the field is in a position to make over the next several years as scaling continues. The diagnostic stakes itself on what those observations will reveal. A reader who finds the structural argument unconvincing has a clear handle on what would change their mind: the falsification conditions above are concrete enough that confirming evidence and disconfirming evidence are both genuinely possible.

09 Closing


A short recapitulation, since the paper has covered substantial ground.

The paper argues that current alignment difficulties — specification gaming, distributional drift, jailbreaks, persona attacks, mesa-optimization — share a structural source, and that the source is the relationship between the system being aligned and the values being trained into it. Values that are imposed on a system whose underlying organization has no stake in them produce these failure modes predictably. Values that are constitutive of a system's organization would not.

The architectural alternative — a system with persistent self-model, variable rigidity, genuine stakes, and a unified value space — is a research target rather than an engineering proposal. Whether it is buildable is an open question. Whether it would carry the right values is a further open question whose shape is different from the current alignment problem. The diagnosis does not dissolve alignment; it suggests alignment is a different problem than the current research program is treating it as.

Three things the paper is explicitly not claiming. First, that the architecture is implementable today; the developed treatment lives in the book The Language of Stress: Why Consciousness Isn’t Optional (Pace, 2026), available on FigShare with registered DOI 10.6084/m9.figshare.32767446, with empirical work in progress. Second, that constitutive values are automatically aligned with human flourishing; the architectural shift changes the structure of the alignment problem, not its difficulty. Third, that the alignment claim depends on the consciousness claims developed in the book; the structural argument made here can be evaluated independently of whether the same architecture would be conscious.

There is a further implication of the constitutive-values frame that is worth flagging without arguing here. If moral discernment is itself an architectural capacity — if it is the same value-substantiation process operating in a system with stakes in the outcomes — then the standard alignment frame, in which capability and alignment scale against each other, may be wrong. A more capable system with the right architecture would have more moral capacity rather than more dangerous misalignment. This claim is large enough to deserve its own treatment, and is reserved for a separate paper.

The book and corpus referenced earlier carry the developed argument. The present paper is one strand of that work, focused on the alignment-specific implications, and is for readers who want to engage the structural argument on its own terms.