Codapult
DocsBlogPricingPluginsDemo
Get Codapult
Logo

Build your SaaS with AI. Keep the architecture under control.

Get Codapult

Product

  • Pricing
  • Architecture
  • Modules
  • AI platform
  • Plugins

Developers

  • MCP
  • CLI
  • Documentation

Resources

  • Blog
  • FAQ

Connect

  • Contact
  • GitHub

Compare

  • SaaS template comparison
  • Codapult vs Supastarter
  • Codapult vs Makerkit
  • Codapult vs ShipFast
  • Codapult vs SaaSBold
  • Codapult vs Gravity
  • Codapult vs Nextbase
  • Codapult vs BuilderKit

Featured on

Codapult on LaunchNestCodapult on LaunchNestbetterlaunch.cobetterlaunch.coFeatured on LaunchBuffFeatured on LaunchBuffCodapult on PeerPushCodapult on PeerPushFeatured on LaunchItFeatured on LaunchIt
© 2026 Codapult·All rights reserved·Privacy Policy·Terms of Service
Full source code · One-time purchase · Self-host anywhere
Back to blog
September 29, 2026·12 min read·Codapult Team

The Agent Shouldn't Be Able to Approve Its Own Rules

AI coding agents can make legitimate architectural changes. The harder problem is separating code changes from policy changes, waivers and approval authority.

ai-agentsarchitecturemcpguardpolicy
Software engineering workflow separating code changes, policy changes and approval authority

How to let AI coding agents change architecture without letting them silently change the rules around it.

One of the less obvious problems I've run into with coding agents isn't that they make bad changes.

It's that sometimes they make a perfectly reasonable change that invalidates one of the rules I'm using to check the project.

That's a different problem.

Imagine a project has this rule:

All payment-provider access must go through PaymentAdapter.

An agent is asked to add support for another payment provider.

It looks at the existing code and decides that the current adapter isn't quite right. It wants to introduce a new abstraction.

The resulting change crosses a boundary that the project currently protects.

The checker reports a violation.

So what should happen next?

The obvious automation is:

  1. agent changes the code;
  2. check fails;
  3. agent changes the rule;
  4. check passes.

Technically, everything worked.

But the system has a problem: The thing being checked was allowed to change the conditions of the check.

Diagram separating implementation changes from policy changes

That made me much more interested in the difference between changing the code and changing the rules around the code.

The dangerous loop

The simplest version looks like this:

agent
  ↓
change code
  ↓
verification
  ↓
failure
  ↓
agent changes policy
  ↓
verification
  ↓
pass

There is nothing obviously broken here.

The agent might even have a good reason for changing the policy.

The problem is that the verification boundary has disappeared.

A policy is supposed to tell the system what has to remain true.

Agent change and policy proposal passing through separate verification and approval paths

If the same actor that made the change can also redefine what "true" means, a successful verification doesn't tell you very much.

It's just a moving target.

And this isn't specific to architecture.

You could do the same thing with:

maximum service size = 300 lines

The agent produces a 420-line service.

The check fails.

The agent changes the limit to 500.

The check passes.

Or:

module A cannot import module B

The agent needs the dependency.

It changes the rule.

The import is now allowed.

Again, maybe that's the right architectural decision.

But those are two separate decisions:

  • The implementation changed.
  • The policy changed.

They shouldn't become one operation just because the same agent proposed both.

Not every policy violation is actually a mistake

This is where it gets more interesting.

I don't want an architecture checker that treats every violation as proof that the agent did something wrong.

Sometimes the agent really should change the architecture.

Suppose an application has:

Checkout
   ↓
PaymentAdapter
   ↓
Stripe

And I decide that the product is getting large enough that payment workflows deserve their own domain boundary:

Checkout
   ↓
PaymentService
   ↓
PaymentAdapter
   ↓
Stripe

The change introduces new files.

Some imports move.

Some old boundaries disappear.

New ones appear.

A strict checker could report a pile of violations.

That doesn't mean the change is bad.

It means the current policy describes the old architecture.

This distinction matters.

A policy isn't supposed to prevent architecture from ever changing.

It is supposed to make architecture changes explicit.

Proposal and approval are different things

This led me to a fairly simple rule: An agent should be able to propose a policy change, but it shouldn't automatically be able to approve that policy change.

For example:

Current policy:

Payment provider access must go through PaymentAdapter.

The agent can say:

Proposed change:

Allow PaymentService to depend directly on a new internal
PaymentProvider interface.

Reason:
The current adapter boundary prevents the new workflow
from sharing transaction state correctly.

That's useful.

The agent has done the hard reasoning.

It has identified the existing constraint.

It has explained why the constraint may no longer fit.

But the proposal should remain a proposal.

The important part is that the authority approving the policy change is separate from the agent that authored the change.

Otherwise the system can silently move the goalposts.

This is the same reason I don't want approval in the prompt

A prompt can say:

Never deploy without approval.

That's useful instruction.

It's not much of a control if the same process can modify the configuration that defines what counts as an approved deployment.

The more autonomous the agent becomes, the more these distinctions move out of the prompt and into the environment around it.

The agent should be able to reason about the policy.

It should be able to request a policy change.

It should be able to explain the change.

The actual enforcement shouldn't depend on the agent remembering to follow its own instructions.

A policy change should leave a trail

Once policy becomes a real project artifact, another problem appears.

You need to know what the policy was when the original change was checked.

Otherwise you can end up with a strange situation where today's successful verification only makes sense because yesterday's policy was replaced.

That's why I like keeping policy changes explicit and versioned.

Something like:

project state
    ↓
policy revision 17
    ↓
agent proposes architecture change
    ↓
verification fails under policy revision 17
    ↓
policy proposal
    ↓
approval
    ↓
policy revision 18
    ↓
verify change again

Now there are two separate facts:

  1. The change passed under policy revision 18.
  2. Policy revision 18 was itself approved.

That's much more useful than simply seeing a green check.

The old policy still matters

There's another subtle point here.

Suppose an agent changes the rule and then verifies the same diff against the new rule.

You can no longer tell whether the original change violated the previous policy.

So I want the system to preserve the distinction between:

what the project allowed before the change

and:

what the project allows after the change

This becomes especially useful when investigating a change later.

You can ask:

  • Why did this dependency become allowed?
  • Was it always allowed?
  • Was there an explicit exception?
  • Did a policy revision happen at the same time?
  • Who approved it?
  • What evidence led to the change?

Those questions are much harder to answer when policy is just another mutable config file.

Temporary exceptions are different again

Sometimes the policy is fine.

The violation is temporary.

For example:

All persistence access must go through Repository.

But I'm in the middle of a migration.

I don't want to remove the rule.

I just need one known exception for two weeks.

That's not really a policy change.

It's a waiver.

And I think treating it as a different object makes the whole system easier to reason about.

A useful waiver has at least:

owner
reason
scope
expiry

So instead of:

remove the rule

you get:

waive this finding until 2026-12-28

owner: platform-team
reason: repository migration

The rule stays.

The exception expires.

That's a very different thing from changing the architecture policy permanently.

Baselines solve a different problem

I also don't want to confuse waivers with baselines.

A baseline answers:

This violation already existed.

A waiver answers:

This active violation is intentionally allowed for a limited time.

And a policy change answers:

We changed what the project considers acceptable.

Those are three different states.

That separation might sound overly precise.

Three distinct concepts: baseline, temporary waiver and policy change

In practice, it makes the tool much more useful.

A real codebase can have old architectural debt.

It can have temporary migration exceptions.

And it can deliberately evolve its architecture.

If all three become "ignore this finding", you lose important information.

What the agent should actually do

This doesn't mean agents have to stop making architectural changes.

Quite the opposite.

I want them to do more.

A useful agent flow could look like this:

requirement
    ↓
agent plans change
    ↓
implementation
    ↓
deterministic verification
    ↓
failure
    ↓
agent explains why
    ↓
policy proposal / waiver proposal
    ↓
approval
    ↓
verification against new state

The agent can drive most of that workflow.

It can inspect the repository.

It can understand the requirement.

It can implement the refactor.

It can identify the policy that stopped the change.

It can prepare the evidence for a policy proposal.

It can even tell me that the current architecture appears to be the problem.

What it shouldn't get is an invisible path from:

my change failed

to:

therefore my own change is now allowed

This also changes what "autonomous" means

I've started thinking that autonomy isn't really one switch.

An agent can have permission to:

  • read the repository;
  • modify source files;
  • run tests;
  • inspect dependencies;
  • propose policy changes;

without having permission to:

  • approve those policy changes;
  • remove its own enforcement;
  • extend a temporary waiver indefinitely;
  • change the authority that verifies it.

That gives you a more useful permission model than simply:

autonomous = yes/no

Different operations can have different authorities.

And that matters more once the agent is running for a long time or can delegate work to other agents.

The verifier should not care who made the code change

There's another distinction here that I find useful.

Verification should primarily answer:

Does the current change fit the current approved policy?

It doesn't need to decide whether the author was:

human
Claude
Codex
Cursor
another agent
automation

That's a separate concern.

The verifier checks the state.

The policy layer defines the constraint.

The authority layer controls who can change the constraint.

Keeping those pieces separate makes the system much easier to reason about.

This is what I started building into Guard

This distinction ended up affecting the design of Codapult Guard quite a bit.

Guard already had the idea of:

facts
  ↓
policy
  ↓
verification

But that isn't enough when policy itself can change.

So I started treating policy changes as first-class operations.

Guard can discover project facts and prepare proposals.

A project can explicitly approve them.

For protected projects, policy approval can require a distinct actor rather than allowing the authoring agent to approve its own proposal.

The same idea now applies to waivers.

A waiver is not just "ignore this finding forever."

It has an owner, reason and expiry date.

When the expiry is reached, the finding becomes active again.

That gives the project three useful things:

policy
exception
history

instead of one growing collection of ignored warnings.

It also makes agent integration more useful

MCP makes this distinction especially important.

An agent can ask Guard:

What policy applies here?

or:

What did this change affect?

or:

Why did verification fail?

or:

What policy change would be needed for this refactor?

Those are useful questions.

But I don't want the agent to receive:

Change policy until verification passes.

That isn't really verification anymore.

The MCP layer should expose the information and operations the agent needs without quietly giving it the authority to redefine the system around itself.

That's a much more interesting role for project-aware tooling than simply adding another set of commands to an agent.

The rule can change. That isn't the problem.

I don't think architectural rules should be permanent.

Projects change.

Requirements change.

Teams change.

The shape of the system changes.

A rule that made perfect sense six months ago might become actively harmful.

The mistake is treating policy change as an implementation detail.

Changing:

src/payments/StripeAdapter.ts

and changing:

payment-provider-access

are fundamentally different operations.

One changes the system.

The other changes what the system is allowed to become.

Once an agent can perform both, they need separate boundaries.

I'm less interested in stopping agents than in making their authority explicit

The goal isn't to make agents weaker.

I actually want agents to make bigger changes.

But bigger changes require clearer boundaries.

The more of the implementation I delegate, the more I care about questions like:

  • What can the agent change?
  • What can it propose?
  • What can it approve?
  • What evidence does it have to provide?
  • What policy was active when the change was checked?
  • Was the exception temporary?
  • Who approved the exception?
  • Can the agent change the thing that verifies it?

Those questions aren't really about whether the model is smart enough.

They're about the architecture around the model.

And that's probably the part that becomes more important as coding agents become capable of doing more work without waiting for a human after every step.

I don't want the agent to be afraid of the rules.

I want it to understand them.

I want it to be able to challenge them.

I want it to be able to propose better ones.

But I don't want it to be the final authority on whether its own change should become the new rule.

The code can change.

The architecture can change.

Even the policy can change.

Those changes just shouldn't all happen under the same authority.


Want AI agents to propose architectural changes without approving their own rules? Codapult Guard separates project facts, proposed policy, approval, and deterministic verification. Read the Guard documentation or see how Codapult exposes Guard through its MCP server.