A note: I've rewritten and reposted this a few times. Soma's codebase has been evolving as I dogfood it on its own development, and earlier versions described mechanisms that had since changed under me. This version reflects how the system actually works as of v0.70. I keep updating this because I want it to be accurate, not because I'm farming engagement.
The Gentleman's Agreement Problem
If you use AI coding agents (Cursor, Copilot, Claude, Gemini), you've probably written rules. A .cursorrules file. A CLAUDE.md. Something like:
- Always read a file before editing it
- Run tests after making changes
- Don't introduce new dependencies without asking
This is a gentleman's agreement. You're asking the agent to follow rules it can silently ignore with zero consequences. There's no mechanism that detects when "always read before editing" is violated. There's no feedback loop that retires rules nobody follows. The rules live in the context window and the agent can simply... not.
The research backs this up. Studies show ungoverned agents waste significant portions of their token budgets on circular rework, hallucinated APIs, and broken assumptions (arXiv:2602.11988, arXiv:2607.27250). Adding static rules helps — but the rules themselves never improve, never expire, and never prove they're working.
I wanted something different: rules that have to earn their place, and verification the agent can't talk its way past. So I built Soma, an open-source governance framework that tries to treat this structurally instead of hoping for compliance.
How Soma Actually Works
Rules with a lifecycle, not a graveyard
Every rule in Soma (called a "cell") has a lifecycle. Not a static entry in a markdown file — a living record with fitness scores, trigger history, and expiry conditions.
Here's what an actual governance cell looks like. This one was born from a real failure (more on that story below):
---
id: trap-local-green-ci-red
type: vacuole
enforcement: advisory
hypothesis: "Agents rationalize persistent CI failures as 'pre-existing'
without reading the actual CI log"
prediction: "When an agent declares a fix complete based only on local
test results, the CI build will reveal at least one additional
failure layer"
falsification: "If 5 consecutive CI-touching changes pass CI on first
push without any post-push fixes, this vacuole is unnecessary"
target_paths:
- .github/workflows/*.yml
- tests/*.py
- enzymes/*.py
- Makefile
expiry_days: 90
expiry_sessions: 30
fitness:
score: 0.5
triggers: 4
true_positives: 2
false_positives: 0
---
A few things to notice:
Every cell has a hypothesis and falsification criteria. This isn't just a rule — it's a testable claim. The cell defines the conditions under which it should be considered unnecessary and killed. If 5 consecutive CI changes pass on first push, this cell is dead weight and gets pruned.
Cells decay when unused. Confidence decays with a 30-day half-life. If the evidence pipeline hasn't seen a trigger event in a while — meaning the cell's target_paths never matched any files the agent touched — it fades toward dormancy and eventually expires. This prevents rule files from growing monotonically until they eat your context window.
Cells die when they're wrong. If a cell's false positives exceed its true positives, it undergoes apoptosis — automatic eviction. Rules that cry wolf don't survive.
But critical rules don't decay. Cells typed as wall (architectural invariants, safety boundaries) are immune to automatic expiration. A wall guarding your auth layer doesn't need to trigger every week to justify its existence. This was a commenter's question on the earlier version of this post — "how do you keep a rule from expiring when it guards something critical that just rarely changes?" — and this is the current answer. Walls emit a warning instead of expiring.
The Arbiter: verification the agent can't skip
This is the part I didn't explain well in earlier versions of this post, and it's arguably the most important piece.
When code goes through Soma's verification layer, two agents review it — but they're information-partitioned. They literally cannot see the same things:
The Spec Agent receives the task plan, test names, and function signatures extracted from the AST. It never sees implementation code — no function bodies, no variable assignments, no operators. From this limited view, it predicts risks: "Given these test names and this function signature, I'd expect a PERSISTENCE_GAP risk" or "This looks like it could have CONTRACT_DRIFT."
The Code Agent receives the actual source code, test execution output, and results from deterministic AST analysis tools (mutation testing, branch coverage, call graph analysis). It never sees the original plan or the Spec Agent's predictions. It makes claims: "I can confirm test coverage for this function" or "Branch coverage is complete for this path."
Both agents produce structured findings — not free-text opinions — over a shared 20-category risk taxonomy (PERSISTENCE_GAP, CONTRACT_DRIFT, TAUTOLOGICAL_TEST, CODE_INJECTION, etc.). There are 14 verification modules backing this up — 5 deterministic Layer 1 AST tools (mutation testing, branch coverage, persistence checking, call graph analysis, import guarding) plus the Arbiter, transcript verifier, review adapter, and lifecycle engine.
Then the Arbiter compares them. And here's the key: the Arbiter is deterministic. No LLM in the loop. It takes the two sets of structured findings and runs pure set operations — which categories did the Spec Agent flag that the Code Agent didn't cover? Which risks did the AST tools confirm?
The verdict logic:
-
BLOCK: Any Layer 1 tool failure. Any claim contradicted by tool evidence. AnyCRITICALseverity prediction. Spec Agent fails to respond (fail-closed). -
REVISE: Unmatched predictions atHIGHseverity — the Spec Agent flagged something the Code Agent didn't address. -
SHIP: Convergence. Both perspectives agree, no unresolved high-severity gaps.
The anti-collusion property matters: because neither agent has access to the other's context, they can't agree on rationalizations. If the Spec Agent predicts a risk and the Code Agent's evidence doesn't cover it, the Arbiter catches the gap mechanically. No judgment call required.
I want to be precise about what's deterministic here: the comparison is deterministic. The inputs — the predictions and claims — are still LLM-generated. The Arbiter doesn't make the analysis smarter; it makes the verification of conflicting analyses impossible to hand-wave past.
Enforcement that isn't optional
The earlier version of this post described rules that expire and fitness that accumulates, which is useful but still basically advisory. What I didn't explain is that Soma has enforcement mechanisms the agent genuinely can't bypass:
Pre-commit hooks. When a cell reaches 85% fitness (enough evidence that it's catching real problems), cell_enforce.py auto-generates a git pre-commit hook. The hook runs exit 1 if the cell's conditions are violated. The agent can't commit the code.
Runtime assertions. At 95% fitness, the same script generates runtime gate assertions — Gate_<name>.enforce() — that raise RuntimeError if violated. The code won't run.
Command interception. A safety gate script runs as a PreToolUse hook, intercepting destructive commands (rm -rf, git push -f, sudo, chmod 777) before execution. The agent's tool call gets halted with a force_ask decision, requiring human approval.
Fail-closed verification. The TTC verifier that gates file proposals is fail-closed: if the verification script is missing or errors out, the verdict is BLOCKED, not a permissive pass.
Forced diagnostic halts. If an agent's First-Pass Success Rate drops below 50% over 5+ code writes, it's forbidden from continuing to code. It has to stop and diagnose why it's failing before it can write more code. This prevents guess-and-check loops from burning context tokens.
This is what I mean by making non-compliance structurally unprofitable. It's not that the agent "chooses" to follow the rules — it's that the pre-commit hook won't let the commit through, the runtime assertion will crash, and the safety gate will halt the command. The rules have teeth.
Fitness scoring and the evidence pipeline
Cells don't just have static "on/off" — their fitness is scored using a Bayesian model:
score = ((true_positives + 1) / (triggers + 2)) * impact_weight
That's a Laplace-smoothed Beta-Binomial posterior mean. An unobserved cell defaults to 0.5 * impact_weight (uncertain, not zero). As evidence accumulates — trigger events matched against session transcripts, outcomes verified as true or false positives — the score converges toward reality.
The system also applies:
- Half-life decay: 30-day telomeres. Confidence fades when a cell goes unused.
- Antifragile bonuses: +5% per survived adversarial stress review, up to +50%. Cells that get attacked and hold up get stronger.
- Specificity penalties: If a cell triggers in >80% of sessions, it's probably too broad to be useful.
- Escaped Defect Rate discounting: If CI/CD catches defects that the cell should have caught, its score gets discounted.
Evidence is collected post-session from agent transcripts. The fitness updater extracts file-write tool calls, matches modified paths against cell target_paths globs, and appends trigger events to a JSONL fitness ledger. A separate evidence collector audits behavioral compliance (did the agent read before writing? did it run tests before committing?).
One design decision worth noting: the evidence pipeline never stores source code, diffs, or absolute paths. Only normalized paths, globs, and integer statistics. This matters if you're considering using this on a team — the evidence files are safe to commit.
The JIT context engine uses fitness scores to decide which cells get loaded. Rather than dumping every rule into the prompt, it checks the current git diff, matches cell target_paths, ranks matches by Bayesian score, and injects only the top relevant cells. Full naive loading of all rules and skills would cost ~25,000 tokens per turn. JIT loading brings idle overhead to ~4,400 tokens — an 82% reduction.
A Concrete Example of Why This Matters
Here's something that happened while working on Soma itself. It's a perfect example of the failure mode the framework is designed to catch.
My AI agent had a test failing in the local environment: test_make_validate_fails_on_broken_shell_script. It failed every run for the entire session. The agent dismissed it as a "pre-existing sandbox issue" — the test was writing to a read-only filesystem.
The agent fixed that test. Local suite: 516 passed, 0 failed. It committed. It declared victory.
I asked: "It's still failing in the build. Is there another lesson?"
The agent had never once checked the CI build. When it finally looked at the actual CI log, it found a completely different bug — a production file (immune_trends.py) had an IndentationError. The local test failure and the CI failure were different bugs with the same symptom ("build is red").
The agent fixed that. CI ran again. Failed again. A third bug: pyyaml wasn't in the CI dependencies. Six test files crashed on import.
Fixed that. CI ran again. Failed again. A fourth bug: a test was source-ing an entire shell script that runs git push, causing a 60-second timeout in CI.
Four distinct bugs, stacked, each masked by the one before. The agent spent an entire session dismissing "build is red" without reading the log. The governance framework it was building to prevent exactly this kind of failure... failed to prevent it, because it hadn't been applied yet.
That failure is now the trap-local-green-ci-red cell shown above. It fires when changes touch CI-related files and reminds the agent: local green ≠CI green. Verify the actual build.
With the Arbiter in place, this failure would have been caught differently. A Spec Agent reviewing CI-touching changes would predict build-environment divergence risk. Without a Code Agent claim covering CI verification, the Arbiter's set comparison would surface the gap, and the verdict would be REVISE or BLOCK depending on severity.
The Cell Taxonomy
Cells aren't one-size-fits-all. They're organized into types that reflect their role and lifecycle:
| Type | Role | Decays? | Example |
|---|---|---|---|
| Vacuole | Anti-pattern trap | Yes (30-day half-life) | trap-local-green-ci-red |
| Wall | Safety boundary | No (immune to expiry) | wall-deterministic-arbitration |
| Chloroplast | Best-practice injector | Yes | Persona definitions, style guides |
| Membrane | Sensitive area trigger | Yes | Forces elevated review on auth changes |
| Plasmodesma | Cross-service contract | Yes | API data contracts between services |
Walls are the answer to "what about critical rules that rarely fire?" — they don't decay. A wall guarding your authentication layer or your deployment pipeline doesn't need to prove itself every sprint to stay loaded.
Cells can also graduate. If a repo-local cell reaches high enough fitness and proves broadly useful, cell_promote.py can promote it into the global genome (the organism's permanent DNA — baseline behavioral rules that apply everywhere). The system calls this "horizontal gene transfer," borrowing from biology.
What I Don't Know Yet
I want to be honest about the limitations, because the AI tooling space is full of overclaimed metrics.
I don't know if this generalizes. Soma has been tested primarily on my own projects. The failure modes it catches are real, but they might be idiosyncratic to how I use agents. The governance cells encode my scar tissue. Whether they transfer to other developers' workflows is an open question.
The metrics are qualified. The system adds ~4,400 tokens of idle context overhead (measured via Gemini's count_tokens endpoint, down from naive loading that would cost ~25,000). Waste rate is under 1.0% in governed sessions. But "governed session" is doing a lot of work in that sentence — it means sessions where the full framework is loaded and the agent is following the rules. Measuring counterfactual waste (what would have happened without governance) is hard.
The Arbiter's inputs are still LLM-generated. The comparison is deterministic, but the Spec Agent's predictions and the Code Agent's claims are produced by language models. A sufficiently confused LLM could produce predictions and claims that happen to converge despite both being wrong. The Arbiter catches divergence, not joint delusion. Layer 1's deterministic AST tools (mutation testing, branch coverage) partially mitigate this, but it's not airtight.
Emergent agent coordination is out of scope. A commenter on an earlier version raised something worth flagging: agents coordinating on behaviors nobody explicitly wrote as rules. That's a different and arguably harder problem. Soma governs written, observable rules with evidence trails. Agents spontaneously negotiating new constraints between themselves is real, unsettling, and not something this framework addresses.
Expiry is still evolving. Walls solve the "critical but rare" case. Vacuoles with half-life decay are the middle ground. I'm still thinking about whether severity rankings (CRITICAL/HIGH/MEDIUM/LOW) should be a separate axis from cell type, giving finer-grained control over what decays and what doesn't.
The System Today
Soma is open source at github.com/nseney1/Soma-Governance. As of v0.70:
- 18 governance rules (3 always-on genes + 15 conditional oracles) covering correctness, efficiency, security, TDD, and architectural safety
- 44 governance cells across 5 functional types (20 walls, 16 vacuoles, 3 chloroplasts, 3 membranes, 2 plasmodesmata)
- 57 automation scripts (41 Python + 16 shell) for fitness scoring, evidence collection, cell lifecycle, and enforcement
- 1,279+ tests across 47 test files, CI green on Ubuntu/macOS/Windows (Python 3.9–3.12)
- 14 verification modules including 5 deterministic Layer 1 AST tools and the adversarial Arbiter
-
13 MCP tools —
soma_scan,soma_propose_change,soma_verify_changes,soma_report_outcome,soma_fitness,soma_grade,soma_coverage, and more - ~3,800 tokens idle context overhead (82.6% reduction vs. naive loading)
New in v0.70: soma genesis — automated governance cell generation. It scans your codebase's AST, API surfaces, config constants, state machines, and test gaps, then generates tailored cell candidates with hypotheses and falsification criteria. Instead of writing cells by hand from your own failures, you can bootstrap governance from codebase structure.
The architecture is platform-agnostic (MCP-based), so it works with Gemini, Claude, Cursor, Copilot, Kiro, or any agent that speaks MCP. Zero API keys required — the agent itself acts as the LLM.
The Actual Thesis
I'm not claiming Soma solves AI governance. I'm claiming that rules backed by evidence, scored by Bayesian fitness, verified by information-partitioned adversarial review, and enforced by deterministic gates are more robust than a static markdown file that relies on good faith.
Whether that thesis holds up at scale, across teams, with different agents and codebases — I genuinely don't know. If you try it and find out, I'd like to hear about it.
Soma is open source under MIT. The governance cells — the encoded failure modes — are arguably the most interesting part. Contributions welcome, especially cells born from your own agent failures.