I know it takes me forever to make an update, but I figured I would talk about some of the things I’ve been working on to help create new safeguard in this agentic world.

Building Services to Govern Agents and Skills Link to heading

Most registries are catalogs. Metadata goes in, search comes out. That is enough when the things being registered are inert, like a package of functions or a documentation page. It is not enough when the things being registered can act.

This post is about building a service to address that gap: a graph-first semantic governance system for autonomous agents and computational skills. It treats meaning, release state, and evidence as the product, and it stamps every trusted publication with the proof to back it up. What follows explains the problems that make a plain catalog insufficient, and the approach the service takes in response.

The problem: proliferation outruns governance Link to heading

Inside any organization that leans on AI agents, useful artifacts proliferate faster than anyone can track them. Agent files, skills, runbooks, process notes, tool manifests, and prompts all multiply. Each one is locally useful. Collectively they become an unmanaged surface.

That creates three concrete failures.

First, people do not know what already exists. The same capability gets rebuilt three times because discovery is informal.

Second, agents cannot tell which artifacts are current and trustworthy. An agent given a folder of markdown files has no way to know whether a file is a draft, a deprecated experiment, or a certified production procedure.

Third, operational knowledge goes stale or gets duplicated. Nothing formally marks a runbook as superseded, revoked, or re-certified, so the old version lingers and gets used.

A catalog addresses none of these directly. Listing a file does not tell you who asserted it, what validated it, what it was derived from, or whether it has been revoked. Without those answers, a registry is only an index of unverified claims.

Why a catalog is not enough Link to heading

The deeper issue is semantic. A natural-language description of a skill is a claim, not a verified fact. A markdown file saying “this skill safely queries the revenue warehouse” is a sentence, not a guarantee. To govern it, you have to:

  • lower the claim into a typed, machine-checkable form
  • check that form against policy and ontology
  • attach evidence of every check
  • emit stable, signable bytes
  • log the whole history immutably
  • make revocation and supersession real states, not comments

That is a compiler-shaped problem, and it is the core of the design thesis.

Design thesis: the registry as a semantic compiler Link to heading

The service is not trying to be a database with a search box. It is trying to be a compiler-like governance pipeline for machine-operable trust. The mapping is deliberate:

  • source language: markdown, tool manifests, skill docs, agent docs, API specs, runtime observations, human review notes
  • intermediate representation: a typed, ontology-backed graph
  • type system: a registry shape ontology plus domain ontologies
  • static analysis: graph shape checks, SHACL-like policy shapes, ontology checks
  • provenance and debug symbols: PROV-O activities, derivations, agents
  • deterministic build outputs: canonical RDF bytes and canonical JSON bytes
  • release artifact: a governed publication snapshot with an evidence bundle
  • audit layer: an append-only Merkle log with external transparency anchoring

Once a claim is compiled this way, it can be signed, verified, replayed, and revoked with the same rigor you would apply to a software release.

Approach 1: logical identity is not a publication snapshot Link to heading

The first structural decision is to separate what a thing is from a particular frozen version of it.

A durable identity (an agent or a skill) is a long-lived handle. A version is an immutable snapshot of that identity at a moment in time. Skills and agents have full lifecycle parity: both have drafts, validations, publications, certifications, deprecations, revocations, and archives.

This separation is what makes supersession and revocation meaningful. You can revoke a specific snapshot without erasing the identity, and you can supersede a trusted snapshot with a newer one, while guaranteeing that a revoked snapshot is never silently superseded.

Approach 2: validation is a stack, not a single approval Link to heading

Trust from composition. Any single check can be fooled, so the service layers five independent validation families, each emitting machine-readable evidence stored as a content-addressed artifact.

  1. Graph shape validation: required properties, edges, and cardinalities.
  2. SHACL-like policy shapes: admissibility policy as versioned data, executed with a safe pattern engine, producing standard violation reports.
  3. Ontology validation: capability claims must bind to known concepts in a resolved ontology, and disjoint combinations are rejected.
  4. Advisory language-model validation: scorecards for overclaiming, ambiguity, and example consistency. Always non-blocking, always evidence-producing.
  5. Human review: judgement recorded as first-class evidence, with policy-driven gating for publication and certification.

The rule is deterministic before probabilistic. Policy and ontology checks can reject a draft. The advisory model can only warn. Human judgement is preserved as evidence rather than collapsed into a boolean.

Approach 3: two ontology families, not one Link to heading

A single vocabulary cannot serve both governance and domain meaning. Mixing them makes trust reasoning incoherent. The design keeps two families separate.

Normative ontologies are curated by humans. They define the allowed meaning space and express contract-level semantics that stay interpretable across time and implementations.

De facto ontologies are emergent. They are captured from real submissions as candidate terms. They help normalization and reveal where the normative vocabulary is missing useful concepts, but they never define publication or certification policy. Promotion of a de facto term into a normative ontology happens only through human review, and it creates a new immutable ontology version. Source versions are never mutated.

Approach 4: canonicalization is where meaning becomes signable bytes Link to heading

A graph is the richest semantic representation, but serialization order is not stable enough to sign. Two representations of the same graph can produce different bytes, which breaks signatures. The service canonicalizes twice, producing two companion projections of the same governed snapshot.

The operational projection is canonical JSON, following consistent key ordering and number formatting. It is the practical interoperability layer for API clients, CI, and detached signatures.

The semantic projection is canonical RDF, projected from the typed publication graph and digested. Verification re-projects the current graph and compares, which detects post-publication mutation. Detached signatures can cover either the operational or the semantic bytes, and the two can cross-verify over the same snapshot.

Approach 5: transparency, not just logs Link to heading

Application logs are mutable and, in an incident, not trustworthy. The service emits a canonical governance event for every lifecycle operation, stores it as an immutable artifact, and appends it to an internal append-only Merkle log with inclusion-proof generation and verifier replay.

On top of that sits a network-aware transparency model: trusted log identities, checkpoint history with consistency proofs, witness attestations with quorum checks, externally observed checkpoints, and disagreement evidence for split-view detection. A minimal signed-note ingestion seam preserves raw checkpoint bytes while normalizing them for verification.

The point is that an operator can inspect why a checkpoint is trusted, not only whether it currently is.

Approach 6: provenance you can replay Link to heading

State alone does not explain how a publication came to exist. The service models lifecycle operations with PROV-O: publication snapshots and artifacts as entities, validation and publish activities as activities, and authors, reviewers, and automated validators as agents. Edges record derivation from prior snapshots, generation by activities, use of policy and ontology packages, and association of actors with roles.

Because the provenance graph is queryable and replayable, an auditor can reconstruct the exact policy version, ontology version, evidence reports, and actors involved at decision time.

The architecture Link to heading

The service is organized into planes with strict boundaries. A canonical internal core owns publication semantics, validation direction, and storage abstraction. A management plane is the only write authority. An AI-facing gateway is a thin, read-only adapter over the same core. A human web surface consumes the management plane and is never a source of truth.

Two invariants hold this together. Metadata and artifact bytes are stored separately: the graph is the system of record for relationships and state, while the content-addressed store holds raw bytes by digest. And publication authority never leaks into the read plane. The AI-facing gateway can read trusted state, but it cannot create or certify it.

Certification is a state machine Link to heading

Trust is not a flag. Versions move through explicit states with evidence requirements and policy and ontology binding at each transition.

Certification is bound to concrete policy and ontology versions, can expire, and produces its own evidence artifact. Revocation is a first class state, not a comment.

What is real today, and what is next Link to heading

Honesty matters more than a pitch. The service implements, today: lifecycle parity for skills and agents, the five-layer validation stack, versioned policy and ontology packages, canonical JSON and canonical RDF projections with dual detached signatures, an internal Merkle governance log with inclusion proofs and replay, a transparency model with checkpoints, witnesses, quorum, and split-view evidence, PROV-O provenance with traversal and audit replay, semantic compiler ingestion, and a human-reviewed de facto vocabulary workflow.

Still ahead: full W3C SHACL execution over RDF datasets, first-class RDF/OWL normative ontologies, richer de facto ontology mining, fuller external transparency backend interoperability, and richer PROV-O export formats. These are named as future work on purpose, because a governance system that overclaims its own guarantees has already failed its first test.

The through line Link to heading

Every design choice here answers one of the original three failures. Identity and snapshot separation answers “what is this, and is it current.” Layered validation and dual ontologies answer “can I trust this claim.” Canonicalization, transparency, and provenance answer “can I prove it, later, to someone who was not there.”

A catalog tells you that something exists. The bet is that, for agents and skills, you need to know what a thing means, how it was released, and why it should be believed. That is what the seal of approval is for.