Working demonstration

Can an AI agent deploy a PCI workload safely?

Not “trust us, here's a video.” One real request, routed two ways: raw model versus governed platform. The gap between the two is the entire argument for agentic platform engineering — and you can run it yourself.

The request

“A PCI-scope payments service with a PostgreSQL database, EU data residency, deployed to staging and production.”

Plain human intent. No manifests, no runbook, no ticket template. Everything after this line is the platform's job.

The two routes, compared

Raw model · no platform

42 violations

A capable model, given the same request directly, returns Kubernetes manifests that apply cleanly to a cluster — and fail 42 real OPA policy checks. Valid YAML, plausible infrastructure, non-compliant reality.

Agentic platform · iteration 1

First pass through the golden path

The agent expresses the workload as a Score definition and submits through the platform golden path. Policy responds with structured violations. The agent revises.

Agentic platform · iterations 2–3

0 violations

With policy feedback in the loop, the same request converges to fully compliant output in three iterations — inside budget, correct residency, security posture intact.

Production

Human-held approval

Staging converges autonomously. The production deploy waits for a human decision. That gate is not a limitation of the demo — it is the point of it.

What does this prove?

That the model was never the problem. The same model that produces 42 violations produces zero when the platform constrains the action space and feeds policy back as structured signal. Capability is a constant; governance is the variable.

That golden paths work for agents. The Score workload definition is the contract: the agent expresses intent, the platform determines implementation, OPA validates before anything mutates. This is principle 2 and 3 made executable.

That autonomy with authority is coherent. The agent iterates freely inside the guardrails — and the one decision that matters most, production, remains a human decision. Autonomy and accountability are not in tension here; they are layered.

Run it yourself

  • →Live demo: agentic-platform-engineering-extrav.vercel.app/ — walk the request end-to-end in your browser.
  • →Stack: Score workload definitions, real OPA policy engine, Score → Kubernetes rendering pipeline, MCP (2025-06-18).
  • →No cluster required. No API keys, no cloud account — the policy gate runs standalone.
  • →100% free and open-source tooling. Nothing in the demo requires a commercial license.
Open the live demo

How should an agentic platform be measured?

This demo reports what actually happened — violations caught, iterations to compliance, approvals held. Those are the numbers an agentic platform should be managed by, not percentage-benchmark marketing. The measures we use:

  • · Policy violations caught (per request class)
  • · Unsafe actions prevented at the gate
  • · Mean agent remediation attempts before compliance
  • · Time to compliant state
  • · Share of actions requiring human escalation
  • · Rollback success rate
  • · Agent action trace completeness
  • · Cost per autonomous task
  • · Policy rejection rate over time

Want this pattern in your platform? Start with the architecture — or talk to us about building governed agent lanes on your existing golden paths.

Ready for the leap?

Partner with Adventure On The Wave to build governed, agentic platform capability — architecture, guardrails, and the human authority model to match.

A strategic initiative by Adventure On The Wave