Seven experiments → a governed execution model.
The biggest shift was not better prompting. It was moving from supervising individual steps to designing the system around the work: desired outcome, persistent context, evidence, guardrails, delegation, verification and human decision points.
7 problem typessame operating principles tested across implementation, tools and decision work
Code → decisionthe same principles tested in implementation and knowledge work
Bounded autonomyincreased only where outcomes were verifiable and reversible
Evidence boundary: six cases were personal projects built on my own time. Rovo / Project Sources is the case applied in a real enterprise work environment. Personal experiments are not presented as company production solutions or measured enterprise productivity results.
Projects as learning environments
The number of projects is not the outcome. Their value was forcing the same working principles through very different constraints, tools and quality criteria.
Enterprise context
Rovo / Project Sources
Decision log, evidence/status rules and persistent context in a real enterprise work environment.
Lesson: Source authority has to be explicit. Design intent is not evidence of production behaviour.
External engine + runtime
PDF Checker
veraPDF engine and Docker runtime. A deterministic rules engine became an independent verification layer for a bounded requirement set.
Lesson: Deterministic verification strengthened the workflow because acceptance could be checked independently of the model.
API + data visualization
Ajokeli
APIs, data processing, map visualization and UI. A useful baseline for more conventional full-stack/data work.
Lesson: A conventional full-stack/data workflow provided a useful baseline for testing what AI changed without adding unusual technical constraints.
Full-stack + data
Satsi
Cloudflare D1, schema, auth and domain modelling. The agent had to work across persistent data, application logic and UX.
Lesson: Persistent state made schema, domain rules and acceptance criteria part of the implementation boundary, not just setup details.
Visual / 3D · 3 iterations
Ghostlight
Three development rounds on the same idea. Quality improved through deliberate planning, iteration and visual review, not merely by reaching a working version.
Lesson: A working version was not a useful quality threshold. Planning, iteration and visual review materially changed the result.
Tool use / MCP
Blender + MCP
Connecting an LLM to an external 3D application. Tool access increased capability while making permission, data and trust boundaries part of the design.
Lesson: Tool access increased capability, but permissions, data boundaries and reversibility became part of the implementation problem.
Decision support
Capital investment decision analysis
RFP structure, financing analysis, calculator and decision material. Facts, calculations and interpretation were kept separate.
Lesson: Decision-support work required facts, calculations and interpretation to remain separately traceable.
A concrete workflow shift: repo-aware work brought code, tests and diff review into the same loop. Tool access later allowed the agent to perform bounded environment actions instead of only explaining how a human should do them.
Context > promptCurrent state, decisions and evidence matter more than one perfect prompt.
Verification > self-report“Done” is not done before tests, diff, visual review or another independent signal.
Delegate what is boundedVerifiability and reversibility determine how much of a task can be delegated.
Agentic ≠ defaultOne agent, subagents, conventional automation or a human should be chosen by task.
What this means for organizations
The practical lesson is less about a particular model and more about how work, control and evidence need to be designed around AI.
01 · LeverageMore capacity, not automatic productivityAI can increase how much a small team can explore and execute. The business case still needs a baseline and measured outcomes.
02 · ManagementSupervise the system, not every stepLeadership shifts toward outcomes, context, decision rights, controls, verification and escalation points.
03 · SelectionNot every workflow should be agenticThe strongest candidates combine meaningful manual effort with bounded scope, verifiable outputs and reversible actions.
04 · ScaleEvidence before rolloutBaseline the current process, pilot, compare quality and effort, then scale, modify or stop based on evidence.
Governed execution model
From business goal to governed execution
This is the execution pattern the experiments converged on: a practical way to structure AI-assisted work so that increasing autonomy remains bounded, verifiable and controllable.
1 · OutcomeUser or business result
2 · ContextCurrent state, decisions and evidence
3 · GuardrailsScope, permissions and acceptance criteria
4 · DelegationHuman, automation, one agent or roles
5 · ExecuteResearch, implement, test and iterate
6 · VerifyIndependent tests, diff, evidence or review
7 · Human gateDecision, merge or production action by risk
8 · MemoryUpdate decisions, state and open questions
Enterprise boundary: in production, this execution model would sit inside existing organizational controls for data protection, information security, regulatory obligations, vendor risk and decision rights rather than replace them.
Next step in a real work environment
A bounded end-to-end agentic delivery pilot
The next step would be to test the same governed operating model across the full delivery path of a small, recurring and reversible change type. The pilot case would be selected so that its outcome can be verified through tests, rules or metrics — not reviewer judgement alone.
Selection criterionRecurringEnough comparable cases for a real baseline
Selection criterionReversibleBounded blast radius and a recoverable change
Selection criterionVerifiableTests, rules or metrics — not judgement alone
Scope and autonomy would increase only where evidence supports progression. See the pilot design and measurement model →
Leadership takeaway: the goal is not maximum autonomy, but the highest level of autonomy in enterprise delivery that is safe, verifiable and demonstrably valuable.