Seven experiments → a governed execution model.
The biggest shift was not better prompting. It was moving from supervising individual steps to designing the system around the work: desired outcome, persistent context, evidence, guardrails, delegation, verification and human decision points.
7different problem types tested in one summer
Code → decisionthe same principles tested in implementation and knowledge work
Human controlautonomy increased only where outcomes were bounded, verifiable and reversible
Projects as learning environments
The number of projects is not the outcome. Their value was forcing the same working principles through very different constraints, tools and quality criteria.
Enterprise context
Rovo / Project Sources
Decision log, evidence/status rules and persistent context in a real enterprise work environment.
What this may enable: Managing context and source authority may matter as much as model capability for enterprise AI quality.
External engine + runtime
PDF Checker
veraPDF engine and Docker runtime. A deterministic rules engine became an independent verification layer for a bounded requirement set.
What this may enable: controlled AI workflows become stronger when model output can be checked against deterministic evidence.
API + data visualization
Ajokeli
APIs, data processing, map visualization and UI. A useful baseline for more conventional full-stack/data work.
What this may enable: familiar workflows are good places to measure what AI changes before attempting broader autonomy.
Full-stack + data
Satsi
Cloudflare D1, schema, auth and domain modelling. The agent had to work across persistent data, application logic and UX.
What this may enable: small teams can test service ideas faster when design, implementation and verification share the same context.
Visual / 3D · 3 iterations
Ghostlight
Three development rounds on the same idea. Quality improved through deliberate planning, iteration and visual review, not merely by reaching a working version.
What this may enable: Higher AI throughput may be of limited value unless the workflow also contains a quality loop.
Tool use / MCP
Blender + MCP
Connecting an LLM to an external 3D application. Tool access increased capability while making permission, data and trust boundaries part of the design.
What this may enable: more autonomy requires more explicit control of identities, tools and reversible actions.
Decision support
Capital investment decision analysis
RFP structure, financing analysis, calculator and decision material. Facts, calculations and interpretation were kept separate.
What this may enable: the same principles extend beyond coding into decision support when traceability is preserved.
A concrete workflow shift: repo-aware work brought code, tests and diff review into the same loop. Tool access later allowed the agent to perform bounded environment actions instead of only explaining how a human should do them.
Context > promptCurrent state, decisions and evidence matter more than one perfect prompt.
Verification > self-report“Done” is not done before tests, diff, visual review or another independent signal.
Delegate what is boundedVerifiability and reversibility determine how much of a task can be delegated.
Agentic ≠ defaultOne agent, subagents, conventional automation or a human should be chosen by task.
What this means for organizations
The practical lesson is less about a particular model and more about how work, control and evidence need to be designed around AI.
01 · LeverageMore capacity, not automatic productivityAI can increase how much a small team can explore and execute. The business case still needs a baseline and measured outcomes.
02 · ManagementSupervise the system, not every stepLeadership shifts toward outcomes, context, decision rights, controls, verification and escalation points.
03 · SelectionNot every workflow should be agenticThe strongest candidates combine meaningful manual effort with bounded scope, verifiable outputs and reversible actions.
04 · ScaleEvidence before rolloutBaseline the current process, pilot, compare quality and effort, then scale, modify or stop based on evidence.
Governed execution model
From business goal to governed execution
This is the execution pattern the experiments converged on: a practical way to structure AI-assisted work so that increasing autonomy remains bounded, verifiable and controllable.
1 · OutcomeUser or business result
2 · ContextCurrent state, decisions and evidence
3 · GuardrailsScope, permissions and acceptance criteria
4 · DelegationHuman, automation, one agent or roles
5 · ExecuteResearch, implement, test and iterate
6 · VerifyIndependent tests, diff, evidence or review
7 · Human gateDecision, merge or production action by risk
8 · MemoryUpdate decisions, state and open questions
Scope: six cases were personal projects built on my own time. Rovo / Project Sources is the case applied in a real enterprise work environment. Personal experiments are not presented as company production solutions or measured enterprise productivity results. Enterprise boundary: in production, this execution model would sit inside existing organizational controls for data protection, information security, regulatory obligations, vendor risk and decision rights rather than replace them.
Next step in a real work environment
A bounded end-to-end agentic delivery pilot
The next step would be to test the same governed operating model across the full delivery path of a small, recurring and reversible change type. The pilot case would be selected so that its outcome can be verified through tests, rules or metrics — not reviewer judgement alone.
Selection criterionRecurringEnough comparable cases for a real baseline
Selection criterionReversibleBounded blast radius and a recoverable change
Selection criterionVerifiableTests, rules or metrics — not judgement alone
Needdefined outcome
Impactaffected scopeHuman gate · impact approval
Designsolution directionHuman gate · solution / architecture
Buildbounded execution
Testindependent evidence
Reviewed candidatehuman-owned releaseHuman gate · release decision
AI would operate within a defined task scope, tool set and permissions. The pilot would progress in stages, with explicit human decision gates retained at the points where judgement and accountability matter.
1 · Baselinecycle time, human effort split between production and verification, quality and cost
2 · Bounded workflowcontext, evidence rules, permissions, agent roles and independent verification
3 · Comparethe same measures plus missed impacts, rework, traceability and control effectiveness
4 · Decideevaluate scope and autonomy separately; modify or stop if evidence does not support progression
Leadership takeaway: the goal is not maximum autonomy, but the highest level of autonomy in enterprise delivery that is safe, verifiable and demonstrably valuable.