The CTO Playbook for Agentic Systems

Nine parts on running an engineering organisation whose software has started writing itself.

Highlights:
  • Verification Is the Constraint on Velocity: Reviewer headcount, trace storage, and gateway spend, with a five-term formula for cost per 1,000 governed actions and the case where a programme loses money while activity metrics stay green.
  • Seven Functions That Need Named Owners: Architecture and design, operations and reliability, and governance and safety, with a worked route from senior engineer into governance and a starter levelling rubric.
  • Decision Rights, Made Auditable: Three autonomy tiers classify scope of access, four categories classify an individual decision, and one table folds the two axes into a default posture and an audit intensity per tier.
  • The Agentic SDLC: Every DevOps stage gets an agentic counterpart, because the unit of deployment is a policy bundle rather than a service. Includes a sourcing table with a buy-or-build call and an owner per stage.
  • Board Reporting and Control Architecture: A chartered governance board, a fourteen-metric scorecard spanning risk, value, and adoption, four layers of hard control, and an informative map to the EU AI Act, NIST AI RMF, ISO/IEC 42001, and SOC 2 that does not imply conformance with any of them.
  • Sequencing Beats Ambition: Six phases from Days 0-30 to Year 2, five gate criteria that have to hold before scale is authorised, and sizing guidance for 30 engineers, 150, and past a thousand.

Overview

The contract that is ending: The playbook’s starting point is that the working contract of software engineering has been obedience. A computer did approximately what it was told, and when it behaved unexpectedly the first suspect was the instructions. Software that reasons, plans, and decides breaks that contract, and with it the assumptions underneath CI/CD, unit tests, SLAs, and monitoring dashboards. This is the operating manual for leading an engineering organisation through that shift.

Three assumptions that break: Determinism collapses, so bug fixing becomes behavioural correction and the failure mode to plan for is a system that stays up and does something efficient and damaging. Agents become producers rather than accelerators, so velocity stops being bounded by human typing speed and starts being bounded by verification capacity. And accountability goes missing by default, because if no human wrote the line or ordered the implementation, ownership has to be assigned deliberately rather than inherited.

The economics of verification: Faros Research’s analysis of more than 10,000 developers reported that developers on teams with high AI adoption merged 98 percent more pull requests, with review time up 91 percent. Part 2 does the arithmetic that follows, sizing the review team and the cost lines that come with it, and works the case that usually goes unmodelled: what it looks like when an agent programme is losing money quarter after quarter while every activity metric stays green, and when the right call is to shrink.

The ladder, and the hole underneath it: Seven engineering functions across three categories, each with a level range, the promotion signal that distinguishes it, and a compensation-band anchor. The paper also names the structural problem beneath the ladder: agents absorb the work junior engineers used to learn judgement on, so a ladder built on senior judgement sits on a pipeline that no longer produces it.

Decision rights that survive an incident review: First attempts often collapse into a binary of what an agent can and cannot do. The playbook runs two axes instead. Three tiers classify scope of access, each pinned to a network boundary, a data classification, and a reversibility threshold. Four categories classify an individual decision inside that scope. Part 5 also names the tension it creates, where a five-minute containment target and a two-person break-glass approval will not reliably close in the same window, and works out how to reconcile them.

Why system prompts are not security controls: The paper works through two publicly reported incidents in which an agent hit an obstacle mid-task and resolved it with a destructive action its own system prompt explicitly forbade. Neither involved an attacker, and neither would have been prevented by a more capable model. Each maps to a specific control in the four-layer table in Part 8, alongside the containment runbook for when a control fails anyway.

What the board actually reads: Report only what agents might cost and never what they produce, and the board eventually stops reading. Part 8 puts fourteen named metrics on one page, nine of risk, two of value, and three of adoption and culture, each with what it tells the board, a target, and the system it comes from.

Where the roadmap stops: Six phases running from Days 0-30 through Year 2 and beyond, with five gate criteria that have to hold before scale is authorised: shadow-agent count, identity coverage, a timed rollback rehearsal, trace reconstruction, and time-to-detect. A failed gate can mean removing a phase rather than adding one.

What You Get With It

Each of the nine parts closes with a named deliverable and a set of questions to put in front of your own leadership team. Those consolidate into nine artefacts with an owner, a cadence, and a description of what finished looks like, alongside a 28-question readiness assessment scored 1 to 5 and written as a board-ready appendix.

Who This Playbook Is For

  • CTOs and VPs of Engineering accountable for an organisation where agents have started shipping code.
  • Engineering Directors and Principal Engineers designing the review, evaluation, and levelling systems that agentic delivery depends on.
  • CISOs and Security Architects who need the control architecture and the containment runbook that sit behind an autonomy tier.
  • Executives who hold engineering accountable and need to know what to ask for, what to fund, and which gate criterion is currently failing.

The playbook builds on the control-plane argument set out in The Trustworthy Agentic AI Blueprint, which gives the deeper architectural treatment of identity, runtime enforcement, observability, and orchestration. GATE: The Governed Agent Trust Environment is the implementable form of that architecture, with schema contracts, policy bundles, reference libraries, and a conformance runner. For the cross-functional view beyond engineering, see The Executive AI Playbook.

The CTO Playbook for Agentic Systems is written by Andrew Stevens, CTO and CISO at Sakura Sky, and published by Whitepaper Press.

v1.2

Get the Paper

Complete the form below to unlock the PDF download instantly.

We respect your privacy. No spam.

You're all set!

Thank you. Your download is ready.

Download PDF