First post in a series about how I work at Clarity AI with an AI-native workspace. This one covers the experiment and the way of working; the next ones go deeper into specific pieces.
The forest problem
In April I decided to try to see the whole forest and still work on any tree, in any repo, owned by any team, without giving up on quality or on small, safe, sustainable steps. This post is how that went.
As Head of Engineering and Data Platform at Clarity AI, I have to look at the whole system all the time: dozens of repositories, many teams, data pipelines that depend on other data pipelines. If I only look at one repo, I miss what matters. If I only look at the whole, I lose contact with the code, and with it the ability to judge whether what we’re doing makes sense. The usual way out is a trade-off: you see the forest (meetings, documents, dashboards) and stop touching the trees, or you get into one tree and only hear about the rest secondhand. For years I lived somewhere in between, never comfortably.
The experiment: turning up the dial
Kent Beck described Extreme Programming as taking the practices that already work and turning the dial all the way up. If code reviews are good, review all the time. If testing is good, test all the time.
My experiment has two dials that move together. One is the practices I’ve trusted for years: small safe steps, TDD, trunk-based development, everything versioned, checking changes in the running system instead of stopping at the merge. The other is how much work I delegate to AI agents. What I’ve seen so far is that the second dial can only go as high as the first: delegating more without stronger practices just means breaking things faster. That’s an observation from one person over nearly six months, not something I’ve measured, and the rest of the post is the evidence I have for it, including where it falls short.
It started on 16 April 2026 as a git repository just for me, which I called platform-office, with a first commit called “bootstrap platform-office AI-native workspace”. The idea was modest: a place where an agent could find everything it needed, so I could stop re-explaining context in every session. The agents are Claude Code sessions, and the co-author trailers in the commits record which model made each change, and it changed over those months.
The mantra: everything is in the repo
One rule shaped everything else: if it’s not in the repo, it doesn’t exist.
That covers code, tasks, decisions, knowledge, communications with other teams, incidents, recaps, the skills the agents use, the tools those skills call, and even the engineering practices, packaged as skills. The company’s code repositories are not inside; they’re cloned when a task needs them. What I ended up with is a single repository for all the context around the code.
The first commit already had tasks/, decisions/, knowledge/ and five skills to manage them. Today it holds hundreds of tasks, decision records and knowledge documents, plus tracked communications, incidents and 45 skills, all plain Markdown with YAML front matter, readable by a person and parseable by an agent.
Writing all of that down is what lets an agent pick up any thread from a five-word prompt, and a colleague pick it up without asking me.
Keeping it that way meant going against the defaults. Agent harnesses like Claude Code (the program that runs the agent and enforces its permissions and hooks) save what the agent learns as personal memory, one user at a time. I wanted that knowledge to belong to the team and to anyone using the repo, so I had to be explicit: in August the agents were told to write what they learn into files in the repo, and a check blocks the commit if an instruction file points to private memory. On 10 September I moved the last personal memories worth keeping into the repo. A week later a colleague committed a piece of knowledge with the message “out of one laptop’s memory”.
How the work gets done: the skill families
A skill, as I use it here, is a flow with phases and gates (points where it stops and waits for a human decision), backed by deterministic tools in the same repo. There are 45, and they cluster into a few families: orienting the agent in the current state of the work, driving a task from capture to close, operations and incidents, the data platform, knowledge and decisions, cross-team communications, code quality, and setup and housekeeping. Several come from the practice skills I wrote about in Encoding Experience into AI Skills.
In practice, a session starts with a short sentence. Some kick off a flow: “run task X”, “look at this alert”, “let’s do a recap of what happened this week”. Others are open questions: “explain how system X works”, “let’s look at how we could improve the architecture of Y”. Either way, the agent loads the context it needs from the repo. It also looks up our organization inventory (more on it below) to find the code repositories related to the request, and clones them when it needs to read the code to give a better answer. Then it keeps working until it needs a decision from me.
Three of them are worth describing in more detail.
The task skill, po-do-task, runs a task through 14 phases and stops for a human decision only where it matters: before implementing, committing, pushing or talking to another team. The phases are: identify, take ownership, clarify, propose alternatives, implement with TDD, mutation testing, test desiderata, report, commit, push, watch CI, validate in the running system, close, and a mandatory postmortem, where the agent reviews what went wrong in the run and proposes changes to the skill itself. Those changes go back into the skill, which several people have rewritten many times.
The incident skill takes a single trigger (a Slack link, a failing pipeline, a pasted alert) through hypotheses and validation, works back from the immediate cause to the systemic one, and ends with follow-up tasks and a postmortem.
The summary skills look at a window of time instead of a single alarm. A census of the Kubernetes event stream complements alerting: it looks for steady signals that never cross a threshold, and proposes which ones deserve a closer look. A weekly health check runs five sources in parallel and crosses them. The manual sweep behind it found its best signal in that crossing: Airflow DAG issues clustering in the same afternoon window as memory pressure in a heavy batch process, which neither source showed on its own.
The principles, in practice
Looking back, the principles I wanted to follow show up in the commits.
See the forest
One repo holds the context of the whole platform, and the code repos a task needs are a command away. It didn’t start that way: I began with git submodules, and by late June there were dozens of them, copied again into every worktree, and the repo had become slow and heavy. On 12 July a decision record replaced them with plain clones on demand.
One code repository is the exception: our organization inventory, present in every session. It’s where Clarity keeps its governance data: dozens of teams and the repositories, Airflow DAGs, container images, Kubernetes deployments, S3 buckets, runtime applications and models each one owns, kept up to date automatically. The workspace reads it instead of copying it. So the first questions of any investigation are cheap: which repositories, applications and data pipelines relate to this topic, and which teams work on it. Even the catalogue of repos and the list of people who can own a task come from there.
Work on any tree
A task registered here is ours to finish, even when the fix lives in another team’s repository. The default is to diagnose the root cause, make the change and coordinate with the owners, with a tracked communication so they’re never surprised, rather than to open a ticket for them. The workspace brings the context; each repository keeps its own rules. A change in someone else’s repo follows that repo’s conventions, CI, review process and owners, in the same small steps as everything else.
Small, safe, sustainable steps
Two weeks in, the pre-commit hook was already running the whole test suite, and the workspace’s own tooling was being hardened with mutation testing. Everything goes trunk-based. Some things stay deliberately slow: any destructive change to live infrastructure is run by a person, one command at a time. The agent writes the exact command; a human runs it and reads the output.
Agents have no access of their own. They work with the access of whoever starts the session, and the rules every session loads keep the important steps in human hands: nothing is pushed, or sent to Slack or to another team, without an explicit go, and when someone approves a specific command, that exact command is what runs.
Experimentation and impact
From the first commit, every task declares its why and can be framed as a hypothesis. In May a validating state appeared: a merged change stays open until it’s checked in the running system. In August came discarded, because a well-argued “we’re not doing this” is also a result; 128 tasks have been discarded on purpose so far.
Learning
Corrections become rules or tools, and sometimes the right call is neither. In July I swept the agent sessions for repeated frictions and the conclusion was clear: rules that live only in documentation get violated; rules enforced by the harness, through hooks and checks, get followed. So corrections keep moving from text into checks, as the next section shows.
When the system got it wrong
Few of the failures I found were an agent doing something silly. Most were the system breaking in ways I hadn’t designed for, and more often than not it was the agent itself, in the mandatory postmortem, who noticed. Three patterns keep coming back.
Rules written as text get ignored; rules turned into tools get followed. A rule that said “use this fallback, the coverage plugin isn’t installed” was ignored in three separate tasks, with the rule in context every time. The postmortem concluded that a rule which has failed three times as prose becomes a make target. The same thing happened with links left behind when a task file moves: the skill first got a paragraph with a workaround, and then the task tool learned to repoint the links itself. When an instruction keeps failing, I stopped writing a better paragraph and changed the mechanism.
The system also has to check itself, including its own checks. One postmortem wrote its lessons into the wrong checkout and then said they were in the commit. Nothing contradicted it, because the working tree looked clean. Now it has to confirm that each lesson’s file shows up in the right git status. Mutation testing reported kill counts that were fake, because tests that crashed were counted as kills. A green pytest line turned out not to be the gate’s verdict. And when a postmortem proposed two more rules for a single friction, it rejected them, with the reason written down: a skill that grows on every friction stops being read.
Then there are the checks that weren’t what they looked like. A shortcut that was safe when someone wrote it stopped being safe once things around it changed. A check ran on every change and still stopped nothing, because nothing was wired to listen to it. A guard added to a shared tool was bypassed by skills that called the tool directly. In each case the check existed and was documented, but it relied on an assumption that was no longer true.
Each time a failure moved from a paragraph into a mechanism, I could delegate a bit more.
From a repo for me to a repo others use
In April, other people made 1% of the commits. In May it was 18%, and in August 46%, with ten people besides me contributing that month. Thirteen people have committed so far. Use keeps spreading, slowly, to other teams and individuals, although most of the commits still come from a few of us.
What surprised me is that nobody came to “use Eduardo’s tools”. They came with their own problem: running our data pipelines, Snowflake operations, migrating services between platforms, alert noise per squad. The repo gave them context and a mechanism, and several of them then started improving the system: seven of the 45 skills were written by someone else, and one colleague launched an initiative to make the workspace’s context cheaper for every session.
None of them adopted the whole thing, and none of them had to. You don’t need to file tasks, record decisions or run the 14-phase flow to get value from the repo. The pieces are useful on their own and they compose: you can start by reading the knowledge base, pick up one skill weeks later, and register a task or a communication only when it helps. I suspect that’s a big part of why adoption happened without a mandate.
For people who are new to the repo, those open questions are usually the way in. With the context of the whole platform in one place, anyone can open a session and say “this is my situation, I want to do this, how is it done at Clarity?”, and get an answer grounded in our decisions, runbooks and code. A colleague who started using it recently told me it had given him more onboarding than any person had. That kind of use leaves almost no trace in git, so the charts undercount it.
The first time the tooling broke because I wasn’t its only user came in June, when two of us created different decision records with the same number, three times in a week. Sequential IDs assume a single author. Random suffixes on every ID fixed it: a small change, but one you only make when the tool has to work for more than one person.
Then I stepped away. Between 20 August and 1 September I made zero commits, and the repo got 191 commits from six other people.
On 6 September we recorded what was already true: the workspace is open to anyone at Clarity, and nobody has to go through me to use it or change it. I steward it: I maintain the tooling, review what lands on main and run the nightly loops. The deeper changes of direction still tend to come from me, mostly because that’s the easy path, not because anyone has to ask.
What I haven’t solved yet
Two problems grow at the same pace as adoption, and I don’t have good answers for either yet.
Integration with the company’s task system
The repo is the source of truth for our work, but the company runs on Jira, and keeping the two in sync is manual: a task here, its mirror ticket there, keys copied by hand. Many commits cite a Jira key in the subject line alone, and some exist only to say “mirror ticket created” or “comment posted”. One housekeeping skill exists only to stop GitLab from turning our DR-NNN and INC-NNN references into links to Jira tickets that don’t exist. I haven’t decided which direction the sync should go, or how to avoid double bookkeeping without breaking “everything is in the repo”. I’m running a few experiments on exactly that right now.
Pruning, cleaning and compacting the knowledge base
“Everything in the repo” also means everything accumulates. Since April, 602 knowledge documents have been added and 30 deleted, with a single explicit cleanup in July. The instruction file every agent session loads grew from about 430 words to about 4,600. Disk space is the least of it: every session spends part of its context on what nobody has pruned, and outdated knowledge is worse than missing knowledge, because an agent treats it as true. My idea is a weekly routine, like the health check, that detects duplicates, stale documents and contradictions with the code or with newer decisions, and proposes the compaction as a reviewable change. The agent proposes, a human decides. It’s not automated yet.
What I’m not claiming
Commits are activity, not value, and many are the workspace’s own bookkeeping. This is not a controlled experiment: one person driving it, better models arriving during the same months, people who were already hands-on. I haven’t compared the workflow against a small guide alone, so I can’t tell how much of the result is the system and how much is the model doing well by itself. The 191 commits from six people in the 13 days I made none tell me the repo works without me. They don’t tell me people would use it if I weren’t the one asking. Writing everything down has a real cost I feel but haven’t measured. And none of this is autonomous: a human approves every push and runs every destructive change, on purpose.
What I can say is how it feels from the inside. The feedback from the people using it has been very good, and I’ve never been this productive, above all on cross-cutting problems, the ones that touch many teams and repos at once and usually stall. That’s an impression, not a measurement, but it’s a strong one.
What comes next
In the next posts I’ll take apart some of the skill families and some of the subsystems behind them, like communications and task management. I’ll pick the order as I go.
And if you lead a platform or engineering team: which of these problems are you fighting right now?
Related reading
- Encoding Experience into AI Skills
- My Base Setup for Augmented Coding with AI
- Radical Detachment in the AI Era: Reinventing How We Build Software
- The Art of Small Steps in Software Development: A Lean Vision
- Small Safe Steps workshop
- Kent Beck, Extreme Programming Explained
All numbers come from the platform-office git history, from 16 April to 27 September 2026; the failure examples run to early October.





