We have all read plenty about how AI is evolving the SDLC: it helps us explore, design, implement, test, and review. Across all these stages, the pattern is the same: the AI provides the data and the options, while the judgment, what the output means and what to act on, remains for humans to form. That was always true. What changed is how much weight it carries: when writing code is cheap and an AI reviewer already catches the obvious bugs and conventions, judgment is most of what is left, and the first thing a team can lose without noticing.
That side of AI, working with AI inside the SDLC, is well-covered ground. What I want to talk about today is a different aspect of how AI can help engineering teams: using it deliberately to protect and grow that judgment, and to accelerate the growth and maturity of the individuals and the team. Here, “workflows” means more than writing code. It is how we measure our own quality as contributors and as reviewers, how we run retros, and how we give each other feedback.
As an engineering manager, I look for the repeatable, high-impact moments in the work of growing people and the team, and build a small piece of AI tooling for each, so that friction is eliminated and only judgment remains. I usually build these pieces together with the engineers, to help drive AI adoption on the team.
All of this rests on one competency, the same one that sets a strong engineer apart: judgment. In engineering, judgment means directing and challenging the AI, questioning the design, and knowing what “good” actually looks like. It matters just as much when the work shifts from writing code to growing people, where it means honestly assessing our own quality, identifying the takeaways that actually matter, and deciding what to change. That is the human-centric part, and it is what these workflows are built to grow.
The habit lives where the work already happens
We wanted growth to be a deliberate habit that lives in the daily flow: quality measured as the work happens, our own and each other’s, and the judgment kept with the person. In James Clear’s book Atomic Habits, he explains how small, consistent changes to your environment shape behavior. One line stuck with us:
You do not rise to the level of your goals, you fall to the level of your systems.
A practice that lives only in someone’s head or in a document nobody opens will not be applied consistently. If we want that habit to be the default, it has to live in the daily environment.
For us, that environment is Slack, where the team already spends its day. So we built Klaudia, an AI assistant that lives inside Slack, and gave it a set of skills: small, purpose-built AI workflows, each triggered by a slash command. For example, /eng-cr-report is a skill that reads the data it needs from GitHub and from Slack threads, and posts a report back into the channel. Most skills fire on their own at the right moment, triggered by events, for instance when a new feature ships. The rest just surface a prompt, so no one has to remember to run them.
Each skill removes the friction: the blank page, the data you have to gather first, the step people keep forgetting. What’s left is the judgment. Remove that friction and the habit is easy to keep. Since the team already collaborates and logs its work in Slack, most of the raw material these skills need is already there: the daily updates, the review threads, the QA and communication back-and-forth. It is a byproduct of how we already work, not something anyone has to produce on the side.
The tools reflect; people decide
There is one rule behind all these skills: each one returns data to a person, and the person decides what to do with it. The tool does not decide. A report is a starting point, not a verdict. It might surface an insight or a recommendation, but that is not ground truth. Two engineers can read the same report and disagree about what it means, and that is okay, because deciding what is true and what to act on is a human job.
Shared channels, shared ownership
By design, most skills post into shared channels the whole team can see. When the data is visible, anyone can act on it.
Take, for example, our #klear-this-is-live channel, where every shipped feature gets announced. Because everyone sees each release, anyone can call for a retro: the owner, the mentor, a Project Manager, or an Engineering Manager like me. They trigger /retro-catalyst, the skill that gathers all the data a feature left behind and drafts a retrospective report from it. Whoever calls for a retro says why, so it lands as a reason to look closer, not an order from above. And the owner is the one who runs it and keeps the takeaways, even when someone else spotted the opportunity.
Two things to measure: quantity and quality
If you run a team on the data these skills surface, you are really answering two questions: how much gets done, and how well it is done. Both apply to each person in two roles: as a contributor and as a reviewer. Together they form a two-dimensional view (the two roles across, quantity and quality down), and each of the four cells has its own measurement skill, as in the diagram below. Quantity and quality need very different treatment, and mixing them is where measurement usually goes wrong.
Quantity is the easier one. Every sprint, the skill /eng-sprint-delivery posts what each engineer merged, weighted by PR size (your output as a contributor), and the skill /eng-sprint-cr-load does the same for who carried the review load. Both are retrospective on purpose: someone looks back and decides to take on more next time, or sees they were overloaded and eases off. We do not rank people, and no single sprint is a verdict: the same numbers roll up into multi-month trends, so what matters is the direction over time. Everyone sees their own numbers in the open and can adjust accordingly.
Quality is the harder one, and it took most of the work.
Reviewer quality
Once the AI reviewers handle the surface stuff, counting a human’s comments tells you almost nothing. What matters is the thinking behind them: is the person challenging the design, or just reacting to what is easy to see? That is what the skill /eng-cr-report looks for.
The skill reads every human review comment on a pull request and rates it on a depth scale we defined, D0 to D4:
- D0, non-review: “LGTM”
- D1, surface: formatting, a rename with no reason, a linter’s job
- D2, craft: reuse, naming, missing tests
- D3, a real defect: a concrete bug, race, or data risk
- D4, systemic: questioning the approach, or something that affects many callers
A deep comment only matters if it changes something. So the report also tracks whether it landed: whether the owner acted on it or answered it. It gets that from the replies in the thread, not the resolved flag, because a resolved thread can mean “fixed”, “answered”, or “quietly ignored”. Response time is tracked too, but kept apart from depth: slow but thorough and fast but shallow are different problems, and averaging them hides both.
Two things keep the report fair. A depth score is the skill’s reading of a comment, not an objective measurement, so every score links to the exact comment it came from, and anyone can challenge it. And the caveats travel with the numbers: depth depends partly on the work you were assigned, volume is not quality, and new joiners are read in their own context rather than against everyone else. The report says so at the top: input for a conversation, not a score, and it stays between the engineer and their manager, for growth.
Contributor quality
The /retro-catalyst skill does the same for a shipped feature. It pulls together the feature’s history from GitHub (PRs, review threads, deploy timing) and, less obviously, from Slack: the async daily updates the team writes as it works, and the long QA threads spread across the feature’s channels. From that, it lays out the feature’s timeline visually, then drafts what is worth keeping and what is worth improving, each point linked to the evidence behind it.
/retro-catalyst report, drawn from the feature's GitHub and Slack history.
The owner and the mentor each read it alone and write, from a blank page, the few takeaways that actually matter. Then they compare notes, and when a takeaway reaches past this one task, they bring it to the team.
Growth belongs to the person, and candor keeps it honest
This is the part I care about most. Before a growth conversation, I want engineers to have already looked at their own full picture. That means the sprint reports for quantity, as a contributor and as a reviewer, and the /eng-cr-report skill for quality. The /retro-catalyst skill adds the task-level view, worth running on a meaningful task even when no formal retro happens. They come in with their own takeaways, ready to set their own goals. That ownership is the point.
But self-reflection has a blind spot: people soften their own patterns, or simply cannot see them, and no report will fix that. What is missing is someone willing to say it. Kim Scott calls that radical candor, from her book of the same name: caring personally about someone while challenging them directly.
Challenging someone directly isn’t just my job, and it isn’t limited to quarterly reviews. I give this feedback in career conversations, but more often in the everyday back-and-forth. The team’s multipliers do the same, to borrow Liz Wiseman’s term from Multipliers: senior engineers and technical mentors who help the people around them grow through direct, honest feedback.
That everyday feedback is where the real power lies. Without it, self-review can drift into a comfortable story; without ownership, challenge becomes nothing more than criticism. Both have to be present.
Build judgment into the workday
None of this is a separate management technique. It is the same way we already build software (gather the data and options, decide with human judgment, iterate as evidence evolves), pointed at how we run the team instead of at the product. The tools do not make the calls; they make the judgment easy to reach, put it in front of everyone involved, and keep it tied to real evidence. The people stay in charge, always.
Environment shapes behavior more than any values statement does. So that is what we spend our time building.
