Blog

  • Exploration: What Happens to the Knowledge Worker When AI Removes the Bottlenecks?

    I learned something useful from performance testing many years ago. When you remove a bottleneck from a system, another bottleneck usually appears.

    That does not mean the optimization failed. The performance of the whole system has increased. The next constraint simply became visible because the previous one no longer dominated.

    I wonder if this is a useful way to think about AI and knowledge work.

    Much of knowledge work is built around expensive activities: finding information, reading it, comparing it, writing, analyzing, programming, documenting, communicating and reviewing. Organizations have developed roles, processes, queues and handovers around those constraints.

    AI is now reducing the cost of several of those activities at once. The obvious conclusion is that fewer people will be needed. That may be true in some cases. But another possibility is that the bottleneck simply moves.

    If software implementation becomes dramatically faster, development capacity may stop being the main constraint. Then other constraints become more visible: deciding what is worth building, understanding customer needs, integrating changes, assuring quality, absorbing change in operations, or learning from what actually happens after release.

    The system may still become much faster overall. That matters because a role is not the same thing as the activities currently associated with it. A Product Manager may spend a lot of time managing a backlog because implementation capacity is scarce. A technical writer may spend a lot of time producing documents because creating and maintaining representations is expensive. A Quality Manager may spend a lot of time collecting and summarizing evidence because that work is slow and fragmented.

    If AI removes those frictions, the activity may shrink without the underlying function disappearing.

    That raises a more interesting question:

    If AI can perform many of the activities associated with a profession, were those activities actually the function of the profession, or merely the mechanisms through which that function happened to be realized?

    For a Product Manager, the deeper function may be understanding context and deciding what should change and why. For a technical writer, it may be preserving and making technical understanding available. For a Quality Manager, it may be maintaining enough understanding of quality, risk and evidence for the organization to make informed decisions. AI may therefore push roles upstream toward their underlying purpose.

    There is also another effect: when something becomes cheaper, we often use more of it.

    If software becomes cheaper to build, more software may become economically viable. If analysis becomes cheap, questions that were never worth investigating may suddenly be worth answering. If quality evidence can be gathered continuously, assurance may expand rather than simply become cheaper.

    So the future may not be today’s amount of knowledge work performed by fewer people.

    It may be far more knowledge work being performed, with different constraints determining where human attention is most valuable.

    And that is where the performance-testing analogy becomes useful again.

    The next bottleneck might be judgment. Or trusted evidence. Or customer attention. Or organizational ability to absorb change. Or something AI will also reduce later.

    The important thing is that removing one constraint does not tell us that the system is finished. It tells us where to look next.

    So perhaps the more useful question is not:

    Will AI replace the knowledge worker?

    But:

    Which bottlenecks currently shape this role, what happens when AI removes them, and what function remains when the old mechanisms are no longer necessary?

  • Retrospective #5: When the Body of Knowledge Started Learning

    After the previous retrospective, the pace of new Practice Notes slowed down considerably.

    At first glance, that might look as if the inquiry itself had slowed down. In reality, almost the opposite happened. I was simply spending less time adding new notes and more time working with what had already accumulated.

    The Practice Notes started feeding temporary synthesis notes, more structured theory, standard assessment and development approaches, LinkedIn posts, and articles for my employer’s blog.

    Retrospective #4 had already suggested that the notes might become the primary body of evidence, while papers, models and guidance became derived representations. What I had not yet understood was that the relationship would not be one-way. 

    Those derived representations started giving something back.

    Trying to synthesize several notes exposed gaps and tensions between them. Turning theory into a standard approach forced abstract ideas to become usable. A LinkedIn post forced an argument to survive without all the context around it. And once those compressed ideas entered public conversations, people from other disciplines began challenging them, extending them, or asking questions I had not considered.

    Some of those questions eventually returned as new Practice Notes.

    The flow had changed. Experience still generated notes, but the notes now also generated synthesis, applications and conversations, and those in turn generated new experience and new questions. The Body of Knowledge was no longer merely accumulating. It had begun to behave like a learning system.

    I think that also explains why it has become increasingly difficult to say what the Body of Knowledge is actually about.

    The blog began in a reasonably recognizable territory. My original questions concerned software quality, assessment, evidence, organizational capabilities and the models we use to understand them. Even then, I described the scope as broader than testing alone, extending toward systems thinking and the conceptual models used to understand complex organizations. 

    But every cycle of inquiry seems to expose another relationship beyond the previous boundary.

    Quality led into risk and evidence. Requirements led into intent. Product development led into operations and continuity. Documentation led into the preservation of understanding. Organizational capabilities led into stewardship. AI raised the possibility that future understanding may not need to be stored primarily in predefined documents at all.

    Even an Exploration of Vivaldi’s Winter Largo followed a strangely familiar path. The interesting question was not merely which notes the bass voice played, but what function those notes performed within the music.

    That recurring habit has become more visible to me lately. When I encounter a mechanism, practice or artifact, I tend to move upstream from it.

    Why does this requirement exist? Why is this criterion in the Definition of Done? Why do we create this document? Why does this capability need to exist?

    The question is not simply a search for root cause. It is an attempt to recover the relationship between mechanism, function, intent and outcome.

    This also helps explain why several recent threads seem to converge around the idea of stewardship. A document, requirement, model or report may be useful, but it is still only a representation. Different artifacts may contain different claims about the same product, while different specialists each hold part of the relevant understanding.

    The challenge then shifts. Instead of trying to preserve one perfect artifact, perhaps we need to preserve enough coherence around the subject itself that people can continue to understand, develop and assess it as circumstances change.

    That idea led to one formulation I keep returning to:

    Contribution follows expertise; stewardship preserves coherence.

    Quality has not disappeared from the work as the scope widened. If anything, the latest notes show how strongly it still anchors the inquiry.

    The recent note on Ready and Done began with a very ordinary Sprint-end problem: tests being negotiated away because work needs to become Done. But asking why those controls existed changed the meaning of the situation. Skipping a test was no longer merely a delivery shortcut. It could be understood as accepting additional risk without the evidence that the test was supposed to provide.

    Following that line of thought into refinement produced another observation: refinement is not only about reducing uncertainty about what to build. It is also about reducing uncertainty about what could go wrong.

    That is still recognizably quality thinking, but it no longer feels bounded by QA. It reaches into product management, risk, decision-making, operations and learning.

    Perhaps that is why Quality Management feels closer to the territory I am moving through, while still not quite describing all of it.

    Looking back across the Body of Knowledge, I increasingly see one part concerned with how we intervene in systems and another concerned with how we understand them. But even that distinction is becoming difficult to maintain. We need understanding in order to intervene responsibly. The intervention produces consequences. Those consequences give us new evidence. The evidence changes our understanding, which should then influence the next intervention.

    The two strands keep folding back into each other.

    And perhaps this is why the subject itself seems harder to define now than it did twenty or forty notes ago.

    Not despite the Body of Knowledge becoming a learning system, but because of it.

    Every cycle seems to reveal relationships that were not visible before. That does not necessarily mean that the subject itself is objectively expanding. It means that my understanding of its boundary keeps changing as the inquiry develops.

    That distinction matters. Retrospective #4 already contained an important warning: as the theory became more ambitious, it became increasingly important to distinguish observation from interpretation and speculation, and not to make an experience prove more than it actually did. 

    So I am not ready to give the Body of Knowledge a new grand title.

    Quality Assurance is clearly too narrow. Quality Management gets closer. Systems thinking, organizational capability, assessment, knowledge management, product management and organizational learning all seem to touch parts of it.

    For now, one question seems to describe the territory better than any label:

    How do we understand changing systems well enough to intervene responsibly, learn from the consequences, and preserve that understanding over time?

    I do not know whether that is what the Body of Knowledge is ultimately about.

    But I think it is now a fair question to ask.

    When I started the blog, I wrote that writing was not merely a way to communicate conclusions. It was a way to discover what I think. 

    That still feels right.

    The difference is that the mechanism has become larger than writing. The Practice Notes feed synthesis, theory, approaches and public discussion. Those representations encounter other contexts and other people. What comes back changes the Body of Knowledge again.

    The previous retrospective ended with the observation that the theory had started talking back.

    Now it seems to be asking a more difficult question:

    What is it actually a theory of?

    And perhaps, for now, it is better that I do not know.

  • Practice Note: When AI Removes the Implementation Friction

    Several years ago, I started experimenting with controlling a model railway through a Roco Z21 command station.

    Looking back, I was really trying to solve three problems at once.

    I wanted to show that coding could be easy and fun by using it to control something physical. I wanted to make a complicated German-language protocol specification more accessible to an international audience by explaining it in English. And I wanted to learn Python and refresh my own coding skills.

    That meant implementing the Z21 protocol was not just scaffolding.

    It was part of the purpose of the exercise.

    My first experiment was simple: retrieve the Z21 serial number. But even that meant understanding the protocol, byte ordering and UDP communication, then translating all of it into Python.

    As the project grew, so did the scaffolding. Incoming packets could contain multiple messages, so I needed to split them. Different messages needed dispatching. Some information arrived asynchronously. Feedback modules reported track occupancy. Locomotive control introduced another protocol carried through the Z21 protocol.

    Along the way I also started experimenting with functional programming, refactoring, behavioral examples and automated tests.

    All of that was useful. But I was simultaneously learning the protocol, learning Python, designing the software, debugging it, documenting it and trying to automate a railway.

    Returning to the same problem

    Several years later, I came back to the project.

    The railway has not fundamentally changed. Neither has the protocol.

    The development environment has.

    Today I can give AI the Z21 protocol specification and ask how to retrieve the serial number. It can find the relevant section, interpret it, explain the request and response and produce a candidate implementation.

    But code generation is only a small part of what changed.

    I could also give AI my old implementations. It found bugs in code I had written, compared my interpretation with the specification, suggested improvements and helped reconstruct what I had been trying to do.

    Where I had previously become stuck, it could suggest possible ways forward.

    When I wanted to develop without depending on the physical command station, it helped design a small Z21 simulator over UDP on localhost.

    When I wanted to return to automated testing, it could help translate acceptance criteria into Python unittest tests.

    And when an answer looked questionable, I could challenge it.

    That matters. AI was not always right. I corrected its reconstruction of the history of some of my old experiments, and I questioned its interpretation of parts of the protocol.

    The point is not that AI provides trustworthy answers automatically.

    The point is that it removes a great deal of friction.

    Three problems that largely disappeared

    The contrast with the original project is striking.

    Previously, I wanted to make coding accessible, translate a German specification for an international audience and learn Python while refreshing my programming skills.

    Today, none of those needs to be a prerequisite for the real experiment.

    Roco published the specification finally in English. AI can help produce and inspect Python. If I encounter a language feature or protocol detail I do not understand, I can investigate it at the moment it becomes relevant.

    Technical knowledge has not disappeared. AI can supply implementation proposals on demand, but my understanding of what good software looks like becomes what allows me to direct, challenge and improve those proposals.

    But much of it has moved from being a prerequisite to being an on-demand resource.

    That changes where I can spend my attention.

    Before, the path often looked like this:

    intent documentation → technical understanding → implementation → debugging → experiment

    Now it can look more like this:

    intent → user story → acceptance criteria → automated tests → implementation → experiment → observation and learning

    All AI assisted. The implementation is still there. So is the technical knowledge behind it. But they no longer need to dominate the work.

    That leaves more room for the questions I actually care about:

    What am I trying to achieve?

    What behavior do I require?

    How will I know that it works?

    What assumptions am I making?

    What did I learn from the result?

    Not just programming

    I recognize the same pattern when working with AI on Practice Notes and other documents.

    I bring intent, observations and source material. AI helps analyze, retrieve, compare, synthesize and draft. I inspect the result, challenge it, provide more evidence and iterate.

    The artifact may be prose instead of Python, but the pattern is similar:

    intent → AI-assisted interpretation → AI-generated candidate artifact → judgment → evidence → revision

    For software, some of that evidence can be executable. User stories preserve intent. Acceptance criteria describe expected behavior. Automated tests make those expectations executable. A simulator provides a controlled test environment. Eventually, the physical model railway layout provides another form of evidence.

    AI can help throughout that chain without becoming the authority on whether the result is correct.

    From implementation to intent

    The first time around, implementing the Z21 protocol was part of the exploration.

    I wanted to understand it, explain it and learn how to implement it.

    This time, I do not particularly need to rediscover how to implement a binary railway protocol.

    I want to explore what I can do with it.

    Earlier, I explored how to implement the Z21 protocol.

    This time, I can explore what I want to do with it.

    And perhaps that is the broader lesson:

    AI does not merely make implementation faster. It can move implementation knowledge from a prerequisite toward an on-demand resource, allowing more attention to move toward intent, evidence, judgment and learning.

  • Practice Note: Ready, Done, and the Risk Between Them

    It is Friday afternoon. The end of the Sprint is approaching. Several stories are not finished and the pressure rises. A few tests could perhaps be skipped. A defect may not be serious enough to hold the story back. Regression testing could be completed later.

    Nobody says, “Let’s accept more product risk so that we can avoid Sprint spillover.” Yet that may be exactly what is happening.

    This made me look differently at the Definition of Done. Why are its criteria there in the first place? If testing, review, security checks or other quality controls are part of the DoD, they should not merely be activities the team has agreed to perform. They should be there because they control risks that matter.

    Seen that way, there are two quite different problems with a Definition of Done.

    The first is that it may not have been defined particularly well. It can become a generic checklist of things teams conventionally do rather than a set of controls related to the actual risks of the product.

    The second is that even a meaningful DoD may become negotiable when delivery pressure rises.

    “Can we skip this test so we can finish the Sprint?” sounds like a delivery question. But if the test exists to control a risk, the real question is closer to:

    What risk was this test controlling, and are we prepared to accept the residual risk without it?

    That may be a perfectly legitimate decision. The point is that it should be recognized as a risk decision. This suggests that risk should be made more explicit in the work itself.

    A user story normally contains intent and acceptance criteria. It could also identify the significant quality risks associated with the story: what might prevent the intended outcome, how likely that is, and how serious the consequences would be.

    Acceptance criteria and risk then serve related but different purposes. Acceptance criteria help us judge whether we implemented what we intended. Risk-based evidence helps us judge whether the important undesirable outcomes have been sufficiently controlled.

    That leads to an interesting reinterpretation of Ready and Done.

    The Definition of Ready becomes a shared checkpoint asking whether we understand the work — including its significant risks — sufficiently to start responsibly.

    The Definition of Done becomes the corresponding checkpoint at the other end: have those significant risks been sufficiently mitigated, and do we have appropriate evidence to support that judgment?

    In simple terms:

    Ready: do we understand the important risks sufficiently to proceed?

    Done: do we have sufficient evidence that those risks have been addressed to proceed further?

    But then another problem appears: stories do not exist independently.

    Story A can be Ready. Story B can be Ready. Story C can be Ready. Yet combining them in the same increment may create risks that none of the stories contains individually. They may change the same component. One may invalidate assumptions made by another. Performance, security, resilience or operational complexity may emerge only when the changes interact. So readiness has an increment aspect too.

    It is not enough to ask whether the individual stories are Ready. We also need to ask whether we understand the significant risks created by combining them.

    The same applies to Done. Three stories can individually satisfy their DoD while the integrated increment still behaves unacceptably.

    A set of Done stories does not automatically make a Done increment.

    As work aggregates, risk can emerge at the new level. Assurance therefore has to aggregate with it.

    The same logic can continue toward release. A completed increment can still be unsuitable for deployment because release and operational risks only become meaningful at that level.

    What started as a Scrum mechanism therefore begins to look more like a broader assurance chain:

    Intent → Risk → Control → Evidence → Judgment

    At each level, we are asking what we are trying to achieve, what could prevent it, what controls those risks, what evidence we have, and whether the remaining risk is acceptable.

    And the chain should not end when the release reaches production. Production gives us new evidence. Some risks materialize despite our controls. Some turn out to have been underestimated. Others were never identified at all.

    Eventually, the production defect may come back to the development team. But the more important question is:

    Did the defect come back, or did the lesson come back?

    If an incident reveals that a risk was missed, misunderstood or inadequately controlled, that knowledge should change future stories, risk assessments, controls, and perhaps even what Ready and Done mean.

    The loop becomes something like:

    Ready → Develop → Done → Integrate → Release → Operate → Learn → Ready…

    The product continues, and so should our understanding of its quality.

    The principle that seems to emerge is simple:

    As work aggregates, assurance must aggregate with it.

    That turns Ready and Done from static checklists into checkpoints in a continuous quality-learning system.

  • Practice Note: Reference Model as the Antidote to Understanding Debt?

    This morning I woke up with a rather practical question.

    Since my laptop is getting older but is still considered rather capable, I wondered how suitable it is to run a language model locally on this machine. And I wondered what a local LLM could be used for. A specific use case came up quickly:

    Organizations accumulate enormous amounts of material over time: architecture descriptions, requirements, decision records, source repositories, project documentation, incident histories, assessments, reports, presentations, and wikis. Could a local LLM collect, analyze, and synthesize such material into a body of knowledge that could then be interrogated?

    We had already explored the idea that this might reduce the need to keep producing new documents simply to carry existing insights into another context. Instead of anticipating every future question and writing a document to answer it, perhaps we could maintain the knowledge and construct the answer when somebody actually asks.

    There is an attractive side effect to this. Documents start aging almost as soon as they are published. If the underlying body of knowledge changes instead, the same question asked six months later can produce an answer based on what is known six months later. The emphasis shifts from periodically republishing knowledge to maintaining something that can continuously be interrogated.

    Reconstructing a Product Reference Model

    I had already been thinking about this in relation to Product Reference Models. An existing product tends to leave a substantial trail of evidence behind it: architecture, requirements, code, decisions, incidents, operational information, and many other artifacts. Rather than manually recreating an understanding of the product from scratch, perhaps AI could work across those existing sources and help reconstruct a Product Reference Model.

    The interesting part would not simply be collecting links. A useful model would also need the relationships between those sources and, particularly, the rationale for their inclusion. Why does this architecture repository matter? What does this incident history tell us about the product? Why is an old decision record still relevant? Where do different sources contradict each other, and where are there white spots for which no useful evidence appears to exist?

    A list of artifacts gives us information. The relationships and rationale between them begin to give us understanding.

    Such a Product Reference Model would not necessarily need to become another physical document. It could remain connected to its underlying sources and be synthesized when required. If a production incident exposed a previously unknown dependency, or a new business need resulted in changes to architecture and code, the evidence available to the model would change with it. The next interrogation could therefore reflect a different understanding of the product without somebody first having to rewrite a reference document.

    From products to capabilities

    That thought made me look differently at another Reference Model I had been working with.

    In parallel, I had been developing a concept Quality Assurance Capability Reference Model. Unlike the rather distributed Product Reference Model I was imagining, this one is very tangible: a spreadsheet with multiple views of the capability, including dimensions and maturity characteristics, stakeholder perspectives, an evidence landscape, indicators, anti-patterns, and example practices. It is intended to support both capability assessment and capability development.

    If a Product Reference Model could become interrogable, why couldn’t this one?

    An assessor could ask what stronger capability looks like in a particular dimension, which stakeholder perspectives might be relevant, or what evidence could challenge an emerging interpretation. Someone working with capability development could interrogate the same model from a different perspective. The spreadsheet might remain useful for maintaining the model, but it would no longer have to be the primary way in which people consume it.

    There is an additional twist here. Assessment and development do not only use the Reference Model; they also generate evidence about the subject it describes. An intervention may succeed, fail, or work for entirely different reasons than expected. Repeated assessments may reveal a pattern that the existing model does not explain particularly well.

    So a Reference Model is not simply a stable body of knowledge from which answers are retrieved. It participates in a loop. The model helps us interpret evidence, while evidence can support, extend, or challenge the model. The Reference Model therefore becomes part of a learning loop.

    And that started to connect the local-LLM thought experiment with another idea we had already been approaching: stewardship.

    From management to stewardship

    Management disciplines such as product management and portfolio management necessarily spend considerable attention on the things associated with their subject: decisions, processes, plans, repositories, ownership structures, and other artifacts. But even when those things are managed well, the organization can still gradually lose its accumulated understanding of the subject itself.

    Stewardship suggests a different concern:

    Can the organization continue to understand a subject as people leave, evidence accumulates, and reality changes?

    That question exposes something slightly awkward. Where does persistent organizational understanding actually exist?

    Some of it exists in people’s heads, some in documents and systems, some in evidence, and a great deal of it exists in the relationships between those things. We had already suspected that Reference Models might play a foundational role in stewardship, but I think the reason for that is now becoming clearer.

    Reference Model = Understanding?

    We normally describe a Reference Model as a representation of our understanding.

    Perhaps that distinction is unnecessary.

    If a Reference Model maintains what matters about a subject, how things relate, why they matter, what evidence supports them, where uncertainty remains, and how all of that changes as we learn, then what separate thing are we calling “the understanding”?

    Perhaps the Subject Reference Model is the maintained organizational understanding of the subject.

    That does not mean every fact or artifact needs to be contained inside it. The Product Reference Model demonstrates why. Much of its detailed knowledge can remain in the sources to which it refers. Knowing that a source matters, understanding why it matters, and maintaining its relationships to other parts of the subject can itself be part of the model.

    A Capability Reference Model may encode much more of its domain knowledge explicitly. The physical forms are different because the subjects are different, but the fundamental function may be the same.

    A Product Reference Model is maintained product understanding. A Capability Reference Model is maintained capability understanding.

    More generally:

    Subject Reference Model = Subject Understanding

    And that connects to a problem we have encountered from an entirely different direction.

    Understanding debt

    We have been using the term understanding debt for something that documentation alone does not seem to solve. Organizations are extraordinarily good at accumulating information while gradually losing the context, relationships, and rationale that make the information understandable.

    The architecture document remains, but nobody remembers why the architecture was chosen. The old system remains, but nobody is entirely sure what depends on it. A practice continues long after its original purpose has disappeared from organizational memory. A decision can still be found, but the alternatives that were considered and the reasons they were rejected are gone.

    Nothing necessarily went undocumented. There may be more documentation than ever.

    What disappeared was the understanding that connected it.

    If the Subject Reference Model is maintained understanding, then understanding debt becomes easier to describe. It is the gap between what the organization needs to understand about a subject and what its maintained Subject Reference Model actually allows it to understand.

    Seen that way, Reference Models become rather more important than useful assessment frameworks, maturity models, or knowledge structures. They become a mechanism for allowing organizational understanding to persist beyond the people who happen to hold it at a particular moment.

    AI may make Reference Models more important, not less

    It would be easy to assume that increasingly capable AI makes Reference Models less important. If an AI can read everything, why bother maintaining a model?

    I suspect the opposite may be true.

    Giving AI access to everything gives it information: twenty years of documents, code, incidents, presentations, abandoned decisions, obsolete descriptions, and contradictory claims. Without some maintained conception of what matters and how it fits together, the AI has to reconstruct an understanding from that history every time we ask a question.

    A stewarded Reference Model gives that information meaning and continuity. AI, in turn, may make it practical to reconstruct such models from existing evidence, interrogate them rather than continually publishing new documents, and help keep them aligned with new evidence as the organization learns.

    The interesting role of AI may therefore not be producing more organizational knowledge at all. It may be helping us maintain the understanding underneath it.

    If Subject Reference Model = Subject Understanding, then stewardship of Reference Models is stewardship of organizational understanding.

    And that may make the Reference Model our best antidote yet to understanding debt.