Category: Practice Notes

  • Practice Note: Reference Model as the Antidote to Understanding Debt?

    This morning I woke up with a rather practical question.

    Since my laptop is getting older but is still considered rather capable, I wondered how suitable it is to run a language model locally on this machine. And I wondered what a local LLM could be used for. A specific use case came up quickly:

    Organizations accumulate enormous amounts of material over time: architecture descriptions, requirements, decision records, source repositories, project documentation, incident histories, assessments, reports, presentations, and wikis. Could a local LLM collect, analyze, and synthesize such material into a body of knowledge that could then be interrogated?

    We had already explored the idea that this might reduce the need to keep producing new documents simply to carry existing insights into another context. Instead of anticipating every future question and writing a document to answer it, perhaps we could maintain the knowledge and construct the answer when somebody actually asks.

    There is an attractive side effect to this. Documents start aging almost as soon as they are published. If the underlying body of knowledge changes instead, the same question asked six months later can produce an answer based on what is known six months later. The emphasis shifts from periodically republishing knowledge to maintaining something that can continuously be interrogated.

    Reconstructing a Product Reference Model

    I had already been thinking about this in relation to Product Reference Models. An existing product tends to leave a substantial trail of evidence behind it: architecture, requirements, code, decisions, incidents, operational information, and many other artifacts. Rather than manually recreating an understanding of the product from scratch, perhaps AI could work across those existing sources and help reconstruct a Product Reference Model.

    The interesting part would not simply be collecting links. A useful model would also need the relationships between those sources and, particularly, the rationale for their inclusion. Why does this architecture repository matter? What does this incident history tell us about the product? Why is an old decision record still relevant? Where do different sources contradict each other, and where are there white spots for which no useful evidence appears to exist?

    A list of artifacts gives us information. The relationships and rationale between them begin to give us understanding.

    Such a Product Reference Model would not necessarily need to become another physical document. It could remain connected to its underlying sources and be synthesized when required. If a production incident exposed a previously unknown dependency, or a new business need resulted in changes to architecture and code, the evidence available to the model would change with it. The next interrogation could therefore reflect a different understanding of the product without somebody first having to rewrite a reference document.

    From products to capabilities

    That thought made me look differently at another Reference Model I had been working with.

    In parallel, I had been developing a concept Quality Assurance Capability Reference Model. Unlike the rather distributed Product Reference Model I was imagining, this one is very tangible: a spreadsheet with multiple views of the capability, including dimensions and maturity characteristics, stakeholder perspectives, an evidence landscape, indicators, anti-patterns, and example practices. It is intended to support both capability assessment and capability development.

    If a Product Reference Model could become interrogable, why couldn’t this one?

    An assessor could ask what stronger capability looks like in a particular dimension, which stakeholder perspectives might be relevant, or what evidence could challenge an emerging interpretation. Someone working with capability development could interrogate the same model from a different perspective. The spreadsheet might remain useful for maintaining the model, but it would no longer have to be the primary way in which people consume it.

    There is an additional twist here. Assessment and development do not only use the Reference Model; they also generate evidence about the subject it describes. An intervention may succeed, fail, or work for entirely different reasons than expected. Repeated assessments may reveal a pattern that the existing model does not explain particularly well.

    So a Reference Model is not simply a stable body of knowledge from which answers are retrieved. It participates in a loop. The model helps us interpret evidence, while evidence can support, extend, or challenge the model. The Reference Model therefore becomes part of a learning loop.

    And that started to connect the local-LLM thought experiment with another idea we had already been approaching: stewardship.

    From management to stewardship

    Management disciplines such as product management and portfolio management necessarily spend considerable attention on the things associated with their subject: decisions, processes, plans, repositories, ownership structures, and other artifacts. But even when those things are managed well, the organization can still gradually lose its accumulated understanding of the subject itself.

    Stewardship suggests a different concern:

    Can the organization continue to understand a subject as people leave, evidence accumulates, and reality changes?

    That question exposes something slightly awkward. Where does persistent organizational understanding actually exist?

    Some of it exists in people’s heads, some in documents and systems, some in evidence, and a great deal of it exists in the relationships between those things. We had already suspected that Reference Models might play a foundational role in stewardship, but I think the reason for that is now becoming clearer.

    Reference Model = Understanding?

    We normally describe a Reference Model as a representation of our understanding.

    Perhaps that distinction is unnecessary.

    If a Reference Model maintains what matters about a subject, how things relate, why they matter, what evidence supports them, where uncertainty remains, and how all of that changes as we learn, then what separate thing are we calling “the understanding”?

    Perhaps the Subject Reference Model is the maintained organizational understanding of the subject.

    That does not mean every fact or artifact needs to be contained inside it. The Product Reference Model demonstrates why. Much of its detailed knowledge can remain in the sources to which it refers. Knowing that a source matters, understanding why it matters, and maintaining its relationships to other parts of the subject can itself be part of the model.

    A Capability Reference Model may encode much more of its domain knowledge explicitly. The physical forms are different because the subjects are different, but the fundamental function may be the same.

    A Product Reference Model is maintained product understanding. A Capability Reference Model is maintained capability understanding.

    More generally:

    Subject Reference Model = Subject Understanding

    And that connects to a problem we have encountered from an entirely different direction.

    Understanding debt

    We have been using the term understanding debt for something that documentation alone does not seem to solve. Organizations are extraordinarily good at accumulating information while gradually losing the context, relationships, and rationale that make the information understandable.

    The architecture document remains, but nobody remembers why the architecture was chosen. The old system remains, but nobody is entirely sure what depends on it. A practice continues long after its original purpose has disappeared from organizational memory. A decision can still be found, but the alternatives that were considered and the reasons they were rejected are gone.

    Nothing necessarily went undocumented. There may be more documentation than ever.

    What disappeared was the understanding that connected it.

    If the Subject Reference Model is maintained understanding, then understanding debt becomes easier to describe. It is the gap between what the organization needs to understand about a subject and what its maintained Subject Reference Model actually allows it to understand.

    Seen that way, Reference Models become rather more important than useful assessment frameworks, maturity models, or knowledge structures. They become a mechanism for allowing organizational understanding to persist beyond the people who happen to hold it at a particular moment.

    AI may make Reference Models more important, not less

    It would be easy to assume that increasingly capable AI makes Reference Models less important. If an AI can read everything, why bother maintaining a model?

    I suspect the opposite may be true.

    Giving AI access to everything gives it information: twenty years of documents, code, incidents, presentations, abandoned decisions, obsolete descriptions, and contradictory claims. Without some maintained conception of what matters and how it fits together, the AI has to reconstruct an understanding from that history every time we ask a question.

    A stewarded Reference Model gives that information meaning and continuity. AI, in turn, may make it practical to reconstruct such models from existing evidence, interrogate them rather than continually publishing new documents, and help keep them aligned with new evidence as the organization learns.

    The interesting role of AI may therefore not be producing more organizational knowledge at all. It may be helping us maintain the understanding underneath it.

    If Subject Reference Model = Subject Understanding, then stewardship of Reference Models is stewardship of organizational understanding.

    And that may make the Reference Model our best antidote yet to understanding debt.

  • Practice Note: Attacking Vivaldi’s Winter Largo with a Bass Guitar

    This experiment begins with a problem: I’m trying to learn bass guitar, but much traditional beginner material isn’t speaking the language in which I understand music. Memorizing the notes on each string, working around open strings and the first few frets, and droning through repetitive pop progressions may get me playing quickly, but it doesn’t particularly make me want to play.

    What does fascinate me is a bass that goes somewhere. John Deacon waltzing and pirouetting through The Millionaire Waltz is the obvious example: the bass isn’t merely underneath the music. It has a voice, direction and destination.

    I can’t play like Deacon. Not remotely. But perhaps I can begin to understand why he plays what he plays.

    And I already have a language for doing that. Give me the key of a piece of music and enough time to locate the root on the bass, and I can think in degrees. A bass line becomes not a sequence of unrelated fret positions or note names, but something like:

    Root → 7 → 6 → 5 → 4 → 5

    Those are relationships I can both understand and find geometrically on the fretboard.

    Then an idea came up when listening to Vivaldi. The Largo from Winter lasts only about two minutes. It is slow, technically approachable, extraordinarily beautiful and almost comically soothing: lyrical violin, gently plucked raindrops, a warm harmonic bed underneath.

    Perfect beginner material. Or so I thought. Except the thing is deceptive.

    Start by finding the root: E-flat. There is home. And almost immediately, the music leads away toward 5.

    Then something peculiar happens. It doesn’t give us the satisfying return to the root we might expect. Instead, we spend enough time around 5 that 5 begins to acquire the psychological weight of a new tonal center.

    Now the map has changed. And that means something wonderfully strange has happened to the original root: relative to our newly established tonal center, it has become a 4. The pitch that once meant home is still there, but merely as a passing station on the route from the new tonal center to the new 5.

    Then Vivaldi plays another trick: the second half is largely a repeat of the first, but from a changed tonal perspective. And when we approach its ending, Vivaldi doesn’t simply give us the same destination again. The route changes.

    Now two things conflict.

    Memory says: I know where this goes.

    Harmony says: Apparently you don’t.

    And underneath one of the most soothing pieces of music imaginable, some small part of the listener’s brain is going: Where are we?

    Then Vivaldi finally does it in the last measures: the original tonic becomes tonic again. Not merely the same pitch — the same function. And suddenly the map snaps back into place.

    The ultimate resolution!

    This experiment produced something valuable on its own terms.

    I attacked a soothing two-minute Baroque masterpiece with a bass guitar because I was looking for a more interesting way to learn the instrument.

    I discovered that a technically approachable bass part could teach me about tonal centers, harmonic function, expectation, repetition, reinterpretation and resolution — simply by asking where the root is and listening carefully to what happens when it stops behaving like home.

    I still can’t pirouette around the fretboard like John Deacon. But with exercises like this I may get slightly closer… and learn a lot about music beyond what is usually considered the bass guitar’s domain. That is a win to me!

  • Practice Note: Product Requirements for Operations

    We are used to translating product intent into requirements for development.

    We want to achieve something, so we determine what the product needs to do. Development implements it. Testing gives us evidence that it works. Eventually, we decide that the product is good enough to release.

    But development is only a means to an end.

    We do not build $shiny_new_function because the world needs another successfully implemented function. We build it because we expect something to happen once people start using it.

    Suppose $shiny_new_function has been implemented correctly. It produces the right results, performs well, is secure, handles the expected load and passes all its tests. Development has done an excellent job.

    Now suppose that, according to the original intent, success also meant that the function would be used at least 500 times per day.

    Three months after release, it is being used 17 times per day.

    Is the product good? That is a surprisingly different question from whether it was implemented correctly.

    Before release, we could test whether the function can handle 500 uses per day. We could simulate the load and collect evidence that the implementation is capable of doing what we expect.

    But no test can establish that real users will actually use it 500 times per day. For that, we need reality.

    This suggests that when we translate intent into requirements, we may be doing only half the job.

    We are quite accustomed to asking:

    What do we need to build to achieve this intent?

    Perhaps we should simultaneously ask:

    What would we need to observe in operation to know whether we actually achieved it?

    The first question gives Development something to realize.

    The second gives Operations something to observe.

    If the intent is to reduce customer effort through a new self-service capability, we will probably derive requirements describing what that capability must do.

    But if reduced customer effort is really what matters, we should also know what we expect to see once the capability is operating. Are customers actually using it? Are they completing the process? Where are they abandoning it? Are fewer customers contacting support? Has customer effort actually decreased?

    Those aren’t merely interesting metrics to add to an operational dashboard. They are evidence about whether the product is doing what we created it to do. And that means we may need to think about them before the product reaches production.

    If we want to know how $shiny_new_function is being used, we need to make that usage observable. If we want to understand where users abandon a process, we need to capture the relevant events. If we want to know whether self-service reduces support demand, we need some way of connecting those pieces of information.

    Otherwise, we may arrive in production and discover that we cannot answer one of the most important questions about the product:

    Is it actually working as intended?

    Not whether the software is running. Not whether the implementation satisfies its requirements. Whether the thing we built is actually producing the effect for which we built it.

    This is where production becomes particularly interesting for quality.

    Before release, much of our evidence necessarily concerns what we expect to happen. Requirements, reviews, tests and simulations can give us confidence that our realization is capable of fulfilling the intent.

    After release, we get something we did not have before. We get to see what actually happens. And reality may disagree with us.

    If $shiny_new_function works perfectly but receives only 17 of the expected 500 uses per day, that does not immediately tell us what is wrong. Perhaps users cannot find it. Perhaps they do not trust it. Perhaps another way of doing the same thing is easier. Perhaps we misunderstood their needs. Perhaps the assumption behind 500 uses per day was wrong. Perhaps 17 uses are actually enough to produce the business outcome we wanted, and we chose the wrong success criterion.

    Some of those answers would lead us back to Development. Others would lead us much further back, to our assumptions and even to the original intent.

    That is why a production observation should not merely become another defect or improvement ticket. It should become learning.

    Quality work has traditionally put enormous effort into determining whether something is good enough to enter production. That remains important. We need evidence that the realization is sufficiently sound before exposing it to reality.

    But perhaps we have concentrated too much on the means and not enough on the end.

    A product does not become successful because it passed its release assessment. That only means we had sufficient confidence to let it start doing the job for which it exists. The more important assessment begins when it actually does.

    So perhaps product intent should point in two directions from the beginning. It should tell Development what needs to be realized. And it should tell Operations what needs to be observed.

    One gives us evidence that the product can do what we intended. The other gives us evidence about whether it does.

    And the difference between those two is where some of our most valuable learning may be found.

  • Practice Note: AI and the Succession of Understanding

    Much of the discussion about AI and software development focuses on whether AI will replace developers.

    A different risk may be emerging.

    Organizations may use AI to automate precisely the work through which people traditionally developed the understanding required to become senior.

    Senior developers did not enter organizations as senior developers. Their expertise accumulated over time.

    They implemented small changes. They investigated defects. They read unfamiliar code. They made mistakes. They observed production failures. They asked why something had been designed in a particular way. They discovered that apparently strange implementations sometimes existed for good reasons — and sometimes for reasons that had stopped being good years earlier.

    Through repeated exposure, they acquired more than technical skill.

    They acquired context.

    They learned not only what the system does, but increasingly why it became the way it is.

    The Efficiency Paradox

    AI can remove substantial amounts of routine work from software development. That can be valuable.

    A junior developer may no longer need to spend hours writing relatively straightforward implementation code, searching documentation, constructing tests, investigating logs or performing other work that an AI system can complete much faster. Measured as delivery efficiency, this looks like progress.

    But some apparently inefficient work may have performed another function:

    It created experienced people.

    The work product was not its only output.

    Learning was another.

    If we automate the activity while measuring only whether the immediate deliverable can still be produced, we may overlook the organizational capability that the activity was helping to create.

    Work Has More Than One Outcome

    This suggests a more general distinction.

    An activity may produce an explicit outcome and one or more implicit outcomes.

    For example:

    Explicit outcome:
    – A developer fixes a defect.

    Implicit outcomes:
    – The developer learns how the system behaves.
    – They discover relationships between components.
    – They understand why a particular design decision matters.
    – They encounter a business rule that was previously invisible to them.
    – They become better able to recognize similar problems in the future.

    An AI system may produce the explicit outcome extremely efficiently while reducing some of the implicit outcomes.

    The defect still gets fixed. But who learned from fixing it?

    The Succession Problem

    This becomes particularly important when experienced people leave.

    Organizations already struggle to preserve understanding across personnel changes. Documentation may preserve requirements, architecture, code and decisions while much of the reasoning connecting them remains in people’s heads.

    AI could introduce a second continuity problem. It may not only change how existing knowledge is preserved.

    It may change how new people acquire enough understanding to become future custodians of that knowledge.

    If today’s senior developers retire or leave, organizations cannot simply replace them with today’s juniors plus more capable AI.

    The question is whether those juniors have had opportunities to develop the contextual and causal understanding that made the departing developers senior in the first place. This is therefore not merely a staffing pipeline problem.

    It is a succession of understanding problem.

    Expertise Is More Than Accumulated Information

    It would be tempting to solve this by giving future developers better access to documentation and AI-generated explanations. That will help, but it may not be sufficient.

    Experienced practitioners possess more facts, but they also have developed judgment.

    They recognize when something looks suspicious before they can necessarily explain exactly why. They know which questions to ask. They recognize recurring causal patterns. They understand where apparently local changes may have wider consequences. Much of that ability was developed through interaction with real problems.

    The challenge is therefore not simply to transfer everything a senior developer knows to a junior developer.

    It is also to preserve opportunities through which less experienced people can develop their own understanding and judgment.

    AI Changes the Learning Path

    This does not imply that organizations should preserve inefficient manual work simply because previous generations learned through it. That would confuse the practice with the function it served.

    The important question is instead:

    Which learning functions did the old way of working provide, and how will those functions be preserved when AI changes the work?

    Perhaps some learning can happen faster with AI. AI can explain unfamiliar code, expose dependencies, challenge assumptions, simulate alternatives and make organizational knowledge easier to interrogate.

    Used deliberately, it may accelerate the development of understanding rather than diminish it.

    But that requires designing AI-supported work for more than immediate productivity.

    The objective cannot only be:

    How can AI help this developer produce the result faster?

    It must also include:

    How does this developer become capable of understanding, questioning and eventually owning the results?

    Implication for Continuity of Understanding

    Continuity of understanding therefore has at least two dimensions.

    The first is preservation:

    Can the organization retain enough of its existing understanding when people leave, systems change and structures evolve?

    The second is regeneration:

    Can new people develop enough understanding to extend, challenge and eventually replace the understanding held by today’s experts?

    A reference model can help with the first by connecting fragmented knowledge and preserving relationships, rationale and intent.

    But it may also support the second. If important relationships between intent, requirements, decisions, implementation, evidence and outcomes are made accessible, a less experienced person does not have to reconstruct all of them accidentally over years of work.

    They can explore them deliberately.

    AI may make that exploration dramatically easier.

    The opportunity is therefore larger than preserving the old apprenticeship model.

    It is to create a better one.

    Working Proposition

    AI should not only increase the productivity of today’s experts. It should help create tomorrow’s experts.

    When evaluating AI-supported ways of working, organizations should therefore assess not only:

    • whether work is completed faster,
    • whether fewer people are required,
    • whether output quality remains acceptable,

    but also:

    • whether practitioners continue to develop contextual understanding,
    • whether reasoning and rationale remain accessible,
    • whether people learn from the work AI performs with them,
    • whether judgment and ownership continue to develop,
    • and whether the organization is creating the people who will be capable of carrying its understanding forward.

    Otherwise, AI may solve today’s capacity problem while quietly creating tomorrow’s capability problem.

  • Practice Note: The Recursive Why

    There is a recursive quality to the question why that I find fascinating.

    Suppose we are automating regression testing. Why are we doing that? Perhaps because we want confidence that changes have not broken existing behavior. But that answer can itself become the subject of the next question. Why do we need that confidence? Perhaps because we want to release changes without creating unacceptable risk. Why do we want to do that? Perhaps because the product needs to respond quickly to changing needs.

    Something interesting happens as we move through these questions. The answer to why at one level becomes the what at the level above it.

    Automated regression testing is a What whose Why is confidence in changes. Confidence in changes is then itself a What whose Why is the ability to release safely. The ability to release safely becomes another What, connected to another Why.

    The Why of one layer is the What of the layer above it.

    This suggests that Why is not simply an explanation attached to something. It is a connection between levels. The higher level gives purpose to the level below it, while the lower level is one way of contributing to what the higher level is trying to achieve.

    Perhaps this is also an important part of what we mean by understanding.

    We can know a great deal about something without necessarily understanding it. We can know what a system does, what requirements it has, what decisions were made, what processes people follow, and what technologies are used. But if we can no longer move from those Whats toward the Whys that give them meaning, something important has been lost.

    A requirement without its Why becomes an isolated statement. A design decision without its Why becomes an historical fact. A process without its Why becomes a sequence of activities that may continue long after anyone can explain what capability those activities were intended to provide.

    The information may still be there. The understanding is not.

    This also means that Why has no obvious natural endpoint. Every answer can potentially become the subject of another Why. Eventually we stop, but perhaps not because we have reached some ultimate Why. We stop because we have reached a boundary that is sufficient for what we are currently trying to understand.

    That makes the choice of boundary important. If we draw it too narrowly, the Why may sit outside it. We can then describe the subject in extraordinary detail while remaining unable to explain its purpose. We may even optimize it successfully according to its internal logic while losing sight of whether it still contributes to what matters at the level above.

    Perhaps understanding, then, does not primarily reside in the individual pieces of information we preserve.

    It resides in the connections between them.

    And perhaps one of the most important connections we can preserve is the one that allows us to keep asking:

    Why?

    Come to think of it, Why? may be the most frequently asked question in the universe. After all, every answer gives us something new to ask Why? about.