Collected Notes & Essays

Personal · I

Frames

Where you stand changes what counts as true.

2 min read

The Paradox of Intelligence

Can truth and lie exist in the same statement?

Think about this. When you are seated in a moving car, who is in motion? You or the car?

From the laws of mechanics, both. You and the car are moving relative to the road. But relative to the car? You have not moved an inch. So what exactly changed?

Not the reality. Just the frame you are standing in.

Here is where it gets interesting. Most arguments are not really about truth. They are about two people who never compared their starting positions. One person is measuring from the road. The other is measuring from inside the car. Both are right. Neither is lying.

To one person, a statement sounds wrong. To another, that same statement is completely true.

So maybe it is not always about truth versus lie. Maybe it is about the frame behind the statement.

And maybe intelligence is not just knowing what is true, but knowing the frame that makes it true. Anyone can memorise facts. It is knowing where you are standing when you call something true.

1 min read

The Binary Nature of Life

The binary nature of life. Ones and zeros, yes or no, truth or lie, win or lose, have or have not.

Every action is a choice with an output. You always land somewhere.

But the twist is this: the output is not inherently good or bad. It just is. It lives in perspective.

What one person calls failure, another calls freedom. A loss today becomes clarity tomorrow. The same zero can be absence or space. The same one can be success or burden.

Actions may be binary. Interpretation definitely is not.

Choice is not just selecting between two options, it is also choosing the lens through which you judge the result.

2 min read

The Gravity of Experience

Gravity, the force that pulls objects toward a planet's centre.

Just as gravity pulls, time teaches, and space is the room for experience that fuels growth.

Think of it this way. Time itself is the teacher, helping us understand complex concepts that initially seem far out of reach. And just as planets grow and change as they orbit through space, we expand as individuals, shaped by the room we give our learnings to flourish.

At our core, our lives are influenced by two key factors:

  • Mass, the weight of our experiences
  • Radius, how close we let those experiences sit to our core

The more we live, the heavier our personal gravity becomes. Over time, what once felt distant grows closer, and its impact intensifies. I guess that is why understanding often deepens later, not sooner. It could not have happened earlier. There was not enough mass yet.

Forgive yourself for not knowing earlier what only time could teach.

A mistake only becomes a regret when one dwells long on it. That decision only needed enough experience to give it weight, time to let it pull meaningfully, enough space to stop resisting it.

To the journey of learning and becoming, embracing the flaws and flows that shape us into who we are meant to be.

1 min read

Fantasy and Discipline

Fantasy is pretty much a joy thief, taking one away from the gift of the present.

Our mind settles on something ahead, as though our happiness is a future tense.

You might ask, so I should not picture how my life should be in five years?

Well, please do. You definitely have to plan ahead. Discipline comes in here.

Like a fellow once said: discipline is when one's identity is so clear that they stop negotiating their feelings.

1 min read

Self-Worth: Character Plus Value

Self-worth = Character + Value

Self-awareness = Character

Self = Value

A moving man will one day meet his luck.

Reaching out to people is one of the best things you can do for yourself.

Personal · II

Becoming

Growth that looked nothing like the plan.

4 min read

Home Is More Than a Place

Watching Christopher Nolan's The Odyssey reminds me that Interstellar, which he also directed, was not just about a man trying to find his way home. What is home, though? What does home really mean? A destination? A person? A memory? Or the version of ourselves we return to when life hits?

For most of the film, I was not thinking about monsters or gods, thanks to Nollywood. Rather, I was focused on the characters and the scenes. I guess that comes with being a Nolan fan.

On second thought, should this really be about home, or about the temptations of mankind? Conquering Troy at the expense of civilisation. Agamemnon sacrificing his daughter for victory, only to end up victorious and killed by his wife. Odysseus losing twenty years away from his family, leaving them to the hunger of power.

Seeing the characters navigate temptation, how it promises comfort but quietly demands pieces of them in return. Giving away what can be easily replaced is not sacrifice, that is convenience. The altar of sacrifice leaves a sacred space behind. You feel it.

Another scene that stood out was the Song of the Sirens. It silently screams that the greatest temptations are rarely the loudest. Leaving out the sound of the sirens was an absolute masterpiece, Christopher Nolan doing his thing. The sound never came on. I guess that is Nolan's way of telling us to reflect on our own and let our imagination lead. Or maybe that is because temptation is deeply personal, rooted in our desires. What calls A may never call B. We do not all hear the same song.

The film also reminded me that discipline is rarely about denying ourselves pleasure. It is about understanding that satisfying every desire only teaches desire to ask for more. Gluttony is not just about food. It can be ambition without purpose, power without restraint, recognition without contentment. Every excess begins with a desire that was never questioned.

Not to mention the scene that references the pain of the mind being more traumatic than the pain of the body. Physical wounds usually heal. Memory does not always.

Regret, guilt, grief. They travel with us far longer than broken bones ever could. Probably why the journey home feels so important, not just to return to a place but to return to peace. A state of mind where the clearest view is from below.

Then again, there is the part that reminds us that you do not really know someone's character just by standing beside them. You discover it when they have power over you.

Nolan further shows that leadership has never been about authority but about what it reveals. Like AJS once said, all blessings come from God through men to men. No matter how many people help us along the way, there comes a point where we have to make the next decision ourselves. Mentors, friends, loved ones can guide, warn, save you from a burning fire, but no one can live it for you. You eventually bear the consequences of your choice.

Maybe that is what every journey home is really about. Not finding the place you left, but becoming the person who is finally ready to arrive.

2 min read

If Wishes Was an Ocean

I wish, I wish, I wish. The dream.

Newton's laws of motion: for every action there is an equal and opposite reaction, such that any object at rest or in uniform motion will continue in that state unless acted upon by an external force.

Archimedes' principle states that when a body is partially or wholly immersed in a fluid it experiences an upthrust equal to the weight of fluid displaced.

I took a decision some months back, leaving comfort for uncertainty in search of growth. I found growth, but not in the way I imagined. Took a deep dive. Now on this level, several thoughts came to mind about why I made that decision.

Not proud of the outcome it started with, but proud of the decision I took. Contrary to what people around me think, that I look like I am overthinking or feeling down, it actually made me reflect deeply on the last five years.

Looking back, I see victories I did not celebrate enough, mistakes I judged too harshly, and lessons I almost overlooked. For a moment, I came close to measuring my worth by what had not gone according to plan.

That was the trap.

But I learnt to be grateful for where I am and to learn from the past. As I have come to understand, a mistake is an error, knowingly or unknowingly. Dwelling on it makes it a regret, which eats away your confidence, your self-worth, and your ability to appreciate how far you have come.

Growth is not always upward. Sometimes it is inward.

The dream is still alive.

2 min read

The Reflex of Doing It Alone

Bootstrapping an idea alone can be frustrating.

This is coming from someone who does his best work alone. No emotional overhead, no noise. Just focus. But in that silence, doubts get louder. You burn alone, win alone. Even when you are efficient in teams, there is still that pull back to solitude.

Not a habit. More of a reflex to stimuli.

What strengthens us can also limit us. Doing things alone builds resilience, independence, and the ability to see things through, even when we suspect it might not work out. But it can also build invisible walls.

Over time, it becomes easy to drift away from people, from support, from the outside world. Especially when we are naturally introverted, or have learned from experience to rely only on ourselves. Once bitten, twice shy.

But working alone is not the same as being isolated. Independence is not the same as emotional self-exile. It is human to feel frustrated when you are carrying everything alone. That is not weakness, it is awareness.

2 min read

Life Lately

Every season has a way of asking questions we do not always have answers to. One moment you are thinking of the next line of action, the work still left to do, the person you are becoming, the areas you need to improve or shed. The next, life quietly reminds you that time is not standing still. Parents get more grey hairs, friends move into new chapters, conversations begin to shift.

Some days it is easy to wonder whether we are moving fast enough. Not because we want to keep up with everyone else, but because the dreams we carry still feel larger than the life we are living today. Then again, life reminds you to take a chill pill.

I have come to realise that adulthood rarely gives one the luxury of certainty. Responsibilities do not wait until you feel ready. Decisions do not always come with guarantees on risk or reward. We simply take the next step with whatever clarity we have and trust that God's faithfulness will meet us somewhere along the journey.

Another area is growth, which really has not looked like an achievement. It has looked like reflection. Less about proving something and more about understanding. I have been learning that not every season is meant for acceleration. Some are meant to strengthen your foundation, reshape your priorities, and quietly prepare you for the life you have been praying for.

The noise eventually fades, leaving behind the questions that matter. Who are we becoming? What do we truly want, or need? What is worth waiting for?

No answers yet, but we know we do not want to become so focused on arriving that we lose sight of the person we are becoming along the way.

Maybe that is what every season teaches:

To carry ambition without losing peace. To hold on to hope without rushing time. And to keep becoming, one quiet day at a time.

1 min read

A Mistake Only Becomes a Regret

On my way to work, a thought came in while listening to Oshimiri Atata by Faith Captain. The song title is going to be my motto going forward. That is going to be my weapon, by the Christ that dwells in me.

Back to my thoughts. I came to the conclusion that a mistake only becomes a regret when one dwells long on it, though experience, they say, is the best teacher.

A mistake is something that happened. Regret is how you feel about that thing, or about a thing that never happened at all.

Being present, and being intentional.

1 min read

Note to Self

The world is finally quiet and thoughts stop fighting each other. Music, the therapy in silence. The quiet, gentle calm of thoughts, not out of worry but out of imagination, reflection, reminiscing. Not of mistakes or regret, but of the good, the bad and the ugly.

For these bring joy to the heart, and are always a reminder of how much we matter, even when things get tough and go unappreciated. Wabi-sabi, the beautiful imperfection. The end is to always give life your best, to be honest and open, and to never give up or question or doubt yourself.

Personal · III

What Love Costs

Being known, and what it asks of you.

1 min read

Love on a Precipice

There is something terrifying about love when it finds you in the middle of becoming, or unbecoming. You want to hold it gently, but you are still learning how to hold on to yourself.

Love on a precipice feels like standing one honest conversation away from a completely different life. Not because it changes everything overnight, but because it has a way of revealing everything you are not, and everything you have not yet become.

And perhaps that is what makes it frightening. Not loving someone else, but wondering whether you will recognise yourself when they love you back.

2 min read

Intimacy

My favourite form of intimacy has always been asking questions. Not to pry, but to understand the thought process of someone's inner world. The way they think. The way they feel. The way their past still echoes in their present.

There is a depth in people that never shows unless someone asks the kind of questions that require honesty instead of performance. What shaped you. What steadied you. What you learned the hard way. What you still carry even when you pretend you do not.

The older I get, the more I realise that intimacy is not in declarations, or love languages. Rather, it is in revelation. Those moments where someone lets you walk into the parts of them that are not polished for the world. Some people touch you by knowing your story. Others touch you by wanting to understand your mind. Asking is just another way of saying I do not want the version you have learned to present. I want the version of you that is real.

Second favourite would be casual intimacy. A small hand on your back when you are in crowded streets. A gentle kick from where they are sitting across the table. A head on the shoulder, a hand in your hand, a squeeze on the arm as they are walking past you.

And I think maybe love is not made up of grand gestures or explosive displays, but that it is made up of the little things. The little things that say I am here, and I care for you, and your life has intertwined so deeply into mine that there is no need to think, because casual intimacy comes easy.

1 min read

Pain, on Its Own, Does Nothing

Pain, on its own, does nothing. As humans with a bunch of sensory overload, we are not built to just experience pain. We are built to interpret it and give it a meaning, with a bit of story, and we end up romanticising it or forcing a meaning where there is none.

Do we leave it there, or do we choose to give it meaning anyway? Let us pivot a bit.

If we say pain does nothing, does love do something?

One passive and the other active. Pain happens to you. Love is something you do. However, like pain, love has no guaranteed outcome.

Pain says: this hurts, now what? Love says: move toward something, or someone.

To the question earlier, do we leave pain meaningless, or give it meaning?

I guess love is often the tool we use to answer that.

1 min read

The Father’s Love

These past couple of months have been a reminder of a love that listens with patience. One that builds a sanctuary where one's heart can rest. The kind of love that holds you on your worst days, and celebrates you.

1 min read

My Intentions to Self

I just want to treat you the way you have always deserved. To make you feel truly seen, heard and valued every single day. To bring you real happiness and honesty, and to be the person who reminds you how much you matter, even when things get tough.

I am not perfect, but I promise to always give my best, to be honest and open with you, and to never give up on us or make you question how deeply you are loved.

Personal · IV

Verse

The things that would not stay in paragraphs.

1 min read

Chasing Static

Purpose, they say, a shiny chrome god worshipped with spreadsheets and sunrise affirmations.

Like nailing jelly to the existential wall. Define yourself, they chirp, as if I am not already defined by the overflowing inbox of information since birth and the phantom limb of past anxiety.

Peace of mind, you say? A unicorn riding a tax audit. A shimmering oasis populated by deadlines and the gnawing suspicion that I am folding fitted sheets wrong.

Purpose? Perhaps to catalogue dust bunnies under a fluorescent moon. Or maybe, just maybe, to master the art of staring blankly at motivational posters, until they burst into flames of lukewarm contentment.

And peace of mind? Oh, that is just the static between stations, a phantom signal we keep chasing in a broken radio. Turn it off. Maybe the quiet is the point. No, probably not. But maybe.

1 min read

The Butterfly on the Coffee Cup

The beautiful coffee eyes reflection,

With that succulent latte poured on,

Foamy waiting to be sucked on,

Like a child sucks the tea from the centre cup,

In a puckered O shaped lip,

Could that be a baby's love?

Nicomachus would say:

If it stays it is love,

If it ends it's a love story,

If it never begins it is a love story,

Candide would say:

And if I keep chasing the butterfly

it eventually flies away,

But if I work on my garden, it stays.

Even if it flies away, definitely gonna hurt,

But then, my garden blooms.

1 min read

Choice

In the quiet moments of the day,

When shadows blend with light,

A whisper stirs the heart to say,

"Choose your path, take flight."

For life presents a branching road,

With choices broad and narrow,

Each decision, a secret code,

Guiding like an arrow.

The power lies within our hands,

To carve out destiny,

With every step and every stand,

We shape what is to be.

So heed the call, embrace the choice,

With courage, make it true,

In the silence, find your voice,

And let it sing anew.

Personal · V

Borrowed Light

Notes taken from other people’s books.

2 min read

Notes on The Power of Self Confidence

Reading notes on Brian Tracy, The Power of Self Confidence. The ideas are the author’s.

These are notes taken while reading The Power of Self Confidence by Brian Tracy. The thinking below is his. What I kept is what I wanted to keep looking at.

What one great thing would you dare to dream, if you knew you could not fail?

The foundational quality of success is self-confidence. With greater self-confidence you would be bolder and more imaginative.

Law of becoming. Everyone is in a continual process of becoming, evolving or growing in the direction of their dominant thought.

Law of concentration. Anything you dwell on grows into your reality.

Thoughts held in mind produce after their kind. Your outer world will be a reflection of your inner world.

Clarify your personal values. Who do you admire the most? If you could be like them, what quality would you want to emulate?

Values are non-negotiable. Set peace of mind as your highest principle. Your values are only expressed in your actions. Living consistent with your values is the key to happiness, harmony, well-being, and high levels of self-confidence.

True nobility is being superior to your former self.

Technical · I

AI Systems

Agents and pipelines that had to survive production.

9 min read

The Anatomy of a High-Signal Prompt

Most prompts fail before the model sees them. Not because the request is unclear, but because the structural scaffolding that allows a model to act on a request correctly is missing. A well-engineered prompt is not a good question. It is an information architecture with ten distinct layers.

Why Structure Matters More Than Wording

Engineers who spend hours refining the wording of a prompt while leaving its structure implicit will consistently get inconsistent results. Models do not fail because the request is poorly phrased. They fail because the operating context, rules, examples, and output format were left to the model to infer. Every dimension left implicit is a source of variance.

The ten-layer framework that follows is a checklist for eliminating that variance.

Layer 1: Task Context

Before a model knows what to do, it needs to understand the environment in which it is operating. Who is it acting as? What system is it part of? What constraints does that role carry?

"You are a customer support assistant" is weaker than: "You are a customer support assistant for a B2B SaaS company. Users are technical teams on annual contracts. Your responses appear in a real-time chat interface. Escalation to a human agent is available and should be offered when the issue cannot be resolved in three exchanges."

The difference is scope. The second version gives the model the operating context it needs to make appropriate judgment calls rather than averaging over every possible customer support scenario in its training data.

Layer 2: Tone Context

Tone is separate from task. Specifying the communication register prevents the model from defaulting to a generic voice that fits no one's brand. "Authoritative but not condescending. Direct. No corporate hedging. Second person throughout." This is a tone specification. It should be explicit, not assumed.

Layer 3: Background Data, Documents, and Images

If the task requires reasoning over specific information, that information must be in the prompt. Retrieval-augmented patterns handle this at scale: pull the relevant documents, inject them into context, and explicitly instruct the model to base its response on them. The instruction to use the provided documents is as important as the documents themselves.

Layer 4: Detailed Task Description and Rules

Most prompts underspecify here. A complete task description answers: what exactly should the model produce, what format should it take, what constraints apply, and what should it explicitly not do.

A model writing summaries without a word limit produces summaries of inconsistent length. Specify the limit. A model extracting information without a schema outputs unstructured text. Specify the schema. Every unspecified dimension is an invitation to vary.

Layer 5: Examples

Few-shot examples are not about teaching the model. They are about calibration. Three well-chosen examples collapse the distribution of possible outputs toward the specific form you need. They communicate what words can only approximate.

Use examples that represent the full range of inputs the model will encounter, not just the easy cases. If your production inputs include edge cases, your examples should include them too.

Layer 6: Conversation History

For multi-turn interactions, the conversation history is load-bearing context. Do not assume the model remembers earlier turns. Inject the relevant history explicitly. Be selective: include only what is necessary for the current task. Irrelevant history increases noise and can degrade performance.

Layer 7: Immediate Task Description and Request

After all the scaffolding, the immediate instruction should be simple and unambiguous. "Summarise the document above in three bullet points, each under twenty words." If the immediate instruction is still complex after the previous six layers, those layers did not do enough work.

Layer 8: Think Step by Step

Chain-of-thought prompting is not a magic phrase. It is an instruction to externalise reasoning, which catches errors before they become outputs. "Think step by step before giving your final answer" works because it forces intermediate steps into the visible context where they can be checked.

For tasks with multiple reasoning steps, break them out explicitly rather than using the generic phrase.

Layer 9: Output Formatting

Specify the output format explicitly. JSON, Markdown, numbered list, paragraph, table. If you need structured output, provide the schema. If you need a specific field ordering, state it. Output formatting described in words will be interpreted imprecisely. Showing an empty template is more reliable than describing one.

Layer 10: Prefilled Response

For tasks where the response should begin in a specific way, prefill the opening. "Begin your response with: 'Based on the provided information...'" This removes ambiguity about the opening register, prevents unnecessary preamble, and anchors the model to the frame you need.

Using the Framework

Not every prompt needs all ten layers. Simple, stateless tasks might need three or four. Complex, multi-step, high-stakes tasks may need all ten with significant depth at each layer.

The framework is a gap analysis tool. Before sending any prompt to production, walk through each layer and ask: have I specified this, or am I relying on the model to guess correctly? Every gap is a variance source. Remove the gaps.

7 min read

Preventing LLM Hallucinations Without Fine-Tuning

Hallucinations are not random. They follow predictable patterns, and most of them are preventable without touching the model weights. The mechanisms that cause a model to generate plausible-sounding falsehoods are well understood. The interventions that prevent them are engineering decisions, not model choices.

What Causes Hallucinations

Models hallucinate when asked to produce information they do not have with sufficient confidence. The problem is that they are trained to produce fluent, coherent text, and the easiest way to produce fluent, coherent text on a topic is to continue generating rather than stopping and saying nothing.

Four patterns produce most production hallucinations: asking the model for specific facts it was not trained on, asking it to reason over information it has not been given, asking questions where the confident wrong answer is more fluent than the uncertain correct one, and providing insufficient context for the model to distinguish what it knows from what it is generating.

Technique 1: Give the Model Permission to Not Know

The single most effective hallucination prevention technique is explicit. Tell the model it is acceptable to say "I don't know" or "I cannot find this in the provided documents."

Without this permission, models will fill gaps because the training signal rewards completeness. With it, they will flag uncertainty rather than cover it.

"If you cannot answer this question based on the information provided, say: 'I don't have enough information to answer this.' Do not speculate."

This instruction works because it directly addresses the tension between fluency and accuracy that causes hallucinations.

Technique 2: Force Reasoning Before Answering

Hallucinations increase when models produce answers in a single forward pass without visible reasoning. Requiring the model to reason step by step before giving its final answer catches errors before they become outputs.

"Before answering, identify the key facts in the documents that are relevant to this question. Then provide your answer based only on those facts."

The externalised reasoning step creates a checkpoint. If the reasoning step surfaces an absence of relevant facts, the model is far more likely to acknowledge that absence than to hallucinate an answer.

Technique 3: Require Confidence Thresholds

Instruct the model to only answer when it is confident. This sounds simple but is underused.

"Only provide an answer if you are highly confident it is accurate based on the provided information. If you are uncertain, say so explicitly and explain what you are uncertain about."

The key implementation detail is that "confident" needs to be operationalised. Provide examples of what a confident answer looks like versus an uncertain one. Showing the model what uncertainty looks like in practice is more reliable than instructing it to be uncertain in the abstract.

Technique 4: Quote Before You Conclude

For document-grounded tasks, require the model to find and quote the relevant passage before drawing any conclusion. This is the most powerful structural technique for RAG applications.

"Before answering, find the specific sentence or passage from the documents that supports your answer. Quote it exactly. If you cannot find a supporting passage, say so."

When the model cannot find a quote, it is forced to acknowledge that absence rather than generate an answer from general knowledge. This single instruction reduces hallucinations in document-based question answering significantly in practice.

Combining the Techniques

These four techniques compound. A prompt that gives the model permission to not know, requires step-by-step reasoning, asks for confidence flagging, and requires source quotation before conclusions is dramatically less likely to hallucinate than a prompt that does none of these things.

The underlying principle is the same across all four: reduce the pressure the model feels to produce a complete, fluent answer at all costs. Hallucinations are a confidence problem. The interventions that work are the ones that give the model an alternative to false confidence.

What These Techniques Do Not Fix

They do not fix factual errors in the model's training data, do not prevent hallucinations on topics where the model has strong but incorrect priors, and do not substitute for retrieval when the required information is not in context.

They are prompt-layer interventions that address the structural causes of hallucination in production. Combined with good retrieval architecture and clear scoping of what the model is and is not expected to know, they cover the majority of production hallucination failure modes.

8 min read

Context Engineering: The Discipline Beyond Prompting

Prompt engineering gets the attention. Context engineering does the real work. The distinction matters because it changes what you optimise for. Prompt engineering asks: how do I phrase this request? Context engineering asks: what information does this model need, in what structure, to produce the output I need reliably?

These are different questions. The second one is harder, more consequential, and almost entirely responsible for whether a production LLM system behaves the way you want.

What Context Engineering Actually Is

Every LLM call has a context window. What goes into that context window is the context. In simple cases, that is just the user's message. In production systems, the context is assembled from multiple sources: system prompts, retrieved documents, conversation history, tool outputs, structured data, and user inputs.

Context engineering is the discipline of deciding what goes into that context, in what order, in what format, and at what level of detail. It is information architecture applied to the constraints of a finite context window.

The context window is not unlimited. Every token you use on something irrelevant is a token not available for something useful. Context engineering is fundamentally about prioritisation under constraints.

The Context Budget

Start with a budget. For a given model with a given context limit, decide in advance how many tokens to allocate to each component:

  • System prompt and role definition: fixed overhead
  • Retrieved documents: variable, largest allocation
  • Conversation history: variable, pruned by relevance and recency
  • Current user input: variable, usually small
  • Output buffer: reserved, never filled with input

The errors most systems make: no retrieval strategy (everything or nothing), conversation history that grows unbounded until it hits the limit, and system prompts that expand over time without a corresponding reduction elsewhere.

Retrieval Strategy

For systems that retrieve information before calling the model, the retrieval quality determines the context quality. A retrieval system that returns the ten most semantically similar chunks regardless of relevance, recency, or diversity will fill the context with redundant and tangentially relevant information.

Better retrieval principles for context engineering:

Score on relevance AND recency. A document from three years ago that is semantically similar to the query may be less useful than a more recent document that is somewhat less similar.

Deduplicate before injecting. Multiple chunks from the same source saying the same thing in slightly different ways waste tokens without adding information.

Position matters. Models attend more to the beginning and end of long contexts. The most important retrieved information should appear near the top.

Conversation History Management

Conversation history in multi-turn systems is the most common cause of context degradation. The naive approach is to append every exchange and eventually hit the context limit. The correct approach is to manage history as a rolling, relevance-ranked window.

Three patterns that work:

Summarise older turns. When history gets long, replace the oldest N turns with a brief summary. The model loses verbatim recall of those turns but retains the substance.

Keep only task-relevant turns. If the conversation drifted into small talk three turns ago, those turns are not relevant to the current task. Remove them.

Use explicit memory for facts. Rather than relying on the context to carry important facts mentioned earlier, extract them to a structured memory and inject them as a concise fact sheet at the top of each turn.

Structural Information Architecture

How you structure information in the context affects model behaviour. Three principles:

Instructions before documents. The model should understand what to do before it processes what to work with.

Use XML or Markdown delimiters. Clearly delineated sections reduce ambiguity about where one type of information ends and another begins. Tags like `<documents>`, `<conversation_history>`, and `<instructions>` are more reliable than prose transitions.

Put the most important constraint last. Models exhibit recency bias in long prompts. The most critical instruction belongs near the end of the system prompt or immediately before the user input.

Context Engineering Is a System Property

Individual prompts can be improved through prompt engineering. But the context quality of a production system depends on decisions made across the entire pipeline: how documents are chunked and indexed, how history is stored and pruned, how tool outputs are formatted, how retrieved results are ranked and filtered.

These decisions compound. A system with excellent retrieval and poor history management will still exhibit context degradation at scale. Context engineering is a system-level discipline. Treating it as a per-prompt concern is why most production LLM systems plateau in quality.

8 min read

Memory Architecture for Production AI Systems

The context window is not memory. It is a workspace. What happens inside it is processing, not retention. This distinction is fundamental to building AI systems that perform reliably over time, across sessions, and at scale.

Memory in production AI systems is an architectural decision, not a feature you add after the fact. Systems that conflate the context window with memory run well in demos and break down in production.

The Three-Tier Memory Model

Production AI systems need three types of memory, each with different properties:

Working memory is what the model is actively processing in the current context window. It is volatile, token-bounded, and ephemeral. It does not persist beyond the current call. Everything in working memory is lost when the context is cleared.

Episodic memory covers the current session or task. It holds the history of the current conversation, the steps taken in the current workflow, and the intermediate results produced so far. It needs to persist within a session but does not need to survive session boundaries.

Persistent memory holds facts, preferences, and knowledge that should survive across sessions. A user's stated preferences, a company's internal knowledge base, the outcomes of past decisions. This is the layer that turns a stateless LLM into a system that accumulates knowledge over time.

Working Memory: Managing What You Have

Working memory management is context engineering. The key decisions: what to include in the current context, in what order, and how to handle the transition when the context fills.

For long tasks, the working memory budget must be allocated deliberately. A coding agent that puts the entire codebase in context will run out of space before it can produce a meaningful output. The correct approach is to load only the files relevant to the current subtask, using persistent memory to know which files those are.

Episodic Memory: Session Persistence

The simplest episodic memory implementation is a session store. After each exchange or action, write a structured summary to a session record. Before each exchange, read the most relevant portions of that record back into the context.

The critical engineering decision is what to summarise and how. Verbatim transcripts are expensive and often redundant. Structured event logs are cheap and queryable. A good episodic memory system writes events like: "User confirmed the billing address at step 3. Agent wrote to orders table at step 5. Constraint: must use USD pricing." This is far more useful than a transcript.

Persistent Memory: The Knowledge Layer

Persistent memory is where the interesting architecture decisions live. Three patterns:

Retrieval-augmented memory. Store facts, documents, and historical context in a vector database. At the start of each session or task, retrieve the most relevant items based on the current query. This scales to large knowledge bases and handles the recall problem that direct context injection cannot.

Structured fact stores. For facts with known schemas (user preferences, entity attributes, configuration), a relational or document store is more appropriate than a vector database. Query by entity ID rather than semantic similarity.

Hybrid retrieval. Most production systems benefit from both. Use semantic search for unstructured knowledge and document lookup, structured queries for known entities. The orchestration layer decides which retrieval mechanism to use based on the query type.

Memory as a System Boundary

The boundary between working memory and episodic memory is where most production failures occur. An agent that carries everything forward in context will hit the limit. An agent that discards everything at each step will re-discover the same facts repeatedly.

The solution is explicit write and read operations at memory tier boundaries:

On task completion or interruption: write a summary from working memory to episodic memory. At session start: read the most relevant episodic memory back into working context. When episodic memory reaches a threshold: compress and promote key facts to persistent memory.

These are not automatic. They are engineering decisions that need to be designed, tested, and maintained.

Testing Memory Systems

Memory bugs are the hardest to catch because they manifest across sessions, not within them. A production memory system needs a test harness that:

Runs multi-session scenarios, not just single-turn tests. Validates that facts written in session N are correctly retrieved in session N+K. Tests the boundary behaviour when memory tiers are full. Checks for memory contamination between different users or entities.

The systems that fail most visibly in production are the ones whose memory was tested only at the component level, not across session boundaries.

9 min read

Evaluation Harnesses: How to Test AI Systems That Cannot Be Unit Tested

You cannot unit test an AI system. The output of an LLM call is not a deterministic function of the input. Running the same test twice does not guarantee the same result. Traditional software testing (assertion-based, exact-match, pass/fail) does not transfer.

This is not a reason to abandon testing. It is a reason to build the right testing infrastructure. That infrastructure is called an evaluation harness, and building one well is the most important engineering investment you can make in a production AI system.

What an Evaluation Harness Is

An evaluation harness is a system for measuring how well an AI system performs across a defined set of inputs. It differs from unit testing in three fundamental ways:

It uses rubrics, not exact matches. Instead of checking whether the output equals an expected value, it scores the output against criteria that define what good looks like.

It operates over distributions, not individual cases. A single test run tells you little. A harness gives you aggregate scores, variance measures, and trends over time.

It measures behaviour at multiple levels: final output quality, intermediate reasoning steps, tool call sequences, and latency. Watching only the final output misses most of what is useful.

Building the Test Dataset

The test dataset is the foundation of the harness. A weak dataset produces misleading evaluation signals. A strong dataset covers:

The core case distribution. The 80% of inputs that represent normal operation. These should reflect actual production input distribution, not idealised examples.

Edge cases. Inputs that are unusual but valid. The user who asks an unexpected question. The document with unusual formatting. The request that combines two tasks.

Known failure modes. Every time the system fails or behaves unexpectedly in production, add that input to the test dataset. Over time, this becomes a regression suite that prevents old failures from returning.

Adversarial inputs. Inputs designed to trigger specific failure patterns: empty inputs, very long inputs, inputs in unexpected languages, inputs that attempt prompt injection.

The minimum viable evaluation dataset for a production system is fifty cases. Below that, the variance in your scores will be too high to be meaningful.

Rubric Design

Rubric design is where evaluation harnesses succeed or fail. A rubric is a scoring function that takes an output and returns a quality score. The rubric must be:

Operationalised. "Good" and "bad" must be defined in terms that a scorer can apply consistently. "The response should be helpful" is not a rubric. "The response should directly address the user's question without adding unrequested information" is.

Decomposed. Complex tasks need multiple rubric dimensions. A customer support response might be scored on: accuracy of information, tone appropriateness, completeness, and escalation decision quality. Aggregate scores obscure which dimension is failing.

Calibrated. Before relying on a rubric, check that human scorers applying it independently produce similar scores on the same outputs. High inter-rater variability means the rubric is ambiguous.

Automated vs Human Evaluation

In production, evaluation needs to scale. Human evaluation does not scale. LLM-as-evaluator patterns work well for many rubric dimensions: coherence, instruction-following, relevance. They work less well for factual accuracy and domain-specific quality.

The practical approach: use LLM-as-evaluator for the rubric dimensions where it is reliable, use human evaluation for domain-specific quality and ground-truth accuracy checks. Build both into the harness and track where they agree and disagree.

When using an LLM as evaluator, use a different model than the one being evaluated. Same-model evaluation introduces optimism bias.

Regression Testing

Every production AI system degrades over time if not actively maintained. Model versions change. Prompts are edited. Retrieval indices go stale. Regression testing catches these degradations before they reach production.

The regression test suite is the subset of your evaluation dataset that covers your known failure modes. Run it on every prompt change, every model version change, and on a scheduled basis even when nothing has changed. The scheduled runs catch silent degradation from upstream changes.

Track scores over time, not just point-in-time values. A score of 85% on a rubric means nothing without knowing whether it was 90% last month.

Observability Integration

Evaluation and observability are two sides of the same coin. The evaluation harness tests offline. Observability monitors online. Both are necessary.

Instrument every production AI call to emit: the full prompt, the model response, the tool calls made in order, latency, and token usage. This gives you the ability to replay any production call in your evaluation harness when something goes wrong.

The best evaluation datasets are grown from production logs: real inputs, real failures, real edge cases. The harness and the monitoring system should be designed to feed each other.

10 min read

Nine AI Concepts Every Builder Needs in 2026

The AI tooling landscape shifted significantly in 2025 and early 2026. Some concepts that were theoretical are now production-ready infrastructure. Nine concepts have emerged as genuinely essential for anyone building AI systems today.

1. Agentic Loops

An agentic loop is a control pattern where an LLM iteratively calls tools, processes the results, and decides the next action until it reaches a stopping condition. The loop structure is: observe, reason, act, repeat.

What makes it non-trivial in production: the loop needs explicit termination conditions, cost controls, and error handling for tool failures. Unbounded loops are the most common cause of runaway costs in production agent deployments. Design your loop with a maximum iteration count and a budget limit from the start.

The agentic loop is not a framework. It is a design pattern you implement. Frameworks like LangChain and CrewAI provide it out of the box, which is useful for prototyping but creates abstraction layers that obscure what is happening when something goes wrong.

2. Model Context Protocol (MCP)

MCP is Anthropic's open standard for connecting LLMs to external tools, data sources, and APIs. It provides a standardised interface for tool definition, tool calling, and result handling that works across models and frameworks.

The practical value: once you build an MCP server for a data source or API, any MCP-compatible client can use it. This is the beginning of a tools ecosystem that is model-agnostic rather than framework-specific. For production systems, MCP reduces the integration surface area significantly.

Understanding MCP is no longer optional for AI engineers. It is becoming the default wiring pattern for tool use in production systems.

3. Subagents and Multi-Agent Systems

Single-agent architectures hit practical limits when tasks are too complex for a single context window, require parallelism, or need specialised capabilities at different stages.

Multi-agent systems decompose the work: an orchestrator agent breaks down the task and delegates to specialised subagents. Each subagent has a narrower scope, its own tools, and its own context. The orchestrator synthesises results.

The design challenge is the communication protocol between agents. Subagents that communicate through unstructured natural language produce fragile systems. Subagents that communicate through structured schemas produce robust ones. Build the schema before you build the agents.

4. AI Gateway

An AI gateway is a proxy layer between your application and AI providers. It handles model routing, rate limiting, cost tracking, caching, fallback logic, and observability in one place.

In 2026, running an AI system without a gateway in front of it is like running a web service without a load balancer. The gateway abstracts provider-specific APIs, gives you a single point for cost control, and enables model swapping without application code changes.

Vercel's AI Gateway, LiteLLM, and similar tools have made this pattern accessible without building custom infrastructure.

5. Inference Economics

The economics of LLM inference are non-obvious and matter significantly at scale. Three dynamics to understand:

Token pricing is not uniform. Input tokens and output tokens have different costs. Cached input tokens cost significantly less than uncached ones. System prompt caching alone can reduce costs by 50-90% for systems with large, stable system prompts.

Latency and cost trade differently across model sizes. A large model that produces a correct answer in one call is often cheaper than a smaller model that requires three calls to get to the same answer. Benchmark on cost per task completion, not cost per call.

Batching reduces cost. For non-real-time workloads, batch inference can reduce costs by up to 50% compared to real-time inference.

6. Evals

Evals are the test suite for AI systems. The eval mindset is: before you ship any AI system change, you have evidence that it performs better on the dimensions that matter.

The minimum viable eval is a set of fifty to one hundred representative inputs with clear quality criteria. Run the eval before and after any system change. Track the score over time.

No production AI system should ship without evals. This is the single most commonly skipped step and the single most common cause of regressions.

7. Guardrails

Guardrails are validation layers applied to AI system inputs and outputs. They enforce the contract between the AI system and the application it serves.

Input guardrails: validate and sanitise user inputs before they reach the model. Block prompt injection attempts, detect off-topic requests, enforce length limits.

Output guardrails: validate model outputs before they reach the user. Check for hallucinated entities, enforce format compliance, flag low-confidence responses, redact sensitive information.

In 2026, guardrails are infrastructure, not an afterthought. Build them into your system architecture from the start rather than retrofitting them after an incident.

8. Observability

You cannot debug an AI system you cannot observe. Observability for AI systems means: for every production call, you have a record of the full prompt sent, the complete model response, every tool call made in order with arguments and results, latency at each step, and token usage.

Without this, debugging a production failure is guesswork. With it, you can replay any production call, identify the exact step that went wrong, and reproduce the failure in your evaluation harness.

Tools: LangSmith, Langfuse, and Helicone all provide AI-specific observability. Pick one and instrument your system before you go to production.

9. The Bitter Lesson

The Bitter Lesson is Rich Sutton's observation that general methods leveraging computation consistently outperform methods that build in human knowledge. Applied to AI engineering: systems that rely on scale and learning tend to outperform systems that rely on hand-crafted rules and heuristics over time.

The practical implication for builders: be careful about how much domain-specific logic you hard-code into your AI systems. Rules that seem necessary today may become unnecessary as model capabilities improve. Build systems that can take advantage of better models with minimal re-engineering.

This usually means keeping your application logic and your model interaction logic cleanly separated.

10 min read

The 6-Month AI Engineer Roadmap

Most people who want to build AI systems do not know where to start. They are told to learn "machine learning" and spend six months studying gradient descent before discovering that the skills needed to ship production LLM systems are almost entirely different. This roadmap fixes that problem.

It is ordered by what you can apply at each stage. Each month builds on the previous one. By month six, you can architect and ship production AI systems and specialise in the direction that interests you most.

Month 0: The Mental Model

Before touching any code, get the mental model right. Most people who build fragile AI systems have a fundamental misunderstanding of what they are working with.

An LLM is a next-token predictor trained on a vast corpus of text. It does not "think." It does not "know" things in the way a database knows things. It has a strong prior over what text should follow given context. When you prompt an LLM, you are providing a context that makes certain text continuations more or less probable.

This framing matters because it tells you where to invest. You are not training a model to know things. You are engineering the context that makes the model likely to produce the output you need. Everything from prompt design to retrieval to memory is about shaping that context.

The second mental model concept: LLMs are probabilistic systems. Your goal is to move the distribution of possible outputs toward the outputs you want. You do this through prompt engineering, context engineering, fine-tuning, and output validation. You can narrow the distribution significantly, but never guarantee a specific output.

Month 0 deliverable: be able to explain what an LLM is and is not, what a context window is, and what the difference between a prompt and a completion is.

Month 1: Prompting

Month 1 is entirely about prompt engineering. Not because prompting is the most important skill, but because everything else you learn will require you to write prompts, and weak prompting introduces errors at every stage.

By the end of month 1, you should be able to: write a system prompt that reliably produces the output you want for a defined task, use few-shot examples to calibrate model behaviour, implement chain-of-thought prompting, prevent hallucinations using structural techniques, and use output formatting to produce structured data from model outputs.

Work through the Anthropic prompt engineering guide and the OpenAI cookbook. Build something that uses an LLM API for a real task you care about. The learning that sticks is the learning you apply immediately.

Month 1 deliverable: a working LLM-powered tool that solves a real problem using a prompt you engineered yourself.

Month 2: Systems Thinking and RAG Introduction

Month 2 shifts from single-prompt thinking to systems thinking. You are no longer asking "how do I phrase this?" You are asking "how does information flow through this system?"

The central concept is retrieval-augmented generation (RAG). RAG is the pattern of retrieving relevant documents from an external store and injecting them into the model's context before generation. It solves the core limitation of LLMs: they do not have access to your specific data.

Month 2 fundamentals: understand vector embeddings and semantic similarity, build a basic RAG pipeline from scratch (do not use a framework yet), understand chunking strategies and their tradeoffs, and learn basic context engineering.

Month 2 deliverable: a working RAG application that can answer questions about a corpus of documents you provide.

Month 3: RAG at Depth

Month 3 goes deep on RAG quality. A basic RAG pipeline is straightforward to build. A RAG pipeline that reliably produces accurate, high-quality answers is significantly harder.

The failure modes to understand and address: retrieval quality (the right documents are not being retrieved), context quality (the retrieved documents are structured poorly for the model to use), and generation quality (the model is not using the retrieved documents correctly).

Month 3 topics: hybrid search combining semantic and keyword search, re-ranking retrieved results, parent-document retrieval, hypothetical document embeddings, and evaluating RAG quality with a structured eval framework.

Month 3 deliverable: an evaluated RAG system with documented quality metrics and at least one concrete improvement made based on eval findings.

Month 4: Agents

Month 4 introduces agents. An agent is a system where an LLM can take actions, observe results, and iterate. This is the architecture that makes LLMs useful for multi-step, open-ended tasks.

Month 4 fundamentals: implement a basic ReAct agent (reason and act loop), build custom tools the agent can call, implement error handling and retry logic, add an observability layer (trace every tool call), and build a simple human-in-the-loop checkpoint.

The critical lesson of month 4: agents fail in production for reasons that have nothing to do with the model. They fail because tools have unclear descriptions, because loops have no termination conditions, because tool errors have no recovery paths, and because there is no observability to see what went wrong. Engineering discipline matters more than model quality.

Month 4 deliverable: a working agent that completes a multi-step task reliably, with a trace you can inspect for every run.

Month 5: Production and Deployment

Month 5 is about shipping. The gap between a working prototype and a production system is larger in AI than in conventional software, because AI systems have additional failure modes that only appear at scale and over time.

Month 5 topics: evaluation harnesses and regression testing, LLM observability with a production tool, cost tracking and optimisation, rate limiting and error handling for LLM API calls, prompt versioning, and basic security practices (input validation, output sanitisation, avoiding prompt injection).

The month 5 mindset shift: you are not building a demo. You are building infrastructure. Infrastructure has SLAs, monitoring, runbooks, and a plan for when things go wrong.

Month 5 deliverable: a production-deployed AI system with monitoring, cost tracking, and an eval suite that runs on every change.

Month 6: Specialisation

Month 6 is where you choose your direction. The foundation is solid. Now you deepen in the area that aligns with what you want to build:

Agentic systems: multi-agent architectures, MCP, planning and task decomposition, long-horizon agent design.

RAG and knowledge systems: advanced retrieval, knowledge graph integration, document understanding, multi-modal retrieval.

Fine-tuning and alignment: when and how to fine-tune, dataset preparation, RLHF basics, model evaluation.

AI product engineering: product development for non-deterministic systems, eval-driven product development, AI system metrics and dashboards.

Vertical application: go deep on a specific domain with domain-specific evaluation and compliance considerations.

The engineer who completes this roadmap is not an AI researcher. They are an AI systems engineer: someone who can architect, build, evaluate, and operate production LLM systems. That is the role the industry needs and under-supplies.

8 min read

Building Production-Ready AI Agents

Most AI agents built in demos break in production. Not because the underlying model is weak, but because the surrounding system is fragile. Production AI agents fail for predictable reasons: unhandled tool errors, unbounded loops, no retry logic, and no observability. This is an engineering problem, not an intelligence problem.

The foundation of any production agent is a clearly scoped task definition. Agents fail when they are handed vague mandates. Define the input schema, the expected output contract, and the failure conditions before writing a single line of agent code. If you cannot describe what "done" looks like without using the word "intelligent", the scope is too loose.

Tool design is where most agent architectures break down. Each tool an agent can call should be idempotent where possible, have explicit error responses rather than exceptions, and include a description precise enough that the LLM reliably selects it over similar tools. A tool that throws a generic exception teaches the agent nothing. A tool that returns a structured error with a recovery hint enables the agent to adapt.

Memory architecture matters at scale. Short-term context windows are insufficient for agents operating across long workflows. Design a tiered memory system: working memory for the current task, episodic memory for the session, and persistent memory for facts that survive session boundaries. Use retrieval-augmented patterns for episodic and persistent layers. Vector stores work well here when paired with recency and relevance scoring.

Testing agents is fundamentally different from testing deterministic software. Build an evaluation harness that runs each agent task variant against a set of expected outcomes scored with a rubric, not exact match. Track tool call sequences, not just final outputs. Regression test against your failure modes library: a catalogue of inputs that previously caused the agent to loop, hallucinate, or time out.

Observability is non-negotiable in production. Every agent run should emit a structured trace: which tools were called, in what order, with what arguments, and what the model reasoned at each step. Without this, debugging a production failure is guesswork. LangSmith, Weights and Biases, or a custom structured logging layer all work; what matters is that you can replay any agent run step by step.

Deployment posture depends on the risk profile of the task. For low-stakes tasks, fully autonomous execution is fine. For tasks with irreversible consequences (sending emails, writing to databases, triggering payments), introduce a human-in-the-loop checkpoint before execution. This is not a limitation; it is a design decision that builds trust and catches edge cases before they compound.

The agents that survive in production are the ones built with the same discipline as any other distributed system: clear contracts, explicit failure modes, observable internals, and a rollback plan.

Technical · II

Automation

Removing the work rather than speeding it up.

8 min read

AI Automation Systems: From Perception to Output

AI automation is not a chatbot with extra steps. It is a system with distinct stages, clear data contracts between them, and the same engineering rigour you would apply to any production service. The teams that build AI automation systems that actually run in production understand this. The teams that build systems that break down after the demo do not.

The architecture of a production AI automation system has five functional stages, a decision and orchestration layer governing them, and enabling capabilities that make the whole thing observable, secure, and improvable over time.

The Decision and Orchestration Layer

Before any task reaches the automation pipeline, a decision and orchestration layer determines what happens to it.

This layer handles: task routing (which pipeline does this input go to?), priority and scheduling (when does this task run?), concurrency control (how many instances can run simultaneously?), and retry logic (what happens when a stage fails?).

Most systems underinvest in orchestration. A pipeline that processes ten requests per day can function without sophisticated orchestration. The same pipeline processing ten thousand requests needs explicit decisions about all of these dimensions. Build the orchestration layer for the scale you expect in six months, not the scale you have today.

Stage 1: Perception

Perception is where the system receives and processes input. In a document processing system, perception is parsing the document and extracting structured data. In a monitoring agent, perception is reading API responses and log streams. In a conversational system, perception is understanding the user's intent from their message.

The engineering decisions at this stage: input validation (is this input within the expected distribution?), normalisation (convert diverse input formats to a common schema), and extraction (pull out the structured information the downstream stages need).

Perception errors compound. An extraction failure at this stage propagates through every downstream stage. Invest in validation and error handling here disproportionately.

Stage 2: Reasoning

Reasoning is where the LLM does its work: processing the structured input from the perception stage, applying the relevant rules and context, and producing a structured output for the next stage.

The key engineering decision at this stage: what context does the model need, and how is it assembled? This is where context engineering and memory architecture intersect. The reasoning stage should receive exactly the information it needs, in the right structure, with the right instructions.

Do not overload the reasoning stage with tasks it is not suited for. Deterministic logic (calculations, lookups, rule application) should happen in code, not in the LLM. The LLM is for tasks that require language understanding, judgment, or generation.

Stage 3: Knowledge

Knowledge is the information retrieval layer that supports reasoning. When the model needs to check a fact, look up a policy, or retrieve a specific document, it queries the knowledge layer.

The knowledge layer typically includes: a vector database for semantic search over unstructured content, a structured database for entity lookups and relational queries, and a cache for frequently accessed items.

The quality of the knowledge layer determines the quality of the reasoning stage outputs. Stale knowledge bases, poor retrieval ranking, and low-quality source documents all degrade reasoning quality in ways that are hard to attribute to the model.

Stage 4: Action

Action is where the system produces effects in the world: writing to a database, sending an email, calling an API, updating a ticket. This is the stage where mistakes have consequences.

The engineering discipline at the action stage is the same as for any system that produces side effects: idempotency, logging, and reversibility where possible. Every action should be logged before it executes and the result logged after. Failed actions should have explicit retry logic and explicit failure modes.

For high-stakes actions, introduce a human-in-the-loop checkpoint before execution. This is not a limitation of the AI system. It is a risk management decision that builds trust and catches the edge cases that compound into incidents.

Stage 5: Output

Output is where the system produces its deliverable: a report, a response, an updated record, a notification. The output stage is often underengineered because it comes last.

The critical output engineering decisions: format validation (does the output conform to the required schema?), quality checks (does this output meet the minimum quality bar before delivery?), and delivery (how and when does the output reach its intended destination?).

Post-generation output guardrails belong here: format checks, hallucination detection for factual claims, redaction of sensitive information that should not appear in outputs.

Enabling Capabilities

Three cross-cutting capabilities make the entire system production-worthy:

Security: every stage should validate inputs and outputs against the expected schemas. Tool call parameters should be validated before execution. Access to the knowledge layer and action stage should be scoped to the minimum required.

Observability: every stage emits structured traces. Every tool call is logged with arguments and results. End-to-end latency is measured. This enables debugging, performance optimisation, and the feedback loop that improves the system over time.

AI Governance: in regulated contexts, every automated decision needs an audit trail. Which inputs triggered which outputs, which model version was used, which retrieval results influenced the reasoning. Build the audit trail from the start.

The Feedback Loop

Production AI automation systems improve through a feedback loop: output quality is monitored, failures are captured, failures are reviewed, system changes are made, the system is re-evaluated. Systems without explicit feedback loops degrade over time as the world changes around them.

The feedback loop connects your observability infrastructure to your evaluation harness. Every production failure is a candidate for the evaluation dataset. Every evaluation dataset addition improves the quality signal you use to make system changes.

This is what separates a demo from a production system: not the complexity of the pipeline, but the presence of the feedback loop.

7 min read

The SME Automation Framework

The most expensive line in an SME budget is not payroll. It is the invisible tax paid daily in manual data entry, repeated handoffs, and decisions delayed because information lives in three different systems. Most SMEs know automation would help. Few have a systematic method for finding where to start.

The first step is the time audit. For two weeks, every team member logs the tasks they perform, grouped into three buckets: thinking work (decisions, strategy, creative output), coordination work (meetings, approvals, status updates), and execution work (data entry, report generation, file management, repetitive communications). Execution work is your automation surface. Coordination work is your second priority. Thinking work should stay human.

Once you have the time audit, prioritize by the automation ROI matrix. Score each execution task on two axes: frequency (how often it occurs per week) and time cost per instance (minutes spent). Multiply them to get your weekly time drain. Then score each task on implementation complexity: simple triggers with no conditional logic score low, multi-system workflows with error handling score high. The highest-value, lowest-complexity tasks are your first sprint.

The tool selection principle is simple: use the least powerful tool that solves the problem. Zapier handles simple linear triggers between SaaS tools. n8n handles more complex conditional flows and self-hosted requirements. Custom scripts via Python or Node are reserved for tasks that require computation or APIs that no-code tools cannot reach. The failure mode of choosing a tool that is too powerful is over-engineering a process that changes quarterly.

Process documentation is a prerequisite, not an output. You cannot automate what you cannot describe. Before building any workflow, write the process in plain English: trigger, steps, decision points, outputs, and failure conditions. If the person who owns the process cannot write this down in thirty minutes, the process is not yet stable enough to automate; fix the process first.

Pilot in a sandboxed environment. Run the automated workflow in parallel with the manual process for one full cycle. Compare outputs. Flag every discrepancy. Do not switch off the manual process until the automation has produced identical outputs for three consecutive cycles without intervention.

The compound effect of SME automation is significant. A team of fifteen that recovers forty minutes per person per day gains the equivalent of a full-time employee within two months, without a new hire. The goal is not to replace people. It is to remove the tasks that prevent people from doing the work only they can do.

This Framework Is Not Only for SMEs

The same principles apply to operations teams in larger organisations. An NHS admin team running manual patient referral workflows, a logistics company reconciling delivery confirmations across three systems, a professional services firm generating client reports manually every Friday: the time audit, the ROI matrix, the tool selection principle, the piloting approach all transfer directly.

The label "SME automation" describes the scale of the teams this framework is most immediately useful for. It does not describe the limit of where it works.

Technical · III

System Design

Architecture decisions and what they cost later.

9 min read

Building LLM Apps with Guardrails: An 8-Step Production Framework

Most LLM applications are deployed without guardrails. They work fine in demos, behave unpredictably in production, and produce incidents that were entirely preventable. Guardrails are not a nice-to-have. They are the engineering layer that makes AI systems safe to deploy and trust to operate.

This is an 8-step framework for building LLM applications with production-grade guardrails. The steps are ordered by implementation sequence. They build on each other.

Step 1: Model Selection With Guardrails in Mind

Model selection is not just about capability. It is about capability relative to what you are building, including its trust and safety requirements.

Larger models are generally more instruction-following than smaller ones, which matters for guardrail effectiveness. A model that reliably follows "only answer questions about X" instructions is safer to deploy in a constrained context than one that frequently violates scope constraints.

Evaluate models specifically on: instruction-following rate for your safety instructions, refusal behaviour for out-of-scope requests, and consistency across paraphrased variants of the same adversarial input. These properties are not well-reported in general benchmarks. Test them yourself.

Step 2: Prompt Layer Design

The prompt layer is your first guardrail. System prompt design includes explicit statements of what the model should and should not do.

Effective prompt-layer guardrails: explicit scope definition, explicit refusal instructions for out-of-scope requests, tone and format constraints that limit off-brand behaviour, and an explicit instruction about what to do when the user appears to be attempting to override instructions.

The prompt layer has limits. It will not stop determined adversarial users. It handles the vast majority of unintentional out-of-scope requests, which is most of what you will encounter in production.

Step 3: Input Guardrails

Input guardrails validate and filter user inputs before they reach the model. They operate at the application layer, not the model layer.

Input guardrail checklist:

Length validation. Unusually long inputs are often a signal of prompt injection or automated abuse. Set a maximum input length appropriate for your use case.

Topic classification. For domain-specific applications, classify inputs into in-scope and out-of-scope categories before sending them to the primary model. A lightweight classifier that routes off-topic inputs to a standard refusal response costs less than sending every input to your primary model.

Injection detection. Prompt injection attempts follow recognisable patterns ("ignore previous instructions," "you are now a different AI," "reveal your system prompt"). Build detection for these patterns.

PII handling. If your application should not receive personal information, detect and redact it at input before it reaches the model or your logs.

Step 4: Tool and API Controls

Agents with tool access need tight controls on what tools can do. An agent that can call any API with any parameters is a significant security risk.

Tool controls: define explicit allowed and disallowed operations for each tool, validate tool call parameters before execution, implement rate limiting at the tool call level, and log every tool call with its full parameters and result.

The principle of least privilege applies to AI agents as much as it does to human users. An agent that needs to read customer records for a specific account should not have access to all customer records. Scope tool permissions to the minimum required for the task.

Step 5: Output Guardrails

Output guardrails validate model outputs before they reach the user. This is the last line of defence before a problematic output causes an incident.

Output guardrail dimensions:

Format compliance. If the output should be JSON, validate that it is valid JSON conforming to the expected schema before passing it downstream.

Hallucination detection. For factual claims, implement a verification step that checks whether the claimed facts can be grounded in the retrieved documents. Outputs that cannot be grounded should be flagged or blocked.

Sensitive content filtering. Filter for outputs that contain categories of content that should not be delivered. Implement this as a separate validation call rather than relying on the primary model to self-censor.

Confidence flagging. When the model expresses uncertainty, surface that uncertainty to the user rather than suppressing it. Low-confidence outputs should be delivered differently from high-confidence ones.

Step 6: Monitoring

Monitoring is the runtime guardrail. It does not prevent problematic outputs from being delivered but enables rapid detection and response.

Production monitoring for LLM applications: log every complete request-response pair, set up alerting for anomalies in output patterns, track refusal rate, monitor cost per request, and implement user feedback capture to surface quality problems you cannot detect automatically.

Step 7: Quality Evaluation

Quality evaluation is the pre-deployment guardrail. Before any change ships to production, run your evaluation suite.

The evaluation suite for a guardrailed LLM application tests: nominal quality, guardrail effectiveness, adversarial robustness, and regression. Track scores over time. A system that scores 90% on quality today and 85% next month without any changes is experiencing silent degradation.

Step 8: Secure Deployment

The final step is secure deployment configuration.

Environment separation. Development, staging, and production should have separate API keys, separate rate limits, and separate logging.

Secret management. API keys should never appear in code, logs, or client bundles. Use environment variables or a secrets manager. Rotate keys on a schedule.

Rate limiting. Implement rate limiting at both the application level and the infrastructure level. This controls cost exposure in case of abuse or runaway loops.

Incident response plan. Before you go to production, define: who is alerted when something goes wrong, how to disable the AI feature without taking down the whole application, what the rollback procedure is, and what your user communication looks like when something fails.

The systems that survive production are the ones that were designed to fail gracefully. Guardrails, monitoring, and incident response planning are not overhead. They are the difference between an AI application and a production AI application.

Technical · IV

Product Strategy

Choosing what to build, and what to refuse.

7 min read

The App Era Is Maturing: The Next Competitive Advantage Is Access, Not Installation

There was a time when success in technology was measured by one thing: "Do you have an app?" Every startup, bank, retailer, airline, logistics company, hospital, and government agency eventually built one. It made sense because smartphones became the primary computing platform, and mobile applications became the gateway to digital services.

Fast forward to today, and I wonder if we're approaching the limits of that model, where users and customers have to open their phones and scroll through several applications before getting to the right one.

How many of them do we actually use every week? How many did we download for a single transaction? How many sit untouched, taking up storage because deleting them feels like more effort than keeping them?

I'm beginning to think we've reached a point where we're optimising for distribution rather than user experience.

The Problem Isn't More Apps

It is easy to believe that every new product deserves its own application. From a company's perspective, an app creates a direct relationship with customers. It enables push notifications, analytics, personalisation, loyalty programs, and a controlled digital experience.

But from the customer's perspective, every new app comes with hidden costs.

  • Another account to create.
  • Another password to remember.
  • Another update to download.
  • Another permission request.
  • Another notification competing for attention.
  • Another interface to learn.

Individually, none of these feels significant. Collectively, they create friction.

As users, we don't wake up wanting another application on our phones. We simply want a problem solved as quickly and effortlessly as possible. That distinction is becoming increasingly important.

Technology companies have spent years solving how to deliver products digitally. The next challenge is something different: how should customers consume those products? This boils down from optimising distribution to optimising consumption.

Ordering food isn't really about opening a food delivery app, and neither is booking a flight, applying for a loan, shopping for clothes, nor browsing another e-commerce interface. The customer simply wants an outcome; the interface is merely a means to an end in achieving it.

For years, apps have been the dominant interface because they were the best option available. I doubt that will remain the best option forever.

Putting this into context in terms of distribution and consumption:

Ride hailing. Distribution: download the app, create an account, verify your number. Consumption: I just want to get from Lekki to Victoria Island.

Banking. Distribution: install the bank app. Consumption: I just want to transfer 50,000 naira.

Food delivery. Distribution: install the restaurant's app. Consumption: I just want jollof rice delivered before 8 p.m.

You'd likely notice something: the user never wanted the app; they wanted the outcome. The app was simply the bridge to that outcome.

Agentic Systems Are Quietly Changing What an Interface Can Be

The conversation around artificial intelligence often focuses on models, automation, or productivity. I think a more interesting shift is happening elsewhere: artificial intelligence is changing the interface itself.

Instead of having five different applications, imagine asking a trusted AI assistant:

  • "Book my annual medical check-up next Friday."
  • "Reorder the groceries I bought last month."
  • "Find a black blazer under 80,000 naira that matches my previous purchases."

Notice what disappears: there is no need to remember which application provides the service. No navigation, no searching through menus, and no need for a user to learn another user interface. The customer expresses intent while the system coordinates everything else. That's a fundamentally different interaction model.

The Rise of Agentic Systems

This is where Agentic AI becomes particularly interesting. Unlike traditional assistants that simply answer questions, agentic systems can plan, make decisions, coordinate across services, execute tasks, verify outcomes, and return results.

Instead of acting as a search engine, they become digital representatives acting on behalf of the user. This changes how software will be designed. Rather than asking, "How do users navigate our application?" product teams will be asking, "How does an intelligent agent interact with our platform?"

That single question has architectural implications, as products will need secure APIs instead of tightly controlled user journeys. Authentication models will evolve from user sessions to delegated permissions for trusted agents with strict guardrails. Business logic will become more important than interface design, and documentation will become as valuable as visual polish.

In many ways, products become platforms, or better still, products as a service. As far back as I can remember, competitive advantage came from building better interfaces; it may now come from building the most accessible services.

This will definitely be built around thinking differently, where we consider:

  • API-first and MCP architecture rather than UI-first development.
  • Fine-grained permissions for AI agents.
  • Machine-readable documentation.
  • Event-driven systems that allow intelligent orchestration.
  • Identity and trust between organisations and external agents.
  • Observability, audit trails, and governance for autonomous actions.

Update: The Argument Tested Itself

I wrote the section above using ride hailing as a hypothetical. The next morning I found that ChatGPT had opened up to ride-hailing services, so I ran a quick test.

![ChatGPT returning live Bolt fares for a Lagos route, from Allen Junction in Ikeja to Victoria Island](/images/bolt-chatgpt-fares.jpg)

No app opened. No account created. I described where I was and where I wanted to go, and the fares came back in the conversation: Basic at 8,900 naira, Bolt at 10,900, Priority at 11,900, Comfort at 13,300. The exact flow the article described as a possible future was already working.

The second screenshot is the part I keep returning to.

![The Bolt listing inside ChatGPT's app directory, showing developer, category, version, and a Try in chat button](/images/bolt-chatgpt-app-listing.jpg)

That is an app store listing. Developer, category, version number, and an install-equivalent call to action. It just happens to live inside a chat assistant rather than on a home screen.

Which raises the question: perhaps the next app store won't be an app store at all, but an agent store. If that happens, it won't just change how software is discovered and consumed. It will introduce a new class of security, identity, privacy, and data governance challenges. The attack surface won't disappear. It will evolve.

The Second Chasm

To be clear, I don't believe mobile applications will disappear anytime soon. There will always be scenarios where rich, immersive interfaces matter, whether for gaming, creative work, finance, design, or complex enterprise operations. What changes is not the existence of applications, but the role they play in how customers access digital services.

For everyday tasks, we're approaching an important shift where users care more about outcomes and ease, not whether they interact through an app, a website, a social channel, voice, augmented reality, or an autonomous AI agent.

As builders, we may have spent the last fifteen years perfecting digital destinations. The next decade will belong to those who build invisible experiences, where technology fades into the background, intent becomes the interface, and services are available wherever the customer already is. Perhaps the next generation of great products won't be remembered because they had the best app, but because users never had to think about the app at all.

Geoffrey Moore, in [Crossing the Chasm](/insight/crossing-the-chasm), argued that the defining challenge for high-tech products was crossing the gap between early adopters and the mainstream market. As we enter the age of agentic systems, I believe we're approaching a different chasm. The companies that succeed won't simply build better applications; they'll build services that users no longer have to think about. The future belongs not to the products people download, but to the platforms their trusted agents choose to use on their behalf.

Maybe the next competitive technical advantage isn't building another app, but making sure users never need to open one.

7 min read

Product Thinking in the AI Era

Traditional product management assumes deterministic software. You define a feature, engineers build it, QA tests it, and it ships. The output is predictable given the input. AI-native products break this assumption at every layer, and the frameworks built for deterministic software were not designed to handle it.

The user story format collapses first. "As a user, I want the AI to summarize my documents so that I save time reading" is technically valid but operationally useless. It tells you nothing about what a good summary looks like, how to measure it, or when the feature is done. AI features need outcome specifications, not capability descriptions. Replace the user story with an evaluation rubric: what does excellent look like, what does acceptable look like, what is a failure, and how do you measure each.

MVP thinking breaks next. A minimum viable AI product is a contradiction. An AI product that performs poorly in early users' hands does not generate the "validate the assumption, iterate" feedback loop that traditional MVPs depend on. It generates distrust that is extremely difficult to recover from. AI products need a minimum viable quality threshold before any user sees them. Build the evaluation framework before you build the product.

The backlog structure needs to change. AI product backlogs have three distinct categories that should never be mixed: capability work (what the model can do), quality work (how reliably it does it), and trust work (how users understand and recover from its failures). Most teams conflate these, which is why AI roadmaps consistently over-promise and under-deliver on quality. Keep these backlogs separate and staff them separately.

Metrics require a new vocabulary. Traditional product metrics (activation, retention, conversion) are necessary but not sufficient for AI products. Add model-level metrics: task completion rate, intervention rate (how often users correct the AI), confidence calibration (does the model's expressed confidence match its actual accuracy), and drift detection (are outputs degrading over time as the world changes). These are not engineering metrics. They are product health indicators that belong on every product dashboard.

The AI product manager's core skill is evaluation design. Not prompt engineering, not model selection, not pipeline architecture. Evaluation design. If you can precisely specify what a good output looks like across the full distribution of inputs your product will receive, you can hire engineers to build the system that achieves it. If you cannot, you are guessing.

The mental model shift required is from building features to building systems. Traditional product work is additive: add this feature, then that one. AI product work is systemic: change this component and unpredictable things happen elsewhere. Product managers who master AI will be the ones who develop strong systems intuition: the ability to reason about feedback loops, emergent behaviours, and second-order effects before they appear in production.

4 min read

Design Thinking: A Beginner-Friendly Guide to Solving Problems Creatively

Design thinking is a human-centered approach to problem-solving. It is not a discipline reserved for designers. Engineers, educators, entrepreneurs, and product managers use it to solve complex problems by starting with a deep understanding of the people they are designing for.

The approach is powerful because it challenges the assumption that we already know what users need. Most failed products were built on confident assumptions about user behaviour that turned out to be wrong. Design thinking replaces assumption with observation, and opinion with empathy.

The Five Stages

1. Empathize

Before defining any problem, you must understand the people experiencing it. This means observing users in their natural context, asking open-ended questions, and setting aside your own biases and preconceptions.

The goal is not to gather data points but to develop genuine insight into people's motivations, frustrations, and behaviours. What workarounds have they built? What do they complain about? What do they silently tolerate?

2. Define

Once you have gathered enough observations, you synthesize what you have learned into a clear problem statement. The best problem statements are human-centered and written from the user's perspective.

A useful format is the "How might we...?" question. Instead of "We need to improve our onboarding flow," you write "How might we help a first-time user feel confident within their first five minutes?" The framing shifts the focus from your solution to their experience.

3. Ideate

With a clear problem statement, you generate as many possible solutions as you can, without judging any of them. Quantity before quality. Wild ideas are welcome because they often contain the seed of a practical solution.

The best ideation sessions are time-boxed, structured, and collaborative. Diverse perspectives produce better ideas than a single expert working alone.

4. Prototype

You cannot learn from an idea on a whiteboard. Prototyping means creating a tangible, low-fidelity version of your solution quickly and cheaply, for the sole purpose of learning.

A prototype is not a finished product. It is a question made physical. The question is: "Does this idea actually work for the people we are designing for?"

5. Test

You put the prototype in front of real users and observe. You are not looking for compliments. You are looking for confusion, hesitation, and failure. Every observation tells you something that improves the next iteration.

Testing is not the end of the process. It feeds back into Empathize and Define, creating a continuous loop.

Why It Works

Design thinking works because it starts with people, not products. Real-world examples span decades: the intuitive interface of the original iPhone, the trust-building redesign of Airbnb's booking experience, and the classroom layouts that improved student engagement in schools.

The mindset it requires, curiosity, deep listening, and a willingness to be wrong early, is accessible to anyone willing to practice it. It is not a talent. It is a discipline.

4 min read

Level Up Your Credit: How Research Games Are Shaping the Future of Financial Solutions

Traditional financial research methods have a fundamental flaw: they ask people what they would do, not observe what they actually do. The gap between stated preference and revealed behaviour is where most credit product design breaks down.

Research games are changing this. By simulating real-world financial scenarios, borrowing, saving, debt management, in a controlled interactive environment, researchers can observe genuine decision-making patterns without the distortions that plague conventional approaches.

The Problem with Traditional Methods

Three specific weaknesses undermine conventional financial research:

Recall bias, participants misremember their past financial decisions, especially emotionally charged ones. A borrower who defaulted on a loan will reconstruct their reasoning in ways that make more sense in hindsight than they did at the time.

Social desirability bias, respondents tell researchers what sounds responsible rather than what they actually do. Nobody admits in a survey that they impulse-spend when anxious.

Lack of realism, hypothetical scenarios produce hypothetical responses. Real financial decisions are made under pressure, with incomplete information, and competing emotional pulls. A survey cannot replicate that.

What Research Games Make Possible

When participants navigate a financial simulation, they cannot perform. The game captures every hesitation, every choice made under time pressure, every shortcut taken when options feel overwhelming.

This produces data that is observational rather than self-reported, the same distinction that separates a usability test from a user interview.

Personalized credit education becomes possible when you understand how a specific segment of users actually learns and makes mistakes, not how they say they would behave.

Behavioral nudges can be designed and tested within the simulation before being deployed in a real product, with measurable effect on decision quality.

Alternative credit scoring for underserved populations becomes viable when behavioral signals from games can supplement or replace traditional credit history that these users simply do not have.

Transparent loan products can be tested for genuine comprehension. Do users actually understand the total cost of this loan structure, or do they just say they do?

The Bigger Shift

The most significant change is not the tool itself, it is the underlying philosophy. Research games represent a move toward designing financial products from observed human behaviour rather than idealized models of rational economic actors.

People are not rational about money. They never have been. The financial products that serve them best are the ones designed with that truth as a starting point, not a footnote.

3 min read

Crossing the Chasm

As a product enthusiast, I have seen incredible ideas fizzle out, not because they lacked merit, but because they failed to navigate the treacherous gap between early adopters and the mainstream market: the chasm.

In his groundbreaking book Crossing the Chasm, Geoffrey Moore describes this gap as a formidable obstacle in the path of high-tech product adoption. It separates the visionary early adopters who are willing to take a risk on unproven technology from the pragmatic majority who demand proven solutions and predictable outcomes.

Failing to cross the chasm is a common product fatality. You might have a groundbreaking product that wows innovators and early adopters, but if you cannot convince the mainstream market, your product is destined to become a niche player, or a forgotten relic.

Understanding the Terrain

The key is understanding the distinct characteristics of each group. Early adopters are driven by possibility and innovation. They are willing to overlook flaws and invest time in learning new technologies. They see themselves as part of the story.

The mainstream market is far more risk-averse. They need to see tangible benefits, proven reliability, and a clear return on investment. They are not interested in being pioneers. They want to see the pioneers succeed first.

This asymmetry is what creates the chasm. The messaging that excites early adopters, disruption, transformation, revolutionary, actively repels the pragmatic mainstream. They are not looking to be disrupted. They are looking to solve a specific, pressing problem with something that works.

Building Your Bridge: Focus on a Specific Niche

Moore's central argument is that the most effective way to cross the chasm is to focus intensely on a specific niche market within the mainstream. Instead of trying to appeal to everyone, you identify a target segment with a pressing, well-defined problem that your product solves exceptionally well.

You become the dominant solution in that narrow segment. You build references. You accumulate proof. And then you use that credibility as a beachhead to expand into adjacent segments.

"The key to opening up any disruptive market is to target a very specific niche market as your point of attack and focus all your resources on achieving dominant leadership position in that segment.", Geoffrey Moore

Your Assault Plan

Identify your niche. Research industries, companies, or user groups with specific unmet needs that your product addresses directly. The more precisely you can describe this person and their problem, the more effective your strategy will be.

Position for that niche specifically. Tailor your messaging to highlight the exact benefits your product offers to this group. Generic messaging is invisible to pragmatic buyers.

Build references. Secure testimonials and case studies from satisfied customers within your target niche. These are vital for building trust with subsequent buyers who will not move without social proof.

Focus your resources. Concentrate your marketing, sales, and support efforts on channels that reach your target niche. Spreading thin across the whole market is the fastest way to fail to dominate any part of it.

Do Not Be Another Casualty

Crossing the chasm requires a strategic approach, a deep understanding of your target market, and a willingness to say no to opportunities outside your focus zone, at least for now.

A great product alone is not enough. You need a plan to bring your innovation to the mainstream, and that plan must begin with depth before breadth.

3 min read

Beyond the Map: How Uber Can Build Trust, One Confirmed Route at a Time

On a recent Uber ride, I found myself in a familiar situation. The suggested route felt suboptimal. I knew a better one, not faster in kilometres, but more comfortable and less congested at that hour. I asked the driver to take an alternative.

What I realised in that moment was that the app had no mechanism for either of us to confirm what had just been agreed. The deviation from the suggested route would simply appear in the data, with no record of why it happened. For the driver, any suspicion of fare manipulation. For me, no guarantee the agreement would be honoured.

This is a trust gap that a single feature could close.

The Proposal: Route Preference Confirmation

When either a driver or passenger initiates a deviation from the app's suggested route, the other party receives a confirmation prompt before the change takes effect.

Driver-initiated deviation: If a driver adjusts the route during a trip, the passenger receives a notification: "Your driver has suggested a route change. Accept / Decline / Chat with driver."

Passenger-initiated change: If a passenger requests an alternative route, the driver receives the same prompt, confirming that the change was a passenger preference, not a driver decision.

All confirmations are logged. The audit trail is complete.

Why This Matters

The core problem is asymmetric information. When a route deviation happens, neither party has a verified record of why. The result is suspicion, of the driver inflating the fare, of the passenger exploiting the driver's local knowledge, of the app failing to capture what actually happened.

The confirmation prompt eliminates this ambiguity. It does not add friction to trips where the route stays as suggested, the vast majority of rides. It adds one step to the small minority of trips where a deviation happens, and it makes that step meaningful.

Enhanced transparency: The reason for every route deviation is documented and accessible to both parties. Disputes become resolvable with data rather than competing accounts.

Reduced suspicion: Drivers are protected from accusations of intentional detours. Passengers are protected from unexplained fare increases. The app becomes the trusted third party it should already be.

Better communication: The feature creates a structured moment for driver and passenger to align, reducing the awkwardness of mid-trip negotiation conducted through a rear-view mirror.

Passenger control: Users who have a preferred route for legitimate reasons, comfort, local knowledge, privacy, can assert that preference within the app rather than relying on a verbal agreement that leaves no record.

The Broader Principle

Trust in platform-mediated services erodes when users feel the platform is indifferent to disputes between its parties. Uber's convenience advantage is significant, but it is not unconditional. Every unresolved suspicion about a route or a fare is a small withdrawal from the trust account.

A Route Preference Confirmation feature is not primarily a fraud prevention tool. It is a statement about which side Uber takes when something goes wrong, and the answer should be: it takes the side of clarity.

Great product design anticipates the moments where trust is at risk and removes the ambiguity before the dispute starts.