“Where can we use AI?” is usually the wrong first question. It begins with a technology and searches for justification. The better question is: what must this system do, and what is the simplest reliable way to do it?

That question often leads to ordinary software: deterministic rules, a well-designed form, a database constraint, a scheduled job or a clearer interface. None of those options becomes less valuable because a language model exists.

Use AI where the input is genuinely variable

AI becomes useful when the work contains ambiguity that cannot be economically captured in fixed rules. Documents arrive in inconsistent formats. Operators describe the same condition in different language. Images contain patterns that matter but resist manual inspection at scale. A decision depends on retrieving relevant knowledge from a large, changing corpus.

These are not guarantees that AI is appropriate. They are signals that probabilistic methods may create value. The next step is to define the task narrowly enough to evaluate.

If the desired behaviour cannot be described and tested, it cannot be responsibly improved.

Separate capability from authority

A model may be capable of producing a recommendation without being entitled to act on it. This distinction matters in healthcare, financial crime, defence, infrastructure and any environment where an error has a meaningful consequence.

The surrounding system should define what the model can see, which tools it can call, how confidence is interpreted, when a person reviews the output and what happens when the model is unavailable. Authority belongs to the operational design, not to the model.

Evaluate the whole workflow

A strong benchmark score does not establish operational usefulness. Evaluation needs representative inputs, including difficult and undesirable cases. It also needs end-to-end measures: time saved, exceptions created, review burden, false confidence, latency, cost and the effect on downstream decisions.

For generative systems, a single accuracy percentage is rarely enough. The evaluation set should reflect the categories of failure the operation cares about. Results should be reviewed with the people who understand the domain, not only by the team building the model layer.

FOUR QUESTIONS

Before adding a model

  1. Is the task too variable for deterministic rules?
  2. Can useful performance be evaluated with representative cases?
  3. Is there a safe fallback or review path?
  4. Does the benefit justify the added uncertainty, cost and operating burden?

Engineer for change

Models, providers and prices change quickly. The valuable part of the system is often the stable layer around them: the task definition, evaluation set, domain context, permissions, observability and user workflow. Treat the model as a replaceable dependency where possible.

This does not mean hiding every model behind an elaborate abstraction. It means avoiding accidental dependence on behaviour that has never been specified. A clean boundary lets a team compare alternatives, roll back a change and understand whether a new model actually improves the operation.

AI should make the system better

The bar is deliberately plain. If AI makes the capability more useful, more accessible or more effective—and the risks can be managed—use it. If conventional software provides the better result, build that instead.

AI earns its place by improving the work. Everything else is theatre.