Prompt Engineering Principles for the Workplace.

Three golden rules for using artificial intelligence: be specific, the unknowns and verify what is missing.
We keep making a basic mistake with artificial intelligence.
We talk to it. It talks back. So we begin imagining somebody is in there.
ChatGPT says, “I think.” Claude says, “I understand.” A model apologises, and can sound deeply convinced of something.
The linguistic surface is so human that we instinctively supply the rest:
mind, intention, belief, consciousness.
We should be much more careful.
There is no accepted scientific basis for treating fluent language from today’s large language models as proof of subjective consciousness. A model can produce the language of belief without demonstrating that it possesses beliefs. It can produce the language of sadness without showing that anything is experiencing sadness.
AI does not “believe” a sentence it produces but convinces you it deeply believed in the statement. AI cannot believe. It’s a machine. It’s your own consciousness that projects this onto a machine. Just the way people watch a random cloud formation and say “Oh the cloud is smiling at me”. The cloud is not smiling at you.
This is also the consequence of a “Godless society” which doesn’t believe in the Soul or Afterlife. It is also the consequence of “Bad Government” which hasn’t prioritised “AI education” nearly enough to keep pace with the technology. And the “hype” created by “teary eyed panic faced AI researchers” coming on LinkedIn feeds saying “Ohhh it’s too dangerous… it will kill us all” fuels the fire.
This is “highly irresponsible” – because AI researchers should know better. Especially from the “top labs” like – Anthropic & OpenAI.
They know or ought to have known – they keep spreading “False news” on purpose – but many do this to “market themselves” – and I would stop short of claiming to “sell their product” in the market. It may be ignorance. Which is inexcusable. The problem is “at least a part of the population” take them seriously and immediately start imagining AI as a “Marvel Villian”. It’s not. It’s a tool. It’s simply a machine.
This problem is also “not new”. Weisenbaum at MIT made Eliza in the 1960s. It had a “simple script” – IF-THEN logic of a psychotherapist. People got hooked and wanted to “Speak to ELIZA in Private”. This was not even AI. It was simply 420 lines of code.
Less code than a few emails you send in a day.
Science till date does not back this up at all.
A more useful metaphor is far less romantic.
Think of a large language model as an extraordinarily complicated pinball machine.
The training process builds the machine.
Your prompt launches the ball.
Context positions the bumpers.
Constraints close particular routes.
Examples encourage others.
Tools connect the machine to the outside world.
And prompt engineering is, fundamentally, the art of routing the ball.
A pinball machine can also be used as an analogy of the brain in the way ideas “evoke emotions, inspire thinking in us”.
Nature describes the underlying generation mechanism even more precisely. A 2025 Nature paper explains:
“Their central task is to predict the probable next token, given a sequence of prior tokens.”
There is also an important qualification concerning apparent “reasoning.” A very recent Nature Machine Intelligence paper warns against assuming that readable reasoning text is a transparent window into what the model internally did. It says generated reasoning tokens:
“are not a faithful explanation of the model’s actual reasoning process.”
What actually happens when you prompt AI?
A language model is trained on enormous quantities of information and learns statistical relationships within that information. Those learned relationships are encoded in numerical parameters.
When you type a prompt, your text is broken into tokens. The model processes those tokens together with whatever other context it has been given and generates possible continuations. One token is produced, becomes part of the context, and the process repeats.
Simplified brutally:
Prompt → context → model → possible next tokens → output → repeat.
OpenAI describes its models similarly: they learn patterns from large quantities of information and generate new material from those learned relationships rather than retrieving a prewritten answer from a database. OpenAI Help Center
The result can be Shakespearean prose, computer code, mathematics or a legal argument.
But underneath the prose is computation.
That matters because human speech and AI prompting are not the same thing.
Based on trained human speech, it is simply predicting the next “likely” word.
Why this looks eloquent is because of the amount of “human speech” it has been fed.
A probabilistic match is highly unlikely to be wrong in many but not all contexts. How recent the data is, the quantum of the data available for the query, how you prompted, whether you supplied your own documents to provide examples, all of this feeds into the quality of the answer.
In the end, you remain the Judge of the answer and Owner of anything you choose to use from the answer.
Human Prompted
Suppose I walk into an office and tell a colleague:
“Can you fix that presentation?”
Almost every piece of information required to complete the task is missing.
Which presentation?
What is wrong with it?
Who is presenting?
Who is the audience?
What does “fix” mean?
Yet my colleague may understand immediately.
Why?
Because humans communicate on top of immense amounts of implicit context.
They remember yesterday’s meeting. They know which slides the chief executive hated. They know my preference for sparse slides. They recognise my tone. They understand the organisational politics. They remember what “fix it” meant the previous five times.
Humans are prompted by much more than words.
Memory. Environment. Relationships. Emotion. Culture. Incentives. Body language. Physical experience. Fear. Desire. Habit.
Human speech can therefore be extraordinarily compressed.
AI Prompted
An AI system starts somewhere different.
An application may provide it with conversation history, memory, documents, search results, databases, images or tools. But it does not automatically share your situation simply because its language sounds human.
It receives context.
This produces the first major principle of AI use:
Human communication can tolerate implicit context. AI becomes more reliable when important context is explicit.
Compare:
“Analyse this company.”
with:
“Determine whether operating-margin improvement over the previous eight quarters came primarily from gross-margin expansion or SG&A leverage. Adjust for disclosed restructuring costs. Separate facts from assumptions and identify what evidence would falsify your conclusion.”
The second prompt does not magically make the model smarter.
It makes the problem narrower.
That gives us the first golden rule.
Golden Rule One: Specificity
Specificity reduces the number of reasonable wrong answers.
A weak prompt leaves thousands of legitimate directions open.
A strong prompt establishes:
Objective → Context → Evidence → Constraints → Output → Verification
Suppose I ask:
“Tell me about nuclear power.”
Almost anything can emerge.
Instead:
Compare nuclear, offshore wind and solar for providing dependable low-carbon electricity to the UK between 2030 and 2050. Separate capital cost, operating cost, carbon intensity, intermittency and grid-system requirements. Prefer government, grid-operator and peer-reviewed evidence. State where credible estimates disagree. Give me an 800-word analysis and a comparison table.
The model itself has not changed.
The route has.
This is why expertise continues to matter.
The difference between:
“Analyse this investment”
and:
“Separate revenue growth caused by volume, price and acquisition effects; normalise EBITDA for exceptional items; calculate cash conversion and identify the assumptions most likely to invalidate the valuation”
is not primarily knowledge of prompting.
It is knowledge of finance.
You cannot specify what you do not understand.
You cannot ask for the right evidence if you do not know what evidence matters.
And you cannot recognise an excellent answer unless you possess some conception of excellence yourself.
AI therefore creates what I would call a specificity premium.
The more clearly you understand a problem, the more precisely you can route the model through it.
It also creates a discernment premium, and even more of a knowledge premium than before.
You cannot prompt well if you don’t know, and you cannot generate or use high quality output if you don’t know.
You wouldn’t know what high quality output looks from low quality output if it smacked you in the face, unless you are highly knowledgeable in your subject area.
AI is not a crutch, it’s not an equaliser.
It’s a very sharp blade.
Don’t use it well, and it will cut deeper.
Many of these concepts, I will expand in future articles.
Training Shapes the Machine
But your prompt is only half the story.
The pinball machine existed before you launched the ball.
Imagine two otherwise identical models.
Train one predominantly on trolling, conspiracy theories, ragebait, insults and poorly reasoned internet arguments.
Train the other primarily on Shakespeare, mathematics, legal judgments, philosophy, high-quality journalism and two centuries of scholarship.
Then give both exactly the same prompt.
We would hardly expect identical answers.
Real frontier models are obviously nothing like this simple experiment. Their training data are huge mixtures and are subsequently filtered and subjected to extensive post-training.
But the principle remains:
Distribution in; tendencies out.
What a model encounters affects what it learns.
And we generally do not possess a complete public, auditable inventory of every document and weighting that shaped frontier systems.
So data quality matters twice.
It matters during training.
Then it matters again when you give the model documents, examples, sources and context.
We have relatively little control over the first.
We have enormous control over the second.
That is why prompt engineering is gradually becoming context engineering.
The sophisticated question is no longer merely:
“What should I type?”
It is:
“What information should this model reason over?”
The Two Advocates
Einstein famously used thought experiments (Gedankenexperiment) to expose the underlying structure of difficult problems: chasing a beam of light, standing inside an accelerating elevator.
We can do something similar with AI.
Imagine two rooms outside a courtroom.
In Room A, an AI receives all the evidence from a murder trial and this instruction:
Act as prosecution counsel. Construct the strongest defensible argument that the defendant is guilty.
Room B receives precisely the same evidence:
Act as defence counsel. Construct the strongest defensible argument that reasonable doubt remains.
The doors remain shut.
From Room A emerges an extraordinary prosecution.
Confident. Logical. Persuasive.
From Room B comes an equally extraordinary defence.
Confident. Logical. Persuasive.
Listening outside, we could easily imagine two minds passionately disagreeing.
Now open the doors.
There were not two models.
It was the same model.
Same parameters.
Same training.
Same evidence.
Different prompt.
Which position did the AI actually believe?
Guilty?
Innocent?
Both?
Neither?
Perhaps the question itself is wrong.
We have confused the language of conviction with the possession of conviction.
Now give the same model a third instruction:
Act as the judge. Ignore advocacy. Evaluate both arguments against the evidence and identify what can actually be established.
A new voice appears.
Balanced. Restrained. Judicial.
The AI did not undergo a moral conversion.
We moved the bumpers.
The More Dangerous Problem: AI May Not Know What It Does Not Know
This brings us to the second golden rule.
Language models hallucinate.
They can fabricate names, dates, citations or explanations while sounding perfectly confident. OpenAI explicitly warns that confidence is not reliability and that ChatGPT can produce plausible but false claims. Its research argues that one cause of hallucination is that conventional training and evaluation can reward guessing rather than admitting uncertainty. OpenAI Help Center
But there is a deeper issue.
A model does not possess a perfectly reliable boundary between what it knows and what it does not know.
Research suggests models can sometimes estimate their uncertainty surprisingly well. Anthropic has shown that language models can, under suitable conditions, predict whether answers are likely to be correct. OpenAI has similarly studied calibration—the relationship between a model’s expressed confidence and its actual accuracy. But neither result means the boundary is infallible, particularly on unfamiliar or out-of-distribution questions. Anthropic
This distinction matters enormously.
There are three different states:
I know.
I know that I do not know.
I do not know that I do not know.
The third is the dangerous one.
And Abraham Wald gives us one of the best examples in the history of statistics.
Abraham Wald and the Missing Bullet Holes
During the Second World War, Abraham Wald worked with the Statistical Research Group at Columbia University on improving aircraft survivability.
Military analysts had data from aircraft returning from combat.
Those aircraft contained bullet holes.
The obvious approach seemed to be: find where aircraft are being hit most frequently and reinforce those areas.
But there was a problem.
The dataset contained only aircraft that came home.
The aircraft that had been hit and destroyed were absent.
Wald’s actual work was mathematically more sophisticated than the simplified internet story about “put armour where there are no holes,” but its fundamental problem was precisely this: estimating aircraft vulnerability when the damage data for aircraft that failed to return were unobservable. Taylor & Francis Online
The visible bullet holes could therefore be misleading.
A part of the aircraft covered in holes might actually be relatively tolerant of damage.
A part with very few holes among returning aircraft might be extraordinarily vulnerable—because aircraft hit there did not survive long enough to enter the dataset.
The most important information was, in a sense, missing.
That is the perfect warning for AI.
You can give a model an immaculate spreadsheet.
You can ask an intelligent question.
The model can analyse everything beautifully.
And the analysis can still be wrong because the decisive evidence is not in the data at all.
AI cannot reason over evidence that neither it nor its tools possess unless it first recognises that something may be missing.
And sometimes it will not.
That gives us the second golden rule.
Golden Rule Two: Verify the Invisible
Do not merely ask:
“What does the evidence say?”
Also ask:
“What evidence might be missing?”
Who is absent from the sample?
What assumptions are hidden?
Which alternative explanations were never tested?
What information would reverse this conclusion?
What does the model not have access to?
What is it inferring rather than observing?
Ask explicitly:
“Identify the three most important pieces of missing information that could change your conclusion.”
Or:
“What would have to be true for this answer to be wrong?”
Or simply:
“What don’t you know here?”
That may be one of the most valuable prompts you can write.
Hallucinations Are Not Unique to Artificial Intelligence
We should nevertheless be careful not to treat hallucination as evidence that machines are uniquely unreliable.
Humans produce tremendous quantities of fluent nonsense too.
Not by the same mechanism.
But epistemically, the warning is familiar.
I can listen to a rap lyric and sometimes, for the life of me, fail to reconstruct a coherent literal proposition from it. That does not mean rap is inherently meaningless—some of it contains exceptional storytelling and social commentary. It simply reminds us that rhythm, confidence and linguistic effect are different things from propositional truth.
Political language provides an even cleaner case.
A politician can speak beautifully for five minutes and never answer the question.
A slogan can substitute for causation.
A carefully selected statistic can substitute for a complete dataset.
Emotion can substitute for evidence.
The sentences remain grammatical.
The audience applauds.
Nothing guarantees that the argument is balanced or true.
AI hallucination and human rhetoric are technically different phenomena.
But they teach the same rule:
Fluency is not truth.
Confidence is not evidence.
Eloquence is not reasoning.
Verification remains necessary whether the speaker is silicon or biological.
Prompt Engineering Without the Mystique
There are dozens of prompting frameworks. Most are useful checklists rather than deep theories.
Three are enough for most work.
RTF — Role, Task, Format
Act as an equity analyst. Assess the company’s earnings quality. Return five findings and five risks.
Use it when the task is simple.
CO-STAR — Context, Objective, Style, Tone, Audience, Response
It is useful when communication itself matters: speeches, articles, emails, marketing or executive communication.
RISEN — Role, Instructions, Steps, End Goal, Narrowing
Use it for complex analytical work where the process and boundaries matter.
But acronyms are secondary.
A strong prompt usually answers seven questions:
What do I want?
Why do I want it?
What does the AI need to know?
What evidence should it use?
What must it avoid?
What should the finished output look like?
How will I verify it?
That is most of prompt engineering.
Human Prompted, AI Prompted
Humans are prompted too.
A teacher prompts a student.
A manager prompts an employee.
A road sign prompts a driver.
Hunger prompts eating.
Fear prompts movement.
But humans bring embodied memories, relationships, emotions, biological drives and subjective histories to every instruction.
We should not simply assume an equivalent internal world exists inside a language model because its sentences resemble ours.
The more useful picture is simpler:
Training shapes the machine.
The prompt launches the ball.
Specificity routes it.
Context determines what it can see.
Missing information determines what it cannot see.
Hallucination reminds us that fluency is not truth.
Verification tests where the ball actually landed.
And that leaves two golden rules.
1. Be specific.
Reduce ambiguity before generation.
2. Verify the invisible.
Look not only at what the AI tells you, but at what the AI—and perhaps you—may have failed to see.
The first gets better answers.
The second prevents beautiful answers from becoming beautifully wrong.
The value of human knowledge therefore does not disappear as AI improves.
It moves upstream.
Into defining the problem.
Choosing the evidence.
Recognising missing evidence.
Knowing what to ask.
Knowing when to doubt.
And deciding what any answer actually means.
AI can route through an astonishing landscape of possibilities. Human judgement still has to decide where the ball should go—and whether the map itself is missing something.
So the next time you prompt, think of yourself more as Ronnie O Sullivan at a Snooker Table, rather than Edward Norton speaking to Brad Pitt in Fight Club.
