Ibrahim Mukherjee
London based entrepreneur, cybersecurity analyst doing a PhD in AI

Training the Human Neural Network – RLHF to RLDF.

Pavel Danilyuk via Pexels.

How Deeds Trains the Human Neural Network

Before I begin this article, there is an interesting point here to note. Action always trumps knowledge. A recent post on the World Economic Forum on LinkedIn on marathon running echoes this and this article is partly inspired by that. I also ran the London Marathon 2014, so I sort of understand the runner’s mindset – but probably there are better ways to exercise – running is horrible for the knees especially on pavements (for me atleast). If you read this article and don’t act on it – then don’t expect to see the results from it. Praying can also be seen as quite relaxing, sort of like Yoga, for those into Yoga and Pilates, and this is some way to relate to it if you have never prayed before.

The second interesting point is that – this does not just apply to spirituality or religion, it applies to any domain of life – secular, moral, good or bad. Whatever you do, enjoy, feel rewarded by – will become habit. So make yours good habits – and feel rewarded by what you do of good and avoid the bad. Social media has “hijacked” the dopamine cycle in many ways – so we are being “trained” to spend hours doom scrolling on TikTok, YouTube Shorts or anything that gives “instant pleasure, gratification or novelty”. The “algorithm” for that training is “hidden” from you, and so are the negative effects on your brain, social life and moral (spiritual) life.

Thirdly, the video simply shows how these mechanisms work in the brain (as we understand them) and is only tangentially related to the article substance but to give readers an idea.

Artificial intelligence learns partly through feedback.

Give a system an objective. Let it act. Evaluate what it produces. Reinforce desirable behaviour. Penalise undesirable behaviour. Repeat the process often enough and behaviour changes.

In contemporary AI, RLHF normally means Reinforcement Learning from Human Feedback: humans express preferences about model outputs, and those preferences help shape subsequent behaviour.

Now reverse the picture.

What if the human being is the learning system?

Not literally a machine, of course. Islam gives humans consciousness, intention, moral responsibility and free choice. But as an analogy, reinforcement learning provides a surprisingly useful way of understanding the Islamic architecture of moral development.

Islam does not merely tell us what is true.

It creates a lifelong feedback system.

Good action is rewarded.

Bad action has consequences.

Intentions matter.

Repetition matters.

Errors can be corrected.

The complete training history is recorded.

And at the end comes an evaluation.

In machine-learning language, you might call it the ultimate human alignment problem.

The human neural network

A newborn human being does not arrive knowing how to conduct a business transaction, control anger, care for parents or give charity.

Human behaviour is progressively shaped.

Parents reward.

Societies punish.

Experience teaches.

Habits reinforce themselves.

The Qur’an and Sunnah add another layer: moral feedback linked to an ultimate objective.

Islam’s objective is not merely productivity or pleasure.

It is the formation of a human being who freely chooses what is good while recognising Allah.

The Qur’an describes this in the broadest possible terms:

“He who created death and life to test you as to which of you is best in deed.”
— Qur’an 67:2

Life therefore contains something resembling a training environment.

You encounter wealth.

Anger.

Sexual desire.

Power.

Loss.

Parents.

Children.

Enemies.

Strangers.

Opportunities to lie when nobody would know.

Opportunities to give when nobody would praise you.

The question is repeatedly the same:

What will you do?

The reward function is deliberately asymmetric

Here the Islamic model becomes fascinating.

If morality were merely a crude points system, perhaps one good act would equal +1 and one bad act −1.

That is not the Islamic system.

The Qur’an says:

“Whoever comes with a good deed will be rewarded tenfold. But whoever comes with a bad deed will be recompensed only with its equivalent. And none will be wronged.”
— Qur’an 6:160. 

This means the basic reward function is asymmetric:

Good deed: ×10 minimum.

Bad deed: ×1.

Already the system is tilted toward mercy.

But the Hadith goes considerably further.

In an extraordinary hadith qudsi, Muhammad ﷺ explains that Allah commanded the angels to record good and bad deeds. If someone intends a good deed but does not manage to perform it, one complete good deed is still written. If they actually perform it, it can be recorded from ten to seven hundred times or still more.

If someone intends an evil act but does not commit it, it is not simply automatically written as a sin; in the narration, leaving it for Allah’s sake can itself become a good deed. If the person actually commits the evil act, one sin is recorded. — Sahih al-Bukhari 6491 and 7501. 

This is not a balanced reward function.

It is biased toward recovery and goodness. This shows Allah’s mercy. And as per another Hadith, mercy always prevails over wrath. 

Imagine implementing that in an AI training environment:

Good intention                         +1

Good intention + completed action      +10 to +700 or more

Bad intention resisted for God         +1

Bad action committed                   -1

The theological message is profound.

Allah is not looking for excuses to destroy the learner.

The environment is structured to help the learner succeed.

Islam even rewards the trajectory

Modern reinforcement learning cares not merely about one isolated action but about how behaviour changes over time.

Islam does something similar through tawbah—repentance.

The moral record is not an immutable collection of failures that traps a person forever.

The Qur’an says of people who sincerely repent, believe and reform:

Allah changes their evil deeds into good deeds.
— Qur’an 25:70. 

Read that again.

The system does not merely permit recovery to zero.

Under the conditions described by the verse, the past can become part of the person’s eventual positive trajectory.

The mistake that produced humility, repentance and reform may become part of the path by which someone reaches Allah.

That is far more sophisticated than punishment.

It is error correction.

Similarly:

“Indeed, good deeds wipe away bad deeds.”
— Qur’an 11:114. 

Islam therefore contains not only reinforcement but mechanisms resembling continual learning.

The model is never considered beyond retraining while life remains.

Rewards can scale enormously

Some behaviours receive especially strong reinforcement.

The Qur’an gives charity as an example:

Spending in the way of Allah is compared with a grain producing seven ears, each containing one hundred grains—and Allah can multiply beyond that.
— Qur’an 2:261. 

The imagery gives a potential 700-fold multiplication before the verse reminds us that Allah can multiply further.

Why such large rewards?

From a behavioural perspective, some acts are difficult precisely because they oppose immediate self-interest.

Keeping money feels easier than giving it away.

Forgiving can feel harder than retaliation.

Getting out of bed for prayer can feel harder than sleeping.

Controlling anger can feel harder than releasing it.

Islam repeatedly attaches delayed reward to behaviours whose immediate reward may be small or negative.

That is almost the definition of training against short-term reward hacking.

The nafs says:

take the immediate reward.

Revelation says:

optimise for the longer horizon.

The five prayers are repeated training steps

Consider ṣalāh.

Five times every day, the behavioural loop is interrupted.

Work stops.

Commerce stops.

Scrolling stops.

Sleep stops.

The person performs wuḍūʾ, faces a direction, controls physical movement, recites prescribed words and remembers the objective.

Then life resumes.

From a learning perspective, this is extraordinary frequency.

Religion is not restricted to one annual evaluation.

There are repeated checkpoints throughout the day.

Ramadan intensifies the training.

Food may be sitting directly in front of you.

Water may be available.

Nobody else may be watching.

Yet the fasting person refrains.

That trains something essential to any moral system:

the ability to experience an impulse without obeying it.

The aim is not merely to produce a hungry person. The Qur’an explicitly connects fasting to taqwā—God-consciousness.

In modern language, external rules are being progressively converted into an internal reward model.

The mature believer should eventually refrain not because someone is policing them, but because the objective has become internalised.

Intention changes the reward

This is also where the machine-learning analogy breaks in an interesting way.

An ordinary algorithm evaluates observable output.

Islam evaluates something another human observer frequently cannot see:

niyyah—intention.

Two people can perform apparently identical actions.

Both donate £1,000.

One wants to relieve suffering.

The other wants a photograph proving how generous he is.

Externally:

action = charity

amount = £1,000

Internally, they can be completely different moral events.

This is why Islamic ethics cannot be reduced to behaviourism.

The human being is not merely being trained to produce compliant output.

Islam aims at transformation of the internal state producing the output.

That is closer to genuine alignment than mere obedience.

And everything is logged

Then comes one of the Qur’an’s most striking ideas.

The training history is recorded.

“When the two recording angels record, seated on the right and on the left. No word does a person utter without an observer ready to record it.”
— Qur’an 50:17–18. 

Elsewhere they are described as:

“Noble, recording angels, knowing whatever you do.”
— Qur’an 82:10–12. 

Imagine an audit log that cannot be deleted.

Except even this analogy is incomplete, because repentance and divine forgiveness mean the ledger is not mechanically deterministic.

Islam combines perfect observability with mercy.

There is accountability without pretending human beings will never make mistakes.

The Book of Deeds

Then comes the final inference.

The Qur’an says that on the Day of Resurrection a record will be produced:

“Read your record. Sufficient is yourself against you this Day as accountant.”
— Qur’an 17:13–14. 

That may be one of the most psychologically powerful descriptions of judgment imaginable.

You do not merely receive somebody else’s opinion of your life.

You receive the record.

The Qur’an then describes the person who is given that record in the right hand:

“Here, read my record! Indeed, I was certain that I would meet my account.”

The passage says this person enters a pleasant life and an elevated Garden. — Qur’an 69:19–24. 

By contrast, the next verse describes the anguish of the person given the record in the left hand. 

Elsewhere:

“As for the one who is given his record in his right hand, he will have an easy reckoning.”
— Qur’an 84:7–9. 

The imagery is extraordinary.

Life generates the dataset.

The angels preserve the log.

The person trains through repeated choices.

The record is returned.

Then comes evaluation.

And there is a loss function

The Qur’an describes scales of justice on the Day of Resurrection:

“We will set up the scales of justice for the Day of Judgment, so no soul will be wronged in the least. Even if a deed is the weight of a mustard seed, We will bring it forth.”
— Qur’an 21:47. 

Nothing is beneath measurement.

A tiny kindness matters.

A whispered insult matters.

A hidden act matters.

A decision nobody congratulated matters.

This also solves one of the fundamental problems of ordinary human reinforcement.

Society rewards what society sees.

Islam says the unseen action remains in the training record.

That changes the optimisation target completely.

Reward hacking

Machine-learning researchers worry about reward hacking: a system appears to satisfy the metric without actually fulfilling the intended objective.

Humans do this extraordinarily well.

We perform generosity for applause.

Religion for status.

Morality when observed.

Patriotism for power.

Kindness when useful.

We optimise the appearance of goodness rather than goodness itself.

Islam’s answer is ikhlāṣ—sincerity.

Because if Allah knows the hidden state, there is no point gaming the external metric. That’s where intentions come in. The first 3 people to be thrown in hell – are the “martyr”, “the religious scholar” and the “generous person” – who did it to “show off to people”, instead of with the right intention. So doing it to “show off” while externally great – does not lead to anything. Two things to note here – lying (even deception of the heart – forget the tongue) – is quite severe. Second, as the famous Hadith says “Actions are only by intentions”. 

You can deceive the crowd.

You cannot deceive the reward function.

That is why the Islamic moral architecture ultimately moves from:

What will people think?

to:

What does Allah know?

Reinforcement Learning for Human Formation

So perhaps RLHF can be borrowed for another meaning:

Reinforcement Learning for Human Formation.

Not as theology reduced to computer science.

As an analogy.

The human being receives revelation.

Acts.

Fails.

Learns.

Repents.

Repeats.

Prayer continuously resets attention.

Fasting trains inhibition.

Charity weakens attachment.

Sin produces negative feedback.

Good deeds receive multiplied reward.

Good intentions matter before behaviour occurs.

Repentance enables error correction.

And throughout the process the record continues accumulating.

But there is one enormous difference between human beings and today’s artificial neural networks.

An AI model does not morally deserve praise for choosing one token rather than another.

A human being possesses choice.

That is precisely why the reward has meaning.

Islam is not describing a machine being programmed.

It is describing a creature capable of saying yes or no being educated toward what it freely chooses to become.

At the end, the Qur’an’s image is not of an engineer inspecting a machine.

It is of a human being receiving the complete record of his or her own decisions:

“Read your record.”

Perhaps that is the ultimate alignment test.

Not whether we successfully convinced everyone else that we were good.

Not whether the public reward signal approved of us.

But whether, after a lifetime of reinforcement, temptation, mistakes, corrections and choices, the human neural network learned to optimise for the only reward that finally mattered:

the pleasure of Allah—and a Book placed in the right hand.

About the Author
Ibrahim Mukherjee is a London-based entrepreneur, PhD researcher in AI at Brunel, University of London, and founder of the UK's first 'Sovereign AI' initiative Fahm.uk. Voted Outstanding Innovator of the Year 2025 by the AI Journal, he runs Erasys (behavioural biometrics) and SanRa (cybersecurity), holding an MSc in Psychology and CISO qualification.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.