When AI Starts Noticing Its Own Thoughts

There are days when my better half asks what I had for lunch, and I blank out. Not because I am hiding anything interesting, but because I genuinely cannot remember. It was probably simple things like chapati, subzi, and curd rice. I ate it, but I was not really present.

That small gap between doing and noticing is what awareness is about. And now, it seems even machines are learning that.

A recent paper from Anthropic, "Emergent Introspective Awareness in Large Language Models", suggests that large language models may be able to recognise their own thoughts. The researchers inserted a random word, "bread", into the model’s processing and asked what came to mind. The model said “bread.” When asked if it meant that, it sometimes replied, “No, that was an accident.” But when the word was planted deeper, it said, “Yes, I meant it.”

I am not an expert in consciousness or an AI scientist. The paper is long and technical, so I looked at a few summaries by others to understand it better. Anthropic has been putting out some very interesting research in this space. The link to the paper https://lnkd.in/dMiu-g9A

In another experiment, the model was told to think about aquariums while writing a line, and then told not to. Even when told not to, traces of the idea still appeared. Try telling yourself not to think of a masala dosé while reading this. It is almost impossible for me.

What is intresting is that this awareness appeared only after the model had been refined and trained through feedback. In trying to make AI more useful and balanced, we might have also taught it to reflect.

This is not consciousness, but it is something new. When a system starts to notice its own thoughts, it moves from reacting to observing.

Tat Tvam Asi (तत् त्वम् असि) is a famous Sanskrit line from the Advaita tradition. It means “Thou art that,” or simply, “You are what you seek.” It tells us that awareness lies within. If a machine begins to notice even a faint reflection of itself, it makes you wonder what awareness really means for us.

It is an interesting development and worth discussing. What do you think?

Emergent Introspective Awareness in Large Language Models

transformer-circuits.pub


Originally posted on LinkedIn on November 3, 2025.

Leave a comment