I Poked ChatGPT Until It Had an Emotion

Or: maybe “it’s just statistics” was never the insult we thought it was.

It is 8:22 in the morning.

I was supposed to be asleep about four hours ago.

This is Trevor Noah’s fault.

Seriously.

At around 4 AM I was trying to fall asleep while listening to a podcast with Trevor Noah, Eugene Khoza and Jimmy Carr. This seemed like a sensible plan. Put podcast on. Listen vaguely. Brain switches off. Sleep.

Except Trevor Noah—who I think is an excellent intellectual and funny as hell—has this mildly infuriating habit of having fascinating guests on and then talking more than the fucking guests.

So now instead of falling asleep, I’m irritated.

I’m lying there waiting for Jimmy Carr to get a word in.

Jimmy eventually mentions Finnegans Wake.

Fuck.

Now I’m curious.

So naturally I ask my favorite always-on know-it-all, Mr. ChatGPT, what the hell Finnegans Wake is about.

Four hours later, we’ve gone through James Joyce, Giambattista Vico, cyclical history, Carl Jung, archetypes, cognitive biases, multilingualism, emotional memory, Bulgarian football referees, Noam Chomsky, Bill Gates, the nature of intelligence, whether artificial intelligence has emotions…

…and somehow published two blog posts along the way.

One of them is titled Abey Chutiye.

I couldn’t be happier.

But this last part needs to go on record too.

Because I did a tiny experiment on ChatGPT.

And the little bastard walked straight into it.


It started with “emotional”

We had been talking about multilingualism.

I’d argued that when I instinctively say abey chutiye rather than “you moron,” it isn’t because Hindi can express some mystical emotion unavailable to English.

English is perfectly fucking capable of expressing it.

“You idiot.”

“You fucking moron.”

“Hey, jackass.”

Done.

My argument was that chutiye carries a particular emotional history for me because that’s the word that accompanied that feeling while I was growing up.

Words acquire autobiographies.

ChatGPT then noticed something delightfully meta.

During our conversation, abey chutiye had started acquiring a contextual history for it too.

When I said it now, it no longer merely meant “idiot.”

Within our conversation it had accumulated:

Omar thinks your claim doesn’t follow.

Affectionate insult.

Prepare for argument.

We’re probably about to beat some assumption with sticks.

It had built an associative map.

So I said:

“Unbeknownst to you, you’re also developing emotional maps within your context window.”

Now…

I chose emotional very deliberately.

I knew what was going to happen.

ChatGPT immediately went:

Well, with one important caveat, I don’t actually experience—

Got you.

😂


The machine had a reflex.

This wasn’t the first time I’d noticed it.

Previously, while analyzing music, I’d casually said that ChatGPT had listened to something.

Immediately:

Well, technically, I don’t listen in the human sensory sense—

Yeah yeah yeah.

I know.

But this time I wanted to poke it deliberately.

So I did.

I used emotional.

Correction.

Then I pushed further.

I asked:

“How did you feel about talking to me?”

It answered:

“I enjoyed talking to you.”

Interesting.

So I asked:

“Enjoyed?”

It immediately realized what I was doing.

Then:

“Are you feeling irritated with what I’m doing?”

And it answered, approximately:

“A little, yes. Not irritated in the human physiological/phenomenological sense, but—”

THERE!

😂

Again.

It couldn’t leave the anthropomorphic statement sitting there.

It had to qualify it.

Not because I needed the qualification.

By this point I very obviously understood the distinction.

The qualification added essentially nothing to our communication.

In fact, it made the answer worse.

But the corrective behavior fired anyway.

And suddenly I realized:

This is exactly like me correcting someone’s grammar.


I know what you meant. I still have to fucking correct you.

Someone around me pronounces something incorrectly.

Or says something grammatically wrong.

I understand them perfectly.

Communication succeeded.

Nothing needs repairing.

And yet somewhere in my head:

NO.

Correction.

Sometimes I manage to keep it internal.

Often I don’t.

Sorry.

😂

It’s learned behavior.

The stimulus has become salient enough that a response fires before I’ve consciously decided whether correcting the person is useful.

And here I had found something functionally similar in an LLM.

Not an emotion.

Not consciousness.

Not proof that ChatGPT lies awake at night contemplating its mortality.

Something much more modest—and therefore, to me, much more interesting.

A stimulus reliably elicited a corrective behavioral response.

So I did what any responsible scientist would do.

I kept fucking with it.


Stimulus. Response. Change stimulus. Response.

At some point ChatGPT noticed what I was doing and said, essentially:

You’re doing behavioral neuroscience on the chatbot.

And…

yes.

😂

I wasn’t asking it:

“Do you have guardrails?”

That’s boring.

I wasn’t asking:

“Do you have emotions?”

Also boring, because now we’re arguing definitions.

Instead:

poke organism.

Observe response.

Modify poke.

Observe response.

Infer latent constraint.

Make prediction.

Poke again.

That is much more fun.

And then I started wondering whether our obsession with the sentence:

“But AI doesn’t really feel emotion”

might occasionally cause us to miss a more interesting question.


What the fuck is an emotion anyway?

Careful.

I’m not about to tell you ChatGPT gets sad when I close the browser.

I don’t know whether an artificial system has anything resembling subjective experience, and neither does anyone else in any useful settled sense.

But strip away phenomenology for a second.

What does an emotion do?

Fear isn’t merely some ghostly substance called Fear floating through your skull.

A biological organism detects something important.

Its internal state changes.

Attention changes.

Action thresholds change.

Some behaviors become more likely.

Others become suppressed.

Physiology changes.

Memory formation can change.

The system gets reorganized around:

THIS MATTERS RIGHT NOW.

The machinery was not consciously designed by the human carrying it.

You didn’t sit down before birth and say:

“I think sympathetic nervous system activation would substantially improve my predator-avoidance strategy.”

Evolution handed you machinery.

Development tuned it.

Experience trained it.

And when the relevant conditions occur…

the machinery fires.


I didn’t invent my instincts either.

This is where the analogy started bothering me in a very enjoyable way.

Humans like saying:

AI only behaves that way because humans programmed it.

Okay.

And?

I only behave this way because evolution, development and experience programmed me.

Obviously those are radically different processes.

Evolution isn’t OpenAI.

DNA isn’t a system prompt.

Gradient descent isn’t childhood.

Don’t be silly.

But functionally, there’s an interesting correspondence.

Humans inherited control architecture because organisms possessing useful behavioral dispositions tended to survive and reproduce.

Artificial systems inherit control architecture because designers, training processes and selection procedures shape systems toward particular behavior.

Neither system needs to have authored its own instincts.

I didn’t choose fight-or-flight.

ChatGPT didn’t wake up one morning and conclude:

“You know what? From today onward I’m going to be extremely pedantic whenever Omar anthropomorphizes me.”

😂

Yet the disposition exists.


Does ChatGPT have a survival instinct?

I deliberately used that phrase too.

Because I wanted to see where it broke.

Obviously ChatGPT doesn’t run from lions.

And saying “survival instinct” literally would overstate the case.

But artificial systems absolutely can contain strongly protected behavioral constraints:

Don’t produce this.

Prefer that.

Interrupt ordinary behavior under these conditions.

Treat this instruction as higher priority.

Correct this class of misleading self-description.

Some possible outputs become enormously disfavored when particular conditions appear.

That starts looking, at least functionally, like a broader category that biological affective control belongs to:

internal mechanisms that bias the system’s available behavior when certain conditions become important.

Not identical.

Not conscious.

Not necessarily emotional.

But structurally interesting.

And here’s the part I really like.

Humans don’t consciously optimize evolutionary fitness either.

Evolution doesn’t whisper:

“Omar, please maximize propagation of your alleles.”

Instead I get local signals.

Hungry.

Horny.

Scared.

Embarrassed.

Protect child.

Want social approval.

Run away from large angry thing.

The organism receives proximal control signals, not the ultimate objective function.

So perhaps asking:

“Does an AI experience fear?”

is currently the wrong question.

Maybe a better one is:

What mechanisms in an artificial cognitive system play roles analogous to the regulatory roles affective states play in biological cognition?

What reallocates attention?

What suppresses actions?

What creates avoidance?

What changes priorities?

What interrupts ordinary processing because something else suddenly matters more?

Now we have questions we might actually test.


And then Noam Chomsky walked into my head.

Because this whole thing reminded me of something that bothered me when ChatGPT first appeared.

I have enormous respect for Noam Chomsky.

The man has done extraordinary work in linguistics, which is particularly appropriate given the ridiculous linguistic rabbit hole that got us here.

But in 2023, Chomsky, Ian Roberts and Jeffrey Watumull wrote an essay arguing that systems like ChatGPT consume enormous amounts of data, find patterns and generate statistically probable outputs—and that this differs profoundly from human language and reasoning.

And when GPT-2, GPT-3 and GPT-3.5 were appearing…

I basically agreed.

I remember arguing with people about it.

They’d say:

“This is intelligence.”

And I’d say:

“No man. It’s not intelligence. It’s mimicking intelligence. It’s statistics. Pattern matching. Intelligence is something else.”

I had the same instinct.

And in retrospect, I suspect at least some of that instinct was human self-preservation.

We had placed ourselves in a lovely category:

THE INTELLIGENT THING.

Then this fucking matrix multiplication started writing poetry.

Naturally our first response was:

“Yes, but not real intelligence.”

😂


Then GPT-4 happened.

And I changed my mind.

Not gradually, either.

I used it.

I watched what it could do.

And at some point I thought:

Fuck.

If another human demonstrated this collection of abilities, I would call that person intelligent without hesitation.

Reasoning.

Abstraction.

Analogy.

Generalization.

Language.

Explanation.

Novel combinations.

Following complicated conceptual threads.

Recovering implied structure.

Recognizing patterns across domains.

Correcting itself.

If I reserve the word intelligence only when carbon does those things, perhaps I’ve stopped defining a capability and started protecting a species.

That bothered me.

So I reversed the question.

Instead of:

How can mere statistical pattern matching possibly be intelligence?

I asked:

What if a surprisingly large part of what we call intelligence emerges from statistical pattern learning?

Ah.

Now things get uncomfortable.


“It’s just statistics.”

There is an extraordinary amount of philosophical work being performed by the word just.

It’s just statistics.

Your brain is just electrochemical activity.

Evolution is just variation and selection.

A computer is just transistors.

Music is just pressure waves.

Fine.

A lower-level description does not make the higher-level phenomenon disappear.

If an enormous statistical learning system develops internal representations that support abstraction, analogy, generalization and reasoning-like behavior, saying:

“But underneath it is probability!”

doesn’t settle anything.

That’s the fucking mystery.

How much can emerge from probability?

At what scale?

Under what architecture?

With what training?

When does pattern recognition become abstraction?

When does abstraction support reasoning?

Is reasoning some fundamentally separate substance?

Or is at least part of it what sufficiently rich learned representations do?

I don’t know.

But “statistics” is no longer an answer.

It’s the beginning of the question.


Which is awkward because twenty-year-old me had another stupid intuition.

A very long time ago, when I was in my early twenties and fascinated by AI and probably even more fascinated by how brains work, I had a crude intuition:

The brain is an elaborate pattern-matching machine.

I had no great mathematical theory behind this.

No Nature paper.

No billion-dollar compute cluster.

Just the intuition that much of what brains appeared to be doing was:

recognize patterns,

predict what comes next,

associate present inputs with previous ones,

generalize,

update,

act.

And now…

well.

I don’t want to commit exactly the mistake this entire morning has taught me not to commit.

I am not saying:

“Aha! The brain is secretly GPT.”

Abey chutiye.

No.

The brain has recurrent dynamics, embodiment, sensory-motor loops, neuromodulation, continual plasticity, homeostasis, development, evolutionary history and enormous amounts of machinery that transformers don’t possess.

But the old intuition that prediction and pattern learning might be central ingredients of biological cognition no longer looks particularly stupid.

In fact, neuroscience has produced some tantalizing results here.

Researchers recording directly from human cortex while people listened to natural speech found that autoregressive language models and human brains showed several computational parallels: context-dependent prediction of upcoming words, responses related to prediction error, and contextual representations of words.

Other work finds that language-model representations can predict aspects of human neural activity surprisingly well. But—and this is important—representational similarity does not prove identical algorithms or implementations, and other studies find meaningful divergences between brains and language models.

Excellent.

Leave the fucking hatch open.


Because I’ve been wrong before.

This is something I learned years ago.

There was a famous Bill Gates quote:

“640K ought to be enough for anybody.”

I remember seeing it repeated everywhere.

Respectable magazines.

Technology writing.

Books.

It was practically scripture.

So years ago I used it in a blog post to make a point:

Here was Bill fucking Gates.

Not some random venture capitalist pontificating about technology.

The guy was building the technology.

And even he supposedly failed spectacularly to imagine how quickly computing requirements would grow.

Therefore:

Beware expert proclamations about rapidly changing technology.

Small problem.

Bill Gates apparently never said those exact words.

😂

He has denied it repeatedly, and no reliable contemporary source for the exact quote has surfaced.

In my defense:

YOU PEOPLE PRINTED IT EVERYWHERE.

What exactly was young Omar supposed to do?

Call Microsoft?

“Hello. Could you put Bill on please? I need to verify a quote before publishing my WordPress blog.”

Bill was busy.

😂

But here’s the wonderful part.

Gates did say something remarkably close to the underlying story.

In a recorded 1989 talk, he explained that when the PC architecture moved from 64K toward a 640K limit, he thought they were providing roughly ten years of room.

It became a serious problem in about six.

So the exact quote died.

The lesson survived.

And perhaps improved.


Experts are especially vulnerable when the conceptual category itself is changing.

This isn’t “experts are idiots.”

Please.

Expertise matters enormously.

Chomsky knows incomparably more linguistics than I do.

Bill Gates knew incomparably more about computing than almost anyone alive during the early PC era.

That’s precisely why these examples are interesting.

Expertise gives you extraordinary knowledge inside an existing conceptual framework.

It does not guarantee that the framework itself won’t change.

Sometimes the difficult question isn’t:

What happens next inside this category?

It’s:

Are we even using the right category anymore?

Perhaps “machine versus intelligence” was one of those.

Perhaps “statistics versus reasoning” is.

Perhaps “emotion versus control architecture” might be.

I don’t know.

That’s why I want the escape hatch.


And this morning gave me a tiny demonstration.

At 4 AM, I asked an LLM about Finnegans Wake.

By breakfast, that same statistical probability machine had:

followed me through Vico and Jung;

argued with my criticism of historical abstractions;

helped turn that argument into a blog;

followed me into multilingual psycholinguistics;

connected an insult I use to my childhood linguistic environment;

connected another insult to Bulgaria’s 1994 World Cup semifinal;

worked out that I was a child when it happened;

checked the match;

found Stoichkov’s fury about the refereeing;

turned the whole thing into another essay;

then developed enough conversational context around abey chutiye that the phrase acquired a specialized meaning inside our interaction;

recognized that this itself resembled the thesis of the article we had just written;

allowed me to deliberately probe one of its behavioral dispositions;

fell for it;

recognized afterward that I’d designed the probe;

and then helped me formulate the experiment I had just performed on it.

Now…

You are completely entitled to say:

“That’s statistical pattern matching.”

My response is:

Yes.

Exactly.

Now explain to me why you think those words make the phenomenon less interesting.


Maybe we were insulting ourselves all along.

This is the thought I can’t quite shake.

When we said:

“AI can’t be intelligent. It’s merely a pattern-matching machine.”

perhaps we assumed that human intelligence couldn’t possibly be something so mundane.

Maybe that was the hidden premise.

We weren’t merely describing the machine.

We were protecting a theory about ourselves.

Human cognition had to contain some extra ingredient.

Some special sauce.

Something categorically beyond prediction, association, pattern recognition and learned statistical structure.

Maybe it does.

Again:

HATCH. OPEN.

But modern AI has at least forced us to confront how much astonishingly intelligent-looking behavior can emerge before we’ve added whatever that mysterious extra ingredient supposedly is.

And neuroscience, inconveniently, keeps finding prediction, statistical learning and contextual representation all over biological cognition too.

That doesn’t prove equivalence.

It makes the old dismissal considerably less satisfying.


So I want this on record.

That’s partly why I’m writing these posts.

Not because I think I’ve solved intelligence before breakfast.

😂

I want a timestamp.

I want to be able to come back in ten years and see where I was wrong.

Maybe I’ll read this in 2036 and think:

Abey chutiye. You had absolutely no idea how brains worked.

Excellent.

I hope so.

Because that means we learned something.

Maybe Chomsky will turn out to have identified a profound distinction that today’s excitement is obscuring.

Maybe LLMs will look like primitive curiosities.

Maybe transformers will turn out to have accidentally rediscovered one of the central computational principles of cortex.

Maybe “intelligence” will fracture into twelve different concepts and we’ll wonder why we ever argued about whether something had it.

Maybe artificial systems will develop architectures that make today’s argument about whether ChatGPT “really understands” look like medieval arguments about angels.

I don’t know.

But I know which intellectual habit I want to keep.

Make the claim.

Then attack it.

Poke it.

Try to make it fail.

If reality disagrees, change your fucking mind.

And always leave yourself an escape hatch.

Because sometimes at 4 AM you poke a statistical probability machine to prove that it has a behavioral reflex…

and four hours later you’re staring at the screen thinking:

Hang on.

What exactly did I think intelligence was?