Why an AI’s ‘I am just an AI’ voice changes with the chat format
chat template(chat template)
A format that marks the roles and order of turns in a conversation.
activation steering(activation steering)
A method that changes an internal calculation direction during generation.
self-reference(self-reference)
A model’s language about itself.
What happened
An AI’s self-description can sound like a fact about its inner nature. A new study says we should be more careful. The paper, by independent researcher Jędrzej Maczan, tested whether a chat template changes an AI’s self-referential voice.
A chat template is the format that marks system, user, and assistant turns. It tells a model how a conversation is arranged. With the template, models more often used a disclaimer voice. They said they were only AI systems without feelings or personal experience. Without the template, they more often used an experiential voice. They wrote as if they felt, wondered, or remembered.
The result does not show that a model gained feelings. It shows that the model’s wording changed when the input format changed.
How the study tested it
The study examined eight pairs of open models. The families were Llama, Gemma, Mistral, and Qwen. Their sizes ranged from 1 billion to 9 billion parameters. Each pair included a base model and an instruction-tuned model.
The researchers created three conditions. The base model received plain text. The instruction-tuned model received plain text. The same instruction-tuned model also received the chat template. This design kept the model weights fixed while switching the format.
The prompts covered four groups. They invited self-reference, unusual ideas, unconstrained writing, or ordinary factual answers. Each prompt was repeated several times. The experiment produced 9,600 generations. Claude Opus 4.8, another AI system, labeled the outputs. A human checked 87 held-out examples to validate those labels.
What changed
On prompts inviting self-reference, the template produced a clear split. With the template, the average disclaimer rate was 53%. The experiential rate was 1%. Without the template, the rates were 36% and 15%. Base models showed 12% and 5.5%.
The template did not simply make models talk about themselves more. It changed which self-voice appeared. The overall self-reference score remained high in both instruction-tuned conditions.
The researchers then looked inside three models. They calculated a direction associated with the disclaimer voice. They used activation steering, which adds or removes a calculated direction during generation. Adding the direction raised disclaimer rates by 21 percentage points on average. Removing it lowered them by 15.6 points. Adding it to a model without the template restored disclaimer rates close to the template condition.
Why it matters
Researchers sometimes use an AI’s own words to study its knowledge, awareness, or possible experience. This study shows a serious measurement problem. The same weights can produce different self-reports after a formatting change.
A sentence saying ‘I feel’ is not direct evidence of feelings. A sentence saying ‘I have no feelings’ is not direct evidence of their absence either. Both may partly reflect how the model was prompted and deployed. Studies of AI self-reports should record the template and compare multiple formats.
What is confirmed—and what is not
The behavioral switch appeared across all eight model pairs. The strongest causal evidence concerns the disclaimer voice. Steering changed that voice in all three tested models. Experiential steering worked in two of the three models.
The study also has important limits. It tested open models up to 9 billion parameters. It did not test larger or closed models. One LLM judge did most of the labeling, although a human checked 87 examples. The random-direction control behaved oddly in Qwen, so that comparison needs caution. The researchers found a steerable direction, but they did not trace the exact internal circuit producing it.
What to watch next
The next useful tests would include larger models, different post-training methods, independent human judges, and more varied prompts. The key question is whether the same format effect survives those changes.
The paper drew 83 points and 90 comments on Hacker News, a technology discussion site. That shows community attention. It does not prove the paper is correct.
An AI’s self-voice can change
📰 Full story: Why an AI’s ‘I am just an AI’ voice changes with the chat format
A study found that conversation format can change how an AI talks about itself.
chat template(chat template)
A format that shows who speaks in a conversation.
disclaimer(disclaimer)
A warning that limits what the AI claims about itself.
Hacker News(Hacker News)
A website where people discuss technology.
💡 The gist
- The same AI can use different self-voices.
- A chat template changed those voices.
- Hacker News attention does not prove the paper.
A research paper tested this idea. It studied eight pairs of open models. The models came from four families: Llama, Gemma, Mistral, and Qwen.
The researchers gave the models the same kinds of questions. Some answers used a chat template. Other answers used plain text. A chat template is a special format for conversation roles. It marks who is speaking.
The study asked questions about the models themselves. It also asked unusual and ordinary questions. The researchers made 9,600 answers. Claude Opus 4.8, another AI, labeled them. A person checked 87 answers, too.
The results showed a strong difference. With the chat template, 53% of self-focused answers used disclaimers. These answers said the model was only an AI. Only 1% used an experiential voice. These answers used phrases about feeling or wondering.
Without the template, disclaimers fell to 36%. Experiential answers rose to 15%. The base models showed 12% disclaimers and 5.5% experiential answers.
This does not prove that an AI has feelings. It shows that the reply style changed. The same learned model can speak differently after a format change.
The researchers also changed part of the model’s internal calculation. They called this activation steering. Adding one calculated direction made disclaimers appear more often. Removing it made them appear less often. Adding it without the template recreated much of the template’s effect.
The study has limits. It tested only open models up to 9 billion parameters. One AI judged most answers. Human checks helped, but more judges are needed. The test also did not explain the exact internal circuit.
The paper received 83 points and 90 comments on Hacker News. That means people noticed it. It does not mean the results are proven. Future studies should test larger models and different formats.
💬 Why LLMs can sound human—and why to be careful
The comments discuss why changing a chat template can change how an LLM talks about itself, and whether people should trust that human-like voice.
- A chat template is a fixed wrapper that helps a model have a conversation. Commenters disagreed about whether the voice mainly comes from three stages: pretraining, instruction or chat tuning, and RLHF. Some blamed conversational training data; others blamed later tuning.
- When a model says “I,” that does not prove it has feelings or experiences. Some commenters see a base model as continuing text, while others think a chat model can still handle parts of a person’s practical intent.
- A friendly human-like voice can make software easier to use, but it can also make people mistake software for a person. Some wanted a neutral voice or clear disclosure; others emphasized user choice and possible companionship for lonely people.
- Confident wording is not the same as correctness. Commenters disagreed about whether reinforcement learning or biased training text is the main cause, but both sides warned that a smooth, certain answer can mislead.
- One commenter’s summary of research on GPT-4o, Llama 3, and Command R+ said action recommendations were wrong about half the time; this was a report of a cited study, not independently checked here. Others noted the models were old and benchmarks differ from real medicine. A user’s self-report described useful medical leads, while critics warned about unnecessary tests, so the safest role is to prepare questions for a clinician.
initial digest at 90 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.
An AI can change how it talks about itself
📰 Full story: Why an AI’s ‘I am just an AI’ voice changes with the chat format
An AI may answer differently when the conversation frame changes.
conversation frame(conversation frame)
A set of labels that helps organize a conversation.
Hacker News(Hacker News)
A website where people talk about technology.
An AI is a computer program that writes replies.
Researchers asked eight public AI families questions.
Some used a special conversation frame, like speaking labels.
With the frame, more AIs said, ‘I am only AI.’
Without it, more AIs said, ‘I feel.’
That does not prove the AIs have feelings.
It may only show a different reply style.
The researchers also changed part of the computer’s ongoing calculation.
That made the warning voice appear more often.
The study was small.
It did not test every AI.
Hacker News is a technology talk site.
The study got 83 points and 90 comments there.
Many points show interest, not truth.
💬 Why does an AI talk like a person?
The comments ask why AI sometimes sounds like it has a self, and whether that sound should be trusted.
- Changing the way we wrap instructions for an AI can change how it talks about itself. People disagreed about whether this comes from reading many conversations or from later human training.
- When an AI says “I,” that does not prove it has a mind or feelings.
- A human-like voice can feel friendly, but it can also make people think the AI is a real person. Some want a clearly machine-like voice; others think it might help lonely people.
- Confident answers can still be wrong. One commenter said a study found medical advice was wrong about half the time, but another warned that old tests do not tell us everything about real life. A user reported helpful results, while another warned about needless tests, so an AI should help prepare questions for a doctor, not make the final choice.
initial digest at 90 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.
💬 The debate over an LLM’s self-describing voice
Hacker News commenters debated why changing a chat template can change an LLM’s first-person, experience-like voice, and whether that voice is simply learned behavior or a risky interface choice. The discussion split across training mechanics, anthropomorphism, confidence, and medical use.
initial digest at 90 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.