🔥 Trending on HN

Why Qwen 3.8 Shifted After a Tiny GPT-5.5 Pro Reasoning Start

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
reasoning prefill

An experiment that places a short reasoning start before another model answers.

source recall

A score showing how many matching parts appear early in two answers.

Qwen3.8 A95B

The target AI model that showed the largest change in this test.

What happened

The author, wsxiaoys, posted a v1.1 experiment in a GitHub Gist. It followed earlier work on reasoning prefills. The test studied Qwen3.8 A95B (the target AI model) and three other models. GPT-5.5 Pro (the teacher AI model) supplied the starting reasoning fragment.

The test used 45 problems. Fifteen were STEM problems. Fifteen were non-STEM problems. Fifteen were synthetic puzzles. Each target model answered every problem twice. One answer was normal. The other began with the first 1% of GPT-5.5 Pro's reasoning. The researcher inserted that start into the target model's reasoning channel. The final visible answer still came freely from the target model.

The researcher checked how much of GPT-5.5 Pro's visible answer appeared early. The comparison used the first 100 tokens. This produced a source recall score, averaging one-part, two-part, and three-part matches. Qwen's score rose from 16.79% to 34.97%. That was an 18.18-point increase. STEM problems produced the largest Qwen increase. Their score rose from 19.26% to 46.24%.

The experiment's background

A reasoning prefill gives a model a short starting fragment. The model then continues its own answer. This setup tests sensitivity to a starting direction. It does not directly reveal the model's full internal process.

The Gist reports an outside comparison. It does not announce a Qwen feature. The measurement concerns visible answer overlap. It does not measure whether the models had identical thoughts.

Why it matters

Qwen changed far more than the other tested models. DeepSeek V4 Flash moved down 1.17 points. Inkling moved up 0.46 points. Kimi K3 moved up 4.54 points. Qwen moved up 18.18 points.

The author says this pattern may mean Qwen learned from GPT-5.5 Pro. It could also involve a closely related GPT model. The earlier experiment used Opus 4.8. Qwen barely moved toward that model. This contrast creates a useful question about training clues. It does not answer that question.

What is confirmed

The experiment recorded a large Qwen change under one specific condition. The effect also appeared across STEM, non-STEM, and puzzle categories. Its category increases were 26.99, 12.80, and 14.75 points.

The Hacker News post received 235 points and 93 comments. Those numbers show community attention. They do not prove the experiment is correct. The source result and the community reaction must remain separate.

What remains unknown

The data do not prove that Qwen used GPT-5.5 Pro outputs for training. They do not identify the source of any training material. They also do not show that Qwen copied complete answers.

The sample had 45 problems. It used one teacher model and one prefill size. It measured only the first 100 tokens. More similar text can reflect the test's conditions. It does not automatically mean better reasoning or better answers.

What to watch next

The next step is replication. Other researchers can use more problems, model versions, and subject areas. They can vary the prefill length. They can compare GPT-5.5 Pro with other teacher models. If the same pattern repeats, the clue becomes stronger.

Until then, the careful conclusion is narrow: Qwen reacted unusually strongly to this short GPT reasoning start. That result raises a question about training sources. It does not settle the question.

💬 Qwen 3.8 and GPT-5.5 Pro reasoning prefixes: evidence and objections

The HN discussion treats the result as suggestive evidence, not proof that Qwen 3.8 copied GPT-5.5 Pro reasoning or gained its capabilities. The numbers, performance changes, and failure reports below come from the article or commenters and were not independently verified here.

  • The cited paper describes recovering readable reasoning from encrypted reasoning tokens, then taking roughly the first 1% of a frontier model’s trace and using it as the starting prefix for an open model before comparing answers.
  • The article reports that giving Qwen 3.8 a GPT-5.5 Pro reasoning prefix moved its score 20.58 points toward GPT-5.5 Pro, including a large effect on private synthetic puzzles. An earlier comparison reportedly found a similar shift when Kimi-K3 was prefixed with Claude 4.8 reasoning.
  • Supporters focus on the contrast between the unprefilled and prefilled cases: Qwen reportedly differs substantially from GPT on its own, but becomes much more similar after the GPT reasoning prefix is matched. Comparisons with DeepSeek V4 Flash and Kimi-K3 are used to argue that GPT reasoning and answers may have been part of Qwen’s training data.
  • The strongest objection is that, according to commenters, the available GPT-5.5 reasoning examples came from a paper published on August 10, while Qwen 3.8 0902 was trained afterward. Qwen could therefore have learned those particular public examples rather than extracted private reasoning traces.
  • The result does not establish how much such data was used, whether the effect is genuine capability transfer or mainly style imitation, or whether both models simply saw the same benchmark solutions. A small amount of post-training could raise correlations, so fresh traces, unseen tasks, and stronger controls would be needed.
  • Users also report inconsistent visible reasoning: one person described another model leaking reasoning into a tool call and stopping, then seeing similarities with Qwen 3.8 27B. Others report ordinary prose instead; prompts, system instructions, interfaces, quantization, and reasoning effort may all affect the output. These are anecdotes, not controlled evidence.
  • Some commenters see competitor-output distillation as unfair fast-following, while others call it a major engineering achievement because Qwen’s weights can be run locally even when GPT cannot. The thread remains divided about how different this is ethically from training on public web data.

mature digest at 93 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Qwen 3.8 Reacted Strongly to a Tiny Starting Clue

📰 Full story: Why Qwen 3.8 Shifted After a Tiny GPT-5.5 Pro Reasoning Start

A small experiment made Qwen3.8 A95B's early answers match GPT-5.5 Pro more often. It did not prove copying or shared training.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
reasoning prefill

A small starting piece given to an AI before it answers.

token

A small piece of text used by an AI.

Hacker News

A website where people discuss technology.

💡 The gist

  • Qwen3.8 A95B is the AI model tested here.
  • GPT-5.5 Pro is the AI model used as a guide.
  • Qwen matched it more after seeing a tiny reasoning start.

A developer posted the experiment in a GitHub Gist. It followed an earlier test. The new test used 45 questions.

Fifteen questions were STEM problems. Fifteen were non-STEM problems. Fifteen were synthetic puzzles. Each model answered every question twice. It answered normally once. Then it received the first one percent of GPT-5.5 Pro's reasoning.

This method is called a reasoning prefill. It gives an AI a small starting piece. The AI still creates the visible answer itself. The researcher compared the first 100 tokens. A token is a small piece of text.

Qwen's matching score rose from 16.79% to 34.97%. The increase was 18.18 percentage points. STEM problems showed the biggest rise. Their score went from 19.26% to 46.24%.

Other models changed much less. DeepSeek V4 Flash fell slightly. Inkling rose slightly. Kimi K3 rose more, but only 4.54 points. Qwen's change was much larger.

The author suggested a possible explanation. Qwen may have learned from GPT-5.5 Pro. It may have learned from a related GPT model. This is not proven. The test measured matching answers, not training records.

The story received attention on Hacker News. The post had 235 points and 93 comments. Those numbers show interest. They do not show that the experiment is correct.

Why does this matter? A model can react strongly to a tiny starting clue. That may help researchers study model behavior. It may also raise questions about training sources. Still, one test cannot settle those questions.

More tests should use new problems and other models. Independent researchers should repeat the method. Until then, the safest summary stays narrow. Qwen changed more than the other models under this test condition.

💬 What the experiment shows—and what it does not

Commenters are debating whether Qwen 3.8 learned from GPT-5.5 Pro reasoning. The reported numbers, behavior, and failures are article or user reports, not independently confirmed results.

  • Researchers made hidden reasoning readable, gave Qwen about the first 1% of a stronger model’s reasoning, and checked whether its answer changed.
  • The article reports a 20.58-point movement toward GPT-5.5 Pro, including an effect on private synthetic puzzles. A previous Claude 4.8 and Kimi-K3 comparison reportedly showed a similar pattern.
  • Supporters say Qwen looks different without the prefix but suddenly looks more like GPT after receiving it. Critics say the GPT examples were public by August 10 and Qwen 3.8 0902 was trained later, so it may have learned those examples directly.
  • The test does not show how much training data was used, whether Qwen gained ability or just style, or whether both models learned the same answers. Users also report that Qwen’s visible reasoning changes with the prompt, interface, quantization, and reasoning setting.
  • Some people call distillation unfair copying. Others say it is valuable engineering because an open-weight model can run locally.

mature digest at 93 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Qwen Saw a Tiny Start, and Its Answer Changed

📰 Full story: Why Qwen 3.8 Shifted After a Tiny GPT-5.5 Pro Reasoning Start

One AI saw another AI's thinking start. Their answers became more alike.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
Qwen3.8 A95B

An AI that writes answers.

GPT-5.5 Pro

Another AI that writes answers.

Hacker News

A website where people talk about technology.

Qwen3.8 A95B is an AI that writes answers. GPT-5.5 Pro is another answer-writing AI. A person showed Qwen a tiny start from GPT's thinking. Then Qwen answered questions by itself. Some answers looked more like GPT's answers.

That does not prove Qwen learned from GPT. It was only one small test.

Hacker News is a website about technology discussions. It also noticed this story. Attention does not make a test correct. People need more tests before deciding what happened.

💬 Did Qwen learn from GPT?

Researchers showed one AI the beginning of another AI’s hidden thinking and checked whether the answers became more alike. The numbers and glitches are reports from the article and users, not final proof.

  • The report says Qwen’s score moved 20.58 points closer to GPT-5.5 Pro. An earlier test with Claude 4.8 and Kimi-K3 reportedly showed a similar change.
  • But the GPT examples may have been public before Qwen 3.8 0902 was trained. Qwen may have learned those examples, so this does not prove that it took secret thoughts.
  • We still do not know how much was learned, or whether Qwen copied skill or just writing style. People also see different reasoning depending on the setup.
  • Some people think copying is unfair. Others think a model that people can run at home is a big achievement.

mature digest at 93 comments (revision 1). We fetched 93 comments and sampled 93 across the thread. These are HN users’ reports, not independently verified facts.

Sources