@yaojingang: In Zhang Xiaojun's latest episode, she interviewed a 17-year-old high school student, Su Tinghao, who is currently studying at an international school in Hong Kong. Recently, his independent research paper 'Attention Projection Mixing with Exogenous Anchors' was accepted by ICML…

X AI KOLs Timeline Papers

Summary

This article reports on the independent research paper 'Attention Projection Mixing with Exogenous Anchors' by 17-year-old high school student Su Tinghao, which was accepted into the ICML main conference, proposing the ExoFormer model to address the over-smoothing issue in Transformers.

In Zhang Xiaojun's latest episode, she interviewed a 17-year-old high school student, Su Tinghao. He is a current student at an international school in Hong Kong, and recently, his independent research paper 'Attention Projection Mixing with Exogenous Anchors' was formally accepted as a main conference paper at ICML after double-blind peer review. ICML, the International Conference on Machine Learning, is one of the most important academic conferences in the field of machine learning worldwide. Every year, researchers from top universities, tech companies, and AI labs submit their latest work here, which undergoes double-blind peer review to decide which papers are included in the main conference. What problem does this paper address? When a Transformer processes a sentence, it continuously passes word representations to deeper layers. The deeper the layers, the more complex the semantic understanding, but the original identity information of words may gradually fade, which is called over-smoothing. Existing methods repeatedly recall information from the first layer, allowing deeper layers to随时 'reference the original text.' However, there is a contradiction here: The first layer needs to generate stable information suitable for long-term reuse. The first layer also has to perform its own semantic computation. Su Tinghao calls this the First-Layer Tension, meaning the first layer bears two objectives and easily compromises on both. The paper proposes ExoFormer, which, in addition to normal Transformer layers, separately sets an 'exogenous anchor' module. 1. Extracts a stable representation from input word vectors; 2. Separately generates attention Q, K, V, and gating signal G; 3. Each layer can mix its current computation results with this anchor; 4. The mixing ratio is learned by the model and can also be dynamically adjusted based on current input; 5. The anchor is normalized using root mean square normalization before injection to avoid numerical scale conflicts between different layers. You can imagine it as writing a long article: - The exogenous anchor is responsible for saving 'who the characters are and what the raw materials are'; - The Transformer layers are responsible for analyzing relationships, extracting meanings, and forming conclusions; - Each layer can consult the raw materials, so deep computations don't need to carry all basic information long-term. The author explains this division of labor as the offloading hypothesis: the anchor saves word identity, while the main network focuses on processing high-level features. The strengths of this paper: 1. Discovered a very specific architectural contradiction 2. Extended cross-layer reuse to the full Attention pathway 3. The performance improvement is significant with low additional cost. On six downstream tasks, its average accuracy increased from 48.80% for gated attention to 50.27%, and validation perplexity decreased from 14.64 to 14.09 4. Provided a testable mechanistic explanation Some memorable stories from the interview: 1. The entire paper involved about 200 experiments, costing around 30,000 yuan. He also admitted that there was no pre-designed path for choosing Attention; it took multiple tests to find the direction of First-Layer Tension and exogenous anchors. 2. When deciding to scale up experiments, he calculated it would cost an additional 6,000 yuan. That night, he hesitated. Even if the paper was well-done, it might be rejected due to uncertainties in review, and this money could yield no results. With support from his parents and brother, he proceeded with the larger experiments. The paper has a single author, and the research completion was indispensable without his family's financial and emotional support. 3. About 98% of the work was completed by himself, with no mentor guidance. Few people around him could discuss models, so his most frequent interactions were with ChatGPT, DeepSeek, and online courses. 4. He frankly stated that without ChatGPT, it would have been difficult to complete this paper. For example, he would ask AI about how to submit, how to respond to reviewers, and how to schedule training time. 5. He said he hopes to become someone who can make himself happy and also make those he likes happy. Whether with or without AI, in his view, 'happiness is fundamental to humanity.' Related resources: 1. https://icml.cc/virtual/2026/poster/61175… 2. https://xiaoyuzhoufm.com/episode/6a8472b95aeb2a5712e8de78… 3. https://arxiv.org/html/2601.08131v4…
Original Article
View Cached Full Text

Cached at: 08/24/26, 01:53 PM

Zhang Xiaojun’s latest episode features an interview with a 17-year-old high school student, Su Tinghao.
He is a student at an international school in Hong Kong. Recently, his independently researched paper, Attention Projection Mixing with Exogenous Anchors, was officially accepted as a main conference paper at ICML after double-blind peer review.

ICML, the International Conference on Machine Learning, is one of the world’s most important academic conferences in the field of machine learning. Every year, researchers from top universities, tech companies, and AI labs submit their latest findings here. After double-blind peer review, selections are made for which papers enter the main conference.

What problem does this paper address?
When a Transformer processes a sentence, it continuously feeds word representations into deeper layers. As the layers deepen, semantic understanding becomes more complex, but the original identity information of words may gradually fade—a phenomenon known as over-smoothing.
Existing methods repeatedly invoke first-layer information, allowing deeper layers to “revisit the original text” at any time.
However, this creates a contradiction:

  • The first layer needs to generate stable information suitable for long-term reuse.
  • At the same time, it must also perform its own semantic calculations.

Su Tinghao calls this First-Layer Tension, where the first layer bears two objectives and easily compromises on both.
The paper proposes ExoFormer, which adds an “exogenous anchor” module separate from the standard Transformer layers:

  1. It extracts a stable representation from the input word embeddings.
  2. It separately generates the Q, K, V, and a gating signal G for attention.
  3. At each layer, the current computation can be mixed with this anchor.
  4. The mixing ratio is learned by the model and can also dynamically adjust based on the current input.
  5. The anchor undergoes root-mean-square normalization before injection to avoid numerical scale conflicts between layers.

Think of it as writing a long essay:

  • The exogenous anchor is responsible for preserving “who the characters are and what the original materials are.”
  • Each Transformer layer analyzes relationships, extracts meaning, and forms conclusions.
  • Every layer can refer to the original materials, so deeper computations don’t need to carry all foundational information long-term.
    The authors explain this division of labor as the offloading hypothesis: the anchor preserves word identity, while the main network focuses on processing high-level features.

Highlights of this paper:

  1. It identifies a very specific architectural contradiction.
  2. It extends cross-layer reuse to the complete Attention pathway.
  3. The performance improvement is significant with minimal extra cost. Across six downstream tasks, average accuracy improved from 48.80% (gated attention) to 50.27%, and validation perplexity dropped from 14.64 to 14.09.
  4. It provides a testable mechanistic explanation.

Memorable stories from the interview:

  1. The entire paper involved around 200 experiments, costing approximately 30,000 RMB. He admitted that choosing Attention as a focus wasn’t a pre-planned path; it took many tests before he discovered the direction of First-Layer Tension and exogenous anchors.
  2. When deciding to scale up experiments, he calculated it would cost an additional 6,000 RMB. He hesitated that night. Even if the paper were excellent, review uncertainties could lead to rejection—this money might yield nothing. With support from his parents and brother, he proceeded with the larger experiments. The paper has only one author, and its completion relied heavily on his family’s financial and emotional support.
  3. About 98% of the work was done by himself, without advisor guidance. Few people around him could discuss models with him, so his most frequent collaborators were ChatGPT, DeepSeek, and online courses.
  4. He frankly stated that without ChatGPT, completing the paper would have been very difficult. He consulted AI on how to submit, respond to reviewers, and plan training schedules.
  5. He said he hopes to become someone who can make himself and the people he cares about happy. In his view, whether or not AI exists, “happiness is fundamental to being human.”

Related Resources:

  1. https://icml.cc/virtual/2026/poster/61175…
  2. https://xiaoyuzhoufm.com/episode/6a8472b95aeb2a5712e8de78…
  3. https://arxiv.org/html/2601.08131v4…

Similar Articles

@Phoenixyin13: This is one of the most important reposts I've made. The first author of this paper is someone I deeply admire and a good friend of mine—Guowei Xu, a top student from the Yao Class at @Tsinghua_Uni, who is now conducting AI large model research at @Harvard. Guowei's paper precisely hits the current...

X AI KOLs Timeline

Reposting an introduction to a paper by Tsinghua Yao Class graduate Guowei Xu (currently at Harvard) that accurately points out two critical bottlenecks in LLM search: sparse verification and candidate limitation, which are important for improving reasoning capabilities.

@elliotchen100: Translate the work on MiroMind under Shanda. The next step of post-training might be scientific discovery itself. Simply put, it trains a model to propose research hypotheses across different disciplines. Physics, chemistry, and biology all use one method. The paper was accepted at ICML 2026, code open source...

X AI KOLs Timeline

This paper proposes a scalable supervised fine-tuning method for training language models to propose research hypotheses across disciplines. It has been accepted by ICML 2026 and the code is open source.

@Sxy_Cherotich: Recently I've been talking with quite a few model researchers, and a consensus conclusion is: the importance of data is once again highlighted. A while ago I got to know ex-Kimi's @FanqingMengAI, who is doing a startup in the data direction, and invited him to record a podcast. The biggest non-consensus from our conversation is Fanqing's view on the difference between domestic and foreign models...

X AI KOLs Timeline

A podcast about AI model competition, discussing the importance of data, distillation and pre-training innovation, and an interview with Evolvent AI co-founder Meng Fanqing, covering topics such as synthetic data, RSI, and differences in domestic models.

@tanzhengmc97: https://x.com/tanzhengmc97/status/2066531753762656730

X AI KOLs Timeline

Explained the operating principles of large models in easy-to-understand language, including word vectors, Transformer attention mechanism, next-word prediction training, and emergent abilities, suitable for beginners to understand basic AI concepts.

@AlchainHust: Spent most of a day listening to the 4-hour interview between Zhang Xiaojun and Yao Shunyu. This guy, who just moved from Anthropic to Google DeepMind last year, has worked on Claude 3.7/4.5 and Gemini 3. He offered many candid perspectives from a frontline researcher at a top large model lab. The interview is incredibly information-dense…

X AI KOLs Timeline

This article summarizes Zhang Xiaojun's interview with Yao Shunyu, a researcher involved in developing Claude and Gemini, who shares 10 insightful views on AI code generation, company culture, scaling laws, etc.