@yaojingang: In Zhang Xiaojun's latest episode, she interviewed a 17-year-old high school student, Su Tinghao, who is currently studying at an international school in Hong Kong. Recently, his independent research paper 'Attention Projection Mixing with Exogenous Anchors' was accepted by ICML…
Summary
This article reports on the independent research paper 'Attention Projection Mixing with Exogenous Anchors' by 17-year-old high school student Su Tinghao, which was accepted into the ICML main conference, proposing the ExoFormer model to address the over-smoothing issue in Transformers.
View Cached Full Text
Cached at: 08/24/26, 01:53 PM
Zhang Xiaojun’s latest episode features an interview with a 17-year-old high school student, Su Tinghao.
He is a student at an international school in Hong Kong. Recently, his independently researched paper, Attention Projection Mixing with Exogenous Anchors, was officially accepted as a main conference paper at ICML after double-blind peer review.
ICML, the International Conference on Machine Learning, is one of the world’s most important academic conferences in the field of machine learning. Every year, researchers from top universities, tech companies, and AI labs submit their latest findings here. After double-blind peer review, selections are made for which papers enter the main conference.
What problem does this paper address?
When a Transformer processes a sentence, it continuously feeds word representations into deeper layers. As the layers deepen, semantic understanding becomes more complex, but the original identity information of words may gradually fade—a phenomenon known as over-smoothing.
Existing methods repeatedly invoke first-layer information, allowing deeper layers to “revisit the original text” at any time.
However, this creates a contradiction:
- The first layer needs to generate stable information suitable for long-term reuse.
- At the same time, it must also perform its own semantic calculations.
Su Tinghao calls this First-Layer Tension, where the first layer bears two objectives and easily compromises on both.
The paper proposes ExoFormer, which adds an “exogenous anchor” module separate from the standard Transformer layers:
- It extracts a stable representation from the input word embeddings.
- It separately generates the Q, K, V, and a gating signal G for attention.
- At each layer, the current computation can be mixed with this anchor.
- The mixing ratio is learned by the model and can also dynamically adjust based on the current input.
- The anchor undergoes root-mean-square normalization before injection to avoid numerical scale conflicts between layers.
Think of it as writing a long essay:
- The exogenous anchor is responsible for preserving “who the characters are and what the original materials are.”
- Each Transformer layer analyzes relationships, extracts meaning, and forms conclusions.
- Every layer can refer to the original materials, so deeper computations don’t need to carry all foundational information long-term.
The authors explain this division of labor as the offloading hypothesis: the anchor preserves word identity, while the main network focuses on processing high-level features.
Highlights of this paper:
- It identifies a very specific architectural contradiction.
- It extends cross-layer reuse to the complete Attention pathway.
- The performance improvement is significant with minimal extra cost. Across six downstream tasks, average accuracy improved from 48.80% (gated attention) to 50.27%, and validation perplexity dropped from 14.64 to 14.09.
- It provides a testable mechanistic explanation.
Memorable stories from the interview:
- The entire paper involved around 200 experiments, costing approximately 30,000 RMB. He admitted that choosing Attention as a focus wasn’t a pre-planned path; it took many tests before he discovered the direction of First-Layer Tension and exogenous anchors.
- When deciding to scale up experiments, he calculated it would cost an additional 6,000 RMB. He hesitated that night. Even if the paper were excellent, review uncertainties could lead to rejection—this money might yield nothing. With support from his parents and brother, he proceeded with the larger experiments. The paper has only one author, and its completion relied heavily on his family’s financial and emotional support.
- About 98% of the work was done by himself, without advisor guidance. Few people around him could discuss models with him, so his most frequent collaborators were ChatGPT, DeepSeek, and online courses.
- He frankly stated that without ChatGPT, completing the paper would have been very difficult. He consulted AI on how to submit, respond to reviewers, and plan training schedules.
- He said he hopes to become someone who can make himself and the people he cares about happy. In his view, whether or not AI exists, “happiness is fundamental to being human.”
Related Resources:
- https://icml.cc/virtual/2026/poster/61175…
- https://xiaoyuzhoufm.com/episode/6a8472b95aeb2a5712e8de78…
- https://arxiv.org/html/2601.08131v4…
Similar Articles
@Phoenixyin13: This is one of the most important reposts I've made. The first author of this paper is someone I deeply admire and a good friend of mine—Guowei Xu, a top student from the Yao Class at @Tsinghua_Uni, who is now conducting AI large model research at @Harvard. Guowei's paper precisely hits the current...
Reposting an introduction to a paper by Tsinghua Yao Class graduate Guowei Xu (currently at Harvard) that accurately points out two critical bottlenecks in LLM search: sparse verification and candidate limitation, which are important for improving reasoning capabilities.
@elliotchen100: Translate the work on MiroMind under Shanda. The next step of post-training might be scientific discovery itself. Simply put, it trains a model to propose research hypotheses across different disciplines. Physics, chemistry, and biology all use one method. The paper was accepted at ICML 2026, code open source...
This paper proposes a scalable supervised fine-tuning method for training language models to propose research hypotheses across disciplines. It has been accepted by ICML 2026 and the code is open source.
@Sxy_Cherotich: Recently I've been talking with quite a few model researchers, and a consensus conclusion is: the importance of data is once again highlighted. A while ago I got to know ex-Kimi's @FanqingMengAI, who is doing a startup in the data direction, and invited him to record a podcast. The biggest non-consensus from our conversation is Fanqing's view on the difference between domestic and foreign models...
A podcast about AI model competition, discussing the importance of data, distillation and pre-training innovation, and an interview with Evolvent AI co-founder Meng Fanqing, covering topics such as synthetic data, RSI, and differences in domestic models.
@tanzhengmc97: https://x.com/tanzhengmc97/status/2066531753762656730
Explained the operating principles of large models in easy-to-understand language, including word vectors, Transformer attention mechanism, next-word prediction training, and emergent abilities, suitable for beginners to understand basic AI concepts.
@AlchainHust: Spent most of a day listening to the 4-hour interview between Zhang Xiaojun and Yao Shunyu. This guy, who just moved from Anthropic to Google DeepMind last year, has worked on Claude 3.7/4.5 and Gemini 3. He offered many candid perspectives from a frontline researcher at a top large model lab. The interview is incredibly information-dense…
This article summarizes Zhang Xiaojun's interview with Yao Shunyu, a researcher involved in developing Claude and Gemini, who shares 10 insightful views on AI code generation, company culture, scaling laws, etc.