Continual Learning of Frontier Models for SovereignAI. Tech Report + Open Weights Model [R]

Reddit r/MachineLearning Models

Summary

The report introduces Thomson, a new general-purpose AI model trained through continual learning for sovereign AI, demonstrating competitive performance with efficient compute and addressing forgetting issues.

Paper: https://huggingface.co/spaces/tri-fair-lab/publications/blob/main/Thomson_1_0_Technical_Report.pdf The development of frontier models is commonly perceived to be in the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but often providing little concrete advice on how this can be achieved in the short term under a diversity of funding settings. In this report, we argue that frontier performance can be achieved by a wide range of institutions through Continual Learning on readily available open-weight models. As opposed to existing limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation with a frozen model, our Continual Learning approach takes advantage of the effectiveness of a modern mid- & post-training stack while introducing safeguards preserving both plasticity and stability at each training stage and seeking to make the minimal number of high-impact interventions on the parameters. This strategy results in model improvements comparable to the gains typically seen across multiple successive model generations. Crucially, such results are achievable with compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for a wider range of actors. To demonstrate this, we introduce Thomson, a new general-purpose frontier model trained with an enhanced focus on high-stakes professional work: domains commonly predicted to undergo large productivity improvements through AI. Through a unique focus on Continual Learning, data-centricity, and efficiency, we demonstrate that Thomson performs competitively with recent frontier models on a wide range of domains and capabilities, ranging from agentic tasks to safety, legal, tax & multilingualism, to comprehensive large-scale Deep Research. Thorough evaluations show a distinctive π-shaped pattern: distinct improvements across a wide range of capabilities (including those not explicitly targeted), while almost completely eliminating the forgetting problem common to narrow domain adaptation.
Original Article

Similar Articles

Sovereign AI is not a model, but a supply chain problem (20 minute read)

TLDR AI

The article redefines Sovereign AI as a supply chain realignment challenge rather than a model development race, arguing that countries will need to secure domestic or allied infrastructure for training, inference, and operation of AI, which will drive renewed demand for GPUs, memory, and other hardware.

Open weight has made to frontier

Reddit r/LocalLLaMA

Discusses how open-weight AI models have advanced to frontier-level capabilities, signaling a shift in the AI landscape.

Thomson Reuters Launches Its Own Frontier Model

Hacker News Top

Thomson Reuters has launched its first proprietary large language model, named Thomson, trained on a mix of open-source foundation and proprietary content to provide efficient, domain-specific AI for professionals.

A US directive just switched off two frontier AI models worldwide overnight. Does this actually make the case for "sovereign AI", or just for not single-sourcing your models?

Reddit r/ArtificialInteligence

A US export-control directive forced Anthropic to cut off foreign access to its Fable 5 and Mythos 5 models, sparking debate over sovereign AI and the high costs of training frontier models. The article argues that the real lesson is multi-provider resilience rather than building a national ChatGPT.

Open weights aren't catching up to closed models by copying them, but they're winning because of how the whole AI stack is quietly modularising

Reddit r/singularity

The article argues that open-weight AI models are catching up to closed ones not via distillation but due to the modularisation of the AI stack—stable interfaces (Transformer architecture, OpenAI-compatible APIs, agentic harnesses) allow innovations to diffuse rapidly across the ecosystem, shrinking the capability gap while keeping a massive price advantage, potentially leading to a commoditisation of frontier AI.