Tag
This research paper investigates why pretraining in large language models fails to transfer knowledge across languages, identifies disjoint token spaces as a fundamental barrier, and proposes mapping languages to a shared token space to improve cross-lingual generalization.
DRET is a parameter-efficient knowledge-transfer method that injects biomedical domain knowledge into smaller models like DistilBERT via embedding transfer, achieving performance competitive with larger specialized models.
mimeo is an open-source tool that compiles public expert corpora into agent skills and evaluates knowledge access, persona recognition, and judgment transfer in AI agents.
This paper proposes LT-MKT, a method for multi-domain knowledge tracing that incorporates cognitive load and knowledge transfer using large language models to construct a hierarchical graph, achieving state-of-the-art performance on real-world datasets.
ATHENA is a knowledge-guided agentic neural architecture search framework that automates Transformer-based electronic health record modeling by reusing architecture knowledge across hospitals to reduce manual tuning.
The article highlights common issues with complex AI workflows, such as lack of documentation and knowledge transfer, and suggests practices like versioning and testing to improve team handoffs.
J-Miner recovers executable decision knowledge from fine-tuned language-model classifiers by mining named concepts and learning decision rules, enabling inspection and transfer to lightweight models with high fidelity.
This paper proposes Activation-Prune-Merge (APM), a training-free framework for cross-scale fusion that improves smaller language models using larger donors without semantic alignment, achieving performance gains on multiple benchmarks.
The article explores the historical 'chauffeur problem' in early automobiles, where mechanics gained control over wealthy owners due to technical expertise, and draws parallels to modern computer specialists, highlighting the temporary nature of knowledge-based power.
The paper introduces Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent via hierarchical memory, improving tool-use benchmark performance by 3.4–27.2 percentage points.
This paper studies what transfers between transformer models of different sizes in the same family (Pythia), showing that representations align while weights don't, and that conversion works best via initialization rather than direct weight projection.
Author introduces ISNAD, a trust layer for AI agents inspired by the Islamic isnad system, designed to verify claim provenance across multi-agent chains. The paper is published on arXiv and includes code.
Proposes a conditional diffusion-guided knowledge transfer framework for multi-domain knowledge graph completion, generating domain-general entity embeddings without suppressing domain-specific information, achieving 4.3% average MRR improvement over state-of-the-art methods.
The article describes how a skiing accident incapacitated the Tech Lead on a project, testing the team's development practices and underscoring the importance of documentation and knowledge transfer in software development.
NVIDIA's ASPIRE framework enables robots to build a persistent library of skills from successful experiences, allowing reuse for new tasks and improving learning efficiency over time.
HASTE introduces a hierarchical multi-agent system for ML engineering that organizes cross-competition knowledge into three tiers, achieving 77.3% medal rate on MLE-Bench Lite while reducing compute by 52% and demonstrating that structured knowledge transfer outperforms flat memory approaches.
This paper presents a benchmark for Arabic-Russian scientific translation, including a hybrid parallel corpus of 27,000 sentence pairs and fine-tuned multilingual models (mT5, NLLB, Qwen) using LoRA. The best model achieves BLEU 23.15, and the work aims to lower language barriers for scientific knowledge exchange between Arabic and Russian researchers.
This paper introduces a conversational voice agent system that uses a lightweight on-device 'Talker' model to start responding immediately, then incorporates knowledge from a frontier LLM 'Reasoner' as it becomes available, achieving 7-19x faster time-to-first-response while approaching frontier-level performance on a laptop.
CacheRL trains small agent foundation models for multi-step tool-calling tasks, achieving 92% process accuracy (approaching GPT-5's 94%) with 100x less compute using cached rollouts and hybrid reward shaping, with innovations in knowledge transfer, cache-aware rewards, and iterative SFT/GRPO training.
Endava, a global software contracting firm, uses OpenAI's Codex to codify senior expertise into agents, enabling small teams to deliver massive value quickly and transforming how junior and senior engineers collaborate.