Tag
WikiSkill uses a persistent wiki to evolve and transfer skills, showing that smaller models with such skills can outperform larger models without them in benchmarks, raising questions about what knowledge should persist in AI agents.
@mtasic85 demonstrates that with prompt programming alone, without fine-tuning, LFM2.5 2.6B can behave close to Qwen3.8 27B in tool calling and skill system applications.
The article argues that tool-calling reliability often does not scale with model capability; smaller models can outperform larger ones in schema adherence and format discipline, suggesting that raw capability is not the sole factor in choosing a model for tool use.