llama-3

Tag

Cards List
#llama-3

@TheAhmadOsman: Top 5 moments for Opensource AI - Llama 3 - Qwen 2.5 - DeepSeek R1 - GLM 4.5 - Kimi K3 These are the moments that chang…

X AI KOLs Following · 2026-08-07 Cached

A tweet highlighting the top 5 moments in open source AI, naming Llama 3, Qwen 2.5, DeepSeek R1, GLM 4.5, and Kimi K3 as transformative releases.

0 favorites 0 likes
#llama-3

StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models

arXiv cs.CL · 2026-05-13 Cached

This research paper investigates using preference optimization (ORPO, AlphaPO) on small language models like Llama-3.2-3B and Qwen-3-4B to align them with Stoic philosophy using micro-datasets. The study finds that while 300 examples can effectively encode Stoic virtues, small models still struggle with outward-facing cosmopolitan duties.

0 favorites 0 likes
#llama-3

How difficult is distilling?

Reddit r/LocalLLaMA · 2026-05-08

该文章探讨了模型蒸馏的难度和成本,以DeepSeek R1蒸馏到Llama 3 8b和Qwen 2.5 7b为例,询问为何蒸馏模型不常见。

0 favorites 0 likes
#llama-3

Can We Locate and Prevent Stereotypes in LLMs?

arXiv cs.CL · 2026-04-23 Cached

ArXiv preprint maps stereotype-encoding neurons and attention heads in GPT-2 Small and Llama 3.2, showing biases cluster in small neuron subsets yet ablating them barely reduces biased text generation.

0 favorites 0 likes
← Back to home

Submit Feedback