Tag
The paper examines correctness-gated multi-teacher distillation, finding decision shifts, lost label functionality, and an inconclusive grounding audit, with no incremental benefit over hard filtering.
The paper investigates injecting frozen relational transformer embeddings into an LLM via soft tokens, reporting negative results due to performance instability and sensitivity to serialization formats.
A preregistered stress test on a small transformer shows that while latent causal structure can be localized, releasing it into behavior fails: the gate detector inverts out-of-distribution and linear release directions are bounded below sufficiency, dissociating localization from behavioral release.
A researcher made a video about their first ML research attempt on pruning vision-language models, including negative results, and submitted a paper to arXiv.
This paper presents a combined proof-of-mechanism study of ontology-amplified distillation for sovereign enterprise language models and a contextuality-audit method, using a Qwen3.6-27B student adapted via supervised fine-tuning and DPO. The results are underpowered and negative, showing no superiority over frontier baselines and zero contextuality in routing.