open-llm

Tag

Cards List
#open-llm

From Zero to Hero: An Open LLM Ecosystem for Armenian

arXiv cs.LG · 2026-09-04 Cached

The paper curates and releases open datasets and a model for Armenian, demonstrating that continued pretraining with news and STEM data improves performance and addresses data scarcity in low-resource NLP.

0 favorites 0 likes
#open-llm

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

arXiv cs.CL · 2026-08-12 Cached

This paper from Xiaomi introduces reference-free post-training for multilingual machine translation, applying GRPO with quality-estimation rewards to the MiLMMT-46-v0.1 SFT models, producing MiLMMT-46-v1.0 that improves translation across 46 languages and outperforms open and proprietary baselines.

0 favorites 0 likes
#open-llm

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

Hugging Face Daily Papers · 2026-08-11 Cached

This paper introduces MiLMMT-46-v1.0, a multilingual machine translation model improved via reference-free post-training with GRPO and checkpoint interpolation, surpassing strong open and proprietary baselines across 46 languages.

0 favorites 0 likes
#open-llm

@0x0SojalSec: the pirate bay for open LLM's, model weights downloadable with torrents. Our next option is waiting

X AI KOLs Timeline · 2026-07-02 Cached

A tweet announcing a torrent-based site for downloading open LLM model weights, positioning it as a pirate bay alternative for open models.

0 favorites 0 likes
#open-llm

@bnjmn_marie: For LFM2.5 8B A1B, the MoQ GGUFs are the best They have the best ratio accuracy/size Again, it's interesting to see the…

X AI KOLs Following · 2026-06-08 Cached

The poster states that the MoQ GGUFs of the LFM2.5 8B A1B model offer the best accuracy-to-size ratio, advising against using versions with less than 95% accuracy recovery.

0 favorites 0 likes
#open-llm

Measuring Maximum Activations in Open Large Language Models

arXiv cs.CL · 2026-05-18 Cached

This paper measures maximum activation magnitudes across 27 checkpoints from 8 open LLM families, finding significant variance across families, architectures, and training stages, with implications for low-bit quantization and deployment.

0 favorites 0 likes
← Back to home

Submit Feedback