model-weights

Tag

Cards List
#model-weights

HF exploring sale - impact on open models?

Reddit r/LocalLLaMA · 2026-08-26

Hugging Face is exploring a sale valued at around $13 billion, raising concerns about potential policy changes for open models due to third-party investment and a focus on profitability.

0 favorites 0 likes
#model-weights

GLM 5.3 weights. It might offer the best capacity-to-size ratio.

Reddit r/LocalLLaMA · 2026-08-14

Announcement of GLM 5.3 model weights, noting a potential strong capacity-to-size ratio and upcoming evaluation after quantization.

0 favorites 0 likes
#model-weights

@casper_hansen_: I will fill out the apology form if Muse Spark 1.2 is Apache 2.0 licensed and actually is a great model that is not ben…

X AI KOLs Following · 2026-08-10 Cached

A user reacts to Meta's announcement of open-weights Muse Glimmer and upcoming Muse Spark 1.2, saying he'll apologize if Spark 1.2 is Apache 2.0 and genuinely great.

0 favorites 0 likes
#model-weights

@rohanpaul_ai: New Google DeepMind Paper. Model weights do not have to remain static artifacts that can only be fine-tuned or averaged…

X AI KOLs Following · 2026-08-02 Cached

A new Google DeepMind paper, SkillSmith, treats prefix key-value caches as an input modality, composing textual knowledge and existing weights into a fresh prefix cache for a frozen Gemma 3 4B model at inference time, improving adaptation without a target-specific training run.

0 favorites 0 likes
#model-weights

@omarsar0: New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM read…

X AI KOLs Following · 2026-08-02 Cached

Google DeepMind introduces SkillSmith, which treats model weights as an additional modality that LLMs can natively reason over, enabling instruction-steered parametric synthesis for composing skills at inference time. The approach outperforms text-only and weight-only baselines.

0 favorites 0 likes
#model-weights

Twenty-five years ago it was cryptography, today it's model weights

Hacker News Top · 2026-07-28 Cached

An essay comparing today's AI model weight export controls to the 1990s crypto wars, citing OpenAI's containment breach and the use of Chinese open-weight models by defenders, arguing restrictions harm security.

0 favorites 0 likes
#model-weights

Ernos Labs AI Archive: A free, self hosted archive of open model weights

Reddit r/singularity · 2026-07-28 Cached

Ernos Labs introduces a free, self-hosted archive for preserving open AI model weights, emphasizing decentralization and ownership by no single entity.

0 favorites 0 likes
#model-weights

ModelExpress: Distributing Model Artifacts at the Speed of Light (12 minute read)

TLDR AI · 2026-07-27 Cached

NVIDIA introduces ModelExpress, a tool that accelerates distribution of model artifacts across GPU clusters by using GPU-to-GPU RDMA transfers and optimized streaming from storage, reducing model startup times from 8 minutes to under 2 minutes for large models like DeepSeek-V4 Pro.

0 favorites 0 likes
#model-weights

The second K3's weights drop, I'm downloading the full FP16 and storing them in mattresses

Reddit r/LocalLLaMA · 2026-07-22

A user tweets about downloading the second K3 model's full FP16 weights and storing them, jokingly asking others to leave their doors unlocked.

0 favorites 0 likes
#model-weights

High-Bandwidth Flash offers efficient storage for model weights

Hacker News Top · 2026-07-15 Cached

IEEE Spectrum reports on High Bandwidth Flash (HBF), a new memory technology that stacks NAND flash dies using high-bandwidth memory (HBM) packaging techniques to deliver vastly higher read bandwidth for efficient storage of LLM model weights, with first products expected in about a year.

0 favorites 0 likes
#model-weights

@brianbellx: I removed 423 GB from GLM‑5.2 without changing the model. 1,403 GB → 980 GB. 753B weights. Bit for bit exact. No quanti…

X AI KOLs Timeline · 2026-07-12 Cached

A technique to remove 423 GB from GLM-5.2 (753B weights) without quantization or retraining, achieving bit-exact compression by keeping weights compressed in VRAM.

0 favorites 0 likes
#model-weights

@0xkeenz: Today I verified something I've been pondering for a long time. The official Qwen3.6 27B model has BF16 weights, but some quantized versions, like cyankiwi / Unsloth, convert some key weights to FP16 instead of keeping BF16. BF16 and FP16 have the same storage footprint, so why not just keep the original BF16 weights...?

X AI KOLs Timeline · 2026-07-08 Cached

The author verified that converting the Qwen3.6 27B model weights from BF16 to FP16 does not cause numerical overflow, and pointed out that FP16 has higher mantissa precision, explaining why quantized versions use FP16 instead of keeping BF16.

0 favorites 0 likes
#model-weights

Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions

arXiv cs.AI · 2026-07-07 Cached

HASE is a reinforcement-learning framework that co-evolves model weights, task solutions, and harness components (guidance and evaluation) in a unified agentic process, enabling a single 8B-parameter model to match the performance of much larger systems on text classification and alpha factor mining tasks.

0 favorites 0 likes
#model-weights

LingBot-Vision: masked boundary modeling for self-supervised pretraining (0.296 NYUv2 linear-probe RMSE at 1.1B vs 0.309 for DINOv3-7B, trails on ImageNet); weights in 4 sizes[R]

Reddit r/MachineLearning · 2026-07-06

LingBot-Vision introduces masked boundary modeling for self-supervised pretraining, achieving a 0.296 RMSE on NYUv2 linear-probe with 1.1B parameters versus 0.309 for DINOv3-7B, though it trails on ImageNet; weights are released in four sizes.

0 favorites 0 likes
#model-weights

Longcat 2 model weights have been published

Reddit r/LocalLLaMA · 2026-07-03

Meituan has published the weights for Longcat 2.0, available in INT8 and FP8 formats on Hugging Face.

0 favorites 0 likes
#model-weights

Anthropic vs Open weight Chinese AI

Reddit r/ArtificialInteligence · 2026-07-02

Alex Karp argues that true AI safety for enterprises means control over data and model weights, criticizing Anthropic's strategy of capturing downstream value by releasing products that compete with customers. The article frames the open-weight model debate as a business concern rather than a safety one.

0 favorites 0 likes
#model-weights

@0x0SojalSec: the pirate bay for open LLM's, model weights downloadable with torrents. Our next option is waiting

X AI KOLs Timeline · 2026-07-02 Cached

A tweet announcing a torrent-based site for downloading open LLM model weights, positioning it as a pirate bay alternative for open models.

0 favorites 0 likes
#model-weights

@ClementDelangue: Not your weights, not your brain!

X AI KOLs Following · 2026-07-02

Clement Delangue tweets about the importance of owning AI model weights, emphasizing that if you don't own them, you don't own your brain.

0 favorites 0 likes
#model-weights

Google keeps losing top ai researchers, the moat was never the weights

Reddit r/artificial · 2026-06-26

Google's top AI researchers, including Shazeer and John Jumper, are leaving for competitors, indicating that the real asset is talent rather than model weights. The article advises against reliance on any single AI model provider.

0 favorites 0 likes
#model-weights

2-bit QAT model releases

Reddit r/LocalLLaMA · 2026-06-07

A discussion on the potential of 2-bit Quantization Aware Training (QAT) for larger MoE models, comparing their performance to 4-bit QAT and ternary LLMs, and considering feasibility for consumer hardware.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback