bf16

Tag

Cards List
#bf16

@QuixiAI: LESSON LEARNED: Always use BF16 kv cache. I was using turboquant. Yeah the VRAM consumption sucks - so another trick is…

X AI KOLs Timeline · 2026-08-28 Cached

The author shares a technical lesson on using BF16 KV cache instead of turboquant for AI model optimization and implements CPU offloading in SlimServe for Qwen models to manage VRAM consumption.

0 favorites 0 likes
#bf16

Peak Portable Personal Datacenter

Reddit r/LocalLLaMA · 2026-08-25

A user showcases a portable personal datacenter built for running the Qwen3.8-27B-BF16 model, detailing hardware specs, performance benchmarks, and thermal management for high-context AI inference.

0 favorites 0 likes
#bf16

inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8

Reddit r/LocalLLaMA · 2026-08-04

InclusionAI released the Ling-3.0-flash model weights on Hugging Face with MIT license, including BF16 and official FP8 versions. The model uses a fine-grained MoE architecture with 512 experts and 8 active per token, totaling 127.5B params with 5.1B active.

0 favorites 0 likes
#bf16

@0xkeenz: Today I verified something I've been pondering for a long time. The official Qwen3.6 27B model has BF16 weights, but some quantized versions, like cyankiwi / Unsloth, convert some key weights to FP16 instead of keeping BF16. BF16 and FP16 have the same storage footprint, so why not just keep the original BF16 weights...?

X AI KOLs Timeline · 2026-07-08 Cached

The author verified that converting the Qwen3.6 27B model weights from BF16 to FP16 does not cause numerical overflow, and pointed out that FP16 has higher mantissa precision, explaining why quantized versions use FP16 instead of keeping BF16.

0 favorites 0 likes
#bf16

@SpaceTimeViking: Announcing Orinth 1.0 AEON ULTIMATE UNCENSORED! BF16 and Quantized in NVFP4 for the DGX Spark / Blackwell arch. Preserv…

X AI KOLs Timeline · 2026-06-27 Cached

Announcing Orinth 1.0 AEON ULTIMATE UNCENSORED, a model with BF16 and NVFP4 quantization for DGX Spark/Blackwell architecture, claiming 200-300% performance improvement with working DFlash.

0 favorites 0 likes
#bf16

@jpschroeder: ZERO providers offer GLM-5.2 in native bf16.

X AI KOLs Following · 2026-06-26 Cached

A user notes that no cloud providers currently offer the GLM-5.2 model in native bf16 precision, highlighting a gap in hosting options.

0 favorites 0 likes
#bf16

@steeve: another 5 days later, zml/llmd runs fully on Metal, serving 8 simultaneous requests at full bf16 zml/llmd is our LLM se…

X AI KOLs Following · 2026-06-13 Cached

zml/llmd now runs fully on Apple's Metal API, serving 8 simultaneous requests at full bf16 precision, with continuous batching and other modern features.

0 favorites 0 likes
#bf16

Built an AI Accelerator and opensourced it. [P]

Reddit r/MachineLearning · 2026-05-31

The author open-sourced a custom AI accelerator (atik) implemented on FPGA with native BF16 and attention support, demonstrating significant speedups over PyTorch for various models.

0 favorites 0 likes
#bf16

@charles_irl: another page for the @modal LLMEng Almanac: an explorer for low-precision floats, from bf16 to fp4 https://modal.com/ll…

X AI KOLs Following · 2026-05-18 Cached

A page from Modal's LLM Engineer's Almanac that provides an interactive explorer for understanding low-precision floating-point formats like bf16 and fp4.

0 favorites 0 likes
← Back to home

Submit Feedback