rtx-3060

Tag

Cards List
#rtx-3060

@DogukanUrker: Ornith-1.5-9B at Q5 on a single RTX 3060: 200k context at ~52 tok/s, ~1700 tok/s prefill. 11.8 of the 12GB, zero cpu of…

X AI KOLs Timeline · yesterday Cached

The post details running the Ornith-1.5-9B AI model on an RTX 3060 with 200k context, achieving high inference speeds using advanced quantization techniques.

0 favorites 0 likes
#rtx-3060

@DogukanUrker: Gemma 4 12B on a single RTX 3060: the full 262,144 context at ~100 tok/s. (config below) dense model -> MTP speculative…

X AI KOLs Timeline · 2026-07-21 Cached

DogukanUrker demonstrates running Gemma 4 12B with full 262,144 context at ~100 tok/s on a single RTX 3060 using speculative decoding and KV cache splitting, achieving nearly full GPU utilization without CPU offload.

0 favorites 0 likes
#rtx-3060

@alamin_ai_: OMG, guys, this is unbelievable Please listen to the Levantine Arabic and the seamless code-switching with english, a 7…

X AI KOLs Following · 2026-07-06 Cached

A significant breakthrough in Levantine Arabic speech synthesis and English code-switching, achieving a 76% improvement using a single RTX 3060 in an evening.

0 favorites 0 likes
#rtx-3060

Can Qwen3.6-35B-A3B on an RTX 3060 Replace Google Vision for Receipt-to-JSON Extraction?

Reddit r/LocalLLaMA · 2026-06-26

A developer shares their experience using a local Qwen VL model on an RTX 3060 to parse Japanese receipts into JSON, replacing Google Vision, with results showing accurate extraction of key fields at ~31 seconds per receipt.

0 favorites 0 likes
← Back to home

Submit Feedback