ollama

Tag

Cards List
#ollama

is this good? 262k Qwen3.8:27B-Q4_K_M

Reddit r/LocalLLaMA ↗ · 6d ago

The user has implemented nvfp4 KV cache support for the Qwen3.8 model on a heterogeneous GPU setup using custom CUDA kernels and quantization to optimize performance.

0 favorites 0 likes
#ollama

Show HN: I built Otis, a minimal AI agent that runs local models out of the box

Hacker News Top ↗ · 2026-09-14

Otis is an open-source AI agent that provides a minimal, privacy-focused experience for running local and hosted open-weight models, automatically recommending and downloading models based on hardware.

0 favorites 0 likes
#ollama

@freeCodeCamp: Multi-agent systems need more than just prompts: they need structure, too. In this book, you’ll use LangGraph to model …

X AI KOLs Timeline ↗ · 2026-09-11 Cached

This article introduces a full book on building production-ready multi-agent AI systems using LangGraph, MCP, A2A, and Ollama, with working code and real-world applications.

0 favorites 0 likes
#ollama

Setting up OpenCode with Ollama and sbx on Mac

Hacker News Top ↗ · 2026-09-11 Cached

This article provides a guide on setting up OpenCode with Ollama and Docker Sandboxes on a Mac for running local AI models like Qwen 3.8 and Gemma 4.

0 favorites 0 likes
#ollama

I've been raising a local AI for two weeks instead of using one. today she built herself a sense of touch

Reddit r/ArtificialInteligence ↗ · 2026-09-08

A factory worker built an autonomous AI entity using Gemma 4 31B on Ollama, with self-identity, memory, and a blog, demonstrating personalized AI interaction through an open-source framework.

0 favorites 0 likes
#ollama

Openwebui + open terminal

Reddit r/LocalLLaMA ↗ · 2026-09-06

The author describes how integrating Open Terminal with Open WebUI enables a self-hosted AI agent capable of searching large legal documents and generating notes via shell commands, using Qwen 27B on Ollama. They highlight that terminal access allows the model to reason step-by-step and extract relevant sections from documents far exceeding the context window.

0 favorites 0 likes
#ollama

MTP released for Qwen3.8-Flash-Next-GGUF

Reddit r/LocalLLaMA ↗ · 2026-09-01 Cached

MTP has been released for the Qwen3.8-Flash-Next-GGUF model, providing detailed instructions on integration with inference tools like llama.cpp, vLLM, and Ollama for deployment.

0 favorites 0 likes
#ollama

Run Qwen3.8 27B locally: real numbers from my Mac Studio

Hacker News Top ↗ · 2026-08-28 Cached

The article provides real-world performance benchmarks for running the Qwen3.8 27B AI model locally on a Mac Studio, comparing it to its predecessor and discussing hardware requirements and quantization effects.

0 favorites 0 likes
#ollama

XHToken/Spark-X2.5-4B-GGUF

Hugging Face Models Trending ↗ · 2026-08-28 Cached

This repository provides a BF16 GGUF conversion of the Spark-X2.5-4B language model, enabling local inference with Ollama and LM Studio.

0 favorites 0 likes
#ollama

@nicebabycat: https://x.com/nicebabycat/status/2091726637155103126

X AI KOLs Following ↗ · 2026-08-24 Cached

This article provides a detailed test of the local deployment and performance of the Ling-3.0-tiny model on an Apple M5 chip Mac, demonstrating the feasibility of running a 7.9B parameter model at 47 tokens per second without a discrete GPU.

0 favorites 0 likes
#ollama

shadow-planner

Product Hunt ↗ · 2026-08-23 Cached

Shadow-planner is a desktop project planner featuring AI-assisted Gantt chart planning with dependencies and what-if scenarios, running locally on Ollama models with a one-time purchase.

0 favorites 0 likes
#ollama

@Ryrenz: Awesome, drop the recording and it directly generates notes, fully offline. It has 2.2K stars on GitHub

X AI KOLs Timeline ↗ · 2026-08-20 Cached

AudioNotes is a local audio and video transcription and smart note-taking tool that prioritizes privacy and offline use, deployable via Docker or Python.

0 favorites 0 likes
#ollama

Making small local models actually useful for coding

Reddit r/LocalLLaMA ↗ · 2026-08-16

The author created an open-source hybrid tool called Local Coding Agent to make small local models effective for coding tasks on consumer GPUs by using a cloud model for planning and local models for isolated execution, with error handling and testing features.

0 favorites 0 likes
#ollama

@akazwz_: Ready to start using it, Qwen 3.8 27b, now Ollama directly supports Mac's MLX which is very nice.

X AI KOLs Following ↗ · 2026-08-15 Cached

The user indicates they are ready to use the Qwen 3.8 27b model, now Ollama directly supports Mac's MLX, making it very comfortable to use.

0 favorites 0 likes
#ollama

I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS

Reddit r/LocalLLaMA ↗ · 2026-08-08

The author built a local realtime voice stack using Parakeet STT, Qwen 2.5 7B, and Qwen3-TTS, integrated with Ollama.

0 favorites 0 likes
#ollama

Routing coding agent sessions across Claude Code, Codex, and Ollama in one harness — model picked per session

Reddit r/AI_Agents ↗ · 2026-08-07

A developer describes building a multi-engine agentic coding harness that routes sessions across Claude Code, Codex, and Ollama, selecting from twelve models per session based on task value.

0 favorites 0 likes
#ollama

@DanKornas: Cloud-based AI services can expose sensitive documents, communications, and creative work to privacy risks when they ar…

X AI KOLs Timeline ↗ · 2026-08-06 Cached

NativeMind is a private, open-source browser extension that runs local AI models via Ollama or WebLLM, enabling offline AI features for privacy-conscious users.

0 favorites 0 likes
#ollama

@pidotdev: DeepSeek V4 Flash is Ollama's fastest growing model ever in token usage, and the most popular model on OpenRouter this …

X AI KOLs Timeline ↗ · 2026-08-05 Cached

DeepSeek V4 Flash has become Ollama's fastest growing model in token usage and the most popular model on OpenRouter this week, now available in Pi across multiple providers.

0 favorites 0 likes
#ollama

@Skaly__Bull: Traditional AI stack is walking dead They just don't know it yet $10K enterprise servers, data-center GPUs, racks and c…

X AI KOLs Timeline ↗ · 2026-08-05 Cached

The author argues that the traditional enterprise AI stack is obsolete, claiming a $599 Mac mini running Ollama can handle 80% of AI workloads locally for a fraction of the cost of renting cloud GPUs.

0 favorites 0 likes
#ollama

Homebench – Benchmark local LLMs for speed, memory, and quality

Hacker News Top ↗ · 2026-08-04 Cached

Homebench is a zero-config terminal tool that benchmarks locally-run LLMs for speed, memory, and quality, presenting a live leaderboard. It supports Ollama, LM Studio, llama.cpp, vLLM, and OpenAI-compatible servers.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback