tool-calling

Tag

Cards List
#tool-calling

Better Models: Worse Tools

Hacker News Top ↗ · 2026-07-04 Cached

Newer Claude models (Opus 4.8 and Sonnet 5) exhibit worse tool-calling behavior by inventing extra fields in tool invocation arguments, causing validation failures, a regression compared to older models.

0 favorites 0 likes
#tool-calling

GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking

Hugging Face Models Trending ↗ · 2026-07-03 Cached

MiniCPM5-1B-Claude-Opus-Fable5-Thinking is a compact 1B thinking language model fine-tuned from openbmb/MiniCPM5-1B on Fable 5 data, enhancing coding and instruction-following while retaining the native thinking chat template and tool-call format. It supports up to 128K context and is suitable for local deployment.

0 favorites 0 likes
#tool-calling

I benchmarked PrismML's 1-bit Bonsai-8B against IBM's Granite on CPU tool calling. The 1-bit model won, but only with grammar-constrained decoding

Reddit r/LocalLLaMA ↗ · 2026-07-02

An independent benchmark of PrismML's 1-bit Bonsai-8B against IBM's Granite and other models on CPU tool calling shows that with grammar-constrained decoding, Bonsai-8B achieves a 92% pass rate, outperforming larger models, but fails without constraints. Granite is the best raw model at 72%.

0 favorites 0 likes
#tool-calling

@seclink: LongCat-2.0 Released & New Billing Service Launched ​ Core features of LongCat-2.0: Trillion parameters, 1M ultra-long context: native tool calling and multi-step reasoning, reliably handling long-context Agent tasks. Outstanding coding capabilities: in code generation, code understanding, and automatic...

X AI KOLs Following ↗ · 2026-07-01 Cached

LongCat-2.0 model released with trillion parameters and 1M ultra-long context, supporting native tool calling and multi-step reasoning, with outstanding coding capabilities. Also introduces Token resource packs and pay-as-you-go API billing service.

0 favorites 0 likes
#tool-calling

InternScience/Agents-A1 · Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-06-30 Cached

Agents-A1 is a 35B Mixture-of-Experts agentic model from InternScience that achieves competitive performance against frontier-scale systems like GPT-5.5 and DeepSeek-V4-pro using long-horizon trajectory scaling and multi-teacher multi-domain distillation.

0 favorites 0 likes
#tool-calling

@MiaAI_lab: If you mainly use local LLMs for Hermes-style agentic loops, this might surprise you: Qwen 3.6 35B actually *beats* Dee…

X AI KOLs Timeline ↗ · 2026-06-29 Cached

Qwen 3.6 35B outperforms DeepSeek v4 Flash on tool-heavy and coding-adjacent workflows, according to benchmarks from MiaAI Lab.

0 favorites 0 likes
#tool-calling

@rohanpaul_ai: The model ("Owl Alpha") is designed for agentic workloads: - tool calling - multi-step reasoning - long-context executi…

X AI KOLs Following ↗ · 2026-06-28 Cached

Owl Alpha is a new model designed for agentic workloads including tool calling, multi-step reasoning, long-context execution, code generation, automated workflows, and DevOps tasks.

0 favorites 0 likes
#tool-calling

@MiaAI_lab: Qwopus 3.6-27b Coder I had a lot of requests to test it, so I did. I ran the same tests I’ve done on other models. It s…

X AI KOLs Timeline ↗ · 2026-06-27 Cached

MiaAI Lab tested Qwopus 3.6-27b Coder and found it underperformed compared to Qwen 3.6 27b and 35b in tool-calling and code generation, with broken HTML demos.

0 favorites 0 likes
#tool-calling

@timseyde: Dumbo's first steps — LFM2.5-230M doing multi-step tool-calling over pre-trained skills provided by @nvidia SONIC. Same…

X AI KOLs Following ↗ · 2026-06-25 Cached

Liquid AI's LFM2.5-230M model demonstrates multi-step tool-calling capabilities on a Unitree G1 robot, running entirely on-device on an NVIDIA Jetson Orin, acting as a skill-selection layer.

0 favorites 0 likes
#tool-calling

@AYi_AInotes: A counter-intuitive judgment: 80% of Agent production crashes have nothing to do with model IQ — they're all from context overflow, tool misconfiguration, sub-agent runaway. The real watershed in 2026 is Harness and Loop, not the model. Bro, @wizardly_ai's engineering note...

X AI KOLs Timeline ↗ · 2026-06-25 Cached

This article points out that 80% of AI Agent production crashes are not due to model intelligence, but are caused by context overflow, tool misconfiguration, and sub-agent runaway. The author emphasizes that the watershed in 2026 lies in Harness (office systems, security) and Loop (automatic cycling mechanism), not the model itself.

0 favorites 0 likes
#tool-calling

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

arXiv cs.CL ↗ · 2026-06-25 Cached

This paper identifies and analyzes 'tool suppression' in open-weight LLMs when both tool calling and JSON schema constraints are simultaneously enabled, proposing the Constraint Priority Inversion hypothesis and a mitigation strategy called Transparent Two-Pass Execution.

0 favorites 0 likes
#tool-calling

@cwolferesearch: What is an agent? The definition can be pretty simple: it’s just an LLM that runs within an agentic loop. To make this …

X AI KOLs Timeline ↗ · 2026-06-24 Cached

A clear definition of an AI agent as an LLM within an agentic loop, covering components like LLM backbone, instructions, tools, environment, and additional details like context management and memory.

0 favorites 0 likes
#tool-calling

Qwen3.6 27B more dumb in vLLM compared to llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-06-24

A user reports that the Qwen3.6-27B model performs better and more reliably with llama.cpp than with vLLM, citing tool call errors and 'lobotomized' behavior in vLLM despite extensive configuration.

0 favorites 0 likes
#tool-calling

Jackrong/Qwopus3.6-27B-Coder-Compat-MTP-GGUF

Hugging Face Models Trending ↗ · 2026-06-20 Cached

Jackrong releases Qwopus3.6-27B-Coder-Compat-MTP-GGUF, a GGUF quantization of the Qwopus3.6-27B-Coder model with an expanded chat template for better interoperability with tool-using runtimes and OpenAI-compatible agent frameworks.

0 favorites 0 likes
#tool-calling

@PatrickToulme: I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier …

X AI KOLs Following ↗ · 2026-06-20 Cached

GLM 5.2 is a frontier open-source coding model that performs near Claude Opus quality on coding tasks, with excellent tool calling, planning, and local deployment capabilities, at no cost.

0 favorites 0 likes
#tool-calling

@cevenif: For those running local LLMs on Macs, here's a tool worth watching — Rapid-MLX. It delivers 2-4x faster inference on M-series chips than Ollama, thanks to being built directly on Apple's MLX framework for more thorough utilization of the chip architecture. Key highlights: KV cache pruning plus…

X AI KOLs Timeline ↗ · 2026-06-18 Cached

Rapid-MLX is a local LLM inference tool optimized for Apple M-series chips. Built on the MLX framework, it achieves 2 to 4 times faster inference than Ollama, supports multiple models, tool calling, and an OpenAI API-compatible interface.

0 favorites 0 likes
#tool-calling

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Hugging Face Daily Papers ↗ · 2026-06-18 Cached

LedgerAgent is a method for customer service agents that maintains task states in a separate ledger to improve policy adherence and state management during tool calling. It improves average passk over standard approaches across four domains.

0 favorites 0 likes
#tool-calling

@haider1: GLM 5.2 feels like the opus 4.5 moment for open-weight models what genuinely impressed me was during long, multi-step a…

X AI KOLs Following ↗ · 2026-06-17 Cached

GLM 5.2 marks a significant milestone for open-weight models, demonstrating strong context retention across long multi-step tasks and more reliable tool calling.

0 favorites 0 likes
#tool-calling

Kimi K2.7 Code: 1T MoE, $0.95/M tokens, MIT license, beats Opus 4.8 on MCP tool-calling

Reddit r/AI_Agents ↗ · 2026-06-17

Moonshot AI 发布了专注于编程的开放式权重模型 Kimi K2.7 Code,拥有1万亿参数和384个专家,性能在MCP工具调用上超越Opus 4.8,成本仅为十分之一。

0 favorites 0 likes
#tool-calling

Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery

arXiv cs.CL ↗ · 2026-06-17 Cached

This paper studies how routing accuracy degrades as the number of agents scales from 10 to 110 in an enterprise productivity assistant, finding F1 drops of 16–23 percentage points. It diagnoses retrieval and confusion gaps and shows that embedding-based shortlisting recovers 10–11pp F1.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback