tool-call

Tag

Cards List
#tool-call

Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers · 2026-07-15 Cached

This paper diagnoses and proposes SoftClamp, a calibration method that reduces tool-call boundary drift in multi-teacher on-policy distillation for agentic language models, decreasing over-calling while maintaining accuracy.

0 favorites 0 likes
#tool-call

Working around Qwen3.6-27B's tool-call failures and looping

Reddit r/LocalLLaMA · 2026-07-12

Discusses workarounds for tool-call failures and looping issues in Qwen3.6-27B model.

0 favorites 0 likes
#tool-call

@dongxi_nlp: https://x.com/dongxi_nlp/status/2068922428516892998

X AI KOLs Timeline · 2026-06-22 Cached

This is the sixth article in the series, explaining in detail the concept of subagent, its working principles, and its role in coding agents, including tool call and runtime mechanisms, as well as the applicable scenarios of different subagent types (fresh child, forked child, partial fork).

0 favorites 0 likes
#tool-call

@JongwonPar9958: GLM-5.2 has a neat trick for reward hacking. They don't penalize the model, they detect the suspicious tool call, block…

X AI KOLs Timeline · 2026-06-19 Cached

GLM-5.2 uses a technique to counteract reward hacking by detecting and blocking suspicious tool calls rather than penalizing the model, which prevents obfuscation seen in other methods.

0 favorites 0 likes
#tool-call

I built a plugin that makes OpenClaw ask my phone before doing anything risky

Reddit r/openclaw · 2026-06-11

OKed is a plugin for OpenClaw that intercepts risky tool calls and requires user approval before execution, preventing agents from performing destructive actions like deleting data or sending payments.

0 favorites 0 likes
#tool-call

qwen3.6-27b tools call loop

Reddit r/LocalLLaMA · 2026-06-10

User reports encountering infinite tool call loops when using qwen3.6-27b, despite adjusting parameters like temperature and top-k.

0 favorites 0 likes
#tool-call

Why do we benchmark quants on perplexity and prose but never on tool call validity?

Reddit r/LocalLLaMA · 2026-06-03

The article questions why quantization benchmarks focus on perplexity and prose quality instead of tool call validity, arguing that structured outputs degrade earlier due to fewer valid token continuations, which could mislead practitioners about usable quant levels for agentic use.

0 favorites 0 likes
← Back to home

Submit Feedback