The article outlines nine prompt discipline rules that reduced wasted thinking in coding agents by up to 70% based on A/B tests on GLM 5.3 and GLM 5.3 Flash models, with a self-testable exam available on GitHub.
360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%. The block (shipped to global instructions): ## Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it. How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke. Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam
The paper reveals that certain prompt instructions in coding agents can lead to redundant work without improving success rates, and recommends using bounded instructions to minimize waste.
Microsoft introduces Coding-Agent Skill Distillation (CASD), a prompt optimization method where an off-the-shelf coding agent analyzes agent logs to write optimized prompts in one pass, outperforming previous techniques like GEPA and SkillOpt at a lower cost.
The article discusses prompt optimization as a key technique in event-driven development, mentioning frameworks like MLflow that support algorithms such as DSPy, GEPA, and MIProv2 to improve LLM prompts.
Kent C. Dodds discusses a research paper showing AGENTS.md often doesn't improve task success while adding over 20% inference cost, and asks what should actually be placed in AGENTS.md, noting Claude Code reduced its system prompt by 80%.
The article argues that with prompt caching, longer, stable prompts can be cheaper than frequently changing short ones, sharing insights from running AI agents with high cache hit rates.