@AlchainHust: https://x.com/AlchainHust/status/2103711280364240936

X AI KOLs Timeline Tools

Summary

This article describes how Anthropic tripled the speed of Claude.ai in two weeks, along with the author's experience applying this method to optimize their own website, and the release of the '闪电.skill' tool for public use.

https://t.co/lVBVSVJsK1
Original Article
View Cached Full Text

Cached at: 09/27/26, 05:11 AM

Release of “Flash.skill”! How Anthropic Accelerated by 3x in Two Weeks, Adapted for Your Own Use

A few days ago, Anthropic published an article detailing how they accelerated the core experience of claude.ai and its desktop version by approximately 3x in just two weeks. From page load to interactive typing time, they reduced it from 3.1 seconds to 0.55 seconds (measured at the 75th percentile, meaning three-quarters of users are faster than this number). Across four categories of operations and 13 metrics, they achieved a geometric mean speedup of 3.1 times.

The entire process took place within a Slack channel. Claude identified issues, built tests, and modified code, while humans set goals, made trade-offs, and approved changes. According to their retrospective, over 3,000 changes were merged in two weeks with zero user-impacting incidents and no rollbacks. I think this article is highly worth reading—they even detailed how they communicated with Claude and where they intercepted it. (They used Claude Tag, backed by an internal research model similar in capability to Opus 5.5.)

I followed their method and had an agent optimize four of my own websites, organizing the results into a “Flash.skill”. The lab version of the typesetter now takes 90% less time from page load to interactive typing, while the other sites didn’t see such dramatic improvements—I’ll explain more below.

Original article: https://claude.dev/blog/how-we-made-claude-ai-faster/

First, Measure What’s Slow

For someone like me who can’t code, telling an AI “this page feels slow, help optimize it” allows it to make changes and even explain the techniques used and the performance gains. However, I often wouldn’t know how to verify if it’s actually faster or where the improvement lies.

Anthropic started by listing what users frequently do: opening the app, starting a conversation, loading old conversations, sending messages. They measured from the moment a user clicks until the result appears on the screen, letting Claude identify issues, make changes, and then re-measure.

But time measurements can be unstable—results might vary depending on what else is running on the computer. So they had Claude try counting the number of instructions executed. In a controlled Valgrind plus node –predictable environment, the instruction count for pure JS code could be reproduced reliably, while for browser operations they used other metrics like React commit counts and style recalculation counts.

Claude discovered that in a segment of code assembling the conversation message tree, the same message ID was being queried three times. After an hour, instruction count dropped by 48%, and actual runtime decreased by 78%. They required every new metric to be validated this way: the number must improve, and actual runtime must decrease correspondingly—if not proven, the metric was discarded.

Valid metrics were then incorporated into automated checks. If a change degraded these metrics, it wasn’t allowed to merge. As optimization continued and the numbers dropped, the thresholds were adjusted downward. The article calls this a ratchet—ensuring future feature additions don’t revert the time savings.

I stumbled on this myself before. I once created an “expression fingerprint” for a writing system—a single number measuring how similar AI-written content was to my own handwritten articles. After four rounds of AI optimization targeting this metric, the distance dropped to 0.32; another group, which only read three of my handwritten pieces before writing once, achieved a distance of 1.71. By the metric, the first version was five times better. But when I read them, the first was clearly worse, and the second was much better.

I later removed it from the optimization goals, keeping it only for detecting obvious issues. If you give an AI a number, it can indeed optimize for that number effectively—but whether that’s what you truly want might require separate consideration.

Initially, they estimated most projects would take two weeks, but by the third day, 12 out of 13 goals were completed. They then let Claude continue finding opportunities for improvement, opening new threads for each.

One approach I later applied directly to the typesetter: pre-rendering the typing input field in HTML so users could start typing before the page framework fully loaded, handing off control once the framework initialized. They specifically tested for character loss during the handoff and compared two versions of the page across 14 window sizes.

The sprint ran over 150 threads simultaneously, each with an owner, and every change required at least one approval. High-risk changes were first used internally, then rolled out to 1% of users, and finally to everyone. Once, Claude proposed a 900-line change to save 2 milliseconds per message, but the engineer rejected it—maintaining such a change wasn’t worthwhile. Other times, they had to push Claude; Claude initially suggested a small change within the week, but the engineer encouraged bolder action, saying it would be merged immediately. Within a minute, Claude revised its plan, delivering the change within an hour.

Speaking of which, back in February, I also set up a group of agents to help with a product. I then went downstairs to eat grilled fish, and when I returned over an hour later, the 8 agents had burned through my $200 membership quota and an additional $65. I quickly pulled the plug.

Testing on My Own Websites

So this time, when optimizing my own sites, I first had them clearly define what “load to interactive” meant, preserved old versions, set up checks, and then started making changes. I set aggressive goals—90% reduction in load-to-interactive time—one agent per site.

All comparisons were run on my computer, alternating between old and new versions, clearing cache, throttling requests to Fast 4G, and measuring 10–25 times per version to take the 75th percentile. I also discovered that four agents running tests and builds on the same computer would compete for CPU, and the order of testing could affect results—so I rotated the sequence as well.

The typesetter showed the best results. Originally, opening the page required synchronously downloading five scripts externally, with a blank screen until they finished. Two of these libraries—one for image export, another for pasting rich text—were unnecessary for everyone who opened the page. The agent switched these to lazy-load, moved the essential libraries back to my own domain, and implemented the static input field mentioned earlier to allow typing immediately.

In the lab, load-to-interactive time dropped from 2238 ms to 209 ms, a 90.7% reduction. However, the preview area still waited for the framework to initialize, which improved from 2238 ms to 1037 ms. In production, accessed from my computer, load-to-interactive dropped from 3.25 s to 1.52 s. Since my computer uses a proxy, connection establishment alone took about 1.4 seconds—the lab figure wasn’t comparable.

There was also a near-miss issue. The early-rendered editor field was empty, but returning users’ previous drafts only filled in after the framework loaded. If users saw a blank field and pasted a new article, the old draft could be overwritten by subsequent auto-saves.

The agent modifying the code tested for character loss during typing but missed the draft issue—a separate review agent caught it. They then adjusted the code to pre-fill drafts as soon as the editor appeared and added specific tests.

On bookai.top, a tutorial site, the most-clicked article featured a mobile screenshot 2,000–3,000 pixels wide in the first screen, though it only displayed at 753 pixels. After resizing the image appropriately, on a 1440px-wide desktop screen, load-to-interactive dropped from 7776 ms to 1588 ms—a 79.6% reduction. However, on mobile, where this image wasn’t in the first screen, the same change had little effect on load speed.

On img2046, a compression tool page, delaying the loading of two libraries used only for batch downloads reduced load-to-interactive time by 11% and script execution time by 29%. However, after the change, batch downloads initially failed. Fortunately, this was caught in supplementary tests for multi-file downloads, and after fixing, the compressed archives were byte-identical to the original.

My personal homepage saw little change—load-to-interactive time decreased by only 1.6%, within measurement error. Switching the avatar to WebP reduced the largest contentful paint by 13.4%. Initially, the agent suggested the main bottleneck was the server, as the first byte from my computer took 1.4 seconds. Later, using vercel.com as a control revealed the same delay, pointing to my computer’s proxy.

Of the four sites, only the typesetter achieved the 90% goal. I had set identical targets for all, but the actual savings depended on where each site was initially slow.

The impact of slight delays on users has long been studied. In 2009, Google conducted an experiment intentionally slowing search results by 100–400 ms for some users, which reduced their daily searches by 0.2%–0.6%. The group slowed by 400 ms still searched 0.21% less on average even five weeks after the delay was removed. Of course, search engine metrics don’t directly apply to small sites like mine.

Get the Flash.skill

Changes for all four sites are now live, and the pitfalls encountered are included in the Flash.skill. It comes with scripts for measuring speed, recording baselines, and checking for future performance regressions.

The name comes from “Flash,” the sloth at the DMV in Zootopia.

https://github.com/alchaincyf/huashu-flash

Simply pass this link to your agent to install it, or run in the terminal:

npx skills add alchaincyf/huashu-flash

Then tell it: “Help me figure out exactly what’s making this website slow.”

Similar Articles

@max_ai_max: https://x.com/max_ai_max/status/2060221653259547069

X AI KOLs Timeline

This article shares a practical guide to writing a truly usable Claude Skill, covering the operating mechanism, directory skeleton, frontmatter writing, iteration methods, etc., to help developers efficiently build and debug custom skills.

@ZionFeng3364: https://x.com/ZionFeng3364/status/2062702195750191182

X AI KOLs Timeline

This article introduces how to use Claude Code's Skill feature to build a personal workflow in ten minutes, turning repetitive tasks into triggerable automated processes, and demonstrates three abilities: installing Skills, writing Skills, and combining Skills.