@startupideaspod: https://x.com/startupideaspod/status/2069494373604282771
Summary
GLM 5.2 is an open-source AI model with a 1M token context window and strong benchmark performance, narrowly trailing Opus 4.8. The episode provides a practical setup guide for local or cloud use with tools like Cursor and Codex, and emphasizes chaining models for cost efficiency.
View Cached Full Text
Cached at: 06/24/26, 02:25 PM
GLM 5.2: How to Set Up OpenSource AI (With Cursor/Codex etc)
GLM 5.2 is the model everyone on X is screaming about. People are calling it the ChatGPT moment for local AI. Almost nobody has shown you how to actually run it.
So I brought my friend Amir on the show to fix that. In under 20 minutes he broke down what GLM 5.2 is, why it’s crushing benchmarks, and how to set it up today.
Here is the whole playbook.
The numbers that matter
GLM 5.2 is open source. ZAI built it. Run it locally if your machine can handle it, or run it in the cloud through a provider like OpenRouter.
It ships with a 1 million token context window. It scores 81 on Terminal Bench 2.1. That is about 4 points behind Opus 4.8. It holds up on long-horizon tasks, the kind where the model has to plan across a long sequence of steps.
Big leap from 5.1, and it’s strongest right now on front-end, execution-style work.
A free, open-source model is running 4 points behind Opus 4.8. That gap used to be a canyon.
Forget the benchmarks for a second
Amir was honest about this. He doesn’t obsess over benchmark charts. His test is simpler: build something real, watch how it performs, decide from there.
That’s the right instinct. A score tells you a model is good on paper. Shipping with it tells you whether it’s good for you.
Setup, the short version
Two paths. Pick one.
Cursor + OpenRouter: Go to ZAI and grab an API key. Open Cursor settings and paste it into the OpenAI field. Override the OpenAI endpoint with ZAI’s endpoint. Add a custom model named GLM 5.2. Call it directly.
Codex + OpenRouter: Grab your OpenRouter key. Pull the endpoint from the provider. Create a profile in Codex and install the open-source model. Give it the model name and context window. Switch to GLM 5.2 from the CLI.
Codex supports open-source models out of the box. So does any tool that lets you bring your own model.
The real edge isn’t one model. It’s chaining them.
OpenRouter calls this fusion. You sequence a heavy thinking model and a fast execution model so each does what it’s best at.
Here is how Amir used it on a real task.
He wanted to refine a hero section. GLM 5.2 doesn’t read images, so he started with Opus 4.8. He fed it screenshots and asked it to describe exactly what it saw in the front-end design and lay it out in text.
Then he switched to GLM 5.2 to study that layout and make the changes.
Opus does the expensive thinking. GLM does the cheap execution. You get frontier-level output for a fraction of the price.
Plan with Opus. Execute with GLM 5.2. Review with Composer 2.5 or Codex 5.5. Same result, smaller bill.
The mental model Amir used was free trade. Florida grows oranges. Canada makes maple syrup. Each side trades for what the other does better. Models work the same way.
Why this matters more than it look
Most people are vibe-spending on tokens. Amir used to be one of them. In a past episode he admitted he had no idea what it cost. He just spent.
That changed. His usage capped out faster. The bill climbed. The team grew.
He’s watching companies hit the same wall. The first year was about adoption and token maxing to prove they were AI-native. Now they’re realizing they’re spending too much, and they’re canceling subscriptions to the expensive model APIs.
Even Satya at Microsoft has pointed at human capital plus token usage as the new equation.
There’s a governance problem hiding in here too. When someone in marketing fires up Opus 4.8 on high thinking to format an email, that’s the wrong model for the job. Companies are now asking how to teach their teams to pick the right model for the right task.
You shouldn’t be token maxing. You should be token minimizing and output maxing.
Do you need to buy hardware? No.
X is full of people telling you to buy a Mac Studio. Amir’s answer was blunt. You don’t need the equipment. You can start today.
OpenRouter is credit-based and easy. Load $20 and go. Get a model-agnostic agent tool set up, point it at OpenRouter, drop in a few open models, and start experimenting.
His favorite move is pushing the limits. Plan with Opus. Execute with GLM 5.2. Review with Composer 2.5 or Codex 5.5. Mix and match until the output is great and the cost is small.
The closing argument
Some builders say they don’t care what tokens cost because the upside of building is so big. If you can spend $200 and pull $1,000 out, fine, keep going.
But the cheap-token party runs on subsidies. Sooner or later, that subsidy runs out.
The builders who win the next year won’t be the ones who spent the most. They’ll be the ones who got the same output for less.
Stop maxing tokens. Start maxing output.
Similar Articles
GLM-5.2 is a win for local AI
GLM-5.2, a 753B parameter open-source model with MIT license, offers frontier-level coding capabilities and massive context window. Its distillation potential promises significant improvements for local AI setups.
GLM-5.2: Built for Long-Horizon Tasks
Z.AI introduces GLM-5.2, a flagship model designed for long-horizon tasks with a solid 1M-token context, improved coding capabilities, and an MIT open-source license, showing competitive performance against leading models like Opus 4.8 and GPT-5.5.
@hooeem: https://x.com/hooeem/status/2068752941553476002
A comprehensive guide to setting up GLM 5.2, an open-source AI model that claims to beat GPT-5.5 on coding benchmarks while being cheaper, covering cloud and local setup options.
@mervenoyann: GLM-5.2 is comparable to Opus 4.8 with 1M context > new IS attention reuses one indexer every 4 sparse layers (2.9× per…
GLM-5.2 is a new model comparable to Opus 4.8, featuring 1M context, new IS attention, improved speculative decoding, and flexible thinking-effort levels. It is released under MIT license with day-0 support in transformers, vLLM, and SGLang.
GLM 5.2 vs. Opus
GLM 5.2 is a new open-weights model from Z.ai, compared against Claude Opus in a 3D game coding task. Opus performed faster and cleaner, but GLM 5.2 offers compelling cost and accessibility advantages.