@gkxspace: I spend two to three thousand on AI subscriptions every month, some for TTS, ASR, etc. The mainstream ones are expensive and their API protocols differ. I kept thinking: is there a single plan that covers voice cloning, meeting transcription, AI podcast generation, real-time voice Q&A, voice input, and coding? Finally found a godsend—StepFun's S...
Summary
StepFun launches Step Plan subscription at $6.99/month, integrating LLM, TTS, ASR, image generation, and other AI models. Supports direct OpenAI SDK connection, applicable for voice cloning, meeting transcription, AI podcast generation, etc.
View Cached Full Text
Cached at: 05/20/26, 04:35 PM
I used to spend two to three thousand yuan a month on AI subscriptions, some of which were for TTS, ASR, etc. The mainstream services are quite expensive, and their API protocols are all different.
I’ve always been looking for a single plan that could do it all: voice cloning, meeting transcription, AI podcast generation, real-time voice Q&A, voice input, and code writing.
Finally found a true lifesaver — Step Plan by StepFun. It costs $6.99 per month and I can never use it all up. So I gradually canceled all the others.
One subscription gets you access to top-tier models of all kinds:
- LLM: Step 3.5 Flash — incredibly low latency, and you can also integrate it with Claude / Cursor / Cline.
- TTS: stepaudio-2.5-tts (I checked; its ranking is higher than ElevenLabs).
- ASR: Real-time voice conversations with voice cloning support.
- Image generation: Text-to-image + image editing, generating images in 0.7 seconds.
All accessible directly via the OpenAI SDK — just change the base URL.
Here are some use cases (details in the comments):
- English audio recording → Chinese notes in 54 seconds
- Long English article → dual‑speaker MP3 for commuting
- Same text → TTS with 7 different emotions
- Lu Xun’s Kong Yiji → automatic role‑based audiobook
- English podcast → end-to-end Chinese remake
@StepFun_ai
Similar Articles
@FinanceYF5: AI subscription plan subsidies are much larger than imagined. Claude Max 20x: $200/month, actual usage value about $8,000. ChatGPT Pro 20x: $200/month, actual usage value about $14,000. You spend $200, they lose thousands supporting you. This price war,…
Discusses the subsidy scale of AI subscription plans, pointing out that Claude Max and ChatGPT Pro cost $200 per month but actual usage value is far higher, implying fierce price competition.
@FeitengLi: Next week, after adding speaker labeling and speech generation, it won't be this cheap early bird price anymore.
EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.
@MaxForAI: If you are working on voice agents, you should try this project. A team from NTU, NUS, and Shanghai AI Lab released: Mega-ASR. This fully open-source ASR is built on Qwen3-ASR, aiming to break the long-standing bottleneck of ASR performance in noisy, reverberant, or other impaired real-world environments...
NTU, NUS, and Shanghai AI Lab jointly released Mega-ASR, a fully open-source ASR model built on Qwen3-ASR. Using the Voices-in-the-Wild-2M dataset and progressive acoustic-to-semantic optimization, it achieves up to 30% relative Word Error Rate (WER) reduction in real-world noisy environments. With only 1.7B parameters, it enables efficient inference on consumer-grade hardware.
@yhslgg: Old Yang shares another gem open-source tool—KrillinAI, 10,000 stars on GitHub, a must-see for multilingual audio/video content! In a nutshell: from video download to subtitle translation, AI dubbing, video compositing, the entire pipeline is covered, and it can even auto-generate platform covers, supporting Bilibili, Douyin, Xiaohongshu, YouTube…
KrillinAI is an open-source tool that integrates the entire workflow of video downloading, subtitle translation, AI dubbing, and video compositing. It supports context-aware translation, voice cloning, auto layout, and cover generation, and is compatible with multiple AI models, suitable for multilingual audio/video content creation and distribution.
@cevenif: Bro, it's time to say goodbye to those paid voice tools! The open-source and free Voicebox has arrived, completely crushing paid giants like ElevenLabs and WisprFlow. Features: Voice cloning - instantly become anyone, Global voice input - accessible anytime...
An open-source, free local voice AI studio that supports voice cloning, voice generation, and global dictation. No API key required, runs entirely locally, and serves as a free alternative to ElevenLabs and WisprFlow.