Z.ai too cheap to meter
Since release, Z.ai's GLM-5.3-Flash has made quite the impact. It made a bit of a spark when it scored 57 on the Artificial Analysis Intelligence Index. That's the same as GPT-5.6 Sol (high). Which is basically my main work-horse for tough problems. Heck, for more routine problems, I mostly use GPT-5.6 Sol (medium), which sits 1 point lower than that at 56.
What's so exciting about matching those scores? Let's look at the price per task. GPT-5.6 Sol costs $0.43 and $0.29 at high and medium effort, respectively. GLM-5.3-Flash costs $0.09 per task. Okay, 3-5x cheaper. That's pretty good. That's similar in pricing to OpenAI's heavily discounted GPT-5.6 Luna (max) but with Sol performance! When you consider that GLM was 7-12x more verbose in its answers (I believe it was tested at Max effort) then GLM's value shines through even more. Z.ai shows you get almost the same performance with high effort, but half the verbosity, which would drop the cost by another factor of 2, making it 6-10x cheaper than Sol.
What's even more bonkers is that this model is open source. The weights are up on huggingface. Not only are they available but they're relatively small. Given its 320B total and 18B active parameters, it fits in only 642GB of VRAM. Okay, but nobody has 642GB of VRAM sitting around except for inference providers. Right? Wrong. Apple just announced their Mac Studio with 512GB of unified memory is coming in late October. Are you complaining that 642GB is more than 512GB? Well get off my case, because Unsloth quants offered 8-bit quants at 341GB from day 1. Sure it's probably going to be $10k USD, but that's only 8.5x more expensive than my 6-year-old 3090 according to my E-bay sales data. This will certainly be a fan favorite of all local LLM enjoyers. At least the ones that can convince their work into buying a Mac Studio for them. I've already drafted up the business case.
How much cheaper is running those tokens locally than paying the ridiculously low API rate? We can defensibly discount the hardware cost as having negligible depreciation value if secondary market 3090 prices are anything to go by. So what's left? Only electricity cost. Which, my trusty LLM has informed me is best estimated at 140W (130W-160W range) of sustained marginal power draw at a blended Ottawa hydro rate (we call our electricy hydro because Hydro Ottawa delivers it, even though its from a fungible Ontario power grid that's 50.7% nuclear and 24.5% hydro) of 16.23¢/kWh effective rate on top of my household's existing 815.84 kWh/month. Total monthly cost to run 24/7 is $16.36.
How many tokens does that get me? My trusty LLM estimates it at 9.18 tokens per second. This is based on reported 6.2657 tokens/s on M3 Ultra, scaled up by the 1.465x higher memory bandwidth, which is the known chokepoint (1.2 TB/s vs. 0.819 TB/s on the M3). That gives me 23,794,560 tokens in a 30-day month.
How do we compare that to API pricing? They charge per output and cached and uncached input. My 14-day totals calculated with clanker-analytics show me using 256.366 cached-input and 12.813 uncached-input tokens per output token. The average openrouter user uses 176.6 prompt per input tokens, with a 89.7% cache hit rate. For the same 23,794,560 output tokens, that estimates usage of 304.884M uncached input tokens and 6,100.110M cache-read tokens. At current promotional rates that's US$120/month for me and $95/month for the average user. So roughly 8-10x more expensive than running it locally. And expected to double after the promotional period.
In my first test, GLM-5.3-Flash used 60.7M input and 336k output tokens to fail to submit a single score in all 5 runs of my 4 hour benchmark 🙃
Effective rate
The effective rate assumes 24/7 usage at marginal 140W on top of my actual monthly usage profile for the last 12 months, which averages 815.84 kWh per month. Marginal draw totals 100.8 kWh per month. During 6 winter months (Nov-Apr) all usage would be tier 1 at 12.0¢/kWh, and during 6 summer months (May-Oct) it would be all tier 2 at 14.2¢/kWh. The average hourly rate is then 13.1¢/kWh. On top of that we add 1.0332 line-loss factor, 2.16¢ transmission, 0.007¢ low-voltage, and 0.53¢ regulatory additional charges to arrive at an all-in effective rate of 16.23¢/kWh.
Who is Z.ai?
Z.ai is an independent Chinese public AI company spun out of Tsinghua University, founder-controlled, with six founders all Tsinghua-affiliated, three of whom taught there, and with significant Chinese tech, VC, and state-backed investors including Alibaba, Tencent, Ant Group, Xiaomi, and Chinese state-backed funds.
The six founders are Liu Debing, Tang Jie, Li Juanzi, Xu Bin, Zhang Peng, and Wang Shaolan.
Z.ai is the international brand of Zhipu AI (Beijing Zhipu Huazhang Technology Co.), now listed in Hong Kong as Knowledge Atlas Technology.
Sources:
-
Hong Kong IPO prospectus defines:
- the 30.22% Controlling Shareholders as Beijing Lianpai, Dr. Liu, Dr. Tang, Dr. Li, Dr. Xu, Dr. Zhang, Huihui and Zhideng and Beijing Lianpai as 92.70% held by Liu Debing
- the following pre-IPO ownership shares by state-owned/state-controlled investors: Zhongguancun Science City (2.05%), Zhuhai Huafa New Quality Productivity Fund (2.05%), Shanghai Zhihui Linghang (2.05%), Beijing AI Industry Investment Fund (1.98%), Hangzhou Chengtou Industrial Fund (1.44%), Beijing Daxing Industrial Fund (1.23%), and Chengdu High-tech Orinno (1.23%)
- the following pre-IPO ownership shares by state-backed or public-capital investors: Tianjin Haihe Fuxin Youda Fund (3.90%), Social Security Zhongguancun Innovation Fund (1.66%), and Tsinghua Technology (3.86%)
-
Bloomberg lists Chinese investors as Alibaba, Tencent, Ant Group, Xiaomi, HongShan (formerly Sequoia China), and MeiTuan
-
Tsinghua University sources document the founders’ Tsinghua affiliations. Tang Jie, Li Juanzi, and Xu Bin taught there.
- Tang Jie is listed by Tsinghua as a professor in its Department of Computer Science and Technology, which he joined in 2006.
- Li Juanzi is listed by Tsinghua as a professor in its Department of Computer Science and Technology; she received her PhD from Tsinghua in 2000 and remained there after her postdoc.
- Xu Bin is listed by Tsinghua as a professor in its Department of Computer Science and Technology and a member of its Knowledge Engineering Group; he received all three degrees from Tsinghua.
- Zhang Peng is identified by Tsinghua as a 1998-entry alumnus and CEO of Zhipu; the same Tsinghua source states that Zhipu was created through commercialization of Tsinghua technology.
- Liu Debing’s Tsinghua affiliation is documented in the Hong Kong IPO prospectus, which records that he worked as a senior engineer at Tsinghua University before co-founding Zhipu.
- Wang Shaolan’s Tsinghua connection is documented by Tsinghua sources identifying him with the university’s computer-science department and its innovation-leadership doctoral program; see, for example, Tsinghua Journal.