Rendered at 08:05:22 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Jowsey 1 days ago [-]
for those who skip to the comments: Qwen released an updated checkpoint of this model today (0902) with significantly higher benchmark results that appear to place it much closer to Fable/Sol
Iolaum 1 days ago [-]
But not open weight I guess, right?
entrope 20 hours ago [-]
No (current) mention of weights being released, so apparently not.
esperent 21 hours ago [-]
Everyone talks about how these models are getting closer and closer to OpenAI/Anthropic and how they're much cheaper at API pricing, which is great.
But then I compare it to my $200/month subscription - which I believe is how most developers are actually using these - and it actually looks way more expensive from my quick calculations.
Has anyone else calculated it? Is there any way to get Kimi K3 or Qwen 3.8 Max at a similar cost to what we're all paying by subscription for Claude or Codex?
If not, I'd push back and say these Chinese models are actually more expensive, for most developers.
These subscriptions look far more limited and possibly also more expensive than the equivalent Claude/Codex ones at similar prices.
It's tricky to do a proper calculation because everyone obscures what a token or a prompt costs in subscriptions and on top of that we have the complexity of how much different models think.
For example, GLM5.3-flash runs slightly slower than Qwen3.8-Next-Flash but uses fewer tokens to think and thus finishes tasks faster overall.
On top of that we have complications like Anthropic apparently being extremely misleading about what the 20x plan really means (it isn't 20x the weekly limit of the base plan). I would not be surprised if everyone's doing something manipulative like that. The constant changes to promotional periods, frequent limit resets, harness updates etc make this even more difficult.
esperent 13 hours ago [-]
I don't think it's that tricky - there's tools like ccusage that show what you would have paid for Claude at api rates, and I think codex just straight up shows you how many tokens you've used if you run /usage.
trescenzi 20 hours ago [-]
If by most developers you mean individuals then sure. But Anthropic and OpenAI have gotten rid of the subscriptions for most, if not all, corporate contracts. So companies looking to give their employees access are paying api rates.
zuhsetaqi 20 hours ago [-]
> Anthropic and OpenAI have gotten rid of the subscriptions for most, if not all, corporate contracts.
I can’t easily find a source for ChatGPT Enterprise on my phone, but I’m reasonably confident it’s the same.
kenmacd 20 hours ago [-]
I don't know about kimi, but my napkin-math shows that qwen and glm give you at least an order of magnitude more usage than anthropic, and that's before you factor in any off-peak discounts.
esperent 19 hours ago [-]
Are you calculating based on API prices? Because I'm talking specifically about subscriptions here.
kenmacd 16 hours ago [-]
Subscriptions. From some searching I saw a best case if around 600M tokens as a back of a napkin usage on the Anthropic subscription for $200 and at worst 2.6B for glm-5.3 for $120.
The credit system makes it tricky to compare though so I'm curious how you worked them out to be closer. (K3 does seem much closer, I think)
Tepix 19 hours ago [-]
Are you comparing their subscription to yours?
How do you know how many tokens you're getting for your subscription or for theirs?
I believe the answer is: You don't.
vikramkr 17 hours ago [-]
no the answer is `npx ccusage` or any of the other of the trillion ways to see how many tokens you're getting and what the current subsidization rates are.
vikramkr 17 hours ago [-]
no the $200/mo subs are definitely infinitely cheaper. If you're stuck paying enterprise API prices though that's not the case. So for personal or business premium plan use there's no competition but api rate/enterprise there is. Still a ton of spend happening on e.g. bedrock and via api.
But then I compare it to my $200/month subscription - which I believe is how most developers are actually using these - and it actually looks way more expensive from my quick calculations.
Has anyone else calculated it? Is there any way to get Kimi K3 or Qwen 3.8 Max at a similar cost to what we're all paying by subscription for Claude or Codex?
If not, I'd push back and say these Chinese models are actually more expensive, for most developers.
These subscriptions look far more limited and possibly also more expensive than the equivalent Claude/Codex ones at similar prices.
https://www.kimi.ai/resources/kimi-k3-pricing
https://www.alibabacloud.com/en/campaign/ai-landing-page-tok...
For example, GLM5.3-flash runs slightly slower than Qwen3.8-Next-Flash but uses fewer tokens to think and thus finishes tasks faster overall.
On top of that we have complications like Anthropic apparently being extremely misleading about what the 20x plan really means (it isn't 20x the weekly limit of the base plan). I would not be surprised if everyone's doing something manipulative like that. The constant changes to promotional periods, frequent limit resets, harness updates etc make this even more difficult.
Do you have any source for that?
“Usage is billed as you go at API rates”
I can’t easily find a source for ChatGPT Enterprise on my phone, but I’m reasonably confident it’s the same.
The credit system makes it tricky to compare though so I'm curious how you worked them out to be closer. (K3 does seem much closer, I think)
I believe the answer is: You don't.
https://x.com/Alibaba_Qwen/status/2094968708288680276