I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.
I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.
I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it.
Lately I've been throwing tasks at Qwen and a frontier or recently-frontier model (as well as Kimi, GLM, etc) and the smaller parameter models are not really comparable to Opus when it comes to making intelligent decisions about greyer areas of good software architecture.
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing.
DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.
GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side.
The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.
Compare to the cost of professional-grade tools in other trades and craft hobbies.
Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.
And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.
The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.
> it would likely take years to spend $4000 (plus the real cost of electricity)
Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.
I don't understand how people don't consider this.
Plus you're spec'd out of near-SOTA level in months.
The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged
You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
> You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right?
Absolutely I do. Each generation of open-weight models has come with significant efficiency improvements, and there are significant hardware gains on the horizon: both increasing competition from Chinese chip manufacturers, and new custom silicon from the established players. And unlike Anthropic and OpenAI, most of the pure inference providers aren't massively leveraged - the more hardware they can bring online, the cheaper they can serve tokens.
For $200/mo you either have a SotA model you can’t run on those devices or you have a cheaper model where you pay less than $200 or have a really big amount of tokens without the energy costs and the risk of failing machine
I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plans, you're better off renting.
Before getting the spark, I was just using a google colab account, their $49 dollar plan allows you access to h100's and I can run qwen there in a Jupyter notebook... and if I really need that web front end I can just use cloudeflair/tailscale/the local ssh client to reverse tunnel it.
This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does.
Notice all the comments saying like "omg why so expensive so just use the API??". It's a trick for lockin even with, so called, "open" models. Keep trying to run them locally, keep undoing the lobotomies, mod model behavior so that they work for you and do what you want vs only what someone else says they're allowed to do.
The cloud stuff is definitely a much better economic value, but I would argue:
1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months ago probably needs a rewrite). Which is fine, I don't think you're going to be "left behind" if you're not a hardcore AI enthusiast or anything (I'm not), but as a guy that's always been interested in computer science I want to see how it ticks.
2. You can't really depend on this subsidization lasting forever IMO. I know the financials thing has been beaten to death but I guess I'm in the camp that it's good to be in control of your tools so that you can go elsewhere if the economics change.
I like to check in with ccusage pretty frequently, and honestly like if I were paying API prices for Claude I'd probably be paying thousands a month.
Anecdotally, ~$500-1500/month token spend at API OpenAI/Anthropic pricing seems pretty realistic for full-time engineers at companies with "liberal but not unlimited" LLM spend policies.
This is of course anecdata. I know plenty of outliers, too. I know a principal engineer who uses many multiples of the number I quoted above. I am sure we also know many people making do with much much smaller budgets as well, via all kinds of well-discussed methods.
But, "$500-$1500 per month per full-time developer" is just kind of the personal mental baseline I use when making my decisions with regards to thinking about whether any of this makes any economic sense.
I am also not sure I would choose to use the cheap and easy to run at home model, given a choice. The marketing copy says this is a frontier model, but it's not. Sol and Mythos are the frontier right now. GLM 5.3 Flash simply isn't. I'd rather use the frontier model as they waste less of my time than even Opus.
Compared to pricing from 3 years ago, it's insane.
The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.
Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.
(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)
My bicycle was in the 5-digits brand new (now I paid it 1/5th of that and I do thank the first owner for that: the 8 000 out of 10 K I saved were put into stocks, that's his opportunity cost, not mine).
Or I know a great many a going to cry "audiofool", but I can say with certainty the following does sound better than the stereo setup of those crying audiofool:
(not my setup but I've got those speakers: same thing, 15 K EUR brand new for the pair... Previous owner forked the money to buy these brand new and, well, I didn't... And I just hooked them to a wonderful, cheap, fully-integrated Yamaha amp: amazing sound).
If your hobby is DIY job around the house, the cost of tools can very quickly add up too: having 20 K worth of tools is definitely not unthinkable.
You like old cars? Pricey hobby.
Some here even track their cars: tires and brake pads budget (and overall car budget and depreciation)... Through the roof.
There's a saying that you're not really into computers if your setup doesn't cost more than your car.
Is $16 K ($4 K x 4) a lot? It's six months of rent for me and for many here I'm sure. It's not "crazy crazy".
Can anyone afford that? Definitely not. But there are way more insane things out there.
And thanks to the individuals that go through to all the pain of setting those up, we've got feedback, tutorials, explanation, numbers, etc. as to how to run those at home.
For example I helped my brother set up VMs and GPU passthrough and he's now running uncensored models locally and showing me the different answers between the uncensored models and the commercial, censored, ones.
So to GP who bought four of these: we need more people like you on HN, keep it going, blog about it, be "crazy"!
I mean, I have the same machine and the pricing is only what it is because it has that 1TB nVME in it instead of larger. nVME prices are insane and have been for months.
Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.
I already have a recycled QNAP-now-TrueNAS that has 10G connections into the fleet so the 1TB doesn't bother me at all. I did some rough math and I don't think I'll ever need to load weights off NFS for what I'm doing so far, but the capability is there.
> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.
There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.
Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it
The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.
Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.
NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.
It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.
These are great points. It's a little off topic but what you bring up is why i advise new grads to spend the first couple years of their career in small eat-what-you-kill companies. I think software devs who start out in large companies get this distorted view that their twice a month direct deposit is just magic and comes from the ether no matter what they do. The whole industry would be better off if everyone started out in a "you don't deliver, you don't eat" company and grew from there.
Strong agree, but I also think some roles in big companies (for me, infrastructure) or in certain industries (eg trading/finance) can help build the same understanding without as much of the variance/raw exposure to bottom line.
Now that the role of the ticket-cruncher is on the path towards full commoditization, and individuals can move much more quickly (and even more carelessly!), I think product roles will probably shift towards one where developers are more deeply embedded in the product/business process so that they own/understand what to build without as much separation between the decision-making and prioritization of what to build. Or at least, they should.
It was eye opening to me to run the math of "should X people work for Y months on this project to save Z per year?" and realize that in so many cases, the time and effort it would cost to stop "wasting" money on things is WAY more than you could actually save on it. Even "small" projects can very quickly become $1M+ investments in time and resources, and the diminishing returns add up quickly (but also a good way to justify the value of your contributions, when done). But the job only exists if it saves money or makes money...
To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.
Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.
If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?
If your usage wouldn't change with local inference and you don't have security/privacy concerns then at the currently heavily subsidized pricing, sure.. not economical.
But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".
I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.
There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API.
The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead.
Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.
It cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.
At the current point in time I'd argue it's more about opportunity cost/value.
If I'm a professional photographer chasing the best possible end product, I'm not buying cameras because they're economical. I'm buying the best camera I can get my hands on to get the best product I can produce within reason under the understanding that it doesn't have to equate to the best economic decision to be the _right_ decision.
If you're in a position to be able to take advantage of the local inference - it's a no brainer. If you're not sure how that would be done, then it's not a good move.
All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.
I just bought a Framework Desktop. Would have been nice to get it at the introductory price, or perhaps the new 192gb model refresh they’re now teasing, but I settled and got a 64 gb model. At the time, the 128’s price had already risen again, but the 64’s price was still at a lower price.
64 can still easily do a Qwen 4.8 model, so I’m relatively happy with my purchase… plus, it’s price change has caused it to quickly appreciate in value… so I could sell it if my situation ever turned dire lol
For me it's entirely because I have a bunch of projects with my own personal data that would be tough to do with openrouter/claude or any other cloud.
For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.
Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it.
That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.
They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.
Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology.
Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.
(This isn't a comment on GLM-5.3 Flash as I've not used it!)
I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better.
I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again.
I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a single line of code.
Yes. I have no UI experience, and wanted a model that could produce something good without me telling it how anything should look like.
My prompt was something like: "here's data I have, here's what matters to me, create HTML mockup".
All GPT 5.6 models were laughably bad. And I don't want to downplay it - they were just absolutely, objectively horrible. Every single attempt was what I could probably call "if json was ui".
Claude models produced... "claude look".
GLM 5.3 - somewhere between GPT and Claude.
Kimi k3 - each attempt produced beautiful UIs. It used components that I didn't even know existed and wouldn't even know to ask for. But expensive, very expensive.
ox-alpha (GLM 5.3 flash) was very close to K3. And at this price point, it's already configured as "designer" model in my oh-my-pi.
Even Artificial Analysis has Opus 5 better than Fable in their aggregated "Intelligence Index" which combines 9 benchmarks. Opus 5 is heavily benchmaxxed.
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts
It's what people know. Opus is just the common target.
> Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash
The problem with this and DeepSWE is it goes for a very specific profile. I'm not convinced DeepSWE is any accurate in actual work. It's surely a different signal (compared to some that allow cheating) but it has its own issues, e.g. weak harness.
Luna is great at following instructions but bad instructions or anything not covered = death.
Deepseek is more analytical. Good for bug tracking.
GLM is a better all rounder in some ways. Better at creativity.
And don't forget the coolest part, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?
Exactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.
except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience
Broad and perpetual license over inputs and outputs, and even your name and profile picture.
Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it.
HN's for example
> We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time.
> You acknowledge that Y Combinator may establish general practices and limits concerning use of the Site,
> You further acknowledge that Y Combinator reserves the right to change these general practices and limits at any time, in its sole discretion, with or without notice.
> Y Combinator reserves the right to investigate and take appropriate legal action against anyone who, in Y Combinator’s sole discretion, violates this provision, including without limitation, removing the offending content from the Site, suspending or terminating the account of such violators and reporting you to the law enforcement authorities.
None of the major LLM chat providers (ChatGPT, Claude and Gemini, and I just confirmed this) claim rights over your input.
They also don't claim rights over your output, but because of how copyright law might apply, they explicitly assign all the rights to the generated output.
Not just that but, even if they wanted to claim ownership of the output, courts in the US have deemed that copyright cannot be assigned to machine-generated output.
So yeah, happy to take a look at a counter-example if you have one (aside from GLM 5.3, obv.).
> In choosing to submit, create, generate, record, post, or display Inputs on or through the Service, you grant an irrevocable, perpetual, transferable, sublicensable, royalty-free, and worldwide right to SpaceXAI to use, copy, store, modify, process, adapt, transmit, distribute, reproduce, publish, upload, download, display in public forums, list information regarding, make derivative works of, and distribute such Content, including anything referenced therein, in any and all media or distribution methods now known or later developed, for any purpose, and to aggregate your User Content and derivative works thereof for any purpose, including but not limited to: (i) maintain and provide the Service; (ii) improve our products and the Service and for our other business purposes, such as data analysis, customer and market research, developing new products or features, or identifying or displaying usage or User Content trends; and (iii) perform such other actions to enforce these Terms, comply with our Privacy Policy, comply with applicable law or governmental, court, and law enforcement requests or requirements or keep our Service safe.
> To the extent the User Content includes a person’s image, likeness, voice, or other similar attributes, you grant SpaceXAI the same rights to use those attributes as part of the User Content as described above. You represent and warrant that you have obtained all rights, licenses, notices, permissions, and consents necessary for SpaceXAI to use that User Content.
However, if we check the market share of generative AI providers, ChatGPT+Claude+Gemini make up around 88%; while Grok is 2-4% depending on who you ask.
> Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms.
OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona), had me do it 8 times, just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country).
Their support says they can't look into anything or do anything, and their public spokespersons on X deny everything.
I get tons of cyber refusals now (lots of reverse engineering), so it's only matter of time when my account is going to get banned.
At least Z.AI is being honest here. And no provider other than OAI/ANT had me submit my biometric information just to use Ghidra.
- Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward".
- OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use.
- Google and Meta - I don't need to talk about the practices of these companies.
Yes, the terms of service aren't great. But the alternatives aren't great either. I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind.
All the American companies you mentioned still follow American law and regulation. Skirting that blatantly has big consequences.
Chinese companies do not follow American laws and there are absolutely no consequences for violating it.
Moreover, the average American is not even aware of exactly what the legal/judicial environment is like in China. If your code and data is stolen, you can't fly to China and demand justice in the courts.
> All the American companies you mentioned still follow American law and regulation. Skirting that blatantly has big consequences.
> Chinese companies do not follow American laws and there are absolutely no consequences for violating it.
... lmk when anthropic/openai/spacex/xai are held accountable for anything. Anything at all. Hard to be when you're _writing_ the rules.
you can abliterate any open model like this. This is pretty standard stuff in a TOS. I'd be surprised if you couldn't find the same in OAI or Anthropic's
I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.
Give it a couple days, and there will be plenty of other inference companies hosting it. Don't like z.ai's TOS? Use the model on a provider with TOS that you agree with.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.
whether it's the cost to develop models, cost of hardware, cost of serving ie inference.
I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)
And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.
So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.
I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.
If anything it's custom chips from the labs that threatens Nvidia.
These open models serve as price / performance pressure. Not all tasks require frontier models and cheap open models can be quite good for in-app assistants, if you're building that sort of thing. We also aren't sure the subscriptions will continue to be sustainable. They're currently subsidized to the tune of 50-70x. As someone who is hitting limits weekly that would easily cost me over $10k month per sub.
Google probably serves more tokens then OAI and Anthropic combined, even if many of those tokens aren't from explicit gemini requests, but from AI overviews and other service integrations.
xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.
This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
Another self-inflicted own courtesy of US government policy.
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.
That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.
Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.
This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.
The export controls were not revoked, only reduced, and not before, but after China refused to buy low performing chips. Top gear was and is still sanctioned, as is any EUVL equipment.
And to add to the above: by building their own supply chain for chips, China is helping the unprivileged, those who can't front-run the market with long-term contracts. If China wasn't producing their own chips, the prices for us would be even higher.
Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.
Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.
Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
It's also in this very announcement, in the first paragraph:
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
Not really. Chinese AI companies were never using NVidia AI chips.
This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.
Also, NVidia chips are still sold out and supply constrained.
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different?
Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.
I agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have.
For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.
> 320B total parameters and just 18B active parameters
This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.
@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.
Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is.
What irks me about this is that the harnesses seem to be just an afterthought here.
Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex.
I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"
> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
From a biased source, but would be big if true.
I've had great results with GLM 5.2.
From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
At the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×
│ https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.
Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.
When reading this type of announcements, always have keen eyes on graphs.
e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.
- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)
I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)
Well, Luna debuted with 5x higher pricing than is currently available. With the pace of recent development these models might not be relevant by Thanksgiving.
Of course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are always being played for things like this, but this seems excessive.
I haven't noticed the Chinese models going up in price for the same model. They do release new versions of the models with different prices that are higher. But everybody is doing that. One fine point is that deepseek-v4-flash-0731 is really a different model than deepseek-v4-flash and it's priced higher.
They're obviously in a pickle, nobody is going to continue to pay $15-50 a mm tokens here soon. There's a reason OpenAI stopped training large models last week, and it's not because of "saftey" or "alignment" they know these gigantic models are not worth the squeeze.
Will we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out?
Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).
I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.
We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.
> I don't see how NVIDIA can keep their spot as belle of the ball.
FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.
With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.
I really wish GLM models had vision capabilities. I've worked around that in the past to use a vision MCP in my harness that GLM can call. It is not the same, but it allows the model to query images.
It's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed.
What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.
Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.
Chinese universities are really a huge advantage, even in the US many of the top staff in model development are Chinese. Another big thing is the hardware costs required to train models. Between those two factors it really looks like this will remain a US-China competition for the foreseeable future, although there are some other players like Mistral from France.
I find GLM's idea of fast/flash is not really competitive with the speed DS4 Flash has, and it's hard to see them as being in the same segment for that reason.
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.
> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.
I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.
I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.
I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.
I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.
GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
Wow, if you don't mind me asking. How and where?
They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
$4,000 isn't priced insanely? ye gads
Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.
And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.
The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.
Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.
Plus you're spec'd out of near-SOTA level in months.
The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged
Of course, it's not real unless I sell, and the value will eventually go down, but so far I have significant paper profits.
Also, DeepSeek token prices are continuing to _increase_, not decrease.
One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...
Absolutely I do. Each generation of open-weight models has come with significant efficiency improvements, and there are significant hardware gains on the horizon: both increasing competition from Chinese chip manufacturers, and new custom silicon from the established players. And unlike Anthropic and OpenAI, most of the pure inference providers aren't massively leveraged - the more hardware they can bring online, the cheaper they can serve tokens.
Before getting the spark, I was just using a google colab account, their $49 dollar plan allows you access to h100's and I can run qwen there in a Jupyter notebook... and if I really need that web front end I can just use cloudeflair/tailscale/the local ssh client to reverse tunnel it.
Notice all the comments saying like "omg why so expensive so just use the API??". It's a trick for lockin even with, so called, "open" models. Keep trying to run them locally, keep undoing the lobotomies, mod model behavior so that they work for you and do what you want vs only what someone else says they're allowed to do.
1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months ago probably needs a rewrite). Which is fine, I don't think you're going to be "left behind" if you're not a hardcore AI enthusiast or anything (I'm not), but as a guy that's always been interested in computer science I want to see how it ticks.
2. You can't really depend on this subsidization lasting forever IMO. I know the financials thing has been beaten to death but I guess I'm in the camp that it's good to be in control of your tools so that you can go elsewhere if the economics change.
I like to check in with ccusage pretty frequently, and honestly like if I were paying API prices for Claude I'd probably be paying thousands a month.
Any organisation or individuals not wanting to have their sensitive data flowing away (either because of trade secret or data protection laws)
This is of course anecdata. I know plenty of outliers, too. I know a principal engineer who uses many multiples of the number I quoted above. I am sure we also know many people making do with much much smaller budgets as well, via all kinds of well-discussed methods.
But, "$500-$1500 per month per full-time developer" is just kind of the personal mental baseline I use when making my decisions with regards to thinking about whether any of this makes any economic sense.
The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.
Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.
(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)
It depends.
My bicycle was in the 5-digits brand new (now I paid it 1/5th of that and I do thank the first owner for that: the 8 000 out of 10 K I saved were put into stocks, that's his opportunity cost, not mine).
Or I know a great many a going to cry "audiofool", but I can say with certainty the following does sound better than the stereo setup of those crying audiofool:
https://youtu.be/TQg9FTBMcTQ
(not my setup but I've got those speakers: same thing, 15 K EUR brand new for the pair... Previous owner forked the money to buy these brand new and, well, I didn't... And I just hooked them to a wonderful, cheap, fully-integrated Yamaha amp: amazing sound).
If your hobby is DIY job around the house, the cost of tools can very quickly add up too: having 20 K worth of tools is definitely not unthinkable.
You like old cars? Pricey hobby.
Some here even track their cars: tires and brake pads budget (and overall car budget and depreciation)... Through the roof.
There's a saying that you're not really into computers if your setup doesn't cost more than your car.
Is $16 K ($4 K x 4) a lot? It's six months of rent for me and for many here I'm sure. It's not "crazy crazy".
Can anyone afford that? Definitely not. But there are way more insane things out there.
And thanks to the individuals that go through to all the pain of setting those up, we've got feedback, tutorials, explanation, numbers, etc. as to how to run those at home.
For example I helped my brother set up VMs and GPU passthrough and he's now running uncensored models locally and showing me the different answers between the uncensored models and the commercial, censored, ones.
So to GP who bought four of these: we need more people like you on HN, keep it going, blog about it, be "crazy"!
Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.
Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.
There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.
Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it
The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.
Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.
NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.
It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.
Now that the role of the ticket-cruncher is on the path towards full commoditization, and individuals can move much more quickly (and even more carelessly!), I think product roles will probably shift towards one where developers are more deeply embedded in the product/business process so that they own/understand what to build without as much separation between the decision-making and prioritization of what to build. Or at least, they should.
It was eye opening to me to run the math of "should X people work for Y months on this project to save Z per year?" and realize that in so many cases, the time and effort it would cost to stop "wasting" money on things is WAY more than you could actually save on it. Even "small" projects can very quickly become $1M+ investments in time and resources, and the diminishing returns add up quickly (but also a good way to justify the value of your contributions, when done). But the job only exists if it saves money or makes money...
Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.
But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".
I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.
The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead.
Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.
If I'm a professional photographer chasing the best possible end product, I'm not buying cameras because they're economical. I'm buying the best camera I can get my hands on to get the best product I can produce within reason under the understanding that it doesn't have to equate to the best economic decision to be the _right_ decision.
If you're in a position to be able to take advantage of the local inference - it's a no brainer. If you're not sure how that would be done, then it's not a good move.
All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.
64 can still easily do a Qwen 4.8 model, so I’m relatively happy with my purchase… plus, it’s price change has caused it to quickly appreciate in value… so I could sell it if my situation ever turned dire lol
For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.
https://deepswe.datacurve.ai/
That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.
They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.
Congrats to them!
Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.
(This isn't a comment on GLM-5.3 Flash as I've not used it!)
https://kommodo.ai/i/IFSUUQYT522uZXWvePZ1
it still fells stupid sometimes and it is benchmaxxed for sure. but its good enough that im building all the hobby projects with it.
I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a single line of code.
The flash one?
My prompt was something like: "here's data I have, here's what matters to me, create HTML mockup".
All GPT 5.6 models were laughably bad. And I don't want to downplay it - they were just absolutely, objectively horrible. Every single attempt was what I could probably call "if json was ui".
Claude models produced... "claude look".
GLM 5.3 - somewhere between GPT and Claude.
Kimi k3 - each attempt produced beautiful UIs. It used components that I didn't even know existed and wouldn't even know to ask for. But expensive, very expensive.
ox-alpha (GLM 5.3 flash) was very close to K3. And at this price point, it's already configured as "designer" model in my oh-my-pi.
I haven't tried it but I think Qwen Max is also very good at design.
It's what people know. Opus is just the common target.
> Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash
The problem with this and DeepSWE is it goes for a very specific profile. I'm not convinced DeepSWE is any accurate in actual work. It's surely a different signal (compared to some that allow cheating) but it has its own issues, e.g. weak harness.
Luna is great at following instructions but bad instructions or anything not covered = death.
Deepseek is more analytical. Good for bug tracking.
GLM is a better all rounder in some ways. Better at creativity.
July 16th: The "Kimi K3 moment" - China has caught up to Opus!
4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!
12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
Broad and perpetual license over inputs and outputs, and even your name and profile picture.
Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it.
HN's for example
> We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time.
> You acknowledge that Y Combinator may establish general practices and limits concerning use of the Site,
> You further acknowledge that Y Combinator reserves the right to change these general practices and limits at any time, in its sole discretion, with or without notice.
> Y Combinator reserves the right to investigate and take appropriate legal action against anyone who, in Y Combinator’s sole discretion, violates this provision, including without limitation, removing the offending content from the Site, suspending or terminating the account of such violators and reporting you to the law enforcement authorities.
Also, it only applies to their chat offering, not the api. OpenRouter also offers the API with ZDR.
While shitty, i’d say that its really not that special.
None of the major LLM chat providers (ChatGPT, Claude and Gemini, and I just confirmed this) claim rights over your input.
They also don't claim rights over your output, but because of how copyright law might apply, they explicitly assign all the rights to the generated output.
Not just that but, even if they wanted to claim ownership of the output, courts in the US have deemed that copyright cannot be assigned to machine-generated output.
So yeah, happy to take a look at a counter-example if you have one (aside from GLM 5.3, obv.).
> To the extent the User Content includes a person’s image, likeness, voice, or other similar attributes, you grant SpaceXAI the same rights to use those attributes as part of the User Content as described above. You represent and warrant that you have obtained all rights, licenses, notices, permissions, and consents necessary for SpaceXAI to use that User Content.
https://x.ai/legal/terms-of-service
These are arguably even worse to be honest. Absolute nightmare.
However, if we check the market share of generative AI providers, ChatGPT+Claude+Gemini make up around 88%; while Grok is 2-4% depending on who you ask.
OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona), had me do it 8 times, just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country).
Their support says they can't look into anything or do anything, and their public spokespersons on X deny everything.
I get tons of cyber refusals now (lots of reverse engineering), so it's only matter of time when my account is going to get banned.
At least Z.AI is being honest here. And no provider other than OAI/ANT had me submit my biometric information just to use Ghidra.
Then alternatives are:
- Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward".
- OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use.
- Google and Meta - I don't need to talk about the practices of these companies.
Yes, the terms of service aren't great. But the alternatives aren't great either. I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind.
Chinese companies do not follow American laws and there are absolutely no consequences for violating it.
Moreover, the average American is not even aware of exactly what the legal/judicial environment is like in China. If your code and data is stolen, you can't fly to China and demand justice in the courts.
... lmk when anthropic/openai/spacex/xai are held accountable for anything. Anything at all. Hard to be when you're _writing_ the rules.
Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.
whether it's the cost to develop models, cost of hardware, cost of serving ie inference.
RIP Nivida shareholders
And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.
So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.
I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.
If anything it's custom chips from the labs that threatens Nvidia.
https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...
Because I genuinely can't tell if you mean Google or SpaceX/X.ai lol.
xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
https://z.ai/blog/glm-5.3-flash
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.
This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.
Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.
Zai is on another "export control" list outside the broader 1. Doesn't help.
Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
... I'm at a loss for words here. It was being served for free. To the entire world.
Edit: Ah:
> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)
Boy, do you have a rude awakening in store.
Get a 210 strike put contract and if your thesis is that nvidias current 10 day slide continues you could make some money.
This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.
Also, NVidia chips are still sold out and supply constrained.
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
[0] https://news.ycombinator.com/item?id=49397204
[1] https://news.ycombinator.com/item?id=49431231
Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.
For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.
This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.
@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.
… you’ll still need to splurge, though.
What irks me about this is that the harnesses seem to be just an afterthought here.
Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex.
I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"
* me raises hand.-
Think z code gives a token bonus though
From a biased source, but would be big if true. I've had great results with GLM 5.2.
From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
│ https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.
│ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}
e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.
- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)
I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)
edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.
MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...
I don't have enough metrics to compare those costs but still Chinese models have been cheaper except against Luna for me.
FWIW, Luna does everything so well, I just keep using it for all my agents by default.
How is the business model of Anthropic/OpenAI will sustain?
Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.
FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.
With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.
What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.
Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.
- Input: $0.15 - Output: $0.50 - Cached input: $0.03
https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...
EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.
It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
That's good. Keep going.
Now the US is behind in EVs can you guess what they're doing? [1]
[1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...
If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.
"problem" indeed.
Intellectual property is part of WTO agreements but enforcement is domestic.
US companies do it too, regularly, they simply hire and poach staff from competitors.
Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.
(281 points, 118 comments)
Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?
https://x.com/Zai_org/status/2092616204787626030/photo/1
> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.