How Strong is U.S. AI Really?

Last week saw another upheaval in the tech markets. The release of Kimi K3 seemed to shake confidence in the idea that Open AI and Anthropic had a lock on the best models for the foreseeable future. Kimi K3 and before that GLM 5.2 come close on some benchmarks and you can get access to them at lower prices. Furthermore, they release the weights so large scale users possessing powerful hardware can host them locally, meaning the users pay nothing to any AI developer.

All those facts raise the question of how Claude or Chat GPT can justify the massive valuations of their respective companies. Can they possibly do an IPO at near their recent valuations? Even if the U.S. were to ban use of Chinese models domestically it wouldn’t do nearly enough to support the leading companies at their current valuations. There’s a whole world out there outside the USA after all.

Free,(open weight), innovative Chinese AI appears to be catching up to U.S. AI models which seems like a bad sign for potential shareholders. It’s also a bad sign for current investors in most parts of the AI economy. Huge investments go into developing and refining AI models, but the Chinese makers seem to do it cheaper. That would seem to show the huge investments represent overspending.

The actual situation isn’t that simple. The question of whether Chinese AI companies constitute an actual threat to U.S. AI dominance hinges more on how AI model development works (something we’re all still figuring out.) The idea that Chinese models (Qwen, GLM, Kimi K, Deepseek) could over-take Claude, Chat-GPT and Gemini is probably incorrect. Probably. Figuring out this point is a multi-trillion-dollar question.

To understand if Chinese AI really poses an existential challenge to U.S. AI business you have to know how they have developed their models so far, and how model development actually works. Just looking at Kimi K3 current capability next to Anthropic’s Fable 5 and viewing it as a horserace is misleading – how they got to where they are is apples and oranges.

How to Build an AI Model

Distilling vs original training

There’s strong evidence most of the Chinese models were distilled from U.S. models, mainly Claude and Chat-GPT. An overwhelminmg portion of discussion around that claim centers on moral or legal angles, but those don’t really matter. Debating if Anthropic or Open AI or Moonshot AI (Kimi K) are more in the wrong is a distraction – deliberately, I think. [1] What matters is (1) if it’s true (it is,) and (2) once you have an accurate understanding of the situation, what are the technical and financial implications?

If you can skip the expensive post-training, especially the human feedback part which takes a good deal of calendar time you can release models much faster and cheaper than a developer that has to do all the original work. It’s harder to just staff up super fast, compared to provisioning a bunch of servers for a week. For the original AI developer the human feedback is a crucial step that can’t be rushed too much. For the downstream distiller, they can purchase temporary compute to essnentially reverse extract the value the post-training added to the model.

Part of the motivation for this approach had to be that the most advanced Nvidia GPUs were banned for export to China and they had to make do with less powerful processors. That put even more pressure on Deepseek to not only skip RLHF (reinforcement learning with human feedback) but also use less compute overall.

During distillation you’re extracting information created during the laborious post-training phase. Think of the outputs of an AI like a function: If you get enough examples of inputs and outputs you can work out an approximation of that function, at least for inputs you care about. You capture hundreds of thousands of x = f_chatgpt(y) for a good variety of y.

Although distilling allows you to skip a lot of slow and expensive process, you’re never adding your own original post-training ideas, only distilling an approximation of the source AI. Someone actually has to pay for the original post-training to get fundamentally new capabilities on top of a base model.

How they do it

Drafting off the Frontier Models

So, the fact that several Chinese models appear close in capability – but always a bit behind – U.S. frontier AI models may not present much of a threat to a monopoly on the very top of capability. The question remains if most customers require that extra edge, but in national security terms the U.S. could still expect to stay in the lead.

Business-wise it’s certainly a vulnerability. Once six-month old models perform well enough for 99% of users it may not matter to them that they don’t use the latest and greatest.

On the other hand, if the source for post-training the Chinese models were cut off and U.S. developers could continue for a long time while blocking the Chinese “drafting”, they could soon make the Chinese models irrelavent. In reality that state of affairs would probably be pretty temporary and the Chinese developers would find work-arounds or just invest in proven methods once they’re forced into it.

Cold Start to Recursive Self-Improvement

Even when distillation can overcome the cold start problem, it’s not enough to sustain future development. What actually matters is if the Chinese models can achieve recursive self-improvement or not. If so, the U.S. models can’t possibly maintain a monopoly. They can’t expect to capture nearly all the value of the AI market. (The actual size of that market or the maximum capabilities of near-term LLM based AI is another question.)

Recursive self improvement simply means the ability for an AI system (it would have to be more than a mere model) to get powerful enough to improve itself without the bootstrap phase provided by humans. This would probably look like AI providing substantial help in developing a successor model, and that successor becoming even more capable assisting in the development of its successor, and so on.

If recursive self improvement really is a thing in the near-tterm, even blocking (somehow) Chinese AI makers from U.S. model output won’t do much for long.


  1. There are hundreds upon hundreds of comments in social media and forums of the form “Anthropic stole thousands of books and information in general to make their model, so what’s wrong with Chinese AI stealing from them?”. It’s a deceptive nihilistic tactic since obviously two things can be wrong at once. Plus, Anthropic and Open AI have actually settled and compensated authors a lot and now have policies that they at least plausibly can claim sticking to “fair use.” It’s certainly not all settled law yet though. Still, the attacks count on readers being uninformed.

The bigger question of if it’s okay to read the whole internet and most books and create a art / music / text generation machine that overwhelms all the output of actual creative humans is something else and to be clear I think it’s bad.

There are also a lot of comments praising the services providing access to Chinese models. They say the price for GLM 5.2 or Kimi K3 are way more affordable than Claude or Chat GPT. That may be (it’s complicated) but what they don’t point out is that these prices are probably even more heavily subsidized than Chat GPT or Claude pricing. It can’t go on forever. That’s fine, but the tone of the comments implies the U.S. providers are ripping off their users and you’re a chump to pay full U.S. prices. Whatever else you think of the U.S. AI companies, clearly it isn’t the case that they are extracting huge profits from their users. The prices are too low to pay for inference and all its externalities much less underwrite future model development – that’s mostly done with venture capital still. A big reason the Chinese models can be so cheap is if they are open weight the services selling inference didn’t have to pay for any of the model development. And for the developers selling their service directly, the original development didn’t cost nearly as much if they distilled the U.S. models.

The relentless attacks on Anthropic (now that they have eclipsed Open AI) are simply an obvious campaign to demoralize Anthropic users and investors. That’s it. I don’t know if it’s orchestrated by Open AI, or Chinese companies or the Chinese government or what, but you can see the aim clearly enough. Spend a little time looking at who posts the comments and their similarity and it’s clear.

By flooding social media with this sentiment they steer discussions away from people caring about the actually interesting questions – if Chinese AI developers are taking distillation shortcuts and why that matters economically and technically. There’s plenty of ill-will against the big AI companies and it’s easy to harness it to derail discussion.