DeepSeek came from nowhere, topped the App Store's free charts, matched frontier models on reasoning, coding and maths at a fraction of the cost, and rattled Nvidia's share price on the way. The obvious question is why it caught fire, and the more interesting one is why it, and a growing share of the world's notable models, came out of China at all.
Yitian Xu of Alibaba Cloud breaks down the how and the why: DeepSeek's cost advantage, open-source release and speed, its mixture-of-experts architecture and reinforcement-learning training, and the young, flat, curiosity-driven team behind it. He sets that against China's cloud infrastructure, chip-export pressure and a push for self-reliance, then argues the real winners won't be whoever has the biggest model, but whoever pairs models with domain expertise to deliver value to end users. He closes with how Alibaba Cloud runs an "all-in AI" strategy, using its own Qwen models and a practical workflow for turning repetitive human tasks into deployed applications.
Auto-generated transcript - may contain errors.
Tap a timestamp to jump the video.
Yeah. Thank you everyone for having me. I'm Yitian. I'm head of solution architect of Alibaba Cloud. Been in industries for around twenty years, still look young. And and yes, and today's my topic is about AI surgery in China. So just one question, anyone heard about the DeepSig or have everyone's using the DeepSig? Would you raise your hands?
Well, that's much better than I expect. So yes, beginning of this year, so we can see that the DeepSig actually is causing a social vibe, so it's surpassed the Challenge GBT has become the top free download on App Store. And, yeah, it's amazing and suddenly it stopped because so many people download and the server just crashed.
So Ernie's Ernie actually, I think the Mainland China and the Great China's guys actually can use in DeepSeek. It's not so because so popular. And it's given the reasons why deep seq is so popular, I think here this some dashboard you can see that from the performance wise deep seq companies actually come from nowhere suddenly actually it's come up the larger model and the performance is actually head to head with the strategy BTS, OpenAI, even the even Google.
And because the companies actually come from nowhere and also when it's come up the cost for training to the deep sea to develop the larger models is so so cheap and on that day the Nvidia stock price crash. I'm the Nvidia stock shareholder and also I'm the I'm very proud of DeepSeek so the feeling is mixed.
Yeah, obviously the proud I'm with so yeah but it's actually gave some thought. Okay, Why deepsake so popular? And there must be some reason behind that. So first of all I think why deepsake is so popular? I give you the four answers. The first one is cost effective.
So the deepsake is significantly cheap to run than Compare about that the tokens everyone's know about token. So the deep the deeps the cost of the deepsake, the token is fifteen times cheaper than CharterGPT. So for I think of when I talk with customers when they're using the OpenAI's they always complain that when they put in the product put into the productions sometimes the cost is so high.
It's not that case. So it's very it's very cheap for actually it's suitable for large volume task making it's actually a budget of friendlies for a lot of like a developers, universities and small company. And also this is open sourced, so it actually can deploy any kind of like the hardware if you if it can run.
So actually, I deployed a small size of DeepSeek on my own laptop. It's easy and it's free course. First of course. So the second one is the performance and the deep seq is actually excellent in this task like a coding, mathematics and and the reasoning.
So it's actually often out performance in the Charter GPT because it's got the mixture of experts architectures allows actually can specialize a certain task. I will explain in later slides. And also it's got a MNT license so it's everyone actually can download that free of charge and also the DeepSeek is opened its papers, so can read that, understand the reason behind that or how to develop the larger model.
And the final final one is actually the speed. So the deep secret is when actually it's when you actually talk with DeepSpeed, the design is very sufficient and it's actually respond time. It's faster than than the champion GBTs I do the calculation. So, yeah, this is actually the four reasons.
Suddenly, a company has come from nowhere, developed model larger models become cheap and user friendly, open sourced, it's actually causing a little bit of virus on social media. So from the technical lessons, what do we actually can learn from the DeepSeek? First of all, as I mentioned that it's the mixture of expertise architecture.
So it's actually the first of a larger model, so to come up with this design. So it just give you kind of like explanations. So the the mixture of our expertise is a the larger models is actually combined group as expertise expert. So when you actually ask questions to the larger model, and maybe the one of expert were active and answer the question, it's not actually compared to like a group of the expertise discussed and common answers.
So that's why it's much faster and the and it's it's quick it's quicker response and cheaper. And the second one is the training methodology. So the also the deep sick is the first models to do, like, the the reinforcement training is on only.
It's actually skipped the the SFT, the supervisor fine tune. So it's just give this kind of moment. Oh, the reinforcement can actually is can improve the larger and larger than our models recently. Yeah. It's it's very significant. And the third one, it's it the deep sea have lot of like the mode different size of the model.
So it's quite a lot of like a distillation and it's it's kind of like the the largest model is six is the six seventy one parameters, So it's actually when it's actually running to the hardware is too large and it's actually desolate and to a small size models from point five B and until the seventeen B, but although the size of the larger models becomes more, but the performance is still good, so it's actually come through this kind of like a distillation technology.
The final one is about the open source, all the models, all the different size models for the DIPSEQ is actually is free of available to download, free to use and free of charge. Yeah. So it's actually come from the second lesson, what actually so we can learn from the DIPSEQ from the organizing wise or culture wise.
And did everyone knows what the companies to make a deep seeker? What's their background? Anybody knows? Okay. So, yeah, Quanta. Yes. Exactly. So they actually they actually hedge fund companies before, so it's their Quanta. So that's not actually spin off by experienced engineers from Google's, Amazon's.
They're actually a bunch of fresh graduate, PhD interns, and engineers with just few year experience. So that's actually quite a young team. So they don't have, a legacy bond. So they actually have free full of ideas. And the the SEOs actually encourage the young engineer, the teams to do the new things, to do the innovation to do the innovation.
And that when they actually recruiting the new guys, they deliver it past the experienced ones, want to find the new one but have with a follow-up idea. That's quite key. And also, the organization is actually flat buttoned up, so there's no kind of like formal reporting lines.
Everyone actually can report into the CEO, c CEO or CEO. If you got to have new ideas and they give just everyone have the same authorities and you're just free to go. So that's actually create quite of important creativity environment. Once you got ideas, you actually can told to this let CEO knows and that you can deploy the applications in certain GPUs and just run this, just test it.
It's very it's free to use. And it's also prioritized that the teams don't have actually a formal sales KPI. I think it is very weird because I know in the real world, a lot of peoples have like the short term KPIs, short term sales numbers, but I think DeepSea can actually balance quite well about the long term research target and the short term sales short term sales target.
And also, think a deep sea is a little bit streamed because they are hedge fund company, they got enough money to to burn. But again, this is in the in I think in most of companies, you still think about, okay, how to actually balance to the like the long term goal and short term goal.
So the the lessons from the teams in nutshell, so it's rarely I think the talent that can trump our incubacy. So and the second, the openness can be as radical. So it's because it's the open sourced platform, it's open sourced larger larger mobile and let actually everyone can use that and actually can create a lot of like a communities and to use to use that.
So a lot of people love this and it's cheap and it's getting people the chance to learn larger model to deploy the large model to use that. So it's gave you this kind of like a technical issue. When you use that, you don't feel like, okay, I think that I'll cost that too much money.
And I hate Judge GBT, so when they come out, it's cost me like a twenty pounds per month. I just really hate that. It don't think this is a good way to charge this. And so the next question is okay, so why deep sea come up China?
So actually, deep sea is not actually the only larger models come up, notable larger models from China. See this table, and I just took the screenshot from last month, and this list of top fifteen notable AI models, and the red triangle, the red attribute I highlighted, it's actually come from China.
So apart from the top ones come from career universities, and the rest of the larger models is actually come from China and and the States. Six come from actually, the six come from China, and the rest of nine is actually come from the States.
So that's actually approved. The the DIPSEQ approved Chinese AI lab can produce the front t models even with like the export controls on the constraints like the GPU chips or soft things. So the reason behind that, I think I'm working for the cloud industries.
I first of all, I think it's actually created to the we are in China, so we got a very good infrastructure and the cloud cloud support. So for the large number of companies, they don't need to worry about the infrastructure itself, and they actually can more fully focus on the data on algorithms, and once the the models actually become open source and actually can all the most kind of like mainstream cloud companies actually adopt that, so and available to the end users.
That's also very important. And the second, as you know, when the Trump back to the White House and this geopolitical tensions happens, it's a stop actually imported Nvidia's high end chips to to China. Sometimes it's it's very hard. So the government actually said to emphasize that the teams the Chinas need to self residence to to develop applications.
That's kind of like the the major kind of like a strategic strategic policies for the in China. And the the third one, the Chinese technical companies actually presume the larger model like DeepSeek for the domestic advantage and the global positioning. So the open source, the the powered models actually serve actually as a brand building and software, so it's we know that sometimes the Charge GPT is expensive and that's why it's we open source that and it actually can allow the more kind of like the small media companies even
actually especially in like APAC who don't have like enough kind of like a budget to use. That's actually I build this kind of brandings them to catch up. And the next one is actually collaborations and partnership, so I know the deep deep seek when actually they actually collaborate with a lot of different parties, like with cloud companies, with hardware companies and also with academic institution.
So bring everything together to to develop the to develop models, to develop applications, that's also very important. And I think the this is the at the bottom, so I I also I I will also read some articles and this is confirmed at Stanford University and the US still leads the but they're producing the top AI model, but you can see that the trend that the China is actually is closing the performance gap.
But it's not I think in the future I will say at now it's not actually the larger model competitions. I think the real winners will be the those who actually can combine the model and the domain expertise combined together to deliver the actual results to actual to actual solution, the values to end users, that will win the future.
So the companies who actually can build the agent, who actually can build the I would say this AI solution SaaS, I think we'll win because this actually deliver the real value to the to the end user. So that's actually Alibaba Group also follow this strategy and Alibaba cloud Alibaba cloud this is my sort of question.
Anyone heard about Alibaba cloud? Okay, cool. Yeah. So we are the cloud providers, and surprisingly, so we are the are the third of course cloud providers or hypervisors around the world. So we're actually neck to neck to the GCP, to Google, and it's and underneath, we actually not actually provide the infrastructure, and also we're proud of we also have like a platform like a PaaS.
And since two years ago, we also provide this called a model service. It's not actually so we provide to the customer, but also internal usage as well. So from the model service, we provided in our environment that we actually can train the models to do the model inference, to a foundation model development, and also provide the API to the to the committee.
So in the nutshell, so we actually can accelerate this kind of like AI developing development for the for for ourselves and to the user. And since the the QUIN model, Alibaba Cloud, we also build the larger model by ourselves. And it's just the QUIN three just released two two weeks ago, and it's actually achieved quite a competitive result in the reasoning, coding, mathematics compared with another top tier tier model.
So it's also open source. And then once if you want to actually try, so you actually can go to the hungry to download to to try for free. So yeah. And why the I think it's the QUINS-ray is so is the significance of the QUINS-ray, and it's actually better with the language.
It support like more than a hundred language, especially for like small language in the Asia Pacific area, good at coding, scientific, and reasoning. And the actually, the Queen three is the first model can seamlessly switch between the sinking model and no sinking model.
To give you example, so when you are using the strategy strategy BTS, you find like a bottle called the Reason. If you want to actually let the models to think something deeply, so you need to enable that. But the Queen's Ray don't need to do that.
So it's a it's it's more like human. So when I ask someone some guys, what's the result? So one plus one, everyone can immediately answer two. And if I ask the models, okay, there are some guys to plan trip from London to maybe Edinburgh, so they maybe think, oh, you need to take by train, take by car, take by walk, whatever.
It's actually to sync, but the models, the queens actually can do that from can can do the sync illustration, so you don't necessary to to deliberately enable this kind of like a rhythm model. And now it's actually the top kind of like open source model globally, and it's accumulated quite a lot of download download, three hundred million download and a hundred also got a hundred K of the derivative models as well.
So I just encourage everyone to try that. It's quite amazing. So come back to the model service and Alibaba Cloud, we we actually eat our own dog food, basically. So underneath layers, we had our operating system, the infrastructure to develop the model, and also we have like a foundation model.
It's even yourself, we're actually focused on the the **** model ourselves, but also we are using DeepSeek, any account of like open source model as a suit for the application purpose, even the Lama or whatever. And the ZIMs, we also made some the the offshore applications to for the different for the different team.
So in Alibaba Cloud, we also we've got a code corporate. So I think forty percent or fifteen percent of code actually developed by the corporate. So it's very significant. And also we had like a chatbot for the customer service and also some legal advice, some translator.
This is actually is already made to the applications being being powered by the different mode. So any teams, anyone in our organizations want to develop a new things, actually it's a case cherry pick what what tools they want, what application they want. So I think some guys know about the Duolingo and it's it's their CEO just published the articles I think two also the two weeks ago two weeks ago And it's actually I think this is quite important that it's become this the companies actually realized
that using the AI is a is like a paradigm shift. It's a new norm. So when the c the CEO of Duolingo, they mentioned that they are not actually as a recruiter contract unless the AI can't replace it. And all the AIs will be in the in the process of like a talent hirings and the performance review, and they need to actually look at the order kind of like a workload involved, like a human repetitive work to see that if AI can replace that.
So that's actually we Alibaba Cloud, also Alibaba Group, I fully agree with that. I think this is some golden guidelines when you implement the AI to check that, to seek that to seeking that if any kind of like point the AI can replace, not totally replace the human, but improve the human's work.
So Alibaba Cloud, I think last year or last year, and we adopted all AI strategies. We're not calling AI first, we're all AI. So the first war, we think that we need to get to the order organizations that is AI readiness capability. So we actually separate that in the organization, we have like a different shareholders like the business leader, AI leader, engineering, data technical lead, and we we have different pillow and to understand about AI, like AI foundations value in engineers and governance.
So the different roles of the people that they have like a different need to achieve certain readiness of the different AI readiness payroll. So the whole organizations actually can have shared the same vision, same capabilities, and they actually can work on the same strategy.
So the AI strategy and here the the top ones is the is the roadmap and the first of all from strategy levels we need to actually define the AI strategy ambition and doesn't create a kind of like the first user case and then to access that any kind of like AI maturity.
So which model, which tools, and then to develop the the road map and then to implement this to the execution level. And the second second picture is just kind of like an illustration, the roadmap of how to actually develop the AI and it's actually illustrating like the one years, maybe it's much faster than that.
So first of all start up like the user case then it's getting the pirates and there's understanding that and I didn't don't forget I also want to mention that when you actually are using the AIs you need to do like end to end monitoring because sometimes it's the cost will go very high if you not actually check that very carefully.
So and then when actually you've been proved that in the the test environment, in the priority environment, you actually can deploy that into the production environment and finally and keep improve that. And this is the detail, so for the AI adoption work workflow, first of all, and for organizing teams, need to deploy the established user case scenario and to define the work then define the workflow because when you for example, so when you want to do make a product, it's actually have involved different teams, different step, so you need to define that first.
Then it's need to build your proprieties and enterprise level knowledge base to help the AI understand the work process, and then to decompose the workflow into different work tasks to have this starting stage, a preparing stage, and then finishing stage, whatever. And then we need to focus on what tasks actually involve some human repetitive work.
So that's the easiest one actually can help involve AI to applications to replace improve the human efficient work. So that's kind of like a whole process and we did that in our organizations quite oftenly. And the second second pictures and we normally suggest that as a starting point we have like the pirate, it's it's all gathered like a hundred page of documents within knowledge base and they have like maybe ten ish users as like a as kind of like a pioneer to test the applications with single user case
minimum risk security risk and even don't don't have like a mandatory SLA. But when this actually the applications roll up, grows, and it actually can be in the productions, it can actually roll out to the thousand users, multiple user cases, and we need to also define the enterprise SLA because this is a production application.
So this is actually how we have like a three layers how to develop the ALM applications. I think the first the first lines, it's very easy just direct direct call the the base model, just calling that using the pump engineerings and do some like a performance test as in okay, the API is ready.
So you actually for you, you you're ready to go. The second one may be slightly complicated, you actually seem to build applications on the base model and before you need to focus on the integrations, for example, to integrate with the rack to build the data knowledge base and also the plug ins and even recently is the MCP server to let the larger models to call different third party ultimate data and finally you can come up the assistant API service invocation.
And the final ones, it's the most complicated and not actually just using the ready made offshore larger model, but you need to train in the model, preparing the data, and to fine tune the application that's actually involved with data engineering, model expertise to build these models that are dedicated for the company, and with this kind of like excluded models with data, and you actually can integrate with third party APIs, and finally actually can come as like the applications for the productions.
And it actually will involve like the leader of the enterprise for organizations, and you also need to support for the ecosystems and also the larger model providers as well, so this is all the parties and all the work process we define that in our team.
And this is some user cases and Alibaba Group and we just did this kind of few user cases and we published that to the to the end users. It's actually in different stream like the marketing, the customer service operation and the transaction. So we it's it's all applications listed here, actually it did it kind of like in the few few months.
So it's accelerated this kind of like the application I will say the go to market speed and yeah it's for me as I'm background is is a software engineer I never saw that before so when the actuators created applications, then deployed into the into the product environment, it's so it's so fast, it's definitely is for me, it's kind of like a changing moment, so yeah that's it, and all the whole organization is actually in embrace this kind of like the spirit, and we actually, everyone's actually can apply the our model service,
and to use that and to try that and it can realize that the application evaluations deliver to the customer. So, yeah, that's my final slide. And if anyone's want to have my slides or want to talk to me, this is my email. Yeah. Thank you very much. Thank you for having me here.