Are we on the verge of an AI revolution driven by an underdog? DeepSeek, a relatively unknown Chinese startup founded in 2023, is making waves in the global AI community with its cutting-edge, open-source models and staggeringly low inference costs.
DeepSeek first rocketed to the top of the app charts in January 2025, fueled by DeepSeek R1, a model users called “shockingly good.” Eighteen months later the company is three model generations past it: R1 has been retired, and the current lineup is V4 Flash and V4 Pro. This article covers DeepSeek’s backstory, the technology behind its rapid ascent, every model it has shipped, and the challenges it faces now.
DeepSeek’s app raises serious privacy and security issues by transmitting user data, including chat logs and keystrokes, to servers in China. If you are weighing the trade-off, see how the two stack up in our DeepSeek vs ChatGPT comparison. This data is subject to Chinese laws, which may compel companies to share information with the government.
For a safer alternative to DeepSeek’s app, users can host its open-source models locally or use 3rd party platforms which keep data within Western data centers, avoiding Chinese data risks and censorship.
The Key Takeaways
- DeepSeek was founded in May 2023 by Liang Wenfeng and is funded entirely by the High-Flyer hedge fund, with no outside investors.
- The current models are V4 Flash (284B total / 13B active) and V4 Pro (1.6T total / 49B active), both MIT-licensed with a 1M token context window.
- API pricing is $0.14 / $0.28 per million tokens for V4 Flash and $0.435 / $0.87 for V4 Pro.
- V4-Flash-0731, shipped July 31, 2026, beats the much larger V4 Pro preview on every agentic benchmark DeepSeek publishes.
- R1 and V3 are retired. Their API endpoints were switched off after 24 July 2026, and DeepSeek’s docs now list exactly two model IDs.
The Rise of DeepSeek
DeepSeek was established in May 2023 by Liang Wenfeng, who previously headed the High-Flyer quantitative hedge fund. Since High-Flyer fully underwrites DeepSeek, the startup is free to pursue ambitious AI research without the usual pressure to generate short-term returns. Located in Hangzhou, China, the company has gathered a young team of top-tier graduates from Chinese universities, emphasizing strong technical skills over conventional work experience.
From day one, DeepSeek has been guided by two core objectives:
- Pushing toward Artificial General Intelligence (AGI) in a transparent, open-source manner
- Making advanced AI more accessible through aggressive pricing and cost-efficient technology
This open-source spirit and disruptive pricing have rattled incumbents, prompting AI powerhouses like OpenAI, Meta, and major Chinese tech firms, including ByteDance, Tencent, Baidu, and Alibaba, to reevaluate their own costs, strategies, and research approaches.
DeepSeek’s Milestones
Since its founding in 2023, DeepSeek has been on a steady trajectory of innovation, launching models that not only compete with but often undercut their bigger competitors in cost and efficiency. From its early focus on coding to its advancements in general-purpose AI, each release has pushed boundaries in a unique way. Here’s a closer look at the milestones that have shaped DeepSeek’s journey so far.
DeepSeek Coder
Launched in November 2023, DeepSeek Coder was the company’s first significant release, targeting developers with an open-source coding model. At a time when commercial code-generation tools were becoming increasingly expensive, it offered a free and effective alternative. The model could generate, complete, and debug code, quickly gaining traction among independent developers and startups. Its open-source nature encouraged customization and experimentation, further boosting its popularity.
This release set the tone for DeepSeek’s mission to democratize AI access. While relatively simple compared to later models, DeepSeek Coder proved that accessible AI tools could deliver strong performance without high costs, laying the groundwork for future innovations.
DeepSeek LLM (67B)
Following the success of its coding model, DeepSeek released a 67B-parameter general-purpose language model. Despite its smaller size compared to competitors like GPT-4, this model excelled at tasks such as summarization, sentiment analysis, and conversational AI. By optimizing for parameter efficiency, it matched or exceeded larger models in many tasks while maintaining a lean computational footprint.
The DeepSeek LLM demonstrated the company’s ability to develop versatile AI tools that prioritized cost-effectiveness without compromising quality. It also solidified DeepSeek’s reputation as an innovative disruptor capable of delivering competitive models on a budget.
DeepSeek V2
Released in May 2024, DeepSeek V2 was a turning point for the company, sparking a price war in the Chinese AI market. By delivering a high-performing language model at a fraction of the cost of its competitors, DeepSeek forced major players like ByteDance, Tencent, and Baidu to lower their prices. This move made advanced AI accessible to a broader range of businesses and developers.
Technically, V2 improved significantly over its predecessors, offering enhanced capabilities for text generation, sentiment analysis, and more. Its combination of performance and affordability caught the attention of the global AI community, proving that smaller firms could compete with heavily funded tech giants.
DeepSeek-Coder-V2
In late 2024, DeepSeek returned to its roots with DeepSeek-Coder-V2, an advanced coding model boasting 236 billion parameters and a context window of 128K tokens. This upgrade enabled it to tackle complex programming tasks, such as analyzing extensive codebases or solving intricate debugging challenges, with impressive accuracy.
What made Coder-V2 stand out was its pricing. Starting at just $0.14 per million input tokens und $0.28 per million output tokens, it became one of the most cost-effective coding tools available. The model cemented DeepSeek’s reputation for providing high-quality AI solutions at a fraction of the cost demanded by competitors.
DeepSeek V3
The launch of DeepSeek V3 in late 2024 marked the company’s most advanced step yet, introducing 671 billion parameters and two groundbreaking innovations:
- Mixture-of-Experts (MoE): Activates only 37 billion parameters per task, drastically reducing computational costs while maintaining high performance.
- Multi-Head Latent Attention (MLA): Enhanced the model’s ability to process nuanced relationships and manage multiple inputs simultaneously, making it highly effective for tasks requiring contextual depth.
While overshadowed by high-profile releases from OpenAI and Meta, DeepSeek V3 quietly gained respect in research circles for its combination of scale, cost efficiency, and architectural innovation. It also laid the technical foundation for DeepSeek’s most significant achievement to date: DeepSeek R1..
DeepSeek R1
DeepSeek took its boldest step yet with DeepSeek R1, launched on January 21, 2025. This open-source AI model has become the startup’s most serious challenge to American tech giants, owing to its formidable reasoning power, lower operating costs, and developer-friendly features.
🚀 DeepSeek-R1 is here!
— DeepSeek (@deepseek_ai) January 20, 2025
⚡ Performance on par with OpenAI-o1
📖 Fully open-source model & technical report
🏆 MIT licensed: Distill & commercialize freely!
🌐 Website & API are live now! Try DeepThink at https://t.co/v1TFy7LHNy today!
🐋 1/n pic.twitter.com/7BlpWAPu6y
Key Features
- Mixture-of-Experts Architecture (MoE)
R1 expands on the MoE concept first seen in V3, activating only the sub-networks required for a specific query. This allows for high performance on demanding tasks without devouring hardware resources. - Pure Reinforcement Learning (RL)
While many competing AI models rely heavily on supervised fine-tuning, R1 incorporates a strong RL pipeline, learning to reason through constant iteration and feedback, rather than relying solely on labeled datasets. - Massive Context Window
Capable of processing up to 128,000 tokens in one request, R1 easily handles extended tasks like complex code reviews, legal document analysis, or multi-step math problems. - High Output Capabilities
The model can generate up to 32,000 tokens at a time, making it ideal for writing in-depth reports or dissecting extensive data sets. - Unprecedented Cost Efficiency
DeepSeek R1’s inference cost was estimated at a tiny fraction, around 2%, of what organizations paid for comparable OpenAI models at the time. For solo developers and enterprises alike, that gap was the whole appeal.
Performance Benchmarks
DeepSeek R1 has logged remarkable scores on math and logic tests, surpassing OpenAI’s o1 Preview with a 91.6% score on the MATH benchmark and 52.5% on AIME. Although it matches OpenAI’s o1 in many coding tasks, it still falls slightly behind Claude 3.5 Sonnet for certain specialized code scenarios. However, R1’s ability to show detailed step-by-step reasoning stands out as a major benefit, particularly for debugging, educational uses, and research.

Perhaps most telling about its success is user adoption. R1 propelled DeepSeek to the top of the App Store on January 26, 2025, and it quickly reached a million downloads on the Play Store. Users cite the recently introduced “DeepThink + Web Search” feature as one of its standout attributes, an area where even OpenAI has yet to fully catch up.

DeepSeek V3.1 and V3.2-Speciale
Through late 2025 DeepSeek iterated on the V3 line rather than jumping to a new number. V3.2-Speciale was the notable release, a reasoning-tuned checkpoint aimed squarely at the frontier closed models on maths and code. It kept the open-weights posture and the aggressive pricing that had defined the company, and it is covered in full in our DeepSeek-V3.2-Speciale breakdown.
DeepSeek V4 Pro and V4 Flash
DeepSeek V4 arrived on April 24, 2026, one day after OpenAI shipped GPT-5.5, and it split the line in two. V4 Flash carries 284B total parameters with 13B active and is the default chat model. V4 Pro carries 1.6T total with 49B active and is the reasoning model. Both moved to a 1 million token context window with 384K maximum output, a roughly eightfold jump on R1’s 128K, and both shipped under the MIT license with full weights on Hugging Face. Our DeepSeek V4 launch coverage has the full specification.
DeepSeek V4-Flash-0731
The most recent milestone landed on July 31, 2026. V4-Flash-0731 is a re-post-train of V4 Flash rather than a new architecture, and it supersedes the preview outright: calling deepseek-v4-flash on the API now routes to it, at the same $0.14 / $0.28. What changed is agentic capability, and the numbers are lopsided. On DeepSeek’s own benchmark table the small Flash model beats the far larger V4 Pro preview across the board, Terminal Bench 2.1 at 82.7 against 72.1, NL2Repo 54.2 against 38.5und DeepSWE 54.4 against 12.8. Artificial Analysis, running Terminal-Bench independently, scores it 78.7; both figures are real and neither should be quoted without naming the house that produced it.
The commercial effect is that DeepSeek’s cheapest model became its best agentic one, which is the reverse of how most vendors price their lineups. On Artificial Analysis’s Intelligence Index it scores 50 at $0.03 per index task, the lowest cost-per-task figure on that board.
What happened to R2
Worth stating plainly, because the rumour persists: there is no DeepSeek R2. The reasoning line was folded into the main V-series rather than continued under the R name, and R1’s own endpoint was switched off after 24 July 2026. DeepSeek’s API docs now list exactly two model IDs, deepseek-v4-flash und deepseek-v4-pro. Any tutorial still telling you to call deepseek-reasoner or deepseek-chat is out of date.
DeepSeek’s Innovations
Both DeepSeek V3 und R1 leverage the Mixture-of-Experts (MoE) architecture, which activates only a subset of their massive 671 billion parameters. Think of it as deploying hundreds of specialized micro-experts that step in precisely when their skills are needed. This design ensures computational efficiency while maintaining high model quality.
DeepSeek’s adoption of a pure reinforcement learning (RL) approach further sets it apart. The models learn and improve autonomously through continuous feedback loops, enabling self-correction and adaptability. This mechanism significantly enhances their problem-solving capabilities, particularly for tasks requiring deep reasoning and logical analysis.
Beyond MoE, Multi-Head Latent Attention (MLA) boosts the models’ ability to process multiple data streams at once. By distributing focus across several “attention heads,” they can better identify contextual relationships and handle nuanced inputs, even when processing tens of thousands of tokens in a single request.
DeepSeek’s innovations also extend to model distillation, where knowledge from its larger models is transferred to smaller, more efficient versions, such as DeepSeek-R1-Distill. These compact models retain much of the reasoning power of their larger counterparts but require significantly fewer computational resources, making advanced AI more accessible.
Reactions from the AI Community
Several prominent figures in AI have weighed in on the disruptive potential of DeepSeek R1:
- Researchers at the United Nations University argued that R1 challenges the idea that high-performance AI requires immense computational resources, and that delivering top-tier results at a fraction of the cost opens the door to democratizing access to advanced AI across industries.
- Independent analysts covering the release singled out R1’s reinforcement learning framework and search capabilities as markers of a new standard in training methodology, arguing the approach would push the industry to rethink how AI models are trained and optimized.
- Alex Zhavoronkov, CEO of Insilico Medicine, praised the biological inspiration behind DeepSeek R1’s reinforcement learning structure. He described it as a significant step forward in logical self-assessment and adaptability, with implications that extend far beyond current AI research paradigms.
- Marc Andreessen, co-founder of Andreessen Horowitz, described DeepSeek R1 as “AI’s Sputnik moment” and one of the most amazing and impressive breakthroughs he has ever seen. He also praised its open-source nature, calling it a “profound gift to the world.” This level of enthusiasm from a leading tech figure underscores the model’s significance and its impact on the industry.
Deepseek R1 is one of the most amazing and impressive breakthroughs I’ve ever seen — and as open source, a profound gift to the world. 🤖🫡
— Marc Andreessen 🇺🇸 (@pmarca) January 24, 2025
At the same time, there are skeptics. Concerns have been raised about potential biases in training data und geopolitical implications due to DeepSeek’s Chinese origins. While its open-source ethos is widely praised, some worry about regulatory constraints and the impact of Chinese censorship on global adoption.
Business Model and Partnerships
DeepSeek’s funding strategy is unlike most AI startups. The company is financed entirely by High-Flyer, a successful quantitative hedge fund founded by DeepSeek founder Liang Wenfeng. This unique arrangement allows DeepSeek to operate without the pressures of shareholder demands or meeting aggressive Series A milestones.
Freed from the typical constraints of venture-backed startups, DeepSeek can prioritize long-term research and innovation over immediate commercialization. So far, the company has shown no urgency to pursue large-scale commercial opportunities, instead focusing on refining its AI models and driving innovation.
One of DeepSeek’s standout features is its incredibly low API pricing, making advanced AI far more accessible. R1 launched at $0.55 input / $2.19 output per million tokens in 2025, and the current models are cheaper still: V4 Flash at $0.14 / $0.28 und V4 Pro at $0.435 / $0.87, with cached input reads at $0.0028 und $0.003625. Note that the R1 rate is historical, since its endpoint no longer exists. These are rates significantly cheaper than offerings from OpenAI or other American AI labs. This affordability has helped DeepSeek carve out a niche among cost-conscious developers, startupsund small businesses who might otherwise struggle to afford cutting-edge AI tools. By offering such budget-friendly solutions, DeepSeek has positioned itself as a viable alternative to more expensive, proprietary platforms.
DeepSeek’s partnership with AMD has also played a critical role in its success. By utilizing AMD Instinct GPUs und open-source ROCM software, DeepSeek has been able to train its models, including V3 und R1, at remarkably low costs. This collaboration challenges the industry’s reliance on NVIDIA’s high-end GPUs or Google’s TPUs, proving that efficient training doesn’t require access to the most expensive hardware. The partnership is a testament to DeepSeek’s focus on cost-effective innovation and its ability to leverage strategic collaborations to overcome hardware limitations.
Together, these factors underscore DeepSeek’s ability to balance affordability, technical excellence, and independence, allowing it to compete effectively with larger, better-funded competitors while keeping accessibility at the forefront.
Competitive Landscape
DeepSeek has positioned itself as a disruptor in the AI market, taking on both the world’s largest American AI labs und China’s tech giants.
Taking on OpenAI, Google, and Meta
OpenAI, Google, and Meta boast vast resources, established reputations, and access to some of the world’s top AI talent. These companies operate on billion-dollar budgets, allowing them to invest heavily in hardware, research, and marketing. DeepSeek, in contrast, adopts a more targeted approach, focusing on open-source innovation, longer context windowsund dramatically lower usage costs.
DeepSeek’s models, like R1, deliver comparable or superior performance in specific areas like math and reasoning tasks, often at a fraction of the cost. This makes DeepSeek an appealing alternative for organizations that find proprietary AI tools overly expensive or restrictive. By emphasizing accessibility and transparency, DeepSeek challenges the narrative that only big-budget players can deliver state-of-the-art AI solutions.
Disrupting China’s Tech Giants
DeepSeek’s rise has also disrupted Chinese tech leaders such as ByteDance, Tencent, Baiduund Alibaba. These companies are deeply entrenched in China’s AI ecosystem, often backed by state-level computing resources. However, DeepSeek’s open-source philosophy und aggressive pricing strategy have allowed it to carve out a unique niche. By providing cost-effective and efficient models, DeepSeek has forced these firms to reevaluate their own pricing and development strategies.
DeepSeek’s ability to compete with these heavily funded giants underscores its status as a formidable challenger both within China and on the global stage.
The Open R1 Initiative
One testament to DeepSeek’s growing influence is Hugging Face’s Open R1 initiative, an ambitious project aiming to replicate the full DeepSeek R1 training pipeline. If successful, this initiative could enable researchers around the world to adapt and refine R1-like models, further accelerating innovation in the AI space.
While this highlights the impact of DeepSeek’s open-source strategy, it also exposes potential vulnerabilities. By making its models open to the AI community, DeepSeek invites competition from those building on its breakthroughs. However, this openness is a deliberate move to democratize AI development and foster collaboration, a philosophy that sets DeepSeek apart from more proprietary-focused players.
Through its disruptive pricing, open-source commitment, and competitive capabilities, DeepSeek has managed to thrive in a market dominated by tech giants, proving that innovation and efficiency can rival even the largest budgets.
What’s Next for DeepSeek
DeepSeek’s rapid rise comes with challenges that could shape its future. U.S. export controls still shape what hardware it can buy, though the picture is no longer a blanket ban: through 2026 the Commerce Department cleared H200-class chips for a short list of approved Chinese buyers under per-customer volume caps, while the top Blackwell tier still requires export licences. That leaves a compute gap at the frontier rather than across the board. DeepSeek’s answer has been architectural, and V4-Flash-0731 is the clearest evidence it works: a 13B-active model that outperforms a 49B-active one on agentic tasks.
DeepSeek also faces hurdles in market perception. To gain international trust, it must consistently prove its reliability, especially for enterprise-grade deployments. Meanwhile, the fast-evolving AI landscape means competitors like OpenAI or Meta could outpace it with new innovations. Additionally, operating under Chinese regulatory frameworks imposes content restrictions that may limit its appeal in open markets.
Despite these challenges, DeepSeek’s focus on its DeepThink + Web Search feature, which enables real-time lookups, is positioning it as a unique competitor. The company could also enhance reinforcement learning fine-tuning, develop industry-specific models, and forge new global partnerships to expand its capabilities. If it can navigate these obstacles, DeepSeek has the potential to remain a disruptive force in AI.
Abschließende Überlegungen
In just a few short years, DeepSeek has gone from being an unknown research-driven startup in Hangzhou to a global disruptor in AI, shaking up industry giants like OpenAI, Meta, and Google. By combining open-source collaboration, innovative architectures like Mixture-of-Experts (MoE), and fiercely competitive pricing, DeepSeek has redefined how we think about AI development. Models like DeepSeek V3 and the groundbreaking DeepSeek R1 prove that success in AI doesn’t always require billion-dollar budgets. Instead, efficiency, adaptability, and strategic partnerships can deliver results that rival even the most expensive models.
What makes DeepSeek’s journey even more extraordinary is the sheer shock it has generated within the AI community. Industry experts and researchers have been vocal about their amazement at how a smaller player has managed to compete with, and even outperform, some of the most advanced models developed by vastly better-funded organizations.
DeepSeek is showing no signs of slowing down. Its recent launch of DeepThink + Web Search, which enables real-time online lookups, places it ahead of even OpenAI in some capabilities. Looking forward, the company is likely to focus on:
- Refining reinforcement learning pipelines to further enhance reasoning capabilities.
- Developing industry-specific models tailored for fields like healthcare, finance, and education.
- Forging new partnerships with global hardware providers to overcome the compute gap created by export restrictions.
As user adoption of DeepSeek R1 continues to soar, the company is forcing established AI players to adapt. It has proven that efficiency and innovation can rival raw computational power and immense budgets, setting a new precedent for what’s possible in AI.
Whether DeepSeek can sustain this momentum amid challenges like geopolitical restrictions, intense competitionund market trust issues remains to be seen. However, one thing is clear: DeepSeek has already proven itself as a force to be reckoned with, pushing the boundaries of AI while empowering smaller businesses, researchers, and developers around the globe.
For anyone tracking how far low-cost innovation can reshape AI workflows, DeepSeek is a name worth watching. Start with our DeepSeek V4 coverage for the current models, or the DeepSeek pricing guide if you are costing out an API workload.
Häufig gestellte Fragen
What is DeepSeek’s newest model?
DeepSeek V4-Flash-0731, released on July 31, 2026. It is a re-post-train of V4 Flash and now serves every deepseek-v4-flash API call. The other current model is V4 Pro, the larger reasoning model. Those are the only two model IDs DeepSeek’s docs list.
Is DeepSeek free?
The web chat at chat.deepseek.com is free with no subscription tiers. The API is paid, billed per token against a balance you top up, at $0.14 / $0.28 per million tokens for V4 Flash. The model weights are also free to download and run yourself under the MIT license.
Is DeepSeek R1 still available?
Not as an API model. The deepseek-reasoner und deepseek-chat endpoints were retired after 24 July 2026. The open weights for R1 remain downloadable on Hugging Face, so you can still run it yourself, but new API work should target V4 Flash or V4 Pro.
Who owns DeepSeek?
DeepSeek is funded entirely by High-Flyer, the quantitative hedge fund founded by DeepSeek’s own founder Liang Wenfeng. There are no outside investors, which is why the company can price aggressively and publish open weights without answering to shareholders.
Is DeepSeek safe to use?
The app routes data to servers in China, where it falls under Chinese law, and the models enforce Chinese content moderation on political topics. For non-sensitive personal use that is acceptable to many people; for work data it is a real risk. Running the open weights locally avoids both issues. See our full safety analysis.




