Cointime

Download App
iOS & Android

The DeepSeek-R1 Effect and Web3-AI

Cointime Official

From coindesk By Jesus Rodriguez|Edited by Benjamin Schiller

The artificial intelligence (AI) world was taken by storm a few days ago with the release of DeepSeek-R1, an open-source reasoning model that matches the performance of top foundation models while claiming to have been built using a remarkably low training budget and novel post-training techniques. The release of DeepSeek-R1 not only challenged the conventional wisdom surrounding the scaling laws of foundation models – which traditionally favor massive training budgets – but did so in the most active area of research in the field: reasoning.

The open-weights (as opposed to open-source) nature of the release made the model readily accessible to the AI community, leading to a surge of clones within hours. Moreover, DeepSeek-R1 left its mark on the ongoing AI race between China and the United States, reinforcing what has been increasingly evident: Chinese models are of exceptionally high quality and fully capable of driving innovation with original ideas.

The artificial intelligence (AI) world was taken by storm a few days ago with the release of DeepSeek-R1, an open-source reasoning model that matches the performance of top foundation models while claiming to have been built using a remarkably low training budget and novel post-training techniques. The release of DeepSeek-R1 not only challenged the conventional wisdom surrounding the scaling laws of foundation models – which traditionally favor massive training budgets – but did so in the most active area of research in the field: reasoning.

The open-weights (as opposed to open-source) nature of the release made the model readily accessible to the AI community, leading to a surge of clones within hours. Moreover, DeepSeek-R1 left its mark on the ongoing AI race between China and the United States, reinforcing what has been increasingly evident: Chinese models are of exceptionally high quality and fully capable of driving innovation with original ideas.

Inside DeepSeek-R1

DeepSeek-R1 was the result of introducing incremental innovations into a well-established pretraining framework for foundation models. In broad terms, DeepSeek-R1 follows the same training methodology as most high-profile foundation models. This approach consists of three key steps:

  1. Pretraining: The model is initially pretrained to predict the next word using massive amounts of unlabeled data.
  2. Supervised Fine-Tuning (SFT): This step optimizes the model in two critical areas: following instructions and answering questions.
  3. Alignment with Human Preferences: A final fine-tuning phase is conducted to align the model’s responses with human preferences.

Most major foundation models – including those developed by OpenAI, Google, and Anthropic – adhere to this same general process. At a high level, DeepSeek-R1’s training procedure does not appear significantly different. ButHowever, rather than pretraining a base model from scratch, R1 leveraged the base model of its predecessor, DeepSeek-v3-base, which boasts an impressive 617 billion parameters.

In essence, DeepSeek-R1 is the result of applying SFT to DeepSeek-v3-base with a large-scale reasoning dataset. The real innovation lies in the construction of these reasoning datasets, which are notoriously difficult to build.

First Step: DeepSeek-R1-Zero

One of the most important aspects of DeepSeek-R1 is that the process did not produce just a single model but two. Perhaps the most significant innovation of DeepSeek-R1 was the creation of an intermediate model called R1-Zero, which is specialized in reasoning tasks. This model was trained almost entirely using reinforcement learning, with minimal reliance on labeled data.

Reinforcement learning is a technique in which a model is rewarded for generating correct answers, enabling it to generalize knowledge over time.

R1-Zero is quite impressive, as it was able to match GPT-o1 in reasoning tasks. However, the model struggled with more general tasks such as question-answering and readability. That said, the purpose of R1-Zero was never to create a generalist model but rather to demonstrate it is possible to achieve state-of-the-art reasoning capabilities using reinforcement learning alone – even if the model does not perform well in other areas.

Second-Step: DeepSeek-R1

DeepSeek-R1 was designed to be a general-purpose model that excels at reasoning, meaning it needed to outperform R1-Zero. To achieve this, DeepSeek started once again with its v3 model, but this time, it fine-tuned it on a small reasoning dataset.

As mentioned earlier, reasoning datasets are difficult to produce. This is where R1-Zero played a crucial role. The intermediate model was used to generate a synthetic reasoning dataset, which was then used to fine-tune DeepSeek v3. This process resulted in another intermediate reasoning model, which was subsequently put through an extensive reinforcement learning phase using a dataset of 600,000 samples, also generated by R1-Zero. The final outcome of this process was DeepSeek-R1.

While I have omitted several technical details of the R1 pretraining process, here are the two main takeaways:

  1. R1-Zero demonstrated that it is possible to develop sophisticated reasoning capabilities using basic reinforcement learning. Although R1-Zero was not a strong generalist model, it successfully generated the reasoning data necessary for R1.
  2. R1 expanded the traditional pretraining pipeline used by most foundation models by incorporating R1-Zero into the process. Additionally, it leveraged a significant amount of synthetic reasoning data generated by R1-Zero.

As a result, DeepSeek-R1 emerged as a model that matched the reasoning capabilities of GPT-o1 while being built using a simpler and likely significantly cheaper pretraining process.

Everyone agrees that R1 marks an important milestone in the history of generative AI, one that is likely to reshape the way foundation models are developed. When it comes to Web3, it will be interesting to explore how R1 influences the evolving landscape of Web3-AI.

DeepSeek-R1 and Web3-AI

Until now, Web3 has struggled to establish compelling use cases that clearly add value to the creation and utilization of foundation models. To some extent, the traditional workflow for pretraining foundation models appears to be the antithesis of Web3 architectures. However, despite being in its early stages, the release of DeepSeek-R1 has highlighted several opportunities that could naturally align with Web3-AI architectures.

1) Reinforcement Learning Fine-Tuning Networks

  1. R1-Zero demonstrated that it is possible to develop sophisticated reasoning capabilities using basic reinforcement learning. Although R1-Zero was not a strong generalist model, it successfully generated the reasoning data necessary for R1.
  2. R1 expanded the traditional pretraining pipeline used by most foundation models by incorporating R1-Zero into the process. Additionally, it leveraged a significant amount of synthetic reasoning data generated by R1-Zero.

As a result, DeepSeek-R1 emerged as a model that matched the reasoning capabilities of GPT-o1 while being built using a simpler and likely significantly cheaper pretraining process.

Everyone agrees that R1 marks an important milestone in the history of generative AI, one that is likely to reshape the way foundation models are developed. When it comes to Web3, it will be interesting to explore how R1 influences the evolving landscape of Web3-AI.

DeepSeek-R1 and Web3-AI

Until now, Web3 has struggled to establish compelling use cases that clearly add value to the creation and utilization of foundation models. To some extent, the traditional workflow for pretraining foundation models appears to be the antithesis of Web3 architectures. However, despite being in its early stages, the release of DeepSeek-R1 has highlighted several opportunities that could naturally align with Web3-AI architectures.

1) Reinforcement Learning Fine-Tuning Networks

4) Reasoning Data Provenance

One of the defining features of reasoning models is their ability to generate reasoning traces for a given task. DeepSeek-R1 makes these traces available as part of its inference output, reinforcing the importance of provenance and traceability for reasoning tasks. The internet today primarily operates on outputs, with little visibility into the intermediate steps that lead to those results. Web3 presents an opportunity to track and verify each reasoning step, potentially creating a "new internet of reasoning" where transparency and verifiability become the norm.

Web3-AI Has a Chance in the Post-R1 Reasoning Era

The release of DeepSeek-R1 has marked a turning point in the evolution of generative AI. By combining clever innovations with established pretraining paradigms, it has challenged traditional AI workflows and opened a new era in reasoning-focused AI. Unlike many previous foundation models, DeepSeek-R1 introduces elements that bring generative AI closer to Web3.

Key aspects of R1 – synthetic reasoning datasets, more parallelizable training and the growing need for traceability – align naturally with Web3 principles. While Web3-AI has struggled to gain meaningful traction, this new post-R1 reasoning era may present the best opportunity yet for Web3 to play a more significant role in the future of AI.

Note: The views expressed in this column are those of the author and do not necessarily reflect those of CoinDesk, Inc. or its owners and affiliates.

Comments

All Comments

Recommended for you

  • Robinhood Launches Perpetual Contracts for U.S. Traders

    On September 30, Robinhood launched perpetual contracts (Perps) for U.S. traders. These contracts have no expiration date, support round-the-clock trading, and allow users to customize leverage up to 10:1. With this product launch, Robinhood becomes the first mainstream brokerage to offer trading services across all six major asset classes on a single platform, including stocks, options, cryptocurrencies, futures, prediction markets, and perpetual contracts.

  • BTC Briefly Drops Below $83,000

    Market data shows that BTC briefly fell below $83,000, currently priced at $83,010.77, with a 24-hour decline of 0.01%. The market is experiencing significant volatility, so please ensure proper risk management.

  • Trump: Iran is Facing Severe Defeat, Oil Prices Will Drop Significantly

    On September 29, U.S. President Trump stated that Iran is facing severe defeat and will soon come to an end. Oil prices will drop significantly.

  • Trump: To Sign Strong Executive Order on Artificial Intelligence

    U.S. President Trump: To sign a strong executive order on artificial intelligence. (Jin Shi)

  • SEC Issues No-Action Letter to Tesla, Clearing Path for Retail Shareholder Voting Plan

    On September 29, Vlad Tenev, co-founder of Robinhood, announced that the U.S. SEC issued a no-action letter to Tesla on the same day, removing obstacles for its plan to empower retail investors in shareholder voting. Tenev stated that when millions of individuals hold shares in a public company, they should be able to exercise their voting rights more conveniently. His team is collaborating with Tesla to implement the project and has already provided features within the app for shareholder information access, retail voting, and real-time Q&A during earnings calls, with more tools to be launched in the future.

  • Trump Administration Launches AI Tool to Create a 'Digital Steward' for Federal Government

    On September 29, The Wall Street Journal reported that the Trump administration has launched an AI-driven tool, which developers claim will serve as a 'digital steward' for the federal government, aimed at modernizing the relationship between citizens and the government. The tool was deployed on the America.gov website, which officially went live on Tuesday, with the goal of becoming a unified entry point for all federal public services. Joe Gebbia, head of the White House National Design Studio and co-founder of Airbnb, stated that this portal will set the 'standards that a superpower should have.'

  • OpenAI's Annual Recurring Revenue Approaches $70 Billion

    On September 29, Axios reported that OpenAI's annual recurring revenue is nearing $70 billion.

  • U.S. 30-Year Treasury Yield Reaches Highest Level Since May 2004

    On September 29, the yield on U.S. 30-year Treasury bonds reached 5.587%, the highest level since May 2004.

  • Bitwise NEAR ETF (NRR) Launches on US Stock Market as First Spot NEAR ETP

    On September 29, Bitwise announced that the cryptocurrency asset management firm Bitwise (with client assets of $9 billion) has listed the Bitwise NEAR ETF (ticker NRR) on the NYSE Arca, making it the first spot NEAR exchange-traded product (ETP) in the United States. The fund began trading on September 29, 2026, with a management fee of 0.75%, providing exposure to spot NEAR and plans to stake the NEAR tokens held by the fund. NEAR serves as a blockchain infrastructure for the AI economy, with block times of approximately 1.2 seconds and transaction costs of less than one cent. Bitwise Chief Investment Officer Matt Hougan stated that NEAR is at the intersection of two major technological trends: AI and cryptocurrency.

  • Nasdaq Opens Up 0.33%, Meta and AMD Rise Over 1%

    On September 29, U.S. stock markets opened with mixed results. The Nasdaq rose by 0.33%, the S&P 500 increased by 0.14%, while the Dow Jones fell by 0.04%. AMD saw a 1.5% increase after announcing an $8.2 billion acquisition of World Labs, with Fei-Fei Li set to become AMD's Chief Scientist. Meta rose by 1.3%, officially launching a corporate platform to provide AI technology to businesses and developers. SpaceX's stock increased by over 1% as its Starship successfully entered Earth's orbit for the first time, deploying 26 V3 Starlink satellites.