Cointime

Download App
iOS & Android

Levels of the game: the psychology of RetroPGF and how to build a better game

RetroPGF is a unique kind of repeated game. With each round, we are iterating on both the rules and the composition of players. These things matter a lot. To get better, we need to study whether the rules and player dynamics are having the intended effect.

This post looks at the psychology of the game during Round 3, identifies mechanics that might have caused us to deviate from our intended strategy, and suggests ways of mitigating such issues in the future.

Disclaimer: I was a voter and had a project in Round 3. I also made a lot of Lists.

Learning to outperform “spray and pray”

It is impractical to thoroughly review more than 100 projects, let alone 600+.

Here’s what Greg did:

Spray and pray is a natural reaction to cognitive overload and limited bandwidth.

However, our focus shouldn't be solely on creating more efficient tools for spraying. A new feature like a CSV upload button would make the work go faster, but it still encourages us to spray.

What we actually need are better ways of designing, iterating, and submitting complete voting strategies.

A voting strategy is a way of testing a hypothesis that funding certain types of projects (X) in the past will result in more desirable outcomes (Y) in the future. We want to use our voting power to incentivize agents to work in areas they know will be rewarded.

There are at least four components to designing a voting strategy:

  • Impact Vector: Identifying the type of impact you want to amplify over time.
  • Distribution Curve: Determining how flat or skewed the eventual distribution of tokens to projects should be.
  • Eligibility Criteria: Establishing parameters to determine which projects qualify.
  • Award Function: Creating a formula or rubric to equate impact with profit.

Currently, badgeholders don’t have much scaffolding to build voting strategies. It’s unrealistic to expect badgeholders to do this work from scratch.

The initial step of defining the impact vector is often the most challenging. To assist in this, we should provide badgeholders with comparative choices to help them identify their preferences. Examples could include:

  • Focusing on attracting new developers versus enhancing the experience for existing ones.
  • Prioritizing virtual community-building efforts versus more in-person events.
  • Concentrating on impacts specific to the Superchain as opposed to impacts affecting the entirety of Ethereum.

Performing these kinds of assessments at the beginning or even during the planning stage of the next RetroPGF round would not only aid badgeholders in formulating their voting strategies but also ensure alignment with the broader intents of the Citizen’s House.

Giving domain experts a platform

Popularity and name recognition are seldom reliable measures of true impact.

When we look for a restaurant on Google or Yelp, we usually don’t want the one with the most reviews. Instead we look for restaurants with a high proportion of detailed five-star ratings. In RetroPGF, the List feature aimed to highlight impactful projects for voters to add to their ballots. However, most Lists showed a conservative distribution of tokens, with little variance between top and median recommendations. This is the equivalent of going onto Yelp and only finding average three-star reviews.

Most Lists were radically un-opinionated

Katie Garcia made a similar point on the forum about low-quality Lists being a drag on the process.

To get stronger opinions, we need to encourage experts to step up and make Lists.

As with any independent review, the expertise, affiliation, and reputation of the List maker is just as important as the List itself. Voters should be able to quickly discern the credibility of a List creator.

While "liking" a List provided some measure of quality, it was unclear how many badgeholders actively utilized this feature. Another problem was that the likes only accrued to the List, not the List maker. Because we could not edit a List, a V2 of the same List started off at zero likes.

Finally, most Lists were generated towards the end of the round, giving the creator very little time to develop a reputation and voters little time to determine if it was useful to their voting strategy.

Most Lists were generated towards the end of the round

Experts should be encouraged to play a larger role in reviewing projects and crafting Lists within their areas of expertise. This includes both helping categorize projects effectively at the start of the round and offering ratings across the full quality spectrum.

Experts should be evaluating projects on multiple impact dimensions, not just issuing absolute scores. For example, a List focused on user growth will have different project rankings than one centered on security. It’s common to see product reviews that consider multiple features and then weight those features to arrive at a final score.

Experts need to be placed in a position where they don’t feel pressure to conform to social pressure. The same is true for whisteblowers. It’s hard to combat the herd mentality that forms in large groups with only certain types of people speaking.

In academia, there's an awareness that today's reviewer could become tomorrow's reviewee. To address this awkwardness, academics use blind reviews. The identities of the reviewers are kept hidden from the participants. A double-blind review extends this anonymity further, concealing the identities of both parties involved. For RetroPGF, implementing blind reviews could mean adopting pseudonymous reviews and List creation. A double-blind system, while more complex, could be realized through the use of standardized impact reports. Another possibility is to organize reviewers so that they are grouped in a domain where there’s no potential for conflict of interest.

While expert-driven systems can be criticized as technocratic, Optimism can address these concerns by promoting transparent and replicable project metrics, supporting its community of grassroots analysts (e.g., Numba Nerds), and requiring experts to publish their evaluation criteria along with their recommendations.

Using data to improve decision-making

The site growthepie played a vital role during the voting period by providing easily accessible and filterable data about projects, such as their presence on ballots/Lists and funding history. This addition, as well as other sites like RetroList, RetroPGFHub, and hopefully our own, Open Source Observer, gave voters more data points to assist with voting.

Projects like growthepie helped surface data to voters

Across the board, voters demonstrated a strong preference for having data in the loop when voting or coming up with their strategies.

As we saw during the round, the best time to request relevant data from projects is during the application phase. However, while projects submitted numerous links to their contributions and impact metrics, the numbers were all self-reported. The nature of self-reported metrics raises concerns about accuracy and data entry errors.

Moreover, the data was cumbersome to analyze without individually examining each project's page. There was no way of comparing similar metrics about similar projects unless you cleaned the data and did the analysis on your own.

Although some badgeholders are committed to conducting their own research, the process of gathering and synthesizing information from various public sources is time-consuming. Streamlining access to comparable impact metrics would significantly aid in the efficient filtering and ranking of projects.

A better approach than self-reporting would be for projects to directly link all relevant work artifacts, such as GitHub repositories, in their applications. This would enable comparable impact metrics to be surfaced automatically from any source where data integration is possible. After linking their artifacts, perhaps projects could get the option to include different “impact widgets” on their profile to highlight the ones that are most relevant to assessing their project.

Data doesn’t tell the whole story. And not all forms of impact are well-suited to some kind of metrics integration. Those are places where tools for making relative comparisons can really shine.

Encouraging relative comparisons over absolute allocations

Another thing we learned from all our “three-star Lists” is how difficult it is to quantify impact in absolute terms. It’s much easier to assess impact in relative terms.

This is what algorithms like pairwise matching are good for. I experimented with the Pairwise app during RetroPGF, but found it difficult given the sheer number of projects and the nature of comparing two projects objectively. A more targeted approach, asking voters to make a more specific comparison, e.g, “do you think Project A or B had more impact on X”, could be more effective.

Another twist on pairwise is to rate a project's recent impact against its past performance. This is reminiscent of the traditional budgeting process of setting this year’s budget in relation to last year’s, only now we’re doing it retroactively and from the vantage point of a community not a CFO.

There’s also setwise comparison, where we analyze how often a project appears in different sets. The most impactful governance project might be the one that appears on a large number of lists independently. Quadratic voting harnesses the power of setwise comparisons. However, effort needs to be taken to prevent sets that are purely based on popularity or name recognition.

In the final 48 hours of the vote, we saw a good example of setwise comparison. Badgeholders made a push to identify borderline projects that needed just a few more votes to meet the quorum.

When looking at a borderline project, the voter just had to make one decision: do I want this project to be in the above or below quorum set.

Yet, as Andrew Coathup noted, this well-intentioned approach might have inadvertently skewed our true preferences:

Badgeholders who "like" a project rather than "love" an impactful project are likely to bring down a project’s median.

In retrospect, it would have been more effective to focus on removing bad apples at the start rather than searching for good apples at the end.

Understanding how game mechanics affect distributions

The experience above is just one of numerous examples of how the rules of the game will affect the distribution.

If you give the same voters the same projects, but vary the rules, you’ll almost certainly get different results. It’s not clear which type of curve the Collective wants to design for.

The results from RetroPGF 2 were pretty exponential. The top project received 20 times more tokens than the median project. Voters didn’t seem upset about this outcome. A comparison with RetroPGF 3’s outcomes and the community’s reaction will be insightful, and may surface if there are strong preferences for one form of distribution pattern over another.

Results from RetroPGF 2 were pretty exponential

Exponential distributions are generally better for tackling complex problems with low odds of success. They channel more resources into fewer, higher impact projects. Conversely, flatter distributions encourage a broader range of teams to pursue more achievable objectives. Both have their place.

The other big distribution question is how meritocratic it is. A meritocratic system would consistently elevate the highest-impact projects to the top of the token distribution, independent of voter identity or the number of competing projects. If the distribution has a long tail, then there will always be a high degree of randomness and variation in it.

The distribution patterns that emerge through repeated games will inevitably affect the mindset of projects and builders who continue to participate. If the model feels like a black box, then the best players will lose the motivation to keep playing.

Given the significance of these outcomes, it’s crucial to simulate and understand the implications of different rule sets before implementing them in upcoming rounds.

Navigating questions around prior funding sources

There was a lot of debate within the badgeholder group and on Twitter about how a project’s funding history and business model should affect its allocation. Many community members including but not limited to Lefteris had strong views.

Like most political issues, once you dig deeper, you realize there’s a lot of nuance. People’s opinions fall along a spectrum. We can get a sense for where the community is on this spectrum by giving them hypothetical comparisons like the following:

  1. Should a ten-person team receive more funding than a two-person team?
  2. Does a full-time project deserve more support than a part-time effort or side hustle?
  3. Should a team in a higher GDP region receive more funding than one in a lower GDP area?
  4. Does the absence of venture capital funding merit more allocation compared to projects with such backing?
  5. Should teams with limited financial resources be prioritized over those with substantial funds?
  6. Does a team generating no revenue deserve more support than one with recurring revenue?
  7. Should projects offering solely public goods receive more funding than those mixing public and non-public goods?
  8. Do teams without prior grants merit more support than those frequently receiving grants?
  9. Should a project never funded by Optimism be prioritized over one that has received multiple grants?
  10. Should a protocol on Optimism with no fees receive more funding than one with fees?

Currently, guidelines state that voters should only consider points 9 and 10. Yet many voters struggle to disregard the other aspects.

To prevent these financial factors from skewing voting, the Collective might need to reevaluate its approach. This type of challenge is another reason why academics use double-blind reviews: the reviewer is forced to only consider the quality of the paper, not how deserving is the team behind the paper.

A potential solution for RetroPGF could involve segregating projects and funding pools. This could alleviate some of the pressure on voters to choose between projects with vastly different funding requirements and access to capital.

While this issue is important, it's regrettable that the focus on reviewing a project's funding sources may have gotten in the way of reviewing its actual impact.

Playing the game better

I’ll end with a recap of my recommendations for how Optimism might get better at playing the game both next round and over the long run.

Our primary goal should be to surpass the 'spray and pray' approach. To achieve this, we need to dig deeper into understanding voters' value preferences and let this knowledge shape the game's dynamics. This encompasses everything from the types of impact voters care about to the financial aspects they consider relevant in reviewing projects. We can start by gathering insights from surveys, forum discussions, and post-round retrospectives. However, acquiring more substantial data before and during the round is essential. This data will be instrumental in differentiating between UI/UX challenges (like updating ballots and lists) and game mechanics issues (such as whether voters should review every project). Critically, we should monitor and simulate how closely the actual outcomes align with voters' expectations regarding funding allocation and distribution patterns.

Improving project comparison methods, reducing bias, and encouraging independent thinking are crucial. We must avoid turning this into a popularity contest by empowering domain experts to have greater influence and build their reputation. Since quantifying absolute impact is challenging, we should focus on better understanding relative impacts and making more effective comparisons. To combat herd mentality and the pressure to conform, exploring blind review mechanisms could be beneficial. Having objective data readily available will assist in making informed decisions and identifying both the strongest and weakest projects.

The voting experience should be optimized to support voters in testing and implementing a well-defined strategy. At the end of each round, voters should have a clear understanding of the process and feel confident about their participation and choices. Speaking from personal experience, I am uncertain about my own effectiveness in the game. Despite investing time in identifying worthy projects and avoiding less desirable ones, I'm unsure if I maximized the potential of my top picks. There's a need for a system that helps voters articulate and evaluate their strategies clearly, as this is key to getting better at a repeated game.

This work won’t be easy. But the upside is huge.

Comments

All Comments

Recommended for you

  • USD/JPY Breaks Above 163 Level

    On July 21, USD/JPY broke above the 163 level, reaching a new high since 1986, up 0.34% on the day.

  • Moonshot AI Reportedly in Pre-IPO Funding Talks at $50 Billion Valuation

    July 21, market news: Moonshot AI is reportedly in Pre-IPO funding talks at a $50 billion valuation.

  • Crypto Bank Augustus Completes $180 Million Funding Round, Led by Tiger Global

    On July 21, Augustus, a startup building a federally chartered clearing bank, announced it has raised $180 million to expand its dollar payment infrastructure as stablecoins reshape the global financial system. The funding round values Augustus at $1 billion. The round was led by Tiger Global Management, with participation from Hummingbird Ventures, QED Investors, and founders of Nubank, Ramp, Circle, and Deel, among other investors.

  • U.S. Trade Representative Greer Says U.S. Is Preparing New Tariffs

    July 21, according to the Wall Street Journal: U.S. Trade Representative Greer said the U.S. is preparing new tariffs.

  • US-listed crypto concept stocks surge; Circle jumps over 10%

    On July 21, Bitcoin returned to above $66,000, and US-listed crypto concept stocks collectively surged. Circle jumped over 10%, Coinbase rose over 9%, Robinhood gained over 6%, TeraWulf and Strategy increased over 4%, while Riot Platforms and CleanSpark rose over 2%.

  • Kalshi Applies to CFTC for Gold-Linked Perpetual Futures

    On July 21, prediction market platform Kalshi submitted an application to the U.S. Commodity Futures Trading Commission (CFTC) to launch gold-linked perpetual futures. (Jin Shi)

  • Official Meeting Between Interior Ministers of Iran and Pakistan Has Begun

    On July 21, according to the Iranian Students' News Agency, the official meeting between the Interior Minister of Iran and the Interior Minister of Pakistan began a few minutes ago.

  • Beijing: In-depth Implementation of 'AI+' Action Plan in the Second Half of the Year with Special Support Policies for Embodied Intelligence Enterprises

    On July 21, according to the Beijing News, the Beijing Municipal Bureau of Economy and Information Technology announced that in the second half of the year, it will focus on the annual key tasks of building a benchmark city for the digital economy and deeply implement the 'AI+' action plan to fully unleash new momentum for the intelligent economy. The plan aims to advance the comprehensive empowerment of artificial intelligence. It will promote AI-enabled new industrialization, establish R&D and pilot platforms for AI in the industrial sector, and enhance the application level of AI in core industrial production processes. Special support policies for computing power and datasets for embodied intelligence enterprises will be introduced to accelerate the iteration of core technologies such as embodied large models and motion control, leveraging embodied intelligence to drive industrial transformation and improve quality of life. The construction of a national (medical) AI application pilot base will be expedited, linking top-tier hospitals, research institutions, and technology companies to address the bottlenecks in the translation of medical intelligence from the laboratory to clinical application. The application of intelligent translation and simultaneous interpretation in the cultural and tourism consumption sectors will be promoted, expanding the range of languages and achieving hardware-software synergy. Additionally, a non-site intelligent supervision system for food safety will be improved to form a complete regulatory chain of intelligent early warning, remote inspection, and rapid response.

  • Beijing to Establish Token Factories in the Second Half of the Year

    On July 21, the Beijing Municipal Bureau of Economy and Information Technology reported that in the first half of the year, the city's digital economy grew by 7.8%, with the core industries of the digital economy increasing by 9.8%, significantly contributing to the city's GDP growth. In the second half of the year, Beijing will promote comprehensive empowerment through artificial intelligence, advancing applications such as intelligent translation and simultaneous interpretation in the cultural and tourism consumption sectors. The Bureau will focus on the annual key tasks for building a benchmark city for the digital economy, implementing the 'Artificial Intelligence +' action plan, and fully unleashing new momentum for the intelligent economy. Policies for the development of the Token economy will be formulated, focusing on key aspects such as Token production, distribution, and application, with plans to establish Token factories and distribution platforms, promoting innovative applications in key areas such as industry, education, and cultural tourism. An 'Innovate for the Future' OPC special roadshow event will be held to stimulate the creative energy of super individuals in AI applications. High-quality training courses on AIGC aimed at OPC will be developed to unlock the value of AIGC technology. Leveraging the Open Source Chip Research Institute and the Beijing Tongminghu Information Technology Application Innovation Center, a 'RISC-V + AI OS' open-source ecosystem will be created, building a full-stack autonomous technology system from chip instruction sets to operating systems to intelligent applications. (Xinjingbao)

  • Changxin Technology: Online Investors Abandon Subscription of 6,586,227 Shares

    On July 21, Changxin Technology (688825.SH) announced the results of its initial public offering on the STAR Market. The issue price was 8.66 yuan per share, with an initial offering of 6,688,088,608 shares, accounting for approximately 10% of the total post-issuance share capital. The final strategic placement was 1,667,064,720 shares, and the final online issuance lot-winning rate was approximately 0.47141739%. Online investors subscribed and paid for 3,844,517,273 shares and abandoned 6,586,227 shares; offline investors subscribed and paid for 2,173,101,821 shares and abandoned 31,567 shares. The joint lead underwriters underwrote 6,617,794 shares, with an underwriting amount of 57.3101 million yuan. Before the exercise of the over-allotment option, the issuance expenses were 281 million yuan.