Is Guessing Really Not Judging? A Reply to Eidenmüller and Hochgürtel

Is Guessing Really Not Judging? A Reply to Eidenmüller and Hochgürtel

Is Kleros really judging... or just guessing? Our new Oxford Business Law Blog article takes on one of the oldest critiques of decentralised justice.

By Federico Ast, William George and Facundo Trotz* 

A shorter version of this post was previously published at the Oxford Business Law Blog and can be accessed here.

Horst Eidenmüller and Anna-Sophie Hochgürtel have recently written a post on the Oxford Business Law Blog critiquing decentralised arbitration protocols. In their post, the co-authors argued that decisions produced by these systems reward participants for predicting what the crowd will decide rather than for reasoning toward the most justified answer, and should therefore not be treated as arbitral awards. We are grateful to the authors for bringing Kleros into scholarly conversation and for their thoughtful critique. However, their argument rests on a few assumptions, which in what follows we hope to clarify.

The Focal Point Is a Strategy, Not an Outcome

In a well-known passage from The Strategy of Conflict (1981), Thomas Schelling argued that strangers would usually choose noon at the information booth in Grand Central Station as the most common meeting point in New York City. Nothing about that place and time made it objectively superior to any other; it worked only because everyone expected everyone else to guess it too. This is what game theorists call a focal point. Eidenmüller and Hochgürtel assume that for a Schelling point to work, a winning verdict must stand out the same way. In a genuinely hard dispute, no outcome ever does. But this misidentifies the coordination problem crowdsourced systems ask their evaluators to solve.

Jurors do not vote in a vacuum; rather, they rely on a shared body of information. This includes objective rules governing how to resolve the dispute (“dispute policies”), typically established by the parties beforehand; rules enacted by the Kleros community to structure general reasoning across different case types (“court policies”); and the evidence submitted by the parties. This shared informational framework anchors the juror community in situations where common norms might otherwise be lacking.

For example, Kleros’ collaboration with the Consumer Protection Office of Junín (Argentina) illustrates how jurors rely on objective rules when making decisions. In this setup, Kleros provides internal advisory opinions via certified courts reserved for lawyers holding identity credentials. Under the dispute policy, these accredited professionals must rule and justify their decisions according to Argentine consumer protection law. Every ruling follows this statutory pattern, and all written rationales are publicly verifiable on the blockchain (or as those in the industry say, ‘on-chain’.)

Because every participant expects their peers to read the same record and follow the same rules, the strategy of honest evaluation itself becomes the safest equilibrium for coordination. If jurors follow the strategy of honest evaluation and arrive at different decisions, this is similar to split decisions on traditional appellate panels reflecting genuine adjudication, not a failure of law. Moreover, the Kleros community has an incentive to ensure that the protocol remains a reliable form of dispute resolution. Hence, if conditions evolve so that the equilibrium strategy begins to deviate from honest evaluation in systematic ways, the community will be motivated to update the court policies so that the strategy of "judging" remains the Schelling point.

It seems to us that concerns that Schelling point mechanisms can lead to coordination around “phony focal points” are similar to the criticisms of financial markets expressed by Keynes, where he likened financial markets to newspaper beauty contests where readers attempted to guess which images would be voted prettiest by other readers. Such effects can certainly add noise. However, just as one sees that financial markets maintain a noisy but correlated relationship with economic fundamentals, when well structured, it is possible for Schelling point mechanisms to provide a signal that is meaningfully correlated with the truth. 

What has to be true for the mechanism to work is not that a uniquely salient honest outcome exists, nor that jurors act as completely independent evaluators in a pure Condorcet jury theorem sense. Rather, where a legal or contractual standard is provided, that standard acts as a shared informational anchor. This prevents the mechanism from collapsing into a Keynesian beauty contest: expected convergence on the applicable rule makes honest, standard-applying judgment the focal strategy, even if individual assessments of the facts carry noise.

The authors contrast mechanisms based on Schelling point convergence which “reach conclusions because one outcome attracts more support than another” with traditional legal settings where “the task is to get the answer right”. However, this objection simplifies both the economic incentives of the system and the nature of legal reasoning itself.

Whether a judge is "really" motivated by "delivering justice" or by the reward for matching the majority is hard to answer, even in traditional legal settings. Traditional courts in many jurisdictions attempt to shield their judges as much as possible from experiencing any personal consequences from their decisions, such as by having lifetime appointments. However, as Richard Posner famously argued in his 1993 account of judicial behaviour, even life-tenured judges respond rationally to personal incentives such as reputation, promotion, and prestige. Often these personal incentives depend on the judge's predictions for what their community will expect from their rulings.

The need to consider how the rulings will be viewed by the community is certainly more present in Kleros than in most other forms of dispute resolution, but this is a difference of degree and not of kind. Decentralised justice mechanisms make these incentive structure explicit: jurors stake financial collateral and are rewarded for converging on the correct application of the rules. This alignment is backed by economic risk. A juror who attempts to guess blindly without reading the record faces a direct hazard: if a lazy initial vote contradicts the evidence, the losing party can appeal to a larger panel, leading to the reversal of the decision and the forfeiture (slashing) of the original juror’s stake.

These ideas are not completely alien to arbitration, given the widespread practice of litigation funding. To be sure, third-party litigation funders operate as external actors rather than adjudicators. Yet the multi-billion-dollar litigation funding industry relies on the exact same underlying logic: assigning financial value to the accurate forecasting of how a governing legal standard applies to a set of facts. In litigation funding, “skin in the game” does not turn legal analysis into speculative gambling; it forces funders to conduct rigorous, objective evaluations of the merits. 

Decentralised justice internalises this predictive mechanic within the adjudicative panel itself. Rather than relying on external market actors to bet on a judge’s decision, the mechanism aligns jurors' financial incentives directly with predicting how a reasonable, rule-applying peer panel will resolve the record under threat of appellate review.

In his celebrated 1897 essay The Path of the Law, Oliver Wendell Holmes Jr. captured this predictive essence, observing that "The prophecies of what the courts will do in fact, and nothing more pretentious, are what I mean by the law." Kleros structures this alignment explicitly: the incentive does not reward speculative popularity contests, but accurate foresight of how the applicable standard applies to the record.

Persuading, Not Just Counting

The authors celebrate traditional tribunals for exchanging arguments and deliberating through reasons, contrasting this with decentralised justice systems where "votes are counted rather than arguments weighed." But this argument conflates a particular procedural preference with an essential requirement of arbitration.

Arbitration does not require a multi-member tribunal, let alone intra-panel deliberation. Under the UNCITRAL Model Law and major national arbitration frameworks worldwide, sole-arbitrator tribunals are routine. A sole arbitrator evaluates evidence and renders a binding award without deliberating with anyone. If internal panel deliberation were a prerequisite for a valid arbitral award, decisions by sole arbitrators would fail to qualify, a position clearly incompatible with global arbitration law.

The authors argue that the justifications provided by sole-arbitrator tribunals provide a function that is similar to deliberation. Kleros jurors are also invited to provide justifications with their votes. In the event of an appeal, the appellate jurors see these justifications. As the incentives for the juror in the initial round depend on how the appellate jurors decide, early stage jurors have an incentive to provide justifications that will persuade the appellate evaluators. Indeed, in addition to the parties to the dispute, the jurors themselves can also trigger appeals if they feel that the panel they were on incorrectly ruled against their position.

For example, in case 1170, an individual who had purchased a hedging product that functioned as a form of “decentralised insurance” against smart contract bugs argued that they deserved compensation for improper functioning of a “blockchain bridge”. This case was appealed five times and was ultimately decided by a randomly drawn panel of jurors representing 20 distinct blockchain addresses. This panel was sufficiently large that one can show that it is statistically likely to represent the entire pool of Kleros jurors signed up to resolve this category of disputes. Thus, as each of these appeal rounds saw the justifications of the jurors of previous rounds, one can take the perspective that those justifications contributed to persuading this community in how it ultimately ruled. 

Moreover, within decentralised justice systems, parties submit evidence and arguments to persuade decision-makers, just as a legal brief persuades a sole arbitrator. When a close case receives majority support, that outcome prevails not because vote-counting replaces reasoning, but because one party's arguments persuaded a majority of independent evaluators. Majority voting is simply the mechanical aggregation of multiple independent adjudications, not a departure from the arbitral function.

Is Guessing Different from Judging?

Eidenmüller and Hochgürtel's critique ultimately seems to rely on a strict ontological divide: a decision is either "pure judging" or "pure guessing." But adjudication in the real world does not operate as a binary switch. It exists along a continuum.

At one extreme sits pure guessing: flipping a coin, picking numbers out of a hat, or speculating on a Keynesian beauty contest with zero external reference points. At the opposite extreme sits the idealized archetype of pure judging: Ronald Dworkin’s mythical Judge Hercules, an adjudicator of superhuman intellect and infinite time, capable of discovering the single, perfectly seamless legal answer in every hard case.

Neither extreme describes real-world dispute resolution. To measure decentralised justice against Judge Hercules is to commit a classic Nirvana Fallacy: comparing a practical, real-world mechanism to an impossible philosophical ideal. 

Traditional courts and arbitral tribunals themselves operate squarely in the middle of this spectrum. Indeed, legal systems explicitly acknowledge the impossibility of absolute certainty by relying on probabilistic decision-making heuristics and standards of proof such as "preponderance of the evidence" or "beyond a reasonable doubt". Adjudicators are required not to discover absolute metaphysical truth, but to evaluate whether a claim crosses a given threshold of probability under incomplete information, cognitive biases, and practical (even budgetary) realities.

Kleros simply sits at a different point along that same continuum. When anchored by a clear contractual policy, evidence, and appellate risk, a crowdsourced system operates far from the "pure guessing" pole. It functions as a mechanism of Bayesian convergence: independent evaluators, incentivized to anticipate how a reasonable peer will read a rule, update their probabilities to align on the most plausible interpretation of the record under the governing standard of proof.

Whether this process reaches Hercules's mythical standard is the wrong question. The relevant question is whether it sits close enough to the middle of the spectrum to deliver consistent and cost-effective outcomes for the disputes it is engineered to resolve.

Valid Criticisms and Possibilities for Improvement

Nonetheless, the authors do raise valid concerns. If a dispute genuinely lacks any anchor (no framework of applicable policies and no standards that can be broadly expected to be shared by the community of jurors), then it is true that correlation among honest jurors can be too weak to lean on.

As we have noted above, however, jurors in Kleros are typically provided a framework of dispute policies and court policies on which to base their decisions. We agree that the quality of these policies is fundamental in shaping the equilibria that determine the quality of this dispute resolution process. 

Shared norms and standards can be facilitated by restrictions in the protocol that guarantee that jurors that judge certain types of cases belong to a given community or have certain backgrounds. Such restrictions are not present in first-generation decentralised justice systems. However, this does not reflect an intrinsic limitation of decentralised justice as a mechanism but rather the result of specific design choices that were made in the early days of the field under very specific technological constraints.

First-generation decentralised justice systems, first deployed in the 2010s, were built with the available technologies at the time, relying on a lightweight model: jurors were anonymous, selected primarily by staking capital, and where a single user could have more than one vote. Such a model was necessitated by the lack of a proper identity infrastructure for blockchain applications.

Second-generation infrastructure, such as Kleros 2.0 which is currently deployed in a beta version, is designed to incorporate advances in blockchain identity infrastructure in ways that allow for selection of jurors that places expertise over capital and is more cognisant of the groups and communities to which they belong. Tools such as soulbound identity credentials (building on frameworks proposed by Vitalik Buterin, Glen Weyl, and Puja Ohlhaver) and Sybil-resistant systems like Proof of Humanity allows panels to be drawn based on verified expertise rather than token stakes, tightening correlation even in unanchored disputes.

The increased modularity of second-generation systems also allows for different dispute logics to be used in different categories of cases, even allowing for mechanisms where secondary cases with different juries and different incentives can be created to consider procedural questions about some underlying dispute. For example, a mechanism like the following:

  • Jurors that decide a dispute between parties are explicitly required by the court policies to provide a written rationale for their vote. 
  • These jurors receive rewards regardless of how they vote and are not penalised for votes that disagree with the majority, per se.
  • Rather, if another participant (such as a party to the dispute, an appellate juror, or a third party) reads the juror’s justification and deems it to be inadequate, that participant can place a deposit to “challenge” the justification.  
  • This challenge creates a secondary dispute to judge the quality of the juror’s justification. The jurors in this secondary dispute are incentivized via a Schelling point mechanism, where they are rewarded or penalised based on whether their vote agrees with the majority of jurors, possibly after appeal. 
  • If the secondary dispute rules that the justification was inadequate, the participant who challenged it receives a reward drawn from a deposit lost by the juror who wrote that justification. 

Where would this system fall in the ontological dichotomy between “guessing” and “judging”, if at all? The dynamics of the underlying dispute perhaps seem more familiar in a context of traditional arbitration, and this approach reinforces the incentives of jurors to provide persuasive justifications for their votes. 

However, there are still game-theoretic aspects here as jurors need to anticipate what justifications will be found acceptable by the community. The strategy that should be the focal point goes from “honestly evaluate the case and vote for a corresponding answer” to “honestly evaluate the case and write a correspondingly acceptable justification”. We would argue that this approach would be yet another point on the spectrum between the mechanisms used in first-generation decentralised justice systems and traditional arbitration. 

The Historical Transition of Dispute Resolution

If arbitration is defined strictly through the twentieth-century paradigm codified by the 1958 New York Convention, then crowdsourced mechanisms indeed fit uncomfortably within the category. This is a concession we gladly make to Eidenmüller and Hochgürtel.

But we need to keep in mind that the framework was designed in a post-war environment specifically to address disputes arising from large-scale commercial transactions and cross-border investments between states and multinational corporations. It was the precise institutional answer to the particular economic architecture of its era.

Today, however, we inhabit a vastly different economic landscape. Traditional international arbitration faces a well-documented crisis of cost, speed, and accessibility. As proceedings grow increasingly formalised and expensive, vast swathes of disputes native to the internet economy (low-value, high-volume transactions, cross-border freelance contracts, and automated protocol interactions) are left completely unserviced by the New York Convention framework. Modern global commerce demands an arbitral infrastructure engineered for this new reality.

As internet technology and artificial intelligence transform international economic life, there are calls for expanding our understanding of what constitutes arbitration. Systems like Kleros look less like twentieth-century commercial tribunals and more like ancient Athenian dikasteria, relying on large-scale civic participation, statistical convergence, and random selection, where adjudicative roles may increasingly be augmented by AI agents alongside human participants. For specific categories of digital disputes where conventional arbitration is economically unviable, these models offer a scalable and accessible solution.

Whether that counts as 'judging' in some philosophically pristine sense is, we suspect, the wrong question. The one that matters is whether the mechanism is reasoned enough, consistent enough, and resistant enough to manipulation to serve the disputes it was built for. And on that question, we hold, the answer is closer to yes than Eidenmüller and Hochgürtel's critique allows.

* Federico Ast is the Founder and CEO of Kleros. 

William George is Research Lead at Kleros, focusing on game theory, mechanism design, and decentralised justice protocol architecture.

Facundo Trotz is an Attorney and Legal Researcher at Kleros, specializing in international arbitration, legaltech innovation, and decentralised dispute resolution frameworks.