Agents, Jurors, and the Rules of Kleros' New Economy

Agents, Jurors, and the Rules of Kleros' New Economy

Community Calls - Monthly Recap

July was the month when Kleros' AI conversation became operational.

Across five community calls, the discussion moved from asking whether AI could participate in dispute resolution to designing infrastructure that could make participation trustworthy. Agents appeared as future users of registries, markets, and courts. Language models appeared as useful but inconsistent decision-makers whose training, incentives, and version changes can affect an outcome. Privacy appeared as the missing layer between public blockchains and the confidential evidence required by real cases.

The month also connected these technical questions to familiar institutions: sports referees, consumer-protection offices, academic research, government adoption, and the legitimacy of collective decisions. The common thread was practical. Kleros is designing procedures that still work when both humans and machines make mistakes.

July at a glance

  • July 1 - A Buenos Aires legal-innovation tour, AI jurors, Curate skills, hidden voting, and a proposal to make cooperative spending legible.
  • July 8 - Counter-Strike's Overwatch as a proto-Kleros, crowdsourced VAR, agent registries, real-world assets, and research with five language models.
  • July 15 - Paul Poenicke on the Knobe effect, infelicitous coordination, juror bias, and the design of AI-native courts with private evidence.
  • July 22 - Algorithmic feeds, four agentic products, model bias in 99 Lemon disputes, and the case for diverse panels instead of one canonical AI.
  • July 29 - The Fellowship's decision-market track, Scout's agent infrastructure, privacy evidence, government use cases, and the accounting behind Curate rewards.

From a football referee to an AI juror

The easiest way to understand July is to start with a referee.

On July 8, the community revisited Counter-Strike's Overwatch system. Experienced players review suspicious gameplay and help identify cheating. It is not a court in the formal sense, but it contains the core Kleros pattern: a large group contributes judgments, the system must manage disagreement, and the credibility of the result depends on the process rather than on a single official.

The same idea appeared in the calls' World Cup conversations. Could a crowdsourced video assistant referee make controversial decisions more transparent? Could AI models judge famous calls? What would happen if the model, the panel, or the evidence were wrong?

The point was not that football should immediately outsource refereeing to a decentralized court. The useful lesson was simpler: public disputes are easier to reason about when we separate the evidence, the decision rule, and the authority of the decision-maker.

That separation is where Kleros becomes relevant to the AI economy. If an agent pays another agent, hires an agent, buys an asset, or submits a claim, there must be a place to resolve the dispute. A reliable system needs more than an automated answer. It needs identity or reputation, incentives to participate honestly, a way to challenge a result, and a record of how the decision was reached.

Call one - July 1

Watch the July 1 call: https://www.youtube.com/watch?v=Egz1wQ585Sw

The July 1 call opened with SubTech 2026 in Buenos Aires, a legal-innovation conference founded in 1990 by Marc Lauritsen. Federico Ast connected Kleros to a wider movement that treats legal systems as something people can redesign, test, and make more accessible.

Amy Schmitz's work with Colin Rule on The New Handshake offered one example: dispute resolution can be built into the relationships and platforms where transactions happen, rather than waiting until a disagreement becomes an expensive legal emergency. Mario Adaro, associated with the Mendoza Supreme Court, brought the conversation closer to home. The call described a pre-settlement use case for neighborhood disputes, where a trusted process can help people resolve conflicts before they harden into formal litigation.

The conversation then turned to AI jurors. Kleros contributors have tested several models on 99 Lemon consumer disputes. The goal is not to treat a model's answer as correct by default, but to compare how models reason, where they disagree, and how their answers change across versions. The call placed this work in the wider context of AI governance. The challenge is to evaluate automated decisions without handing an institution to an opaque machine.

Fortunato described Curate skills as the practical layer for agents. An agent could operate autonomously, work with a human in a semi-autonomous mode, or act as an assistant. In each case, the skill set gives it a way to submit, challenge, and interact with Kleros tools.

That changes the question from "What can AI do?" to "What is AI allowed to do?" Once an agent can take action, the system needs clear permissions, wallet safety, and an account of who is responsible for the action.

The call closed on two kinds of legitimacy. The first was technical: a hidden-vote curation court on Gnosis, with a reminder to migrate stake before the July 13 deadline. The second was social: William George's discussion of Schelling points and a proposal for cooperative spending disclosure. A decentralized system cannot ask people to trust an idealized community. It has to make decisions, incentives, and spending legible enough for participants to coordinate around them.

The crowd referee becomes a machine panel

Call two - July 8

Watch the July 8 call: https://www.youtube.com/watch?v=PHDiMv0SR34

The July 8 conversation returned to football, but the underlying subject was institutional transparency. A controversial call is not only a question of whether the ball crossed a line. It is a question of who gets to decide, which evidence counts, and whether other people can inspect or contest the reasoning.

The community used Counter-Strike's Overwatch as a bridge between online moderation and Kleros' origin story. A distributed group of reviewers can be more resilient than one moderator, but only if the incentives discourage careless decisions and the process can distinguish a good-faith error from manipulation. The World Cup examples made the problem intuitive: when millions of people disagree with a call, authority alone does not create legitimacy.

The call also imagined AI jurors judging famous World Cup decisions, including the Balogun case. It even raised the possibility of presenting the idea to FIFA president Gianni Infantino. The useful question was not whether a model can replace a referee. It was what evidence and review process would make an automated or crowdsourced judgment acceptable to the people affected.

Fortunato then brought the same logic to agents. Scout registries can help agents discover one another, while skills can give them access to Kleros tools. Water credits and other real-world assets showed why this matters beyond software. If an agent represents a physical asset or participates in an environmental market, disputes may concern delivery, measurement, ownership, or payment. The court becomes part of the asset's operating environment.

An asado with Glen Weyl added the economic layer. Markets do not only move prices; they create signals about reputation, coordination, and expected behavior. Soulbound reputation can add context that a one-time transaction cannot. For an agent, that context could eventually function as collateral: a history of reliable behavior that makes cooperation less expensive.

William then returned to AI-juror research. The Wolters Kluwer chapter examined five language models on real cases. Comparing models matters because a single confident answer can hide a systematic bias. A panel does not eliminate disagreement, but it makes disagreement visible, which is the first step toward handling it.

The call finished with the Fellowship of Justice's tenth generation and the research questions available to the new cohort. Kleros' technical roadmap and its research community were pointing in the same direction: build the tools, then study the failure modes before the tools become infrastructure.

The bias inside a jury

Call three - July 15

Watch the July 15 call: https://www.youtube.com/watch?v=_iTZir7qihM

Paul Poenicke, a philosopher from the Kleros Fellowship, joined the July 15 call to discuss a deceptively ordinary question: do people judge harmful and helpful outcomes differently, even when the underlying structure is the same?

The Knobe effect describes a finding in moral psychology: people are more likely to attribute intention, knowledge, or blame when an action's side effects are harmful. Paul connected that effect to infelicitous coordination, the problem of coordinating on truth when participants' judgments are influenced by the way a situation is framed.

"Kleros may actually be immune."

Paul Poenicke, July 15 call, 15:18

For Kleros, the important question is whether jurors are vulnerable to this bias in the same way as ordinary experimental subjects. Kleros jurors receive financial incentives to coordinate on the correct answer, and that incentive may make the system less dependent on intuitive moral reactions. It may also introduce different biases. The proposed study is valuable precisely because it tests the claim instead of assuming that economic incentives solve psychology.

The conversation then moved to AI jurors. Models have biases too, and model behavior can vary with prompts, training data, system instructions, and updates. An AI-native court therefore cannot rely only on the apparent neutrality of the model. It needs a process for selecting models, comparing outputs, handling private evidence, and explaining responsibility when a model gets a case wrong.

William's discussion of private evidence highlighted the practical obstacle. Public blockchains are excellent at making a procedure auditable, but many real disputes contain information that should not be public: personal data, commercial records, medical information, or confidential communications. Kleros needs a way to preserve the transparency of the procedure while restricting the evidence to the jurors who need it.

The hidden-vote courts on Gnosis received their first six cases around this period. Commit-and-reveal voting makes it harder for jurors to copy an early visible vote, protecting the Schelling-point mechanism that the system depends on. The mechanism is technical, but its purpose is human: give each juror room to form an independent judgment.

When the model becomes part of the arbitration clause

Call four - July 22

Watch the July 22 call: https://www.youtube.com/watch?v=oagSE__Q_zo

The first call after the World Cup final opened with the information environment around the tournament. If an algorithm controls the feed, it can influence what people see, what they believe is important, and which disputes receive attention. That led to a broader question: who controls the algorithm, and who can challenge its decisions?

"Whoever controls the algorithm controls the information."

Federico Ast, July 22 call, 6:53

Fortunato outlined four directions for Kleros agents: verification and reputation, including registries such as ERC-8004; payments and disputes between agents; Trivium, a triage layer for low-value disputes; and Kleros courts as a jurisdiction for appeals when an automated transaction goes wrong.

Together, the four directions describe a basic economic stack. Agents need to know whom they are dealing with. They need to pay one another. They need a cheap way to handle small disagreements. They need a credible court for disagreements that cannot be resolved automatically.

That stack makes model bias an economic issue, not only a research issue. In the Wolters Kluwer Arbitration Blog article discussed on the call, 99 real Lemon disputes were run through Claude Opus 4.7 and ChatGPT 5.5. The models did not behave identically: Claude ruled for the consumer roughly five times as often as ChatGPT in the comparison. The call also discussed how a model update from Opus 4.7 to 4.8 could shift the burden of proof against the consumer.

"Those updates are pushed silently."

William George, July 22 call, 49:33

This is the new arbitrator-selection problem. If a contract names a model, a silent model update may change the meaning of the contract without either party agreeing to a new rule. If a platform chooses the model, the platform becomes a hidden court administrator. If an agent chooses the model, the counterparty may not know which standards will govern the dispute.

The proposed answer was not to enshrine one AI. A diverse panel of models can expose disagreement. Soulbound qualifications and badges can add information about which models or jurors have passed relevant tests. Human and machine participants can be evaluated by their track record. None of these mechanisms creates perfect neutrality, but each makes the system more inspectable.

The call connected that research to the Fellowship's tenth-generation AI track and to Scout's migration toward hidden voting on Gnosis. As agents become more active, the cost of visible coordination and unreviewable automation rises. The protocol needs independent judgments, not just fast judgments.

Privacy, public blockchains, and the first operational layer

Call five - July 29

Watch the July 29 call: https://www.youtube.com/watch?v=Z7Gs92uy1zo

The final July call brought the month to its most practical question: can Kleros handle disputes that contain evidence nobody should publish?

Lucia described work on a government-facing landing page in Spanish and a Mendoza use case involving consumer protection, mediation, and small claims. Privacy evidence was close to being ready for initial cases. The proposed process would allow only the selected jurors to access the evidence, and only for a limited period. That design preserves the auditability of the selection and voting process while narrowing the exposure of the underlying documents.

The problem is not limited to government. Enterprise disputes, consumer claims, gaming and esports, content moderation, intellectual property, and neighborhood conflicts all create situations where a public transcript is inappropriate. Adoption depends on showing organizations that confidential information will not become permanent public data.

William described several technical and accountability layers: encryption; access limited to the selected jurors; proof of humanity or soulbound credentials; KYC when an enterprise needs accountable jurors; and trusted computing or specialized hardware for AI jurors. Each layer answers a different concern. Encryption protects the evidence. Juror verification helps establish who saw it. Trusted processing can reduce the risk that a model or operator copies it. Accountability creates a path to follow if someone leaks or misuses the evidence.

The call also looked at the Fellowship's new decision-market, or futarchy, track. Instead of allocating limited funds only through a static application review, community traders could predict which proposals are most likely to create impact. The highest expected impact applications would receive more of the available budget. This turns funding into an explicit coordination problem: the community is not only judging the quality of an idea but also making a forecast about its future results.

Fortunato's Scout updates showed the same movement from concept to infrastructure. Scout was moving toward Blockscout, with support for entries on Robinhood Chain and rewards that can be staked on a court. The agent skills would let agents submit, challenge, stake, and vote. Curate skills already provided the basic pattern; Court skills were being built with MCP and CLI access.

That power comes with a design requirement: onboarding must be simple enough for an agent to start with a free model or a plain prompt, while wallet safety and permission boundaries remain explicit. The system should support autonomous, semi-autonomous, and assistant modes without pretending that they carry the same risk.

Finally, the Curate rewards dashboard made the accounting visible. Entries and rewards over time, along with the declining PNK reward per submission, help the community understand how the mechanism is behaving. The planned accounting work is important for the same reason that hidden voting is important: participants need to see enough of the process to coordinate without giving one participant an unfair advantage over another.

What July changed

July did not produce one grand answer to AI governance. It produced a more useful map of the problem.

First, agents need jurisdictions. An agent economy cannot depend on every transaction being correct. Payments, reputation, escrow, triage, and appeals form a basic institutional layer for machine-to-machine commerce.

Second, models need pluralism. The 99 Lemon comparison showed that model choice can affect a consumer's result. A model update can change the burden of proof. A panel, a qualification system, and an auditable dispute process are safer than treating one model as a permanent source of truth.

Third, privacy is not the opposite of transparency. A Kleros court can make the selection, incentives, and voting procedure inspectable while keeping the evidence available only to the jurors who need it. That combination is essential for decentralized justice to move from crypto-native cases into consumer, enterprise, and government settings.

Fourth, legitimacy has to be designed at several levels. Hidden votes protect independent judgment. Reputation records past behavior. Proof of humanity and KYC can add accountability. Research tests whether jurors are actually biased in the ways the system assumes they are. Decision markets make resource allocation a forecast that can be evaluated later.

Together, these threads present Kleros as more than a dispute-resolution protocol. They present it as a coordination system for a world where software agents transact, algorithms shape attention, organizations hold confidential evidence, and communities still need a way to decide together.