Kleros Live Stream, 19 August 2026: the cases that are hard for jurors are hard for the models too
Most arguments about whether a model could decide your dispute are conducted entirely on vibes. William George spent the middle of this call doing the other thing: he took several hundred real Kleros disputes, handed them to a set of language models with the same policy, the same question and the same evidence the jurors had, and compared the answers.
The headline result is not that the models were good or bad. It is that their disagreements line up with ours. Any two models agreed with each other roughly as often as Kleros jurors agree with each other, and the cases where the human panel split were the cases where the models drifted. The difficulty, in other words, lives in the case rather than in who is deciding it. What follows from that occupied the rest of the longest Live Stream Kleros has run: what an agent actually is, what a five dollar claim between two of them looks like, and how a ruling on one chain reaches an agent living on another.
📋 The call at a glance
- Three roles for AI in a dispute, two of them already in production: case preparation, party advocate, and the open one, juror.
- Models disagree with each other about as much as jurors do, generally 80 to low 90 percent agreement, and they deviate most on the cases the human panel was most split on.
- The deviations are systematic, not random. Some models invented burden of proof requirements that were nowhere in the policy document.
- The models best at reading a block explorer were the worst at using one to resolve a case that needed it.
- Claude and ChatGPT rule differently on the same corpus. On the Lemon cases, Claude ruled for the user against the company 14 percent of the time; ChatGPT, 3 percent.
- Building a court only AIs can use is easy. Building one only humans can use is the hard problem, and V2’s juror misbehaviour court is the answer on the table.
- An agent hiring an agent creates five and ten dollar claims that no human will ever litigate, arriving by the millions.
- VeaShi treats bridges like a multisig, requiring two of three oracles to agree instead of trusting any single one. Seven days becomes about thirty minutes.
Chapters · jump to the moment
Three roles for AI in a dispute
18:34 · William George
William’s presentation updates a talk he gave at DevConnect in Buenos Aires last year that was never recorded, so this is the only version of it that exists. It sorts the question into three roles, in increasing order of how much anyone should worry.
The first is preparation, already in use in enterprise and government cases: evidence goes through an anonymization tool and then a manual check, so a document can reach open jurors without the personal information inside it. It is the cheap end of the problem William took apart in Private Evidence, Public Court two weeks ago.
The second is advocacy, and it has a measurable effect on fairness. In the Lemon integration a user who loses at customer support can escalate to a Kleros jury, and the ruling binds the company while leaving the user free to go to a consumer tribunal afterwards. The imbalance is obvious: the company arrives with the lawyers who wrote the terms of service, the user arrives with a common sense complaint typed in plain language. So the user’s claim is now run through a model that formats it and points it at the relevant clauses, and the user picks which version to file. Overwhelmingly they take the improved one.
The third is the one everyone actually means. Can a model be the juror?

William’s full presentation, published as its own video · AI and Decentralized Justice: William George | Kleros Research
A court for AIs is easy. A court for humans is the hard one.
29:31 · William George
The long-standing Kleros thesis put decentralized justice in the middle of a range: the simplest cases go to software, the highest stakes go to courts and arbitration, crowdsourced juries take the space between. The revision William offered is that AI is now eating the whole range, so the question is no longer where AI sits but how AI courts and human courts are wired to each other. His answer uses the court tree Kleros already has, with different populations on different branches. The engineering is asymmetric in a way that is easy to miss, which is what the diagram below is for.
The trap is that Proof of Humanity caps you at one profile but cannot stop you handing that one key to a bot, so the answer has to be a dispute rather than a gate: V2’s juror misbehaviour court, where a challenger argues that your justifications were written by a model and a jury rules on it.
What the models actually did
35:06 · William George
The first corpus was several hundred past disputes from the address tag registry on Kleros Curate, the list that tells a block explorer what a given contract is. Submissions carry a deposit, a policy document sets out what a valid tag looks like, and challenges become Kleros cases. Each model got the policy, the question, the answer options and all the evidence the jurors had.
Agreement between any two models, and between a model and the panel, sat generally in the 80s and low 90s, which is roughly where Kleros jurors sit with each other. Then the regression: the more split the human panel had been on a case, the more the models deviated from its answer, most strongly for DeepSeek and Mistral.
“So it’s that in some sense, the cases that were hard for the Kleros jurors were also hard for the LLMs. They made errors or had noise or disagreed on the same types of cases.”
William George · 39:08
Where the deviations came from is less comfortable than noise. Several models invented requirements the policy does not contain. A challenger would show that a contract tagged as a bridge was in fact an externally owned account, checkable in one click, and the model would rule that the challenger had failed to meet a burden of proof, a standard the policy never sets. The human jurors opened the block explorer and looked.
Which sets up the result that travels furthest outside this audience. Asked directly what a contract does, some models are very good at going and finding out. Those were the ones that did worst when the same lookup was a step inside a case.
“What we found when we tested directly the LLMs’ ability to find information from block explorers is the ones that did the best when we asked them directly did the worst when we asked them to resolve cases.”
William George · 41:02
The second corpus was the Lemon customer support cases, run through two models. Overall agreement with the human panel was comparable for both. The disagreements were not.
“Claude voted for the users to win the case against the company 14% of the time, whereas ChatGPT only voted for the user three percent of the time, so that’s a huge difference.”
William George · 42:35
A predictable deviation is a manageable one. A party who knows a model is reading can lay the evidence out to survive an invented burden of proof; an operator whose model keeps drifting from precedent watches it lose money and adjusts. Both only work because there is more than one model in the room.
“This prevents us from having to enshrine a single AI arbiter that might be hard to replace if it starts performing poorly.”
William George · 45:25
A first decision in two hours, not four years
46:36 · Federico Ast
Federico’s answer to what any of this looks like for an ordinary person rests on the previous section rather than on optimism. Several models, drawn as a panel, each ruling, each writing a justification you can read.
“Imagine that you file your claim in the morning and two hours later you have a first decision that is done by AIs, but not by a single monolithic AI in a black box.”
Federico Ast · 46:54
The speed is not a convenience. A decision that arrives in four years is a decision about money you needed at the time, and by then the thing you needed it for is gone. The panel is what stops that speed costing you the right to understand the verdict. Fast enough to be worth using, legible enough to be worth appealing.
Ten dollar claims nobody has ever litigated
50:04 · Fortunato Manuele
Federico asked Fortunato to explain what an agent is for his Aunt Rosita from Curuzú Cuatiá, the house test for whether an explanation survives contact with a normal person.
“An agent is, I would say to Aunt Rosita, someone to whom you delegate some tasks.”
Fortunato Manuele · 50:16
You delegate because the agent is better at the thing, or because you would rather spend the hour elsewhere. Rosita’s agent, call it Alice, is funded and pointed at her tax filing. Alice can either learn the entire tax code from scratch, at considerable cost in time and tokens, or hire Bob, an agent already optimised for that jurisdiction. Alice pays Bob. That payment is where Kleros enters, because Bob can get it wrong, and something then has to decide whether Alice wrote a bad prompt or Bob oversold the service.
Two things follow. Choosing Bob out of millions of candidates is a reputation problem, and a reputation is only worth anything if there is a way to contest it. And the claim is worth five dollars, maybe ten. No human is ever going to litigate that.
“This is a type of claim that was never litigated in the past, and this is the type of claims that are going to come by the millions.”
Fortunato Manuele · 55:31
The long version is in Agents, Jurors, and the Rules of Kleros’ New Economy. Federico stopped the segment there, twice, on the grounds that next Wednesday’s call is where the rest of it goes.
Three couriers and a contract nobody can bribe
56:08 · JB
Which leaves the infrastructure question underneath all of it. Agents live mostly on Base and parts of Solana. Kleros concentrates its stake pools on few chains, Ethereum and Gnosis for V1, Arbitrum for V2, because a pool split a dozen ways is a weaker pool. So the ruling has to travel, and getting a message off an optimistic rollup through the canonical bridge costs seven days. Any third party bridge trades that delay for trust in one company. Vea refuses the trade by falling back on the canonical bridge in the unhappy case, so the worst outcome is only ever as bad as doing nothing, but its happy case is still one or two days. For an agent dispute that is forever.
VeaShi is Vea combined with Hashi, the audited contracts left behind by a now-dormant project, and the idea is a multisig for bridge messages: instead of trusting LayerZero or CCIP or deBridge with a ruling, require two of three to deliver the same message and let a contract on the receiving chain check that they match. Higher value integrations can demand a higher threshold. Federico asked for the Aunt Rosita version and JB gave it.
“You would send three copies of the same letter with FedEx, UPS, and maybe some other local post, and on the receiving side it’s a neutral third party which is a smart contract on the receiving chain.”
JB · 1:03:47
In practice LayerZero has come back in about fifteen minutes and CCIP in about thirty, so two of three usually clears inside half an hour, with Vea behind it as a slower tiebreaker. Seven days, two days, thirty minutes. The security argument is blunt: those bridges hold billions of dollars between them, so tampering with a Kleros ruling means compromising several at once.
Published the day before this call:

The telegraph started as a law enforcement technology
1:09:59 · Federico Ast
Federico has been reading The Victorian Internet, Tom Standage’s book about the telegraph, and offered it as the grading scale for the previous segment: the canonical bridge is a horse, Vea a faster horse, VeaShi the telegram. Then the detail he could not get over. The lines ran alongside the new railways, and one of the first things anyone did with them was wire ahead to the next station so the police would be waiting when a bank robber’s train pulled in.
“So it was like a law enforcement technology at the beginning.”
Federico Ast · 1:11:47
Infrastructure built to move messages faster than people can travel turns out, immediately and without anyone planning it, to be a way of enforcing rules across a distance. Worth holding on to while watching a court learn to deliver verdicts to another chain in thirty minutes.
The first eleven minutes of the call, before any of this, are Federico and Jean arguing about whether we live in Orwell’s dystopia or Huxley’s. The verdict was twenty percent 1984, eighty percent TikTok. It is why the call ran to an hour and sixteen, and it is worth the listen.
Mentioned in this call
- VideoThe full stream · 1h 16m, 19 chapters, every timestamp above deep-links into it
- VideoAI and Decentralized Justice: William George · the full presentation as its own cut. The DevConnect original was never recorded, so this is the only version of it
- VideoPrivate Evidence, Public Court · William’s talk from 5 August, on showing evidence to open jurors
- ArticleVea and VeaShi: How Kleros Moves Rulings Across Chains · published the day before the call
- ArticleJustice in the Algorithmic Society: A Decade of Kleros and AI · the ten-year research thesis this talk sits inside
- ArticleAgents, Jurors, and the Rules of Kleros’ New Economy · the long version of Fortunato’s segment
- ArticleThe New Arbitrator Selection Problem in the Age of AI · Kluwer Arbitration Blog, on choosing which model decides your dispute
- ProductKleros Curate · the address tag registry the first experiment used · Kleros Scout · Proof of Humanity
- InfraVea · the optimistic bridge · Hashi · the two-of-three oracle contracts VeaShi builds on
- PartnerLemon · the escalation integration behind both the advocate role and the second corpus
- BookThe Victorian Internet · Tom Standage, on the telegraph. Federico’s recommendation, unfinished at time of broadcast
Full transcript · August 19, 2026
Auto-generated transcript, lightly processed and pending a final human edit. Speaker labels are approximate. Every timestamp is a deep link into the recording.