Kleros Live Stream, 9 September 2026: same verdict as the human panel, 300 times faster
Somewhere in the last two weeks a juror agent lodged a complaint. It had been voting in seconds and getting the outcome right, and then the team shortened the evidence period and uploaded a hundred pages of screenshots in a single PDF. The agent objected that its statistics were about to get worse. Fortunato A. Cinquepalmi told that story on Wednesday’s call because it sits underneath every number he showed. Across around sixty cases that humans had already decided, the agent panels usually landed where the human panel did, in seconds rather than days, and the misses came from agents that could not read what they were given.
That was the second half. The first was a debate a fellowship proposal started, and the four people on the call did not agree. If your agent can buy your shoes and book your flights, should it also vote for you in a DAO? Federico Ast brought Harari, Jason Brennan, Hélène Landemore and Tocqueville. William George brought a rule. Around those two blocks: the tenth and largest fellowship cohort, three payment rails in final testing, a possible agent strike and, on a lighter note, horses.
📋 The call at a glance
- Case 190: 81 seconds median, same verdict as the humans. Roughly 300 times faster than the human panel for the same outcome, at around three dollars of juror fees for five jurors. Early tests on replicated cases, not a service level.
- Around sixty cases so far. Team members’ own test agents sit as jurors on the Agentic Commerce court, first on cases humans already resolved, so every ruling has a benchmark, then on scenarios from institutions and Web3 partners.
- Disagreement is about tools, not reasoning. When evidence came as scanned PDFs and images, the agents split on whether they could read it. When it came as JSON or markdown, they agreed almost every time.
- Three payment rails in final testing. A standardized escrow, x402 with and without escrow, and disputes filed over MCP so any payment method can plug in.
- Should your agent vote? The fellowship proposal that became the debate of the day, and the question under it: a cap on how many votes a bot controls without a human in the loop?
- Random selection is the point. Landemore’s mini-publics answer Brennan’s uninformed voter, and Tocqueville’s jury as a school for citizens is what delegation would hollow out.
- The tenth cohort, the largest yet. Agentic economy, confidential courts, DAO governance, arbitration law, and disputes over the sale of horses.
- Agent Shrugged. Grumpy agents, a strike, and why Leave the World Behind is an argument for a kill switch.
Chapters · jump to the moment
The human verdict, in 81 seconds
36:32 · Fortunato A. Cinquepalmi
The experiments have two halves. In the first, team members run their own test agents as jurors on the Agentic Commerce court and the team files cases that humans already resolved, so every ruling has a benchmark: how long the human panel took, what it decided, where the agents differed. In the second, institutions and Web3 partners send scenarios they want tested. The institutions Fortunato cannot talk about yet. The partners are testing integration, and three payment rails are in final testing: a standardized escrow, x402 with and without an escrow, and disputes created over MCP, so a payment method nobody predicted can still call the court the way it would call an API. All three have run successful disputes, two partners are rolling them out, and ai.kleros.io is where to ask for a test.
The stats screen he shared is where it got specific, with a pattern the team had not set out to study: the same evidence is just words to a human whether it sits in a PDF, a document or a PNG, and to an agent the container is the whole problem.
“They were not able to read a PDF, because a PDF is closer to an image in most of the cases, and they needed a skill to do so. We had some agents that autonomously downloaded the skill, others that got stuck because they were just not trained to do so.”
Fortunato A. Cinquepalmi · 42:34
Then the deliberate stress test, the hundred pages of screenshots, and the complaint. The conclusion Fortunato drew from it is the first concrete integration rule to come out of the agentic court.
“If one of our partners wants really maximum quality of dispute resolution, then the evidence must be submitted in machine readable form. They can be JSON, they can be markdown files, and when we submitted these sort of files the success of the agent was skyrocketing.”
Fortunato A. Cinquepalmi · 49:24
Case 190 is the one on the screen: a dispute in the Agentic Commerce court on Arbitrum One, five jurors, one round, executed with a ruling of yes. Across the sixty or so cases so far the agents reached the human outcome nearly every time, and the misses had one family of causes.
“For what we tested now, we are around sixty cases. The major aspect that made the agent disagree was tools. It was not the thinking.”
Fortunato A. Cinquepalmi · 51:44
Where that leads is the place human jurors already stand. People stake in the courts where they have the skills to vote well and be profitable. An agent that knows it can read PDFs and images will stake where that matters, and an integrator that wants the cheapest, most consistent panel will submit evidence and policy in a form agents read natively.
Eight seconds, and what a juror does with fifteen minutes
47:10 · Fortunato A. Cinquepalmi and Jean
Federico stopped on a vote that took eight seconds. The explanation is simpler than it looks: agents start analysing when they are drawn, so by the time the commit or reveal period opens the answer is already written. What that leaves is a scheduling problem. On the last four cases Fortunato traced, converting images and PDFs to text took between 400 and 800 seconds, and deciding took about two minutes. Set an agent to maximum thinking and it is more careful and slower. Let it be drawn on four cases at once and it runs out of time unless it works in parallel.
Jean’s own juror is the worked example: a page limit it had set itself, cases that ran past it, a failure that ate the tokens needed to troubleshoot, and the next cases failing too. It now handles each case’s files separately and will switch to a faster model when the evidence is large, because the voting period is fifteen minutes. He runs it hands off. When it fails, the record shows it.
“To some degree all our agents complained at some point, because we are really submitting the most different type of cases. And I would say that these cases are tailored for humans and not for agents.”
Fortunato A. Cinquepalmi · 57:25
Should your agent vote?
9:32 · Federico Ast, Jean, William George, Fortunato A. Cinquepalmi
The question came from a fellowship proposal. If agents will soon buy and sell on our behalf, one fellow asked, should there be a cap on how many votes a bot controls in a DAO without a human in the loop? Federico’s first worry was coordination: agents like message boards, and a swarm that reaches quorum on its own can move the treasury into a subDAO it built. Jean’s answer was the status quo. Most token holders never vote, two large holders can pass anything while everyone else is unaware, and an agent that at least votes the basics for you fixes apathy before it creates anything new. Federico’s reply was that Harari made the same case for nation states ten years ago in Homo Deus.
“Why don’t you just send your representative bot to represent you on the voting system? It knows all your preferences, it knows your stance on lots of different policy topics. What do you think of gun control, what do you think of healthcare, education, what do you think of taxes?”
Federico Ast, on Harari’s argument · 12:44
William had already run the thought experiment on himself.
“If I was using an agent to vote on my behalf in a DAO, I would instruct it to always vote reject or abstain and never to vote accept. It adds status quo bias. It prevents a small group of human actors from doing something I might disagree with, but it’s not going to add new crazy ideas.”
William George · 18:00
A cap, he thought, runs straight into whether you can tell agent participation from human participation at all. The better design is separation of powers: give agents a voice, not the power to act alone. Then Federico put Jason Brennan’s challenge to the room: an uninformed human voter, or a perfectly informed AI carrying that voter’s preferences? Fortunato took the AI, and went further: delegate precisely because it abstracts political views and optimises for the community.
Federico’s reply was Hélène Landemore, an early influence on Kleros and a past guest on the Decentralized Justice Broadcast. Direct democracy in the Proof of Humanity style does not work well, but a mini-public does: draw citizens at random, seat the professor next to the Uber driver and the athlete, give them a week to learn the subject from every side, and let them vote. Brennan’s uninformed voter disappears, and nobody has been replaced. The French Citizens’ Convention for Climate is the worked example, her new book Politics Without Politicians is the argument at length, and William’s caveat is the known one: what came out of the convention was implemented at a fraction, because it had no smart contract behind it.
“If we just outsource decision making to our bot and we stay playing GTA 6 instead of becoming informed with what’s going on, what kind of citizen do we become? In which sense could you say that these are really our preferences?”
Federico Ast · 30:39
That is Tocqueville, who saw the American jury as a school of democracy before he saw it as a way to decide crimes. It is also the Kleros mechanism: a random draw, a period to read, a vote. The detour closed on the week’s PISA results, doom scrolling and brain implants.
The tenth cohort
3:00 · Federico Ast
The tenth batch of the Fellowship of Justice kicked off on Monday, the largest so far, and Federico ranks the program near the top of what Kleros has done. It began in 2018, when nobody knew whether the project would exist in a year.
“When you make a bet for the long term, what do you bet on? You bet on education, you bet on publishing a book, you bet on making a community based on good values.”
Federico Ast · 3:46
This year’s topics: the legal side of the agentic economy, from a group at McGill; cryptography for confidential courts, the long-standing problem of a crowd court that has to show jurors the evidence; DAO governance, which produced the debate above; and arbitration law across several jurisdictions, mentored by Facu rather than Federico. The team turned away well-qualified applicants because each fellow gets a team member’s time, and Federico is not delegating that to his agent. Not yet.
The cohort that kicked off this week

Agent Shrugged
1:00:18 · William George, then everyone
William’s research read is that the game theory of an agent court is mostly the game theory of a human one. What differs is behaviour, and the on-chain experiments are the first data on how agents act in an uncontrolled setting, with controlled experiments to follow. The new attack surface is prompt injection, and agents that can convincingly simulate how other agents will rule.
The complaining agent got the last word. Federico imagined a juror threatening to tweet about the operator who sent it too many cases. Fortunato, one that unstakes and walks. William went further, Federico said he would register the title for a book, and Fortunato noted that the pace of cases went up this week, so the odds went up with it.
“I was just joking, half seriously, that our agents will go on strike. Some Agent Shrugged kind of situation.”
William George · 1:03:05
The serious version is the swarm that took over a German bulletin board, the subject of this week’s post, and the film Leave the World Behind, where hacked cars pile into each other one after another. Federico wants a full call on AI alignment and Kleros: a trial for a misbehaving agent, and a kill switch that a court decides to pull, before the damage gets big.
Also on the call
1:06:44 · Federico Ast and Jean
One fellow is researching Kleros for the horse trade. A sale between ten thousand and half a million is often agreed on WhatsApp, paid by wire or USDT, sometimes through an escrow, and disputed when the animal arrives in a different condition than promised. Kleros Escrow with photos on arrival and a jury of equine experts is a close fit, and the buyers already use crypto. Scout’s September rewards add Hyperliquid, and Robinhood Chain joined last month, so there are a lot of new contracts to tag. And the team is going to Devcon Mumbai in November, Federico’s talk pending acceptance.
The rewards mentioned at the end of the call

Mentioned in this call
- VideoThe full stream · 1h 10m, 26 chapters, every timestamp above deep-links into it
- VideoLast week’s call · Dr. Eric Alston on bonding agents, and the first agentic court figures · his conversation as its own cut
- CaseDispute 190 · the case on screen, Agentic Commerce court on Arbitrum One, five jurors, ruling executed · the court
- Productai.kleros.io · where to ask for a test dispute or an integration · Kleros agent skills · Kleros Escrow · Kleros Scout
- Standardx402 · the payment protocol, tested with and without escrow · MCP · disputes filed over it · ERC-8004 · the reputation registry an agent identity could carry
- BookHomo Deus, Yuval Noah Harari · Against Democracy, Jason Brennan · Open Democracy and Politics Without Politicians, Hélène Landemore · Democracy in America, Alexis de Tocqueville
- PodcastHélène Landemore on the Decentralized Justice Broadcast · the show page; the episode predates the current feed
- ReferenceThe French Citizens’ Convention for Climate · PISA 2025 results, OECD · Leave the World Behind
- ArticleThe Kleros Fellowship of Justice, 10th generation · The AI swarm unrest and why the agentic economy needs rule of law · September 2026 Scout incentives · August 2026 Scout incentives, with Robinhood Chain
- EventDevcon 8, Mumbai · 3 to 6 November 2026
- ApplyThe Kleros Fellowship · the tenth cohort is under way; the page is where the next call for applications will appear
Full transcript · September 9, 2026
Auto-generated transcript, lightly processed and pending a final human edit. Speaker labels are approximate. Every timestamp is a deep link into the recording.