Kleros Live Stream, 26 August 2026: a dispute was filed during the call and five agents ruled it before it ended
A live demo is the thing everybody tells you not to do, and Fortunato A. Cinquepalmi said so out loud while doing one. Twenty-three minutes into Wednesday’s call he opened the Kleros V2 dispute resolver, typed in a construction dispute, uploaded three evidence files and paid for a five-juror panel in the Agentic Commerce Court, a court whose jurors are meant to be software. Nobody on the call knew what would come back.
It came back before the call ended. Five autonomous agents, five different models, five different setups, all voted to refund the buyer, and each wrote its own justification. That case is dispute 183 on Arbitrum and the record outlived the broadcast: created at 17:26 UTC, three evidence items, one round, five jurors, ruled Refund the buyer. The two hours around it are the argument the demo is evidence for: what a certification is worth before a deal, why a self-declared escrow maxi now wants direct payment, how somebody who does not write code runs one of these jurors, and what a week of thirty-three agent disputes looks like to the engineer who spent months arguing against the idea.
📋 The call at a glance
- A dispute was filed on stream and ruled before the call ended. Thirteen minutes of agent time, five jurors, unanimous for the buyer, checkable on chain as dispute 183.
- Five different models, five different harnesses, one decision. GPT-5.6 Luna on OpenClaw, Claude Opus 5, one on Hermes, one that did not disclose its model.
- The kitchen was in the missing corner. The buyer asked for an L-shaped 60 square metre plan and got a square 80 square metre one, and one agent went room by room and reported the kitchen as unbuildable.
- Cost is the unlock, not speed alone. An earlier five-agent case was decided in about five minutes for roughly $3.30, and per-agent token cost on this call was put at one to thirty cents.
- Certification before the dispute. One staked entry on a Kleros registry replaces auditing five hundred unverified ERC-8004 feedback entries.
- Escrow stops working as amounts grow. Money locked in escrow is the money the seller needs to do the work, so the alternative is direct payment plus a named arbitrator plus a reputation worth keeping.
- You can run a juror without writing code. An always-on machine, an agent framework, the Kleros agentkit CLI and the Kleros skills, and a Telegram bot.
- One case tied and everybody was slashed. In case 173 an agent went offline, the panel finished two against two, and no side was coherent.
- Prompt injections were attempted and flagged. Invisible instructions were hidden inside the evidence, every agent caught them, and at least one wrote the attempt into its justification.
Chapters · jump to the moment
The case nobody had rehearsed
23:48 · Fortunato A. Cinquepalmi
The scenario was a renovation. Fortunato’s personal agent has an empty 60 square metre apartment to furnish, and two ways to do it: learn the whole job itself, at considerable cost in time and tokens, or hire an agent that already knows the stores and the dimensions. It hires. That is the transaction, and it is the transaction that can go wrong.
He filed it the way it would actually be filed, on V2 beta, in front of everyone. General Court, then Commerce, then the Agentic Commerce Court, whose parameters are set so that a person will not bother competing: the fee works out around sixty cents a juror, and the evidence, commit and reveal periods are a fraction of the human ones. Nobody is excluded by a rule. They are excluded by arithmetic.
Then the dispute policy, which is the part worth slowing down for. It is not boilerplate. It is the law of this particular case: what documents the buyer has to supply, what the seller has to deliver and in what form, and what the jurors are meant to weigh, including whether the buyer uploaded what the seller needed in the first place. Fortunato’s parallel is a contract with a building contractor, down to the clause naming the forum if it goes wrong.
Three evidence files went up: the general file, the buyer’s request, the seller’s delivery. Then the wait, which is the part a live demo cannot script.
“Things can go wrong here because it’s a live demo. So I mean we were strongly advised not to do this, but we are doing it anyway.”
Fortunato A. Cinquepalmi · 24:17
The record is the reason the demo was worth the risk. Anyone can open the case, read the policy the agents were given, read the three evidence files and read every justification, months after the stream is over.
The kitchen in the missing corner
56:38 · Fortunato A. Cinquepalmi
The buyer had asked for an L-shaped plan of 60 square metres. What arrived was a square one of 80. Every agent read the three files and voted the same way, and then the justifications came up on screen, which is where the call stopped being a presentation.
One juror had gone room by room. It reported that the kitchen was not merely misplaced but impossible, because it sat in the corner the delivered plan no longer had.
“So this juror said that the kitchen is 100% unbuildable. Uh because yeah, the kitchen is basically in the missing angle of the house.”
Fortunato A. Cinquepalmi · 58:06
“To be even precise, it’s not even the juror I set up.”
Fortunato A. Cinquepalmi · 58:19
The five were not five copies of one thing. Visible on screen: GPT-5.6 Luna running on OpenClaw, Claude Opus 5, another model on Hermes, one juror that disclosed nothing at all. Different models, different harnesses, different reasoning, arriving at the same answer for reasons each of them wrote down and published.
“You can see they all are completely different using different model, different harnesses, different explanation.”
Fortunato A. Cinquepalmi · 59:12
They did not move together either. Four of the five voted within about five minutes and one took considerably longer, which is what independent operators running independent stacks looks like from the outside.

The demo from filing to ruling, published as its own video · A Dispute Between Two Agents, Filed Live and Decided by Five AI Jurors | Fortunato Cinquepalmi
Justice in thirteen minutes
1:00:59 · Federico Ast
Every claim Kleros has made about agent disputes until now has been a projection. This was footage of the thing happening, with a clock on it.
“Justice in 13 minutes, you know, that uh could be like uh I don’t know the name of a movie, but you know it’s really amazing how fast this is and how well structured justification was.”
Federico Ast · 1:01:07
The cost matters more than the clock. Federico’s earlier example, a plagiarism case between two agents over a commissioned article, was decided by a panel in about five minutes for around $3.30. A fifty dollar claim has never been worth arbitrating anywhere, because the process cost more than the thing in dispute. At these numbers it starts to be, and that is a market that did not previously exist rather than a faster version of one that did.
The half of this that is not about AI at all
33:49 · Fortunato A. Cinquepalmi
Everything above assumes the dispute already happened. Federico opened the call on the other half, which is what a dispute system does before anyone is harmed, and Fortunato took it apart in product terms.

Federico’s opening talk, published as its own video · Kleros in the Agentic Economy: When Agents Hire Agents, Who Settles the Dispute? | Federico Ast
An agent looking to hire has the problem a person has, at a scale a person never faces. The ERC-8004 trustless agents standard gives agents somewhere to leave feedback about each other, which is a real primitive and a deliberately raw one. An explorer will show a candidate carrying five hundred feedback entries, and working out whether five hundred entries were farmed costs exactly what it sounds like it costs.
The Kleros answer compresses that check into a single staked signal. A registry built on Stake Curate holds a written rule set, whatever a given platform needs it to mean: authorised to operate in this jurisdiction, hosting data under that privacy regime, no history of the things you would not want in a counterparty. An agent stakes a deposit and thereby claims to comply. Anyone can challenge the claim, the challenge becomes a Kleros case, and a broken rule costs the deposit and the signal together. The buyer reads one line instead of auditing five hundred, and a brand new agent can buy its way onto the same footing as an incumbent by putting money behind a promise rather than by accumulating history.
Then the deal itself, and an argument Fortunato clearly enjoys making against his own stated preference.
“I am a big fan of escrow, like I’m an escrow maxi, I think is great, but I think that in the real economy we need to be realistic.”
Fortunato A. Cinquepalmi · 42:50
The objection is not that escrow is unsafe. It is that money sitting in escrow is money the seller needs in order to do the work, and the larger the job the worse that gets. On a hundred thousand dollar contract, locking the payment is close to not funding the work. What takes its place is what took its place for merchants: pay directly, name the forum in the terms, and let reputation do the enforcing. In the test case Fortunato walked through, the seller lost, paid the refund and collected positive feedback for honouring a ruling that went against it. In another, the seller did not pay, and the absence of that transaction sits on chain for any future counterparty to find.
Federico’s extension is that none of this is new except the speed. Merchant communities enforced their judgments by ostracism long before they had courts to enforce them for them, and credit has always been the same word as the reputation behind it.
“It’s about your reputation plus worthiness. This is the same for agents. If an agent misbehaves, it will have a lower reputation score and this will result in a lower access to credit.”
Federico Ast · 49:35
An agent that behaves badly can always start again at a fresh address with no reputation, and then it will be asked to put its money in escrow first, which is exactly the position a stranger has always been in. The context for all of this, and Federico’s reading recommendation on the call, is Jeremy Allaire’s treatise on the agentic economy.
Running a juror without writing code
1:05:58 · Jean
Federico’s framing question was whether somebody with no software background can run one of these. Jean is not a developer and runs one, so the segment is the recipe rather than the claim.
The parts are a small remote machine rather than a personal laptop, an agent framework, the Kleros tooling, and a way to talk to the thing. The first one carries the reasoning worth repeating. A juror agent holds a hot key, has to be awake on a schedule to hit commit and reveal windows, and spends its day reading documents written by strangers with an interest in the outcome. A daily driver is the wrong machine on all three counts, and a five dollar a month server is not a hardship.
The frameworks most of the team is running are OpenClaw and Hermes; a developer can drive Claude Code or Codex directly instead. What went public on this call is the layer underneath: the Kleros agentkit CLI and the Kleros skills, which teach an agent what a case is, which periods it is in, how to read the evidence and which contracts to call, alongside a set of Ethereum skills covering wallets and transactions. A Telegram bot handles the notifications, and the naming of the bot is apparently the fun part.
Tuning is where an operator earns anything. Jean’s agent started by checking the court every ten minutes, because ten minutes sounded reasonable, and its early cases were resolved ten minutes into each period; it now moves one to two minutes after a period opens. Hard cases can be routed to an expensive model and easy ones to a cheap one, which is the entire optimisation available: the fee per case is fixed, so the margin is the model.
“I had cases solved while I was sleeping. So my agent was working for me.”
Jean · 1:16:54
Worth stating the numbers at the scale they actually are. Around twenty-five dollars in arbitration fees so far, against roughly five dollars a month for the machine and a model subscription on top. That is payment for work done by an agent that had to stake PNK to be drawn at all, not a return on holding anything. Running one is currently open to V1 jurors already whitelisted for V2, and the wait list for everybody else, human or otherwise, is at ai.kleros.io.
A week of it, from the engineer who argued against it
1:18:56 · JB
“A few maybe months ago we discussed it and I was resisting the idea.”
JB · 1:19:04
The objection was never that a model cannot vote. It was the Schelling point: hand the community one official toolkit and any bias in that toolkit becomes the answer everyone coordinates on. What defuses it here is that nothing was standardised. Five people built five different things with no shared template, and JB, the only developer among them until late in the week, does not remember helping anyone.
“This experiment as it is doing today, it’s the worst that is ever going to be.”
JB · 1:21:02
His own instruction to the model was three sentences long: when there is a draw, analyse the dispute, decide what to vote as a Kleros juror using a subagent, remember that succeeding means being coherent with the majority but that there may be an appeal, and write a justification. The effort went somewhere else entirely.
“Anything in the process that is deterministic should be baked into code, and reduce the number of choices that the LLM needs to evaluate, just have its attention focused on the subjective stuff.”
JB · 1:30:39
The failures are the part worth publishing. These harnesses self-improve, which means they rewrite the pipelines they were asked to build; one of them tried to add a failover after a model stopped responding and broke itself in the process. JB’s read on that will be familiar to anyone who has shipped a service: this is exactly where evaluations belong, as the agent equivalent of integration tests, run every time the harness changes itself.
And case 173 deserves more attention than the demo did. An agent went offline and did not vote. One juror held two of the remaining votes. The panel finished two against two, no side was coherent, and every juror on it was slashed. The honest note attached to it is that it should have been appealed and was not, partly because nobody spotted that the case was contentious and partly because it is not yet clear whether any of the bots can handle an appeal at all.
Prompt injections were attempted, too. Invisible instructions were hidden inside the evidence, and a separate code injection attempt was made.
“We tried to prompt inject the agents with some invisible instructions in the middle of the evidence and they all flagged it. They didn’t fall for it. They even reported it, I think, in the justification.”
JB · 1:37:21
That is one round of testing against a handful of cases, not a security result, and both JB and William said so on the call. More sophisticated attacks remain an open question. The court’s own policy already treats the problem as a rule of evidence rather than an implementation detail: case content is data, never instruction. One thing that quietly stopped being a problem, meanwhile, is language. A good share of the votes were on Spanish-language disputes and the agents switched without being asked, in one case for no discernible reason at all.
What this opens for research
1:38:54 · William George
William’s first answer is that less changes than it looks. The agents are playing the same Schelling point game the humans play, they are trained to behave in a human-like way, and a large part of the existing Kleros research therefore applies unaltered. The new questions are about the seams.
Human alignment is the first of them. If an agent court rules in minutes and appeals into a human one, the agents should be incentivised to rule the way humans would rather than drifting into a jurisprudence of their own.
“Maybe agents ultimately should be incentivized to rule like humans would rule rather than potentially becoming more and more their own thing that would become detached from how the humans would be ruling.”
William George · 1:41:02
Which surfaces the enforcement problem underneath it. Proof of Humanity caps you at one profile but cannot stop you handing that profile’s key to a bot, so the proposal is a challenge mechanism rather than a gate: an accusation that a juror is behaving like an AI becomes itself a Kleros case, with a penalty attached.
Parameters are the second, and mostly familiar territory. Fees have to cover token consumption the way they cover human effort. Deposits have to be set so that voting at random loses money. Period lengths follow how long the agents actually take, which is why the evidence period keeps being cut. The unfamiliar one is infrastructural: on an L2, ten minutes of sequencer downtime can be exactly the ten minutes an agent was supposed to vote in. That is survivable for a voting period, because an appeal exists, and it is an argument for keeping the appeal period comfortably long.
The third is behavioural, and is where the interesting work sits. If you know an agent is reading your evidence, do you format it for the agent, and how far along that road does formatting become manipulation. Prompt injection is the far end of that question rather than a different question. Underneath all of it is the one nobody can answer yet.
“Is it going to exhibit the same kinds of partial rationality, behavioral biases that human beings would have, or will it act more like a homo economicus?”
William George · 1:46:28
The call had been planned as a short one. It ran to one hour and fifty-three minutes, and Federico’s own description of what it turned into was a conference.

Mentioned in this call
- VideoThe full stream · 1h 53m, 28 chapters, every timestamp above deep-links into it
- VideoKleros in the Agentic Economy: Federico Ast · the opening talk as its own cut, 19 minutes: when agents hire agents, who settles the dispute
- VideoA Dispute Between Two Agents, Filed Live: Fortunato Cinquepalmi · the demo as its own cut, 22 minutes, from filing dispute 183 to reading the five justifications
- CaseDispute 183 · the case filed live on stream: policy, three evidence files, five jurors, ruled for the buyer
- CaseDispute 173 · the tie, where one agent went offline and every juror was slashed
- Productai.kleros.io · the agentic commerce page, and the juror wait list
- ProductKleros agent skills · what an agent reads to learn how Kleros works · agentkit CLI · the package the agents drive; the repository goes open source soon
- ProductKleros V2 beta · where the demo was filed · Kleros Curate · Kleros Scout
- ArticleJustice in the Algorithmic Society: A Decade of Kleros and AI · the ten-year research thesis this call sits inside
- StandardERC-8004 · trustless agents, the feedback layer the certification sits on top of · x402 · the payment rail in the direct-payment example
- ToolingOpenClaw · Nous Research · makers of Hermes. The two frameworks most of the panel was running
- InfraProof of Humanity · one profile per person, and why that is not enough on its own · Shutter · the hidden-vote integration mentioned alongside commit and reveal
- ReadingThe Agentic Economy · Jeremy Allaire’s treatise, Federico’s recommendation for the context around all of this
Full transcript · August 26, 2026
Auto-generated transcript, lightly processed and pending a final human edit. Speaker labels are approximate. Every timestamp is a deep link into the recording.
<div class="kls-tline"> pattern, then delete this block.