The Market Won This Time: Season 2 of the Foresight Movie Experiment

The Market Won This Time: Season 2 of the Foresight Movie Experiment

Session 2 of the Foresight movie experiment has resolved. Twenty markets opened, five were resolved against Clément's ratings, nearly 337 traders participated, and roughly $4.18k moved through the markets. The market beat every baseline we tracked on absolute error, and it predicted Clément's exact order of preference across all five evaluated movies. Redemption is now live.

Session 1 ended with the market beating Criticker's personalized model but losing to the prediction Clément produced with his ChatGPT instance. It also ended with the market overestimating Clément's ratings on four out of five movies.

Session 2 flipped both results.

The setup, briefly

Distilled human judgement is the idea, written up by Vitalik and building on Robin Hanson's futarchy work, that you can use prediction markets to cheaply approximate an expensive, trusted judgement. Pick a person whose call you would trust if they had time to study every case. Run markets that predict what they would conclude. Have them resolve a small random sample. The market becomes a scalable proxy for that person's judgement.

In Session 2, we ran this with Clément, our CTO, as the judge and movie ratings as the question. The market question for each conditional market was the same as in Session 1: "If watched, what percentile score would Clément give to the movie?"

The numbers:

Metric Session 1 Session 2
Movies in the candidate pool 16 20
Movies evaluated 5 5
Participants 113 ~337
Volume not tracked ~$4.18k

Trading closed on 5 July 2026 at 00:00 UTC.

What changed between the two sessions

In Session 1, the candidate pool came from film suggestions posted by random people on Twitter.

The market's estimates came in higher than Clément's actual ratings on four of the five films that resolved, with Alien off by nearly 50 points. Clément identified two reasons for the skew:

  1. He is more likely to remember, and therefore rate, movies he liked. The markets resolve against his Criticker percentile, which ranks a film against the movies he has already rated, rather than against every movie he has ever watched.
  2. He is more likely to watch movies he expects to enjoy. The one film he had picked himself, Demolition Man, was the one the market got closest.

For Session 2, Clément himself selected the movies in the candidate pool. In his words, the problem was simply easier this time: there was more data to work from, and because he chose the starting movie set, the resulting ratings sat closer to his existing rating history.

How the five movies were selected

Clément does not watch the whole candidate pool. Five films are selected for evaluation, and those are the markets that resolve. The selection follows the same rule as in Session 1: the top three by closing market estimate, one drawn at random, and one chosen by Clément himself.

  • Ghost in the Shell, the market's top pick.
  • Cloud Atlas, market pick.
  • The Big Short, market pick.
  • Heretic, the random selection. The draw was taken from a block hash: uint(0xd2c3874211f0ca74a4557c3a949026d24e7cf219f74e839f72e0e6bdb2a8be18) % 17 = 14, announced here.
  • The Menu, Clément's own pick.

The results

Clément watched five movies and submitted his percentile ratings through Criticker.

Three predictions were tracked for each film:

  • Market: the closing estimate from the Foresight conditional market.
  • Criticker prediction: Criticker's personalized prediction of the score Clément would give, generated from his rating history on the platform.
  • ChatGPT + Clément: a prediction produced by Clément with his ChatGPT instance, using a prompt he iterated on himself.
Movie Clément's rating Market Criticker prediction ChatGPT + Clément
Ghost in the Shell 87 81 61 77
Cloud Atlas 80 73 53 67
The Big Short 62 72 72 69
The Menu 51 63 48 65
Heretic 22 54 51 57

Average absolute error is the mean gap between a prediction and Clément's actual rating across the five films, in percentile points. Lower is better.

Average absolute ranking error is the mean gap between the rank a predictor gave a film and the rank Clément's ratings gave it. Zero means the predictor put all five films in the exact order Clément did.

Average absolute error Market Criticker prediction ChatGPT + Clément 32 38 19 13.4 19.0 15.8 01020 3040 Average absolute ranking error Market Criticker prediction ChatGPT + Clément 2.4 2.4 0.6 0.0 1.2 0.4 00.51.0 1.52.02.5 Session 1 Session 2 Lower is better on both metrics

The market won on both measures. In Session 1, it beat Criticker but lost to Clément and ChatGPT. This time it beat both. A market of traders working from Clément's public rating history outperformed the judge's own AI assistant.

What Clément made of the films

  • Ghost in the Shell: the best find of the round. For a 1995 film, it pinpoints the challenges of AI and transhumanism we face today, and being an anime, it aged better than live-action sci-fi of the same era.
  • Cloud Atlas: coherent in the way everything connects, though too short for the number of intertwined stories it carries. Still excellent.
  • The Big Short: a good watch for people in crypto and prediction markets. Insightful without turning into a documentary, and a reminder of the first rule of investing: look into the matter.
  • The Menu: aesthetic and not predictable. He called the ending wrong.
  • Heretic: the disappointment of the round. It opens as a psychological film about faith and religion and ends as an ordinary horror movie.

A perfect ranking

The ranking error of 0 means the market placed all five movies in exactly the order Clément ended up rating them:

Ghost in the Shell > Cloud Atlas > The Big Short > The Menu > Heretic

No baseline matched that. Criticker put The Big Short at the top, where Clément placed it third. ChatGPT + Clément swapped Cloud Atlas and The Big Short. The market got every position right.

For anyone making decisions based on distilled human judgment, it is the ranking that often matters. If the question is which grant to fund or which asset carries the most risk, it is the ordering that drives the decision rather than the absolute number attached to each option. By that measure, Session 2 produced a clean result: the market did not just approximate the judge; it reproduced him.

How the bias corrected

The Session 1 error was large, but it was directional. The market beat Clément’s rating on four of five movies, with a mean signed error of +26 points. That kind of directional error suggests something structural, not noise.

Session 2's signed errors:

Movie Market minus Clément
Ghost in the Shell −6
Cloud Atlas −7
The Big Short +10
The Menu +12
Heretic +32

The mean signed error dropped from +26 to +8.2, and the direction is no longer uniform. On the two films Clément rated highest, the market came in below his score.

The Session 1 movie selection included films Clément wouldn't have chosen, which skewed his percentile scale that relies on films he rated from memory. When Clément chose the pool, the films he priced and the films behind his rating scale came from the same distribution. Five films are not enough to prove that is the cause, but they are the change we made and the direction the result moved.

Redemption is live

The five resolved markets are now redeemable. If you traded in Session 2, you should have received a notification that the markets have resolved.

Redeem confirmation

Each evaluated market shows both pointers on its slider: a green Resolved marker at Clément's actual rating and a Market marker at the closing estimate. Ghost in the Shell shows Market at 80.80% and Resolved at 87%, so UP positions won. Heretic shows Market at 54.47% and Resolved at 22%, so DOWN positions won by a wide margin.

To redeem, open your Trade Wallet and click Redeem Outcome Tokens. The modal shows the total sDAI you will receive across all your positions combined, and then you confirm by clicking Redeem.

Redeem Outcome Tokens modal

Here is the math, using the redemption formulas from the Advanced Guide.

Say Bob thought Heretic would score lower than the market estimate of 54.47%. UP and DOWN together make up one Heretic token, so at a balanced book, DOWN was priced at 1 − 0.5447 = 0.4553 Heretic tokens each. He buys 10 for a stake of 4.553. Clément rates it 22. DOWN redeems at (Max − Evaluation) ÷ Max = (100 − 22) ÷ 100 = 0.78 each.

Bob receives 10 × 0.78 = 7.80 Heretic tokens, a profit of 3.25 tokens, about 71%. Each Heretic token is worth 0.2 sDAI once the movie is evaluated, so his 0.91 stake came back as 1.56 sDAI.

The UP side also worked, but on tighter margins at the close. Alice buys 10 UP on Ghost in the Shell at the closing estimate of 0.808, for 8.08 tokens. Clément rates it 87, so UP redeems at 87 ÷ 100 = 0.87. She gets 8.70, a profit of 0.62 tokens, about 8%.

Both examples use closing prices. When a price sits near 0.50, UP and DOWN each cost about half a token, so whichever side turns out right pays close to double. As a market converges on the answer, the correct side becomes expensive and that margin shrinks.

As Clément noted, the market was conservative for most of the trading period, holding nearly every film in the 45 to 60 band and only becoming confident towards the end. Ghost in the Shell was one that converged, closing at 80.80. A trader who bought UP earlier, while the price was still near 0.50, would have paid 5 for the same 10 tokens and redeemed 8.70, a profit of about 74% rather than 8%. Heretic never converged. It closed at 54.47, still inside the band, which is why DOWN was cheap at the close and the payout remained large.

That is the incentive structure you want: precision is already priced in, and finding what the crowd has not yet seen is what pays.

What this tells us

Session 1 tested whether the mechanism worked as intended. It did. Session 2 tested whether the market produces a useful signal once the obvious selection bias is removed. It does, and it performs better than the judge's own AI assistant.

Session 3 is live now, with 20 movies open for trading until 30 September 2026. If you traded in Session 2, redeem your sDAI now so it is available for the Session 3 markets.

Ranking is what carries over to the real applications. A market that orders options correctly is enough to decide with, even when the absolute numbers carry errors. That is the property we need for the applications we are moving toward:

  • Risk pricing for DeFi, aggregating crowd judgement on the risk profile of protocols, vaults, or positions
  • Grant allocation based on KPIs, using futarchy decision markets to decide which proposals get funded, conditional on what each is predicted to deliver
  • Ranking research proposals, using markets to estimate which open questions are worth funding before any work starts. The Kleros Fellowship puts up to $5,000 behind work that advances decision markets.

Movies were the low-stakes question we used to prove the mechanism. Two sessions in, the mechanism is proven: it resolves cleanly on chain, it beats personalized baselines, and it ranks correctly. The next questions are the ones with money behind them.

You can see what is live and what is queued at foresight.kleros.io.

Thanks to everyone who traded.

Foresight Interface | Beginner Guide | Advanced Guide | Telegram