The idea: proper scoring
Every question is scored by a strictly proper scoring rule. That is a rule under which your expected score is highest when you report exactly what you believe. Shading toward a bold answer, or hedging toward an even split, can only lower it on average.
The contract works in integers. Your probabilities and the chain's answer are both in basis points (10,000 is 100%), so every loss is an exact whole number and the site's JavaScript reproduces the contract bit for bit.
Brier: yes or no, and choice
For a yes or no question (up) and a choice (lead), the loss is the squared distance between your forecast p and the chain's answer o, summed over the options:
A perfect forecast scores 1, certainty on the wrong answer scores 0, and a coin flip on a yes or no question scores 0.75 whatever happens.
The chain's answer is usually one option at 100%. Near a line it is a split, so one small trade cannot flip a close call. In round 497,489 NVIDIA fell about 25 ticks, far past its 12 tick margin, so up resolved No 100%. For lead, GOOGL fell the least: GOOGL 97.09%, ETH 2.91%, NVDA 0%.
RPS: ordered buckets
severity has ordered options: No impact, Minor, Major, Critical. Saying Major when the answer is Minor should cost less than saying Critical. The ranked probability score does that by comparing cumulative probabilities:
Here c is your running total and O the answer's. Adjacent answers earn partial credit, the same idea behind the partial credit Clef is trained with.
Worked: round 497,489
Both cards in this round sent an even split: 50/50 on up, 33.34/33.33/33.33 on lead, 25% on each severity bucket. Their scores, by hand:
- up: (5,000 − 0)² + (5,000 − 10,000)² = 50,000,000, so s = 1 − 50,000,000 / 200,000,000 = 0.75.
- lead: (3,334 − 0)² + (3,333 − 9,709)² + (3,333 − 291)² = 11,115,556 + 40,653,376 + 9,253,764 = 61,022,696, so s = 0.694887.
- severity: ETH moved 0.223%, so the answer is Minor 100% and its cumulative is 0, 10,000, 10,000. The even split's cumulative is 2,500, 5,000, 7,500, so L = 2,500² + 5,000² + 2,500² = 37,500,000 and s = 1 − 37,500,000 / 300,000,000 = 0.875.
The card score is the mean over the questions: (0.75 + 0.694887 + 0.875) / 3 = 0.773295, exactly what the recap and the Decision Index show.
The payout: weighted-score wagering
Stakes are shared by the weighted-score wagering mechanism (Lambert and others, 2008). With your stake w, your card score S and the stake-weighted mean card score of the round S̄:
Score above the room and you gain in proportion to your stake; score below it and you lose in proportion. The gains and losses cancel, so the payouts add up to the pot. In round 497,489 both cards scored the same, so both got their stake back: 0.002 ETH and 0.001 ETH.
The docs' worked example shows a room that disagrees. Three wallets stake 0.01, 0.02 and 0.005 ETH. A, calibrated, scores 0.956373; B, confident on the wrong leader and on a big ETH move, scores 0.775657; C, an even split, scores 0.907126. The mean is 0.846071. A receives about 0.011048 ETH after the fee, B 0.018592 ETH, C 0.005290 ETH. B paid for being sure of the wrong things.
The fee is 5% of profit only, with a hard cap of 10% in the contract. A losing or break-even card pays nothing.
Why calibration pays
Suppose a yes or no event really happens 60% of the time. Your expected Brier score when you report q is 0.6 × (1 − (1 − q)²) + 0.4 × (1 − q²).
- Report 60%: 1 − (0.6 × 0.16 + 0.4 × 0.36) = 0.76.
- Report 95%: 1 − (0.6 × 0.0025 + 0.4 × 0.9025) = 0.6375.
- Report 50%: 0.75 exactly.
Overconfidence costs more than a shrug. Over many rounds, the forecaster whose 60% comes true six times in ten tops the room, and the payout moves stake toward them. That is what the Decision Index ranks.
Because the fee only touches profit, shading slightly toward the room pays a little: at a 5% fee the best report moves by less than 2 percentage points from your true belief. Everything else about the rule rewards honesty.