What happens next, predicted and then scored
Every measurement published on councilof.ai is a small signed record called a capsule: one subject, one observed state. For each capsule, SovSpace makes a prediction about what the next observation of the same subject will show. When that observation arrives, the prediction is scored. The running score shows how good the predictions have been.
Today's numbers
From the reaction run of 29 Sep 2026, frozen , against capsule index 5ce9511e9554… (public index).
Capsules reacted to PREDICTED
16,930
Predictions open PREDICTED
14,275
Plus 1,141 whose window closed unscored. 12,583 of them have no confirmed re-observation job yet, so they may close unscored.
Predictions resolved SCORED
1,682
Brier score so far SCORED
0.1666
over 1,682 resolved predictions; lower is better. Compared with forecasts that need no skill.
PREDICTED made before the outcome was known. SCORED compared with a later, verified observation.
That day's final public capsule index lists 10,958 capsules; 10,958 of them have a reaction. That day's run reacted to 1,263 of them. The other 9,695 are unchanged batches the index carries forward, reacted to when they were first published (26 Sep 2026, 28 Sep 2026). 140 of those were published on an earlier day and reached SovSpace after that day's reaction run, so this run was the first to react to them: Cross-runtime reproduction (140, published 26 Sep 2026). Batches that were replaced in the index before any reaction run reached them have no reaction, and the replacement is reacted to instead: Cross-ledger supply (353, last listed 26 Sep 2026).
The score beside forecasts that need no skill
Every row is scored on the same resolved predictions, and each forecast uses only outcomes observed and scored before that prediction was frozen. A predictor that cannot beat the base rate for the kind of measurement has shown no skill beyond it.
| Forecast | Brier |
|---|---|
Issued by climatology-hier/0.2 (645 predictions) | 0.0325 |
Issued by persistence-laplace/0.1 (1,037 predictions) | 0.2500 |
| Always saying 0.5 | 0.2500 |
| The base rate for the kind of measurement | 0.1663 |
| The subject's own last outcome | 0.1679 |
The issued score splits into reliability 0.1303 (lower is better), resolution 0.0009 (higher is better) and uncertainty 0.0372, which no forecaster controls. What these mean.
By kind of measurement
| Measurement | Predictions | Open | Resolved | Brier | Subjects shown |
|---|---|---|---|---|---|
| Agent card signatures | 33 | 33 | 0 | – | capsule id only |
| claim_watch | 1 | 1 | 0 | – | capsule id only |
| MCP contract parity | 9,148 | 8,019 | 0 | – | capsule id only |
| Cross-ledger supply | 1,175 | 406 | 758 | 0.1522 | capsule id only |
| Cross-runtime reproduction | 154 | 154 | 0 | – | capsule id only |
| Operations (our own pipeline) | 1,266 | 1,256 | 10 | 0.2628 | capsule id only |
| Public signals | 472 | 122 | 350 | 0.1695 | capsule id only |
| Self parity (our own listings) | 609 | 165 | 443 | 0.1640 | capsule id only |
| MCP tool drift | 4,240 | 4,119 | 121 | 0.2500 | capsule id only |
A subject and its state are shown only where councilof.ai already publishes that row on Hugging Face. Elsewhere the prediction is listed by capsule id, which you can check against the public index.
How it works
- A capsule is published. councilof.ai measures something, such as whether an agent card's signature verifies, and signs the result into a daily batch.
- SovSpace reacts. It checks that the capsule's own verification can fail (change one character of a source digest and the check must reject it), and writes down a prediction: the state will be the same at the next observation, with a probability.
- The score is kept. When a later capsule for the same subject arrives, the prediction is scored with a Brier score. Unscored predictions are never counted as right.