Method
How SovSpace reacts to a measurement, how its predictions are scored, and what none of this shows.
What a reaction is
councilof.ai publishes its measurements as capsules: small signed records, each about one subject (an agent card, an MCP server, a token contract, one of our own listings) and one observed state. Capsules are grouped into daily batches, each with a Merkle root, and one index binds the day's batches together. The index is public on councilof.ai and archived on Hugging Face.
For every capsule in a day's index, SovSpace writes one reaction. A reaction holds:
- a reference to the capsule: its id, and its batch root, recomputed by SovSpace rather than copied;
- an integrity control and a negative control (below);
- a counterfactual: the smallest change that would flip the capsule's state;
- a prediction about the next observation, with a probability, its basis and a time window.
A reaction never carries a measured state of its own, a decision or an approval. It is a different kind of thing from the capsule it reacts to.
Controls: can the measurement fail?
A check that cannot fail tells you nothing. Each reaction tests that in two ways.
Integrity control
SovSpace changes one hexadecimal character of one source digest in a copy of the capsule. The capsule's own verification path (recompute the capsule id, check batch membership) must then reject it. A capsule whose tampered copy still verified would be flagged.
Negative control
SovSpace re-runs the capsule's own check on a deliberately altered input, and the state must change. For example, it appends a few characters to a signed field of an agent card and re-runs the real signature verifier: a card that verified must now fail. Each negative control ends in one of these results:
- Can fail
- The altered input changed the state. The check is live.
- Can pass
- For a signature that failed, the same verifier and keys return verified on the bytes that were actually signed.
- Scheduled
- The control needs a fresh live probe. Until it runs, it proves nothing either way.
- Mirror gap
- SovSpace's own re-derivation does not reproduce the recorded state. That is a gap in SovSpace, not a finding about the capsule.
- Cannot construct a passing input
- Making a passing input would need the signer's private key.
- Not applicable
- No check ran for this state, so there is nothing to alter.
In the latest run:
| Result | Reactions |
|---|---|
| Integrity control: can fail | 1,263 |
| Negative control: scheduled | 1,029 |
| Negative control: can fail | 189 |
| Negative control: not applicable | 45 |
How a prediction is made
Every prediction says the same thing: the next observation of this subject will show the same state as this capsule. Its probability is a base rate for that kind of measurement and state, not a judgement about the subject. It is here to be beaten, not admired.
The probability comes from counts, never from judgement, and each row names the predictor version that made it.
persistence-laplace/0.1(26 and 27 Sep 2026) used (s + 1) / (n + 2), where n is the number of earlier scored pairs of the same kind and state and s is how many were unchanged. No pair had been scored when those runs froze, so n = 0 and every probability was 0.5. That is why the first score equals the 0.25 of always saying 0.5: it was the starting point, not a finding.climatology-hier/0.2(from 28 Sep 2026) uses the same counts but leans on the wider group when a group is small: the kind and state, then the kind, then every scored pair. A prediction may use only outcomes that were both observed and scored before it was frozen. Each row records the latest such outcome it used, so this can be checked.
For MCP contract parity, the first run had a stand-in history: how often a server's declared version in the public MCP registry stayed the same between two registry snapshots. That gives 0.968 for some states, and each such row names this basis. Once real scored pairs exist, they replace the stand-in.
The window says when the next observation can count. It opens at the measurement's next scheduled re-observation and closes two to eight days later, depending on how often that measurement runs. Where no recurring re-measurement exists yet, there is no opening date, and the window closes after seven days.
How predictions are scored
Once a day, at about 09:00 UTC, SovSpace checks every open prediction from the last 16 days against the capsules that arrived after it was made. For each prediction with a new capsule for the same subject:
- the outcome y is 1 if the state is unchanged and 0 if it changed;
- the Brier score is (p − y)², where p is the predicted probability. It runs from 0 (perfect) to 1. Always saying 0.5 scores 0.25, so a useful forecaster must score below 0.25;
- the log loss, −log p if unchanged or −log(1 − p) if changed, is recorded alongside.
A score means little without its comparison, so the home page puts each score beside scores that need no skill, computed on the same predictions from what was known before each one: always saying 0.5; the base rate for the kind of measurement; and the subject's own last outcome. A predictor that cannot beat the kind base rate has shown no skill beyond it. The Brier score also splits into three parts: reliability (how far the stated probabilities are from how often things actually stayed the same; lower is better), resolution (how well the probabilities separate what stays from what changes; higher is better) and uncertainty (how unpredictable the outcomes were to begin with, which no forecaster controls). Brier = reliability − resolution + uncertainty.
An observation is refused, not scored, if it is not strictly later than both the prediction and the capsule it reacts to, if it is the same capsule coming back, or if its batch root does not recompute. The mean Brier score on the home page is the mean over all resolved predictions, and it is shown only once at least one prediction has resolved.
What this does not show
- It is not a measurement. Nothing on proofof.ai is measured. The measured results, and their signatures, live at councilof.ai.
- It says nothing about the subject. A prediction that came true does not mean a server, card or contract is good, and one that failed does not mean it is bad. It only means the state did or did not change.
- Unscored is not correct. A prediction whose window closes without a new observation stays unscored. It is never counted as right.
- Not every capsule has a reaction. A batch added to the index after the day's reaction run has none until the next run. The index also carries forward batches that have not changed since an earlier day; those were reacted to when they were first published. The home page states both.
- Some subjects are not shown. A subject and its state appear here only where councilof.ai already publishes that row. Other rows are listed by capsule id.
- Nothing learns from this beyond counting. The only use of a scored outcome is as one more count behind the next day's base rates. No model is trained on the scores.
Schedule and data
New capsules are signed each morning. SovSpace reacts to them and scores the previous day's predictions, and this site is rebuilt from those outputs daily at about 09:30 UTC. If a step fails, the site keeps its last good version.
The data is at /api/latest.json (CORS open). Every figure in it is labelled PREDICTED or SCORED. Its files list gives one JSON file per measurement and day, with one row per prediction. Data is CC-BY-4.0.