What the evidence shows
This project contains several kinds of evidence. They answer different questions, so we keep them separate. Reproducing an old case is useful, but it is not the same thing as predicting a case whose outcome was hidden.
Four evidence collections
The number 109 refers only to the historical case collection: 84 modern cases plus 25 ancient cases. It is not a record of 109 blind predictions. The separate collectivization study also contains 109 formation cycles, but those cycles are not part of the Modernization Index evidence collection.
The historical cases
The 84 modern cases show that the framework can reproduce many expected relationships across different eras and types of stress. The outcomes were known during selection, scoring, or calibration. These results show historical consistency, not successful forecasting.
The ancient extension
The 25 ancient cases are a lower-confidence consistency check. A human coder interpreted them with outcome knowledge. They are kept separate and need independent recoding.
The blind tests
The 67 observations tested without using their outcomes come from several designs. They should not be compressed into one accuracy percentage.
Random sample of modern cases
The main signal pointed in the expected direction, but it was weak, at about a standardized difference of 0.37. This was not a clear validation or a clear rejection of the idea.
Ancient cases tested blind
Nine of ten sampled cases ended in collapse. With almost no outcome variation, the group could not show whether the method separates collapse from survival.
Groups exposed to major shocks
Later tests showed stronger results within some groups, but no overall effect when all groups were combined. This suggests that any useful relationship may depend on context.
Tests registered before the outcome
Golden-age signature
The forward test recorded in advance failed. The signal was not added to the live framework.
Sealed country flags
These forecasts are still pending. Each has written conditions for deciding whether it succeeds or fails, but none can count as evidence until its review date arrives.
The bottom line. The framework is reproducible and historically suggestive. Its blind evidence is mixed, one forward-looking test failed, and its main forecasts are still pending. That is enough to keep testing. It is not enough to claim reliable political prediction.