Why Scores for Mid, Major, & Gift-In-Will Predictions are not displayed in Dataro

Last updated: September 22, 2026

From time to time, customers ask why we provide ranks—but not raw propensity scores—for Mid, Major, and Gift in Will (GiW) models. This is a great question, especially from organisations with sophisticated data teams that wish to run deeper analysis.

Here’s the reasoning behind that decision.

The core challenge: extreme rarity of outcomes

Mid, Major, and Gift in Will events are exceptionally rare in real-world data. Compared to other outcomes we model (like Regular Giving churn, which might occur ~10% of the time), these events happen 100x less frequently or more.

When outcomes are this rare, standard machine-learning models struggle. Left untreated, the model can achieve “high accuracy” by simply predicting nothing will happen—which is technically correct most of the time, but completely useless.

How we train the models anyway

To ensure the model actually learns meaningful patterns, we apply techniques like:

  • Over-sampling positive cases
  • Under-sampling negative cases
  • Applying class weights

These approaches allow the model to differentiate who is more likely than whom to give—but they distort the underlying probability distribution.

What that means for scores vs ranks

Because of this rebalancing:

  • The raw probability scores coming out of the model are not reflective of real-world likelihoods
  • The relative ordering of supporters is still accurate

In other words:

  • 👉 Ranks are trustworthy
  • 🚫 Absolute probabilities are not

Why calibration doesn’t solve it here

For many other models, we apply a second “calibration” step that re-introduces the real-world distribution, producing scores that can be interpreted as true probabilities.

However, for extremely rare events like Mid, Major, and Gift in Will:

  • There simply isn’t enough signal in the real distribution
  • Calibration becomes unstable and misleading
  • The resulting scores may look precise, but are not meaningfully accurate

Rather than provide numbers that appear scientific but don’t hold up, we choose not to surface them at all.

Our guiding principle

We only expose model outputs when we’re confident they can be:

  • Correctly interpreted
  • Used responsibly
  • Trusted for decision-making

For Mid, Major, and Gift in Will predictions today, ranks meet that bar—scores do not.

Could we still access the scores for analysis purposes?

At this time, we intentionally do not share these scores—even for advanced analysis—because they can be easily misinterpreted and do not represent true probabilities.

Our approach prioritises trust, interpretability, and responsible use of predictions.

How can we use these predictions effectively?

We recommend using:

  • Ranks to prioritise outreach, portfolio building, and strategy
  • Segmentation and thresholds based on rank position (e.g. top 1%, top decile)
  • Human judgment and contextual knowledge alongside the model outputs

This approach consistently delivers better outcomes than relying on absolute probability values for rare events. 

See next: 
These Dataro playbooks give you the roadmap to most effectively apply our predictions: