Validating a new risk signal when the book has no loss history

You've got a new signal you want in the scoring model. Maybe it's a rooftop condition index, a vegetation encroachment score, or something pulled straight from imagery with no predefined attribute list behind it. The problem is the peril class or territory it's meant to help with doesn't have enough developed losses yet to run the usual lift test. You can't wait three renewal cycles for the loss triangle to mature before deciding if the feature earns a place in the model.

This is the catch-22 every portfolio risk quant hits eventually: the newer and more interesting the signal, the thinner the loss experience behind it. New admitted territory, a recently added line, a peril the book has only started carrying meaningful limits on. The feature might be exactly what the model needs. You just can't prove it the normal way yet.

Why you can't wait for the losses to catch up

The standard workflow (score the book, wait for claims to develop, check correlation) assumes you have a year or two of outcomes sitting there ready to validate against. For a brand-new signal on a thin book, that data doesn't exist yet, and won't for a while. Sitting on the feature until it does means losing a renewal cycle or two where it could have been adding separation to the model. Validate what you can check now, and shadow-run the rest until the losses develop.

Four checks that don't require a loss column

Proxy against what you already know drove losses elsewhere. If the new feature is meant to capture, say, structural deterioration, check it against perils or territories in your book where deterioration-driven losses are well documented. If the feature ranks those known-bad risks the way you'd expect, that's real signal before a single claim comes in on the new book.

Population stability across the renewal file. Score the whole book on the new feature and look at the distribution by segment, by territory, by construction class. A feature that's wildly unstable release to release, or that clusters in ways that don't track anything you'd expect from the underlying risk, is telling you something before it's told you anything about losses.

Monotonicity against your existing loss-based score. Bucket the book into deciles on the new feature and check whether those deciles line up in roughly the same order as your current, loss-validated score. You're checking whether the two agree often enough that adding the new feature is a sane move, and disagree in the handful of places where a second opinion would actually be useful.

A shadow run through at least one loss-development cycle. Score the book, hold the feature out of pricing, and let it sit next to the model for a year. When the next batch of losses develops, you finally get the real answer: did the feature's ranking predict anything, or did it just look plausible.

None of this requires labeled training data. It requires a feature you can score across the whole book consistently, a prior score to check it against, and the discipline to shadow-run before you commit weight to it. That's also the right way to treat a vector-per-site feature that wasn't built from a hand-picked attribute list in the first place: you're not confirming a definition, you're testing whether the number holds up under the checks above.

If you're trying to get a new imagery-derived signal into the model without first building and maintaining a bespoke extraction pipeline just to have something to test, a site-level vector keyed to your own book is one way to get a feature in front of these checks without that upfront build.

Run the proxy and stability checks on it before you decide whether it earns a line in the model.