Learned embeddings vs hand-engineered attributes for risk scoring

Every portfolio risk model lives or dies on its feature list. You pick roof condition, vegetation setback, distance to the flood line, maybe a proxy for structure density, and then someone has to go populate those columns for every site in the book. That's the part nobody budgets enough time for.

The question underneath "learned embeddings vs hand-engineered attributes" is really a question about where you want to spend your modeling budget: on defining what to measure, or on validating a model that already learned what correlates with loss.

The cost of hand-engineered attributes

A hand-engineered attribute pipeline starts with a hypothesis: roof age predicts hail loss, canopy overhang predicts wildfire ignition risk, lot clutter predicts contents loss after a flood. Each hypothesis becomes a tagging task. Someone, an analyst, a vendor, a computer vision contractor, has to define the attribute precisely enough that it can be measured consistently across tens of thousands of sites, then build or buy a pipeline that extracts it from imagery, then QA the output before it goes near the scoring model.

That's fine for the first three or four attributes your book needs. It gets expensive fast once you want to test a tenth hypothesis, because the marginal attribute still needs its own definition, its own extraction logic, its own QA pass. Add a new peril and you're often starting over: the attribute list built for wildfire exposure doesn't transfer cleanly to hail.

And every attribute nobody thought to define is a signal your model never sees. If debris density around a structure matters for wind loss but nobody wrote a spec for "debris density," it isn't in the feature set. A hand-tagged pipeline is only ever as complete as the list of things someone thought to ask for.

What a learned embedding gives you instead

A learned embedding skips the attribute-definition step. Instead of deciding in advance that roof condition and vegetation setback are the things that matter, a model trained on imagery produces a vector per site, a few hundred numbers summarizing what's visually present on that parcel, learned from the imagery itself rather than from a list written down first.

What comes back is coordinates in a learned space: no built-in label like "roof condition: fair," just numbers, with sites that look visually alike landing near each other for whatever reasons the model found useful. That's the tradeoff: you give up interpretability of any single dimension in exchange for not having to enumerate every attribute you might eventually want.

For a portfolio risk quant, the practical upside is testing speed. Dropping a site vector into your existing scoring model as an extra feature block takes a join, keyed to your site list. If it adds lift once blended with your existing rating factors, you keep it. If it doesn't, you've spent an afternoon, not a quarter and a vendor contract.

Where each approach fits

Hand-engineered attributes still win when you need something defensible attribute by attribute, where you have to show a regulator or a treaty partner exactly what "defensible space score" means and how it was computed. If the attribute is already a rating factor with its own actuarial history, keep measuring it the way you always have.

Learned embeddings fit the earlier, messier part of the work: you suspect imagery carries signal for a peril or sub-peril you haven't built a feature for yet, and you want to find out whether it's worth the investment before committing to defining and tagging anything by hand. Representation learning is cheap to test and hard to explain; feature engineering is the reverse. Most books end up using both: embeddings for exploration and for the long tail of signal you'd never get around to hand-coding, named attributes for the handful of factors that have to stand on their own in a rate filing.

The underlying bet with an embedding approach is that imagery-derived features, built without manual tagging, correlate with the same loss drivers your named attributes were trying to approximate in the first place. You can't prove that from first principles; you test whether it actually holds against your own loss history, the same way you'd test any new rating variable.

If you want to see what a site-level vector looks like before it touches your model, Imagery Embeddings lays out the delivery format and cadence without asking you to commit to a pipeline first.

Worth a look if your next renewal cycle has a peril you don't have a feature for yet.