What is a feature vector in a portfolio risk scoring model?

If you build or feed a scoring model across a book, you already work with feature vectors even if nobody on the team calls them that out loud. A feature vector is just the row of numbers a model sees for one risk: TIV, distance to coast, construction class, year built, maybe a wind or flood peril score pulled from a CAT model. Stack those rows for every site in the book and you've got the matrix that goes into your GLM, your gradient boosting model, or whatever scoring approach your team runs.

The word "feature" is doing a lot of work here. Each column is a feature: something somebody decided was worth measuring and could actually get measured across the whole portfolio, not just a handful of flagship risks. That second part is the hard part. A feature that's available for 200 of your 40,000 sites isn't a feature, it's a sample.

Where the features in most books come from today

Most portfolio feature sets are built the slow way. Someone defines an attribute list (roof condition, surrounding land use, proximity to a wildfire interface, whatever the peril committee flags as relevant this renewal cycle), then a team works out how to populate it. For a handful of attributes that means a manual tagging pass over imagery or site photos. For others it means buying a third-party dataset and hoping the match rates against your SOV are decent. Either way, every new attribute is its own small project: define it, source it, QA it, maintain it as the book turns over.

That approach works. It's also why a lot of risk teams have five or six hand-built features per site and a long backlog of ones they'd like to add but haven't gotten to.

What an embedding vector is, and how it differs

An embedding vector is a feature vector too, but nobody defined its columns by hand. It comes out of a model trained to compress an image (or any other dense input) into a fixed-length list of numbers that preserves what's visually distinctive about the site: roofline, density of structures nearby, vegetation, surface materials, the general texture of the built environment. You don't write down "column 47 means tree canopy density." The model learned that structure from the data; the columns don't carry human labels, they carry statistical regularities that happen to be useful.

For a scoring model, that's not a downside. Your gradient boosting model doesn't care whether a feature has a name, it cares whether it has signal. An embedding vector derived from VHR imagery, at 0.5 m resolution or better, captures a lot of the same visual information a human reviewer would use to eyeball a site, just compressed into something you can concatenate onto your existing feature matrix and test like any other input. Drop it in, hold out a validation set, see if it moves your Gini or your lift curve. If it doesn't help for a given peril, drop it back out. No attribute list had to survive a committee meeting first.

Why this matters for feature engineering across a large book

The practical case for embeddings in a reinsurance context isn't that they replace your CAT model outputs or your engineering-grade attributes. It's that they let you test a new risk signal without first deciding what that signal is called. For a book with tens of thousands of sites, the bottleneck usually isn't modeling capacity, it's the feature-extraction pipeline: the manual tagging, the QA, the annual refresh that someone has to own. An embedding vector keyed to your site list sidesteps that pipeline entirely, because the "attribute definition" step never happens. You get a vector per site and you decide, empirically, whether it earns a place in the model.

That's the specific gap Imagery Embeddings is built for: a vector per site derived from imagery, delivered as a file keyed to your book and refreshed annually, meant to sit alongside the features you already maintain rather than replace them.

If you've got a backlog of risk signals you'd like to test but no appetite for building another bespoke pipeline, that's worth a look.