Explaining a black-box imagery feature to model governance
Model governance has signed off on opaque inputs before. Underwriting models have used credit scores and catastrophe model output as black-box features for years, because reviewers could document where the number came from, how stable it was, and what it did to model performance. That's the actual bar for a learned imagery vector too: a documentation trail that survives a validator asking "what is this and how do you know it's doing what you think."
What the reviewer is actually asking
Strip away the SR 11-7 language and a model risk reviewer wants four things from any new feature, embeddings included.
First, provenance: what data went into it, at what resolution, how often it's refreshed. For a site-level imagery embedding this is usually straightforward to state plainly. VHR imagery at 0.5m or better, RGB, one vector per site, refreshed annually. No hidden sensor fusion, no undisclosed third-party enrichment.
Second, what the dimensions represent. An embedding dimension doesn't map to "roof condition" or "vegetation encroachment" the way a hand-built feature would; the vector captures visual structure the model found useful for distinguishing sites, and you don't get a clean label per dimension. Say that plainly instead of assigning each coordinate a story you haven't tested. Reviewers have seen plenty of credit and fraud models where individual coefficients aren't individually meaningful either; the standard is conceptual soundness at the feature-set level, not interpretability of every single coordinate.
Third, stability. Does the feature mean the same thing this year as it did last year? With an annual refresh, this is a real question worth answering before anyone asks it. A new imagery batch processed through the same pipeline should produce vectors that sit in a comparable space to prior years, but you should check that directly, not assume it. Track correlation between a site's new-year vector and its prior-year vector for sites that haven't visibly changed. If that correlation drifts, your downstream model's calibration drifts with it, and governance will want to see that you're watching for it.
Fourth, outcomes evidence. Does including the embedding improve out-of-sample loss prediction versus the model without it? Run the comparison the same way you'd run it for any other candidate feature: holdout set, same splits, same metric you already report to governance. Improving that metric is what earns the feature its place on the model, regardless of whether anyone on the committee can narrate what dimension six is doing.
Building the validation case without overclaiming
A few things make the write-up easier and more credible.
Don't describe the embedding as capturing specific risk attributes unless you've tested for that and can show it. If you haven't run a probe showing dimension clusters correlate with, say, flood exposure or construction density, don't write that they do. Describe what you've verified: the feature improves predictive performance on your holdout, it's stable year over year within a tolerance you define, and the imagery source and cadence are documented.
Treat it like any other vendor-supplied feature with an opaque construction method, because that's what it is. You already have a process for third-party catastrophe model scores and credit-bureau attributes where the underlying model is proprietary. Point your governance process at that precedent rather than inventing a new interpretability standard for imagery specifically.
Keep the monitoring plan simple and recurring: re-run the stability check at each annual refresh, re-run the outcomes comparison when you retrain the downstream model, and flag if a meaningful share of your book shows large year-over-year vector shifts that aren't explained by visible site changes. That's the ongoing-monitoring section most reviewers are looking for, and it's achievable without needing to open up the embedding's internals.
None of this requires the imagery vendor to hand you a feature dictionary. It requires you to treat the embedding the way you'd treat any purchased signal: documented provenance, tested stability, measured lift. If you're scoping a site-level vector that plugs into an existing risk model as an additional feature rather than a replacement for one, that's the version of this that's easiest to get through committee: one input among many, judged on the same outcomes test as the rest.
If a workable embedding feature for your book is still a few steps off, start by seeing what a sample vector looks like against your own site list.