MoldRiskIQ
Methodology

The mold risk model

Mold risk modeling estimates growth conditions from records rather than samples: surface humidity and time drive germination, so the model reads climate, terrain, envelope and documented water history.

The premise

Mold germination is governed by water activity at a surface and by how long that condition persists. It is not governed by room relative humidity, which is a proxy, and it is not governed by spore concentration, which is effectively unlimited in every occupied building.

Common indoor genera — Aspergillus and Penicillium — germinate at a water activity of roughly 0.80, which corresponds to about 80% relative humidity at the surface itself. Stachybotrys requires around 0.90 and therefore a sustained liquid water source rather than incidental humidity. These thresholds are conventions from the building-science literature and are used here as thresholds, not as precise constants.

Time matters as much as the threshold. Brief excursions above 80% surface humidity are ordinary in a bathroom and produce nothing. Sustained ones behind an exterior wall produce a colony. A model estimating from records is therefore estimating persistence: how often does this building put a surface above threshold, and for how long.

Why a room reading is not the variable

Air holds more moisture when it is warm. The surfaces inside a building are not all at air temperature: framing that bridges to the outdoors, the inside face of an uninsulated wall, the floor of a closet on a north elevation. Air touching those surfaces is at a lower temperature and therefore a higher relative humidity than the hygrometer in the middle of the room reports.

This is why a house at a comfortable 45% can have growth behind a wardrobe on an exterior wall, and why interventions aimed at room humidity sometimes do nothing. The model reasons about the conditions that produce cold surfaces — envelope, era, thermal bridging, below-grade exposure — rather than about the reported room figure.

Input class one: regional moisture load

Twenty of the model's hundred points. Four components: average humidity carrying eight, annual rainfall five, flood exposure four and storm exposure three.

This class is predictive because it sets the background pressure every building in a region works against, and because it is dense, long-run and independent of anything being recorded about the property. Its limitation is the mirror of that: it is identical for every property in a postal code, so it can never distinguish two houses on the same street.

Regional figures are seeded per postal code from state-level defaults on first use. Those defaults are planning estimates rather than measurements — adequate as a floor for a single property, and explicitly not the basis for any published ranking.

Input class two: terrain and drainage

Ten points. Elevation relative to the surroundings, slope, and the resulting drainage class, derived from elevation data at the property's coordinates.

This is the class that most often explains why one house is wet and its neighbour is not, and it is measured rather than reported, which makes it unusually reliable for something invisible in a listing. A depression or valley position carries eight of the ten; low-lying ground carries six.

Its limitation is that it reads the ground, not what has been done to it. A property with a working perimeter drain and one with a failed drain present the same terrain.

Input class three: building vulnerability

Twenty-five points. A base figure from construction era, adjusted by the specific materials and details known for the property, then normalised.

The era curve is the part worth arguing with, because it is not monotonic with age. 1990s stock carries the highest base vulnerability in the model, above pre-1940. Older assemblies are leaky, uninsulated and vapour-open, so they dry in both directions; the 1990s tightened envelopes without reliably matching them with mechanical ventilation, during the peak years of barrier EIFS, polybutylene supply lines and complex roof geometry. Drying capacity is half of whether an assembly survives getting wet, and that is the half the era lost.

Specific characteristics are added as weighted factors — barrier EIFS and contaminated imported drywall at the top, then polybutylene, aluminium branch wiring, flat roof drainage, complex flashing geometry, below-grade living space — and each is scaled by how firmly it is established before it is applied.

Input class four: disclosure and event history

Forty-five points, the largest module, because it is the only class that is evidence about this building rather than inference about buildings like it.

It reads state disclosure statements where they can be parsed — Pennsylvania's SPD, New Jersey's SPCDS and Delaware's SDCR — normalising them into risk events with categories and severities. It also reads historical hazard exposure from federal sources.

Documented, unresolved water intrusion sets an evidence floor: a minimum the composite cannot fall below however favourable the other three classes are. Two independent records of the same problem raise that floor further, because corroboration between sources that did not consult each other is the strongest signal available from outside a building.

Evidence weighting

Every factor carries an evidence strength that scales its weight before it is applied. A characteristic confirmed by an inspection counts for a quarter more than one confirmed on paper; one inferred from the building's cohort counts for seven tenths; one merely suspected counts for four tenths.

The effect is that a property with three suspected characteristics does not outscore one with a single confirmed serious one — which is the behaviour you want from a model routinely working with incomplete records.

Provenance is carried through to the report. A construction year from a parcel record or from the owner is treated as verified; one from a census-tract median or a community-mapped footprint is treated as an estimate, and the report says so on the line where it is used.

Limitations, stated first-person

We cannot see current conditions. No input changes if a colony begins growing today. Anything the model says about a building is about its propensity, not its state.

We cannot see undisclosed history. Silence in the record is indistinguishable from an absence of problems, and the two are very different. An investor-owned property discloses almost nothing, lawfully.

We cannot see maintenance. A regraded lot, a replaced roof, a corrected drain — none of these reliably reach a record we read, so a well-maintained property in a vulnerable category will score as its category.

We have not validated against inspection outcomes. There is no measured agreement rate, because the matched dataset does not exist yet. This is stated at greater length on the accuracy page rather than glossed here.

Our disclosure coverage is three states. Outside Pennsylvania, New Jersey and Delaware the largest module is unavailable, and a score there rests on climate, terrain and building data alone.

Data sources

Property and parcel records from ATTOM and Regrid, with OpenStreetMap building footprints and US Census ACS tract estimates as fallbacks. Geocoding via Nominatim and the US Census geocoder.

Terrain from Open-Meteo elevation. Hazard exposure from the FEMA National Risk Index, USGS short-term network high-water marks, and NOAA HURDAT2 storm tracks.

Regional climate profiles seeded per postal code. Where a published dataset is used for a ranking rather than for a single property's fallback, its vintage is stated on the page that publishes it.

Sources

Questions people ask

How is this different from a moisture meter reading?
A meter measures moisture content in one material at one moment, on site. The model estimates how often this building is likely to put surfaces above the germination threshold, from records, without visiting. The meter is evidence; the model is a prior.
What climate framework do you use?
Regional profiles are held per postal code, seeded from state-level defaults. Where a published dataset needs comparable climate across places — the metro index — heating and cooling degree days are derived from thirty years of daily reanalysis rather than from those defaults.
Why is disclosure worth more than construction?
Because it is evidence about this building rather than an inference about its type. A 1990s house with synthetic stucco is a category that fails in a known way. A disclosure describing water in the basement in 2019 is a fact.
Is the model published?
The module weights, the era curve, the characteristic weights and the evidence multipliers are all stated here and on the driver-ranking page. The implementation is not open source.