Knowledge Base

Data Sourcing & Methodology

Where does the data come from?
How does BuildZoom Data collect its data?

Gryd by BuildZoom Data collects data from publicly available government sources, including building departments, municipalities, counties, and contractor licensing agencies. The data is then standardized, enriched, and organized into a consistent national dataset.

Is the data public?

The underlying records are generally public government records. Gryd by BuildZoom Data adds value through aggregation, standardization, enrichment, contractor matching, and ongoing maintenance.

How much geographic coverage does Gryd by BuildZoom Data have?

Gryd by BuildZoom Data has data from all 50 states across the country. We cover over 2500 jurisdictions, which are made up of cities, counties and metropolitan areas.

How much historical data is available?

Our database has over 400M building permits. Historical coverage spans decades in many jurisdictions, with records dating back over 25 years, enabling customers to analyze long-term construction trends and contractor activity.

How often is the data updated?

Gryd by BuildZoom Data receives new data daily and most data becomes available near real-time, depending on when data becomes available from source agencies. We're refreshing over 1.3M records every week and ingesting new data daily.

Data Pipeline

Gryd by BuildZoom Data follows a five-stage data pipeline to transform raw public records into a clean, enriched, analytics-ready dataset:

Stage
Description
1. Identification
A research team maintains and prioritizes a living index of public data sources: state and license boards, building and planning departments, and county clerk offices.
2. Extraction
A variety of methodologies are applied to extract data from each source, accounting for variation in format, cadence, and access method across thousands of jurisdictions.
3. Transformation
Extracted data is stored in a nested JSON format in Postgres, then structured and schematized via a combination of human interpretation and machine-learning methods.
4. Mapping
Data is mapped into a construction knowledge graph, enabling cross-entity relationship building (e.g. linking a permit to the licensed contractor who pulled it).
5. Enrichment
A bundle of ML models develops second-order data points such as entity type classifications and project categories. Data is then aggregated for usage.
Data Quality
How accurate is the data?

Gryd uses automated processing, enrichment, and quality assurance processes to maintain high-quality construction intelligence datasets.

Why might records differ from local government records?

Differences can occur because of source updates, permit amendments, publication timing, or agency-specific reporting practices.