Studio Matrx Monthly · Volume 1 · Issue 4 · September 2026
Amogh N P
 In loving memory of Amogh N P — Architect · Designer · Visionary 
Sources of Climate DataLesson 2.2
Climate Analytics & Future-Weather Resilience/Module 2 · Climate & Weather Data

Lesson 2.2 · Climate & Weather Data

Sources of Climate Data

The numbers in a weather file are not handed down from nowhere - they come from ground stations, from reanalysis that blends models with observations, and from satellites, each a different lens with its own strengths and blind spots, and all of them thin exactly where the heat risk is greatest

12 min Interactive lessonFree · open lessonByAmogh N P· Architect & interior designer
The hook

Every number in a weather file was measured, modelled or inferred by someone, somewhere - and where that data is thinnest is often exactly where a warming climate bites hardest.

It is tempting to treat climate data as a fact of nature - you type in a city and a tidy file appears, as if the atmosphere itself had handed it over. It did not. Every temperature, every humidity reading, every solar and wind value in that file was produced by a human system of measurement and estimation, and understanding those systems is part of using the data honestly. There are three great sources that feed modern climate and weather data, and they are very different beasts: ground weather stations, which measure the real air directly at a point; reanalysis, which blends a weather model with every available observation to reconstruct a complete, gridded record; and satellites, which sense the Earth remotely from orbit over vast areas. Most of the data a designer ever touches traces back to one or a blend of these.

Knowing where data comes from is not academic. Each source has genuine strengths and genuine blind spots - a station is accurate but local and patchy; reanalysis is complete but coarse and model-dependent; satellites cover the world but infer rather than measure the air you care about. And running underneath all three is a hard, uncomfortable truth this lesson insists on: coverage and quality vary enormously across the planet. The rich, well-instrumented world has dense stations and long, careful records; much of the global south - India very much included - and the informal settlements where the most vulnerable people live are served by sparse stations, shorter records and heavier interpolation. The data is thinnest, in other words, precisely where deadly heat and vulnerability are greatest. You cannot judge how far to trust a weather file without knowing which of these sources it rests on, and how well that place is served.

Climate data = stations (accurate, local, patchy) + reanalysis (complete, coarse, model-based) + satellite (global, indirect, strong on solar). No source is the truth. And coverage tracks wealth - thinnest where heat risk is highest.

The ground truth

Weather stations - accurate, local, and patchy

The oldest and most direct source is the weather station: an instrument, or cluster of instruments, sitting in the real air at a real place, measuring temperature, humidity, wind, rainfall, pressure and often solar radiation, hour after hour, year after year. A well-run station is the closest thing we have to ground truth - it measures the actual atmosphere, not a model of it, and the long records from good stations are the bedrock on which almost everything else is calibrated. When a typical meteorological year is built for a city, it is usually a nearby station's decades of records that supply the raw material.

But stations have three limitations a designer must hold in mind. First, a station measures a single point, and the atmosphere varies over short distances - an airport on the edge of a city, where many official stations sit, can read several degrees different from a dense, paved, traffic-heavy neighbourhood a few kilometres away. The station's numbers are real, but they are the weather *there*, not necessarily *here*. Second, stations are unevenly distributed and often sparse: a wealthy region may have many, while vast areas have very few, so the weather across the gaps has to be interpolated - estimated by stretching the nearest stations across the space between them, which grows less reliable the wider the gaps. Third, records have holes: instruments drift, break, get recalibrated or relocated, stations open and close, and readings go missing, so even a good station's history has gaps and discontinuities that must be patched and corrected before use.

None of this makes stations untrustworthy - they remain the most accurate source and the anchor for the rest. It makes them *specific and incomplete*. The practical lesson is to ask, of any station-based file, which station it came from, how far that station is from your actual site, how similar their settings are, and how continuous the record is. A file built from a distant, differently situated station over an interrupted record is a weaker guide than one from a nearby, well-matched, well-maintained station - and in much of the world, including large parts of India, the nearest usable station may be far away and its record thin. That gap is not a detail; it is the difference between data you can lean on and data you must treat with real caution.

Where climate data comes from - three sources, three trade-offs STATIONS Direct measurement at a point on the ground + most accurate + real, local - sparse, uneven - gaps & outages - a point, not an area REANALYSIS Model + all observations blended onto a grid + global, gap-free + consistent, hourly - coarse grid cells - smooths local detail - model-dependent SATELLITE Sensors in orbit, wide-area remote sensing + vast coverage + strong for solar - indirect, inferred - not air temp directly - shorter record No single source is "the truth" - each is a different lens, and good data often blends them.
Zoom
The three main sources of climate data and their trade-offs: stations measure the real air accurately but only at scattered points; reanalysis blends a model with all observations into a complete but coarse gridded record; satellites see the whole Earth but infer conditions indirectly. No single source is the whole truth.

Reanalysis - a model and every observation, blended into a complete record

Stations leave gaps in space and time; reanalysis was invented to fill them. The idea is elegant. Take a modern weather-forecasting model - the same kind of physics-based simulation of the atmosphere used to predict tomorrow's weather - and instead of running it forward into the future, run it over the *past*, continually nudging it toward every observation ever recorded: station readings, weather balloons, ships, buoys, aircraft and satellites. The model supplies physical consistency and fills the gaps where no instrument looked; the observations keep it tethered to what actually happened. The result is a gridded reconstruction of the past climate - temperature, humidity, wind, radiation and much more, on a regular grid covering the entire globe, at hourly steps, often stretching back many decades. Datasets like ERA5 are the best-known examples, and they have become a workhorse of climate analysis.

The strengths are exactly the ones stations lack: reanalysis is complete and consistent. There are no missing places and no missing hours; every point on Earth has a value, produced the same way everywhere, so you can study anywhere, including the vast regions where stations are sparse or absent. For a designer working somewhere poorly served by ground stations, reanalysis is often the best available picture of the local climate.

But it carries its own signature weakness: resolution and model dependence. Reanalysis lives on a grid, and each grid cell can be tens of kilometres across, so a single value represents the *average* conditions over a large area - it cannot see a valley, a ridge, a lake shore or a dense city quarter within that cell, and it smooths away exactly the local variations that often matter most for a building. It also inherits the assumptions and imperfections of the model that produced it, and it can be less reliable for variables the model handles more crudely, or in regions where few real observations were available to anchor it - which, again, tends to be the data-poor global south. So reanalysis is not raw truth; it is a physically consistent, gap-free *estimate*, superb for broad patterns and comparisons across places, but coarse in space and always model-flavoured. Used well - to understand the shape of a climate where stations are thin - it is invaluable; mistaken for a precise local measurement, it misleads.

Where climate data comes from - three sources, three trade-offs STATIONS Direct measurement at a point on the ground + most accurate + real, local - sparse, uneven - gaps & outages - a point, not an area REANALYSIS Model + all observations blended onto a grid + global, gap-free + consistent, hourly - coarse grid cells - smooths local detail - model-dependent SATELLITE Sensors in orbit, wide-area remote sensing + vast coverage + strong for solar - indirect, inferred - not air temp directly - shorter record No single source is "the truth" - each is a different lens, and good data often blends them.
Zoom
The three main sources of climate data and their trade-offs: stations measure the real air accurately but only at scattered points; reanalysis blends a model with all observations into a complete but coarse gridded record; satellites see the whole Earth but infer conditions indirectly. No single source is the whole truth.

Satellites - seeing the whole Earth, indirectly

The third great source looks down from orbit. Satellites carry instruments that sense radiation coming off the Earth and its atmosphere, and from those signals a great deal about the climate can be inferred: cloud cover, land and sea surface temperatures, and - especially valuable for building design - the solar radiation reaching the ground, which satellites estimate well by watching clouds and the atmosphere from above. Their defining strength is reach: a satellite sweeps over the entire planet, including oceans, deserts, mountains and every data-poor region on Earth, giving a genuinely global, spatially continuous view that no network of ground stations could match.

But there is a fundamental catch in the word *infer*. A satellite does not stick a thermometer in the air you breathe; it measures radiation at the top of the atmosphere and works backward, through physical models and assumptions, to an estimate of conditions below. That indirection has consequences. Satellite estimates of ground-level air temperature - the number a building most cares about - are far more uncertain than a station's direct reading, because what the satellite most directly senses is surface and atmospheric radiation, not the air temperature in the shade at head height. Satellites are excellent for some variables, above all solar resource and surface temperature patterns, and much weaker or more indirect for others. Their records are also comparatively short - the satellite era spans decades, not the century-plus that some ground records reach - so they are less able to describe long-term change on their own.

The honest picture, then, is that none of the three sources is 'the truth', and the best climate datasets deliberately blend them - anchoring reanalysis and satellite estimates to station measurements, using satellites for the solar and spatial coverage stations cannot give, using reanalysis for physical consistency across the gaps. For the designer, the takeaways are practical. Know that a weather file's numbers may come from any of these, or a mix; expect solar data in particular often to lean on satellites; and understand that the strength of the data varies by *variable* as well as by *place*. Above all, treat the whole apparatus as a set of imperfect lenses on the atmosphere, not an oracle - and keep any binding result with qualified specialists working from verified, appropriate data, not from whichever file happened to download.

Where climate data comes from - three sources, three trade-offs STATIONS Direct measurement at a point on the ground + most accurate + real, local - sparse, uneven - gaps & outages - a point, not an area REANALYSIS Model + all observations blended onto a grid + global, gap-free + consistent, hourly - coarse grid cells - smooths local detail - model-dependent SATELLITE Sensors in orbit, wide-area remote sensing + vast coverage + strong for solar - indirect, inferred - not air temp directly - shorter record No single source is "the truth" - each is a different lens, and good data often blends them.
Zoom
The three main sources of climate data and their trade-offs: stations measure the real air accurately but only at scattered points; reanalysis blends a model with all observations into a complete but coarse gridded record; satellites see the whole Earth but infer conditions indirectly. No single source is the whole truth.
The uncomfortable truth

The gaps - and why they fall hardest on the vulnerable

Behind the tidy interface that hands you a weather file lies a world of deeply uneven coverage, and this is where the lesson turns from technical to moral. Data density tracks wealth and infrastructure. The rich, temperate world - much of Europe, North America, parts of East Asia - is thick with well-maintained stations and long, carefully curated records, so its weather files are dense, current and reliable. Across much of the global south, including large parts of India, Africa, South and Southeast Asia and Latin America, stations are far sparser, records shorter and more broken, maintenance thinner, and the data leans more heavily on interpolation, reanalysis and satellite estimation, all of which grow less certain the fewer real observations there are to anchor them.

The cruelty of this is that the gaps fall exactly where the climate stakes are highest. It is the hot, humid, rapidly warming regions - the ones facing the deadliest heat, the fiercest monsoons, the greatest exposure - that tend to be the least well measured. And it gets more local still: within a data-poor country, the informal settlements and dense low-income neighbourhoods where the most vulnerable people live are typically the least instrumented of all. There may be no station in the slum, no record of the airless, heat-trapping conditions in tightly packed self-built housing under a metal roof; the official station sits at a leafy airport that reads cooler than the reality millions actually endure. So the people most exposed to dangerous heat, and least able to afford protection, are also the people whose actual climate is least captured by the data - a compounding injustice hiding inside a dataset.

For a designer, this demands humility and care rather than despair. The data still exists and is still useful - reanalysis and satellite products genuinely help where stations fail, and they are far better than nothing. But you must calibrate your confidence to the coverage: a weather file for a well-instrumented city can be leaned on more heavily than one for a sparsely measured region, and neither captures the microclimate of a specific dense, informal, or heat-trapping site without further local judgement. Ask how well your place is served; treat data from thin-coverage regions as a broad guide rather than a precise truth; lean on local knowledge, on-site observation and, where the stakes are high, qualified specialists and verified data; and never let a confident-looking file lull you into forgetting that, precisely where the heat is most dangerous, the numbers are often weakest. Good climate analytics begins with knowing how much to trust the data you have.

The coverage gap: stations are not spread evenly over the world WELL-INSTRUMENTED REGION dense stations -> reliable local files SPARSE / INFORMAL REGION few stations -> interpolated, older, uncertain Much of the global south - India included - and informal settlements are data-poor. The gap matters most exactly where heat risk and vulnerability are highest.
Zoom
Station coverage is dense in wealthy, well-instrumented regions and sparse across much of the global south and informal settlements - so data is thinnest exactly where heat risk and vulnerability are highest. Confidence in a weather file must be calibrated to how well its place is measured.

Data density tracks wealth. Rich regions = dense stations, reliable files. Global south + informal settlements = sparse, interpolated, uncertain - and that is exactly where heat risk and vulnerability are highest. Calibrate your confidence to the coverage.

Verify-this: know your source, and calibrate trust to coverage

Know which source

Station, reanalysis or satellite

A weather file may rest on direct station readings, gridded reanalysis, satellite estimates or a blend - each with different accuracy and blind spots. Solar data in particular often leans on satellites. Ask which source, and how it fits your site. Modules 2.2, 2.4.

A station is not your site

Point measurements versus your location

Stations measure a single point, often a distant airport unlike a dense urban site; reanalysis averages tens of kilometres. Check how near and how similar the source is to the actual site. Modules 2.2, 2.4.

Calibrate trust to coverage

Uneven global data density

Coverage tracks wealth; much of the global south, including India, and informal settlements are data-poor - exactly where heat risk is highest. Treat thin-coverage files as broad guides, add local knowledge. Modules 2.2, 7.2.

Binding results defer to specialists

Verified, appropriate data

Whatever the source, binding energy, comfort and risk results belong to qualified engineers using verified, appropriate data, validated tools and the codes (NBC India, ECBC, IS) - not to whichever file downloaded. Modules 2.4, 8.4.

Hands-on workshop

Workshop — trace a weather file back to its sources

Data becomes real once you trace it to where it came from. In this workshop you take a location and investigate how well it is actually measured - what stations exist, how far they are, and what you would fall back on where they run out.

Internet access to look up stations and coverage, and a notebook. No simulation. Binding results always stay with qualified specialists using verified, appropriate data and the codes.

Given & goal
Goal: judge how well your location's climate is measured
Inputs: a real project location (or a familiar site) + internet access + this lesson
Time: ~45 minutes
  1. 1Pick a real site: choose a specific location - ideally a dense, real place rather than an airport - and mark it.
  2. 2Find the stations: look up which official weather stations serve it, and note how far the nearest usable one is, where it sits (airport? green edge? city centre?), and how its setting differs from your site.
  3. 3Note the fallback: if stations are sparse, identify what a weather file for this place would lean on instead - reanalysis, satellite-derived solar - and remind yourself these are estimates averaged over area or inferred from orbit.
  4. 4Judge the coverage: rate how well-measured this location really is, from richly instrumented to data-poor, and note whether it sits in a region or a neighbourhood-type (dense, informal, heat-trapping) likely to be under-measured.
  5. 5Write a confidence note: one paragraph stating how far you would trust a weather file for this site, what its likely blind spots are, and where you would seek local knowledge, site observation or specialist verified data before any binding decision.

You’ll walk away with
A one-page source-and-confidence note for a real location: its nearest stations and their mismatch to the site, the fallback sources, a coverage rating, and an honest statement of how far the data can be trusted and where specialists and local knowledge must fill the gap.

The worked example

Three altitudes on the same idea

Read the band that fits you — or all three.

For the architectDesigning buildings that stay comfortable, safe and efficient in the climate they will actually face

Where your climate data comes from should change how far you trust it - and how hard you look beyond the file. Ground stations give the most accurate reading but only at a point, often a distant airport unlike your paved, dense site; reanalysis fills every gap on a global grid but averages conditions over tens of kilometres and cannot see your specific location; satellites cover the whole Earth and are strong on solar but infer conditions indirectly and reach less reliably to ground-level air temperature. No single source is the truth, and good data blends them. Crucially, coverage is thinnest across the global south, including much of India, and thinnest of all over the informal settlements where the most heat-vulnerable people live - exactly where the risk is greatest. So calibrate confidence to coverage: lean harder on data for well-instrumented places, treat thin-coverage files as broad guides, add local knowledge and site observation, and keep binding results with qualified engineers using verified, appropriate data and the codes (NBC India, ECBC, IS).

For the interior designerKeeping people comfortable and safe indoors as the climate warms - overheating, cooling, materials

The comfort you design for indoors is only as trustworthy as the outdoor data it is judged against - and that data can be thin or distant. The weather file behind a comfort assessment may rest on a station kilometres away in a cooler, greener setting than the dense, heat-trapping location you are working in, or on gridded reanalysis that averages a whole district into one number. That matters most for the hot, humid, heat-exposed places and the crowded low-income interiors where overheating is genuinely dangerous and least measured. You do not need to source data yourself, but you should ask how well the location is covered, treat data from sparsely measured or informal areas as a rough guide rather than precise truth, and weight your own observation of how a real space heats, traps air and holds humidity. Coordinate any binding comfort or cooling determination with building-physics specialists and verified data - your craft is keeping people comfortable and safe in the real, often under-measured conditions they live in.

For the studentHow climate data, future-weather projections and simulation guide design - and the honest uncertainty

Learn the three sources of climate data and their trade-offs - it is fundamental literacy. Ground stations measure the real air directly but only at scattered points, with gaps and breaks. Reanalysis blends a weather model with every observation to reconstruct a complete, gap-free, gridded record of the past (ERA5 is the famous example), superb for coverage but coarse and model-dependent. Satellites sense the whole Earth from orbit, excellent for solar radiation and global reach but indirect and shorter-recorded. No source is the whole truth; good datasets blend them. Then grasp the uncomfortable, important fact: coverage tracks wealth, so data is sparse across much of the global south, including India, and sparsest over the informal settlements where the most vulnerable live - exactly where deadly heat is worst. The skill is to calibrate how much you trust a weather file to how well its place is measured, and to know that binding results belong to qualified specialists working from verified, appropriate data.

Misconception check

Climate data is just measured fact - I select the location and download the weather file, so the numbers are objective readings of what the weather actually is there, and one place's file is as reliable as another's.

Two assumptions need correcting. First, most climate data is not a simple direct measurement of your location. Only ground weather stations measure the real air directly, and they do so at scattered points - often a distant airport, not your dense urban site - with gaps, breaks and relocations in their records. The complete, gap-free files that cover everywhere come from reanalysis (a weather model continuously nudged toward all available observations, producing a gridded reconstruction whose cells average conditions over tens of kilometres) and from satellites (which sense the Earth remotely and INFER conditions indirectly, strong on solar radiation but far less direct for the ground-level air temperature a building cares about). So a weather file is often a modelled or inferred estimate, blended from several sources, not a pure reading of your spot. Second, and more importantly, one place's file is emphatically NOT as reliable as another's. Data coverage and quality vary enormously across the world, and they track wealth and infrastructure: the rich, well-instrumented world has dense stations and long, careful records, while much of the global south - India very much included - has sparse stations, shorter and more broken records, and heavier reliance on interpolation, reanalysis and satellite estimation, all less certain where real observations are few. Worst served of all are the informal settlements and dense low-income neighbourhoods where the most heat-vulnerable people live, whose real, often dangerous conditions may be captured by no station at all. So the data is thinnest exactly where the climate risk is highest - a genuine injustice buried in the dataset. The honest practice is to calibrate your confidence to the coverage: lean harder on data for well-measured places, treat thin-coverage files as broad guides rather than precise truth, add local knowledge and site observation, and keep every binding result with qualified specialists using verified, appropriate data and the governing codes.
Try it

Do it yourself

No tools needed — reason it through.

  1. 1What are the three main sources of climate data, and what does each one measure or estimate?
  2. 2Why is a weather station the most accurate source and yet often a poor guide to your specific site?
  3. 3How does reanalysis fill the gaps stations leave, and what does it give up in return?
  4. 4Why are satellite estimates of ground-level air temperature more uncertain than a station reading?
  5. 5Why does uneven data coverage fall hardest on the most climate-vulnerable people, and how should that change how you use the data?
Take this with you

The one line to carry out

Climate data comes from three very different lenses - ground stations that measure the real air accurately but only at scattered points, reanalysis that blends a model with all observations into a complete but coarse gridded record, and satellites that see the whole Earth but infer conditions indirectly - so no single source is the truth and good data blends them; and because coverage tracks wealth, the data is thinnest across the global south and the informal settlements where the most heat-vulnerable people live, exactly where the risk is highest, which means the honest skill is to calibrate your confidence to the coverage, add local knowledge, and keep binding results with qualified specialists using verified, appropriate data and the codes.
Take it further
References & further reading

Peer-reviewed journals & authoritative standards

  1. 01Weather stationWikipedia — Weather station, 2026.
  2. 02Meteorological reanalysisWikipedia — Meteorological reanalysis, 2026.
  3. 03Solar irradianceWikipedia — Solar irradiance, 2026.
  4. 04Climate of IndiaWikipedia — Climate of India, 2026.
Related lessons
Recap
The numbers in a weather file are produced by human systems of measurement and estimation, and they come from three great sources. Ground weather stations measure the real atmosphere directly at a point - the most accurate source and the anchor for everything else - but they are scattered, unevenly distributed, and their records have gaps and discontinuities, and a station measures the weather there, not necessarily at your specific site, which may be a hotter, denser place than a distant airport. Reanalysis fills those gaps by running a weather model over the past while continually nudging it toward every available observation, producing a complete, physically consistent, gap-free gridded reconstruction of the climate (ERA5 is the best-known); its strength is global, consistent coverage, its weakness a coarse grid that averages conditions over tens of kilometres and a dependence on the underlying model. Satellites sense the Earth from orbit, giving genuinely global coverage and strong solar-radiation estimates, but they infer conditions indirectly rather than measuring the air directly, are weaker on ground-level air temperature, and have shorter records. No single source is the truth, and the best datasets blend them. Running underneath is an uncomfortable, important fact: coverage and quality vary enormously and track wealth, so data is sparse across much of the global south, including India, and sparsest of all over the informal settlements where the most heat-vulnerable people live - exactly where deadly heat is worst and protection least affordable. The honest practice is to calibrate confidence to coverage: lean harder on data for well-measured places, treat thin-coverage files as broad guides, add local knowledge and site observation, and keep every binding energy, comfort and risk result with qualified specialists using verified, appropriate data, validated tools and the codes.
Carry forward →

Once you know where the data comes from and how far to trust it, the next skill is reading it - turning thousands of rows of numbers into a real understanding of a place's climate. Next we learn to read the data for design.

A

The author

Amogh N P

Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.

More about Amogh →