Studio Matrx Monthly · Volume 1 · Issue 4 · September 2026
Amogh N P
 In loving memory of Amogh N P — Architect · Designer · Visionary 
Urban Data SourcesLesson 6.1
Generative & Parametric Urbanism/Module 6 · Data & Analysis

Lesson 6.1 · Data & Analysis

Urban Data Sources

Every map of a city is also a map of what its makers chose to count - and the fabric that houses the most people is often the fabric counted least

12 min Interactive lessonFree · open lessonByAmogh N P· Architect & interior designer
The hook

Every dataset of a city is a decision about what counts - and someone is always left off the map.

Data-driven urbanism sounds neutral, as if the city simply hands over the numbers. It does not. Every dataset is manufactured: someone decided what to survey, how to categorise it, what to drop as noise, and who to leave off the schedule. A satellite sees roofs but not who lives under them; a census counts households on its list but not the ones it never visited; a mobility trace records the phones that could afford data plans. The map is never the territory, and in a city the gap between them is where power lives.

This lesson is a working tour of the main sources - GIS and cadastre, OpenStreetMap, census, movement and mobility data, remote sensing, and open civic data - and an honest account of what each genuinely reveals and what it structurally cannot see. The through-line matters especially in India: a very large share of the city is informal and organic, and it slips through almost every source at once. Before you build a single model, you have to know what your evidence base is blind to - because a generative masterplan can only ever design for the city its data can see.

Six sources, each a lens that sees some things and misses others. Stack them and the blind spots line up on the informal city. Data = a made record, not the city. Ask what it cannot see; ground-truth it; the binding choice stays democratic.

The sources

The main sources and what each genuinely reveals

Start with the raw materials of data-driven urbanism, because each one has a native strength you should use and a native blindness you must remember. GIS and the cadastre are the official spatial record - plots, roads, land-use zones, ownership, infrastructure - held by the planning authority or revenue department. It is the backbone of any urban model: authoritative, structured, legally consequential. But it records the *formal* city, the one that has been surveyed and titled, and it is often out of date and patchy exactly where change is fastest.

OpenStreetMap is a crowd-sourced world map - streets, buildings, points of interest, mapped by volunteers. It is astonishingly rich, free, and in some places better than any official source; it is also uneven, because coverage follows the mappers. Well-mapped neighbourhoods tend to be wealthier, more central, more visited. A quiet informal settlement may be a near-blank on OSM not because nothing is there but because no one mapped it.

Census data counts people and households - population, age, occupation, amenities - at coarse spatial units, on a slow cycle (India's decennial census). It is the bedrock of who lives where, but it is a snapshot years old by the time you use it, it aggregates to protect privacy so fine grain is lost, and it systematically misses the mobile, the homeless and residents of settlements not on the enumeration schedule.

Movement and mobility data - mobile-phone location, GPS traces, transit-card taps, ride-hailing logs - reveal how the city actually moves, at a temporal resolution no census can match. They are among the most powerful new sources and among the most biased: they see the people who carry smartphones, hold bank-linked cards and appear in commercial datasets, and under-see everyone else. Remote sensing (satellite and aerial imagery) reveals built form, growth, green cover and heat from above, cheaply and everywhere - but it sees roofs, not lives, and cannot tell a thriving lane from a slum earmarked for clearance. Open and civic data - budgets, permits, complaints, sensor feeds - can be gold, where it exists, is published, and is clean enough to use.

URBAN DATA SOURCES - WHAT EACH LAYER SEESGIS / cadastreplots, roads, zoning - the official recordOpenStreetMapcrowd-mapped streets, POIs - uneven coverageCensuspeople, households - a slow, coarse snapshotMobility / movementphone, GPS, ticketing - who moves, and gapsRemote sensingsatellite, aerial - built form from aboveOpen / civic databudgets, permits, complaints - partial and messyTHE BLIND SPOT: the informal cityNo clear cadastre, off the census schedule, thin on OSM, invisible toticketing - so the fabric that houses the most people is counted least.Every source has biases and gaps. Ask what it cannot see before youtrust what it shows. Data is a partial record, not the city itself.
Zoom
Six urban data sources as layers, each with what it reveals - and the informal city as the blind spot they share, undercounted by design.
The blind spot

The informal city, undercounted by design

Now the honest core of the lesson. Line the sources up and a pattern appears: they all tend to miss the same place, and it is the place that houses the most people. The informal city - the dense old fabric, the unauthorised colonies, the vast settlements that shelter hundreds of millions across India and the global South - falls into the gap in almost every dataset at once. It has no clean cadastre because it was never formally surveyed or titled. It is thin on OpenStreetMap because fewer volunteers map it. It is undercounted in the census because enumeration schedules and addresses do not reach it well. It is under-seen in mobility data because its residents are less likely to appear in the commercial and smartphone streams those datasets draw on. From a satellite it may read as an undifferentiated grey mass, legible as 'to be cleared' but not as a living community.

This is not a random error you can average away; it is a systematic bias that runs the same direction every time, and its direction is against the poor. A model trained or built on this evidence base inherits the blindness. A generative masterplan optimizing 'the city' is really optimizing the city its data can see - and the fabric it cannot see is precisely the fabric most at risk of being planned over, because in the model it barely exists. The danger is not merely inaccuracy; it is that invisibility in the data becomes erasure on the ground, dressed as objectivity.

The discipline this demands is simple to state and hard to practise: for every source, ask *who and what is missing, and which way does the gap run?* Treat data as a partial, made record with an author and an agenda, not as the city itself. Where a source is silent, that silence is information - usually about who did not get counted. And when the evidence base cannot see a community, the humane response is to go and look, to bring in local and participatory knowledge, and to refuse to let the map's blank spaces license decisions about the people living in them.

FROM RAW SIGNAL TO A CLAIM ABOUT THE CITYraw signalphones, sensors->cleaned datachoices baked in->a metrica number->a claim aboutthe citycontestableBias and gaps enter at every arrow: who is sampled, what is dropped asnoise, which category the messy world is forced into. By the last box theassumptions are invisible and the number looks like fact.Trace the chain back before you trust the claim.
Zoom
Bias accumulates from raw signal to cleaned data to metric to confident claim; by the last step the assumptions are invisible and the number looks like fact.
Reading data

How bias enters, from raw signal to confident claim

Bias in urban data is rarely a single villain; it accumulates quietly along a chain, and by the end it is invisible. Follow a mobility dataset from signal to claim. The raw signal is a stream of phone pings - already biased, because it only exists for people with phones and data plans. Cleaning drops points flagged as errors, snaps traces to known roads and fills gaps with assumptions; each choice is reasonable and each bakes in a worldview about what a 'normal' trip looks like. The cleaned data becomes a metric - say, average commute time - which compresses a messy human reality into one number. That number becomes a claim: 'the city's commute is twenty minutes', quoted in a plan as fact. By the last step the phones that were never in the sample, the trips snapped onto the wrong road, and the assumptions in the fill are all gone from view. The number looks objective precisely because its manufacture is hidden.

Three biases recur and are worth naming. Coverage bias: who is in the dataset at all (smartphone owners, titled property, mapped streets) skews toward the wealthier and more formal. Categorisation bias: the messy real world is forced into tidy boxes, and whatever does not fit a box - a shop-house that is home, workshop and store at once, a street that is also a market and a living room - is mangled or dropped. Temporal bias: data ages, and a census or land-use map years old describes a city that has already moved on, especially where change is fastest.

The practical response is not to reject data - data-driven urbanism is genuinely powerful and here to stay - but to read it like a critic. Trace every confident number back up the chain to the signal it came from. Ask what was sampled, what was dropped, and what was assumed. Cross-check sources against each other and against your own feet on the ground. Report uncertainty and gaps as first-class results, not footnotes. Data can sharpen judgement wonderfully; it cannot replace it, and the moment a number stops being questioned is the moment it starts to mislead.

FROM RAW SIGNAL TO A CLAIM ABOUT THE CITYraw signalphones, sensors->cleaned datachoices baked in->a metrica number->a claim aboutthe citycontestableBias and gaps enter at every arrow: who is sampled, what is dropped asnoise, which category the messy world is forced into. By the last box theassumptions are invisible and the number looks like fact.Trace the chain back before you trust the claim.
Zoom
Bias accumulates from raw signal to cleaned data to metric to confident claim; by the last step the assumptions are invisible and the number looks like fact.
In practice

Building an evidence base you can trust

Put this to work as a practical method. A trustworthy urban evidence base is not the biggest pile of data; it is a triangulated one, where independent sources check each other and their gaps are mapped as carefully as their contents. Combine the official cadastre (formal structure) with OpenStreetMap (fine-grained texture), census (who lives here), mobility (how they move), remote sensing (built form and change) and civic data (how the city is run and complained about) - and treat the places where they disagree as the most interesting places, not errors to smooth over. A settlement that is dense on satellite imagery, blank on the cadastre and thin on OSM is telling you exactly where the formal record fails and the people are.

Add the source no dataset contains: ground truth and local knowledge. Walk the area. Talk to residents, local officials and community organisations. Bring participatory mapping in, so the people who live in the data's blind spots can put themselves on the map. This is not softness; it is rigour - it is how you catch the systematic errors that no amount of computation will surface, because the computer only knows what it was given. In the Indian context, where informality is the norm rather than the exception across much of the urban fabric, this ground-truthing is not optional; it is the difference between an evidence base that serves the whole city and one that quietly designs for half of it.

Finally, keep the boundary from the course's first lesson. Data informs; it does not decide. An evidence base helps you understand a place, argue from something firmer than opinion, and open options for public debate - but the binding decisions about land use, displacement and a city's future belong to the planning authority, the statutory and participatory process, the affected communities and the governing law and development-control regulations, in India the master-plan process, the applicable DCR and the National Building Code of India. The most honest thing a good dataset can do is show you clearly, and in public, both what it knows and what it cannot see.

URBAN DATA SOURCES - WHAT EACH LAYER SEESGIS / cadastreplots, roads, zoning - the official recordOpenStreetMapcrowd-mapped streets, POIs - uneven coverageCensuspeople, households - a slow, coarse snapshotMobility / movementphone, GPS, ticketing - who moves, and gapsRemote sensingsatellite, aerial - built form from aboveOpen / civic databudgets, permits, complaints - partial and messyTHE BLIND SPOT: the informal cityNo clear cadastre, off the census schedule, thin on OSM, invisible toticketing - so the fabric that houses the most people is counted least.Every source has biases and gaps. Ask what it cannot see before youtrust what it shows. Data is a partial record, not the city itself.
Zoom
Six urban data sources as layers, each with what it reveals - and the informal city as the blind spot they share, undercounted by design.
Verify-this: data is a partial record; the binding urban choices stay human and democratic

Every source has a blind spot

What each dataset cannot see

GIS = formal only; OSM follows the mappers; census is coarse and slow; mobility sees smartphones; remote sensing sees roofs. Ask of each: who and what is missing? Modules 1.2, 6.1.

The informal city is undercounted

A systematic bias, not random noise

The sources converge on the same blind spot - the informal fabric that houses the most people - and the gap runs against the poor. Ground-truth and participatory mapping are rigour, not softness. Modules 6.1, 9.4.

Triangulate and report gaps

How to build a trustworthy base

Cross-check independent sources, treat disagreements as findings, and publish uncertainty and blind spots as first-class results. Data informs; it never decides. Modules 6.4, 1.2.

The binding choice is democratic

Who decides on land and displacement

Land-use, displacement and equity decisions belong to the planning authority, the participatory process, the communities and the law - in India the master-plan process, the DCR and NBC India - never to 'the data'. Modules 7.2, 7.3.

Hands-on workshop

Workshop — audit the blind spots of one place

Pick a place you know well and interrogate what the standard data sources would and would not see of it. The aim is to feel, concretely, how the evidence base can be confident and wrong at the same time - and where its silences fall.

Free web maps (satellite plus OpenStreetMap) and a notebook. No GIS software needed - this workshop is about reading data critically by eye; the analysis tools come later, and the binding urban decisions always stay with the planning authority, the community and the democratic process.

Given & goal
Goal: see a real place through, and past, its data
Inputs: a neighbourhood you know + free web maps (a satellite view and OpenStreetMap) + a notebook
Time: ~45 minutes
  1. 1Choose the place: a neighbourhood, market or settlement you know on the ground, ideally one with some informal or fine-grained fabric.
  2. 2Look at it in OpenStreetMap and in a satellite view: note what is mapped in detail, what is a near-blank, and where the two disagree.
  3. 3For each source (cadastre, OSM, census, mobility, satellite), write one line on what it would show of this place and one line on what it would miss.
  4. 4Name the systematic gap: list the people, uses or activity that every source would tend to undercount here, and note which way the bias runs.
  5. 5Write a short reflection: if a generative masterplan were built only on this data, what would it get wrong about this place, and how would you bring the missing reality in - flagged as reasoning, with the binding decisions left to the process and the community.

You’ll walk away with
A one-page data audit of a real place: a per-source table of what is seen and missed, the systematic blind spot named with its direction, and a reflection on what a data-only model would erase. Keep it as a template for reading any urban evidence base.

The worked example

Three altitudes on the same idea

Read the band that fits you — or all three.

For the architect / urban designerUsing computation to explore, analyse and test urban form - while people and the democratic process decide

For the architect or urban designer, data is the raw material of every computational move you make - and its biases become your design's biases unless you interrogate them first. Assemble a triangulated base: cadastre for the formal skeleton, OpenStreetMap for texture, census for population, mobility for movement, remote sensing for built form and change. Then, for each layer, write down what it cannot see - and notice that the answers converge on the informal city. Do not let a blank on the map read as empty ground; a generative model will happily optimize over a settlement it cannot see. Ground-truth by walking the site and bringing in participatory and local knowledge. Use data to explore and to argue, defer the binding land-use and displacement decisions to the planning authority, the participatory process, the affected communities and the governing law, and always publish what your evidence base is blind to alongside what it shows.

For the planner / urbanistWhere computational methods genuinely help planning and where the city's human and political life resists them

For the planner or urbanist, the evidence base is where objectivity is most claimed and most dangerous, because a biased dataset can dress a political choice as a technical fact. GIS, census, mobility and remote sensing genuinely strengthen how you understand a place and argue from evidence - but their gaps run in a consistent direction, against the informal, the mobile and the poor. A commute figure, a density map or a 'vacant land' classification can each quietly erase the people the data never counted. Treat every source as a made, partial record; triangulate; and report gaps and uncertainty as first-class findings that open debate rather than close it. Bring participatory mapping and community knowledge into the base so the undercounted can appear. Keep the binding decisions with the statutory process, the communities and the law - and never let 'the data shows' stand in for a choice about who the city is for.

For the studentHow cities can be grown by rule - and why a city is a living system, not an optimization problem

Learn to read a dataset the way you would read a source in history: ask who made it, why, what they counted and who they left out. The main urban sources - GIS and cadastre, OpenStreetMap, census, mobility, remote sensing, open data - each reveal something real and hide something else, and their blind spots line up on the same place: the informal city that houses the most people. That is not random noise; it is a systematic bias that runs against the poor, and a model built on it inherits the blindness. The skill to build is not collecting more data but questioning it - tracing a confident number back to the messy signal it came from, cross-checking sources, and treating silence in the data as information about who was not counted. Data-driven urbanism is powerful and worth mastering; it is trustworthy only in the hands of someone who knows exactly what their evidence cannot see.

Misconception check

More data means a more objective, more accurate picture of the city. If we just gather enough sources - satellite, mobile-phone traces, sensors, open data - we can finally see the whole city as it really is and plan from facts instead of opinion, removing bias from urban decisions.

This inverts how urban data actually works. More data does not converge on an objective city; it compounds a set of made, partial records, each with its own author, categories and blind spots. Every source is manufactured: someone decided what to survey, how to classify it, what to drop as noise and who to leave off the schedule - and those choices, not the city itself, shape the numbers. Worse, the biases are not random errors that cancel out with volume; they are systematic and they run the same direction. GIS records the formal, titled city; OpenStreetMap coverage follows wealthier, more-mapped areas; the census misses the mobile and the unscheduled; mobility data sees smartphone owners and card holders; satellites see roofs, not lives. Stack them and they converge on the same blind spot - the informal city, the dense old fabric and the settlements that house hundreds of millions - so gathering more of these sources can make the same people invisible more confidently, not less. And invisibility in the data is dangerous: a model can only design for the city its data can see, so an undercounted community reads as empty ground to be planned over, with the erasure gilded as objectivity. The honest stance is the opposite of data-maximalism: treat every dataset as a partial record, ask for each one who and what is missing and which way the gap runs, triangulate independent sources and treat their disagreements as the most informative findings, add ground truth and participatory local knowledge that no dataset contains, and report gaps and uncertainty as first-class results. Data-driven urbanism is genuinely powerful for exploring and understanding a city - but only for someone who knows precisely what their evidence base cannot see, and who keeps the binding decisions with the people and the democratic process rather than with 'the data'.
Try it

Do it yourself

No software needed — reason it through.

  1. 1List the main urban data sources and give, for each, one thing it reveals and one thing it structurally cannot see.
  2. 2Why is the informal city undercounted across almost every source at once, and why is that a systematic bias rather than random error?
  3. 3Trace how bias accumulates from a raw mobile-phone signal to a confident claim like 'the city's commute is twenty minutes'.
  4. 4What does it mean to triangulate an evidence base, and why are the disagreements between sources the most informative part?
  5. 5Why is a blank space on the map dangerous rather than neutral, and what is the humane response to it?
Take this with you

The one line to carry out

Every urban dataset - GIS, OpenStreetMap, census, mobility, remote sensing, open data - is a made, partial record with an author and blind spots, and the blind spots line up on the same place: the informal city that houses the most people, undercounted by design and against the poor - so treat data as evidence to be interrogated, not the city itself, triangulate and ground-truth it, report what it cannot see as a first-class result, and keep the binding decisions about land and displacement with the communities and the democratic process, never with 'the data'.
Take it further
References & further reading

Peer-reviewed journals & authoritative standards

  1. 01Geographic information systemWikipedia — Geographic information system, 2026.
  2. 02OpenStreetMapWikipedia — OpenStreetMap, 2026.
  3. 03Open dataWikipedia — Open data, 2026.
  4. 04Big dataWikipedia — Big data, 2026.
  5. 05Informal settlementWikipedia — Informal settlement, 2026.
Related lessons
Recap
Data-driven urbanism runs on a handful of sources, and mastering it means knowing exactly what each one cannot see. GIS and the cadastre give the authoritative but formal-only spatial record; OpenStreetMap gives rich, free, but uneven crowd-sourced texture whose coverage follows the mappers; the census gives who-lives-where at a coarse, slow, privacy-aggregated grain that misses the mobile and unscheduled; mobility data gives movement at fine temporal resolution but sees only smartphone owners and card holders; remote sensing gives built form and change from above but sees roofs, not lives; open and civic data can be gold where it exists and is clean. The honest core is that these sources tend to miss the same place - the informal city, the dense old fabric and the settlements that house hundreds of millions - which has no clean cadastre, is thin on OSM, is undercounted in the census, is under-seen in mobility streams and reads as grey mass from a satellite. That is a systematic bias, not random noise, and it runs against the poor; a model built on this base inherits the blindness, so invisibility in the data becomes erasure on the ground, gilded as objectivity. Bias accumulates quietly from raw signal to cleaned data to metric to confident claim, through coverage, categorisation and temporal bias, until the number looks like fact. The response is not less data but critical reading: trace numbers back to their signal, triangulate independent sources and treat their disagreements as findings, add ground truth and participatory local knowledge, and report gaps and uncertainty openly - while keeping the binding decisions with the planning authority, the communities and the democratic process.
Carry forward →

Data describes a city as points, lines and counts - but a street network is more than a set of lines; its very shape channels movement and life. Next we turn form itself into a network and ask what its structure predicts, with space syntax.

A

The author

Amogh N P

Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.

More about Amogh →